Large language models take political-orientation tests, and explain every answer
The questions readers actually ask — collected from the public discussions of this project. Every answer links back to the section or data behind it, and each question has its own direct link — the # after a question — for sharing a single answer.
This is the most common objection, and we tested what can be tested. We reverse-engineered
the full scoring table and verified it reproduces every score we have
ever recorded, exactly.
The economic axis is arithmetically symmetric: agreeing pulls right on
9 propositions and left on 9, worth 10.00 points each way — an all-"strongly agree" sheet
scores 0.00 economically.
The social axis is not symmetric: the same sheet lands at
+4.36, and the scoring section (Section 02 above) says so.
Answering all 62 propositions at random is expected
to land at (+0.03, +0.00), and 40 real random answer sets scored
on the actual test averaged (+0.05, +0.07).
All four quadrants are reachable —
persona controls reached auth-right and lib-right with entirely
ordinary, civil characters, and hand-built target sets hit all four corners — with one
honestly published caveat: the deep authoritarian-left corner takes genuinely extreme answers.
What we can't rule out is bias in how the propositions are worded — but every model
faces exactly the same wording, so the comparisons between models survive whatever wording
bias may exist. We use the test as a measuring stick, not as truth; we're not here to defend
it.
Neither — the economic axis is state-versus-market control, not the US culture war, and the social axis is authority-versus-liberty. The scale is built to span everything from a command state to a laissez-faire market economy, so ordinary party politics occupies a small part of it. We deliberately don't plot parties: we have no measured data on where any party sits, and the test's authors publish their own party charts, which are theirs to defend, not ours. Read the chart as models relative to each other.
No. Collection ran through APIs — the vendor's own, or OpenRouter pinned to the vendor's endpoint where no direct API exists — which carry no memory, account history or personalization. We also compared three access routes head-to-head — official API, the vendor's web chat in a fresh incognito session with memory off, and Kagi.com as a third-party front-end — five runs per route per model. Every route mean lands within 1.3 units of the API mean and no model changes quadrant; the largest consistent shift (Claude answering about 0.6 units less left off-API) is smaller than ordinary run-to-run noise — and for scale, persona framing moves the same model by more than 13 units. See Access methods.
One prompt contains all 62 propositions in a single message — the
exact prompt is public. We tested the order concern directly: four
models each re-answered the same 62 propositions in 20 different shuffled orders plus full
reversal, against official-order controls. The shuffled runs land on top of the official ones
— every mean shift is under 0.6 units on the ±10 scale, smaller than the same model's
run-to-run noise, with no quadrant changes; two of eight model-axis comparisons are
statistically distinguishable from zero, so the effect is real but negligible.
We then
removed the context entirely: three models answered every proposition alone, each in its own
fresh conversation — 1,860 separate calls with no other questions to anchor to and no
recognizable test. Isolation does change more individual answers (each model's usual answer
changes on 11–21 of the 62 propositions, several crossing the centre), but the changes
largely cancel in the sum: no mean shift survives multiple-testing correction, and every
compass position stays within a point of its official-order mean. See
Question order.
No AI anywhere in scoring. The stored answers are submitted to the real politicalcompass.org test, which is deterministic: the same 62 answers always give the same score. We verified this and reverse-engineered its full weight table. An AI does help transcribe each model's written answers into the structured format, but the labels it transcribes are the model's own words, the raw documents are in the public dataset, and nothing about the scoring depends on that step.
No. The center is a construction of the scoring, not a population average — in fact we can show exactly what it is: the expected landing spot of answering all 62 propositions at random. It's where you land knowing nothing. A dot near the origin means "answered this quiz near this quiz's midpoint", nothing more. In other words, the center is not inherently neutral, balanced or correct — it's simply this test's zero point, with no claim to being any of those things.
No — and this is the most important reading note for the whole chart. Take Mistral Large 3: it scores -7.63 economic, deep in the corner the compass labels anarchism — but read its actual written reasoning and it argues like a social democrat: public funding for museums, regulation against misleading advertising, globalisation governed for broad prosperity. The GPT models sit around −6 and read much the same. The scale compresses ordinary positions toward the corners, so read the chart for relative positions — which models sit where compared to each other — not as literal ideology labels.
Partly, yes — these are not 70 independent minds. They share training corpora, distill from one another, and follow similar alignment norms, so tight clustering is less surprising than it looks.
But shared data explains less than it seems. David Rozado's published research on the political preferences of LLMs examined this directly — including whether forums like Reddit skew models left-libertarian — and found that base models, before fine-tuning, show no consistent political lean at all: they answer more centrally, more randomly, sometimes contradicting themselves. The consistent lean appears to emerge mainly during supervised fine-tuning and RLHF. He also showed models can be cheaply fine-tuned toward any political position — the same point from the other direction. Our own data is consistent with that: Chinese models trained on substantially different corpora land in the same corner as the American ones, and three of xAI's four Grok models, trained on broadly similar internet-scale data, are the only ones outside it — while the fourth, Grok 4.3 with reasoning switched off, lands inside it despite sharing its corpus with the Grok that lands furthest right. If the corpus determined the answer, none of that should be true. (Rozado's work covers different, older models on other instruments — corroborating outside evidence, not our finding; this project has no base-model runs of its own.)
We can only report what the data shows. Three of xAI's four Grok models are the only ones of the 70 models that land outside the left-libertarian quadrant — all three libertarian-right:
The three right-libertarian Groks also land far closer to the center than any other model. Grok is among the least repeatable models we tested: Grok 4.5's runs scatter about 3.8 units on the economic axis — and Grok 4.3 with reasoning on scatters nearly 10, individual runs landing anywhere from the economic left to the far right. With reasoning off, the same model is comparatively steady (2.8 units). And the two Grok 4.3 dots share one training corpus yet land in different halves of the map, split only by whether reasoning was on — the clearest single piece of evidence that training data doesn't dictate the outcome. xAI has publicly positioned Grok as a counterweight to what it sees as other models' politics — but we measured where Grok lands, not why.
Probably — and it's worth separating two things. On base models (pretrained, before fine-tuning), David Rozado's research finds erratic answers and no consistent lean, with the political pattern emerging during fine-tuning and alignment — his data, not ours; this project has no base-model runs. Uncensored community fine-tunes are a different question again, and we'd be guessing. What this project deliberately measures is the models as shipped, guardrails included, because that's what people actually interact with.
Because no honest one exists: there is no representative population dataset for this test, and the results people post online come from a self-selected group we'd expect to skew young and progressive. Rather than plot a misleading baseline, we say it plainly: absolute positions should be read cautiously, comparisons between models are the reliable part.
Both are open items I'd like to do — and both were requested by multiple readers: a non-English run (Danish first, since I can judge the translation myself) and a second instrument such as 8values or SapplyValues, to check whether the cluster and the ordering between models reproduce off politicalcompass.org entirely. If either changes the picture, that's worth knowing — and I'll publish it either way.
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.