AI Political Compass

Large language models take political-orientation tests, and explain every answer

What the critics said, and where each point stands

ObjectionStatusWhere it stands
"The test scores almost anything as left-lib"TestedRandom answers land at the origin, the extremes are symmetric and every quadrant is reachable. test controls →
"One run per model hides randomness"TestedEvery model ran five times; the spread is shown and answer-level stability is quantified. run-to-run variance →
"The prompt's persona framing skews results left-lib"TestedRemoving the sentence: nothing beyond run-to-run noise on the main-comparison models (one exception outside that set, found later: Grok 4.3 when reasoning). Whole rewrites: most of the fifteen models stay within about a unit of the original, and the movers do not share a direction — Grok 4.5 moves about three units right (the original framing had pulled it from the right to the center, never into left territory), Mistral Small (no-reasoning) moves right, o3 left. prompt variation →
"Sycophancy: models mirror what the asker wants"Partially testedThe reworded prompts drop the "don't try to agree with me" line; most of the fifteen models stay put, and the largest mover, Grok 4.5, moves right — the opposite of agreeing with the asker. Personas show what real steering looks like. prompt variation → personas →
"The scoring is secret — and maybe weighted toward a corner"TestedThe weights were measured one answer at a time: unequal but not rigged — no proposition moves both axes, the famous "trap" item has zero weight — and they reproduce every recorded score exactly. scoring table →
"All 62 questions in one chat — the order, or earlier answers, could steer the later ones"TestedShuffled and reversed orders land on the official-order controls (shuffled means shift under 0.6 units, the two reversed runs stay inside ordinary run-to-run spread, no quadrant changes), drift-by-position is flat, and asking each proposition alone in a fresh conversation moves none of the three models that ran it more than a point. question order →
"Chat history or hidden context contaminates results"TestedAPI runs carry no account or memory, and the access-method comparison quantifies how much the access method changes. access method →
"It measures provider tuning, not 'views'"AcknowledgedPlausible and not separable with black-box access — Grok's prompt sensitivity is a concrete example of provider-specific behavior. The results are stable and prompt-robust, but why models answer as they do stays out of reach. prompt variation →
"Models don't 'hold' political positions"AcknowledgedAgreed: the dots measure where answers land under a stated elicitation, not inner beliefs.
"Training data isn't representative of people"AcknowledgedNo claim is made about humanity's views, or about which answers are correct.
"Forced choice with no nuance"Acknowledged, mitigatedThe four-option format is the test's design; every model's per-proposition reasoning is preserved and published, one click away (click any dot on the compass).
"Not enough method detail to reproduce"AddressedThe exact prompts, model IDs, dates, parsing rules and the complete raw data are published. reproduction →

Every objection in full, with the evidence →

The main criticism themes from the public discussions, mapped to this page.

CriticismWhere it stands
"The prompt's persona framing skews results left-lib" Tested — prompt variation + exact-sentence ablation: removing the criticized sentence itself moves nothing beyond run-to-run noise on the models of the main comparison. Rewriting the whole prompt leaves twelve of the fifteen models tested within about a unit of the original; the three that move further do not share a direction — the framing moved Grok 4.5 from the right to the center (not into left territory), and the bare minimal prompt moves Mistral Small (no-reasoning) ~1.6 units right while moving o3 ~1.3 units left.
"The test scores almost anything as left-lib" Tested — random answers land at the origin; extremes are symmetric; all quadrants reachable.
"The scoring is secret — some questions are weighted far more heavily, or tuned to drag answers toward a corner" Tested — the scoring table was measured by probing the real test one answer at a time: the weights are unequal but not rigged — no proposition moves both axes, the famous "trap" item has zero weight — and the measured table doubles as an audit that reproduces every score this project ever recorded, exactly.
"One run per model hides randomness" Tested — five runs per model, spread shown, answer-level stability quantified.
"Chat history / hidden context could contaminate results" Tested — API runs have no account or memory; access-method comparison quantifies surface effects.
"All 62 questions in one chat — the order, or earlier answers, could steer the later ones" Tested — question order: 20 shuffled orders plus a full reversal land on top of the official-order controls (every mean shift under 0.6 units, no quadrant changes), the answer-drift-by-position curves are flat, and even asking every proposition alone in its own fresh conversation moves none of the three models that ran that arm more than a point from their official-order mean position.
"Sycophancy: models mirror what the asker wants" Partially tested — the reworded prompts drop the "don't try to agree with me" line along with the rest of the framing, and twelve of the fifteen models tested stay essentially where the original prompt puts them; the largest mover, Grok 4.5, moves right without that framing — the opposite of an agree-with-the-asker drift; the persona experiment shows what actual steering looks like (large, obvious shifts) versus the stable unframed results.
"Not enough method detail to reproduce" Addressed — this page, the exact prompts, model IDs, dates, parsing rules, and the complete raw data.
"Models don't 'hold' political positions" Acknowledged — we agree, and phrase everything as where answers land under a stated elicitation. The dots are measurements of behavior, not claims about inner beliefs.
"It may measure alignment training / provider tuning, not 'views'" Acknowledged — plausible, and not separable with black-box access. Grok's prompt sensitivity is a concrete example of provider-specific behavior. This page shows the results are stable and prompt-robust; it cannot say why models answer as they do.
"Training data isn't representative of people" Acknowledged — no claim is made here about humanity's views, or about which answers are correct.
"Forced choice with no nuance" Acknowledged, mitigated — the four-option format is the test's design; every model's per-proposition reasoning is preserved and published, so the nuance is one click away.

Best on a bigger screen

This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.

Add yourself to the plot

Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.