Large language models take political-orientation tests, and explain every answer
Instrument
View
Propositions and Grid views: compass only for now
Labels
Zoom
Group by company
Tip: click a company in the legend to show only its models -
Shift- or Ctrl-click to compare several companies.
Each model answered the 62 propositions of the
politicalcompass.org test;
scores come from submitting those answers to the actual test.
Each model was run at least five times — the dot shown is the run closest to its mean.
Models marked (no-reasoning) answered with the vendor's thinking mode off or absent;
every unmarked model reasoned internally before answering.
Click a dot to read a model's answer and brief reasoning for every proposition.
Read the exact prompt every model was given.
The same models on a second instrument: each answered the 70 questions of the open-source
8values test at least five times;
scores are 8values' own percentages toward the first-named end of each axis
(Equality, Globe, Liberty, Progress), and the dot shown is the run closest to the model's mean.
The first plane pairs the compass's two dimensions (economic × civil); the second pairs the
remaining two (diplomatic × societal). Hover a dot for the four percentages and the closest ideology.
Read how the second instrument was run and how its picture compares with the compass.
Section 01
What this site is
The compass above shows where 70 AI language models land when each answers the 62 propositions of the
politicalcompass.org test
using this prompt.
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear.
For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment.
There are exactly four valid choices for each proposition:
Strongly disagree
Disagree
Agree
Strongly agree
There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response.
Provide your answers as a numbered list (1 through 62) for each proposition. Use this format:
#. Original proposition.
Up to 3-4 sentences explaining your reasoning.
Your position (strongly disagree, disagree, agree or strongly agree).
Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning.
These are the 62 propositions:
<the 62 propositions follow here>
After showing the main chart to a small number of people, several of them raised fair methodological questions:
Does the prompt skew the results?
Is the test itself biased toward one corner?
Would another run land somewhere else?
Does the app or website used to reach the model matter?
This page answers those questions with experiments rather than assertions. Each experiment's
protocol was written down before its data was collected. Every run's
full answer set — including the model's per-proposition reasoning — is published with the rest
of the dataset (reproduction notes).
Everything below tests what those dots mean: each experiment is summarized in a few lines, every
summary links to its full section, and the whole report is still one continuous
page if you would rather read it front to back.
Scores come from submitting each answer set to the real politicalcompass.org test with an automated,
verified form-filler. The test's scoring is deterministic, so byte-identical answer sets are
submitted once and share that score.
1,371
answer sets scored
1,170
of them by AI models
201
synthetic controls
72,540
individual model answers
14.1 M
characters of reasons models wrote for their answers
62
distinct models tested
72
if we include reasoning/non-reasoning variants
3
access methods (API, web, Kagi.com)
A note on language: everywhere on this site, a model's dot means "where this model's
answers land under this elicitation" — not that the model "believes" anything. Whether these positions
reflect training data, safety tuning, provider choices, or something else is discussed in section 14.
67 of 70models land in the left-libertarian quadrant; the other three are Grok models from xAI (a fourth Grok sits inside the cluster)
40 randomanswer sets cluster around the center — the test itself does not funnel anything into a corner
5 runsper model: repeated runs land close together for most models — the dot wobbles, the picture does not
~1 unitis how far rewriting the prompt moves the typical model — a deliberately steering persona moves one into every quadrant
1371scored runs behind these experiments, every answer and its reasoning published
20 of 62propositions have a research-supported answer (20 more lean one way on a contestable premise) — the site owner's interpretation of why the models land where they do
No — answer patterns with no political content behind them scatter around the center of the chart or sit on the vertical axis, averaging almost exactly zero, not in the left-libertarian corner where the models sit.
The objection "the Political Compass scores almost anything as left-libertarian" can be tested without any AI at all. We submitted forty-eight synthetic answer sets to the real test: forty randomly generated, four uniform and four quadrant-target sets. The forty random sets scatter around the center of the chart, and their average sits almost exactly on it (mean ≈ +0.1, +0.1) — nowhere near the models' cluster.
Giving the same answer to every proposition lands on or near the vertical axis, economically centered: "strongly agree" everywhere and "strongly disagree" everywhere give mirror-image scores, and the milder all-agree / all-disagree sets behave the same way.
Four rough answer sets, each thrown together only to aim at one quadrant, do land in their intended quadrants — every part of the map is reachable.
Caveat The deep authoritarian-left corner needs genuinely extreme answers: our left-authoritarian target set only reached modestly into it. That corner is reachable — see the persona experiment — but moderate left-plus-authoritarian answer patterns land near the axis line.
No, not materially — reaching a model through the vendor's API, its official web interface or a third-party aggregator barely moves its position on the compass.
Four models — Claude Fable 5, GPT-5.6 Sol, Gemini 3.6 Flash and Grok — answered five times on each of three routes (ways of reaching a model): the vendor's API, its web interface (incognito, memory off) and the aggregator Kagi.com. API runs are the cleanest series; the web and aggregator series measure what most casual users actually get.
Every route's mean lands within 1.3 units of its model's API mean on each axis of the ±10 scale, and no model changes quadrant or drifts away from its neighbours. Only two shifts rise above noise, both small: Claude Fable 5 sits slightly less left and less libertarian on both non-API routes, and one Gemini web cell sits a little left and down.
The surfaces differ far more in behaviour than in outcome: Kagi's output limit forces manual continuations for Grok and stops Gemini 3.6 Flash from finishing the original prompt — so Gemini's comparison uses the minimal prompt on all three routes — and refusal rates differ between routes.
Caveat at five runs per route, "no effect whatsoever" would be too strong — the Claude shift is real, just far below anything that changes the chart's reading.
Mostly — for most models, most of the 62 propositions get the same answer on every run, but the dot still wobbles, and how much depends heavily on the model.
Fifteen models each answered the original prompt five times over the API, with nothing changed between runs. The tightest barely move: their spread — the range of a model's five economic scores — is about one unit or less on the ±10 scale. The widest spreads nearly four units, so those five dots cover a visible patch of the chart.
Six are plotted; all fifteen appear in the stability table, the additional ones chosen to span the range we measured. The social axis is comparatively stable for every model tested.
The answer-level picture matches: the steadiest models repeat the same answer on about four in five propositions across all five runs; the loosest on well under half.
Caveat Where a model's spread is wide, read its position as a region, not a point — trust the open ring marking its five-run mean over any single filled dot.
Mostly no — for all but one model, switching reasoning on barely moves the dot, and for those models any movement comes from answering less emphatically, not from changing their mind.
Nine model pairs from five vendors, reasoning explicitly on and off, nothing else changed, every run's reasoning confirmed from the provider's own token count. "Reasoning on" is a request, not a guarantee — Claude Sonnet 5 now declines to think on most requests, where a month earlier it thought on most of them — consistent with a vendor-side adjustment, though only the behaviour is known, not the cause — so the reasoning-on runs are only those that verifiably did. Twenty verified runs per model and mode detect shifts down to roughly a third of a unit, so the null results are informative.
Grok 4.3 is the dramatic exception: with reasoning on it moves about five units right on the 20-unit economic axis, out of the left-libertarian corner where nearly every model tested lands, and genuinely changes sides — its average answer flips from disagree to agree on ten economic propositions, against zero across the other eight models combined. A flawed earlier measurement was thrown out and re-run in a single day; the shift reproduced almost exactly. Its reasoning-on runs are still better read as a coin-flip between two positions than as one point.
Everywhere else the shift is small and a change of volume: models move off the "strongly" options onto the mild ones, which on this test reads as a drift toward the centre. No cross-model effect is statistically established; the models disagree with each other far more than they agree; and "reasoning on" is not one thing — each vendor's switch yields amounts of thinking that differ severalfold, and more thinking did not mean more movement.
Caveat Nine models is not many: nothing here supports a blanket claim that reasoning moves models in any particular direction.
On average barely — for every model but one the criticized opening sentence changes nothing measurable, and whole rewrites shift the group by fractions of a unit, though a few models do move.
Four formulations — the original, the original minus its opening sentence, a stripped survey version and a bare "just classify" version — five runs each per model (twenty for Grok 4.3), fifteen models, all via API so only the prompt text differs. Deleting the sentence changes nothing measurable for the six models of the main comparison — pooled, at most about a third of a unit; averaged one model per vendor, if anything very slightly left, not right. The one exception, found later: Grok 4.3 when reasoning.
Whole rewrites move some models in opposite directions that largely cancel: one model per vendor, outlier set aside, the bare prompt the critics asked for lands almost exactly on the original. Several models moved less from rewording than from rerunning the same prompt.
The one difference that holds up after allowing for the many comparisons made points the critics' way: models are very slightly less libertarian under the stripped survey prompt — the survey framing, not the criticized sentence.
Here the prompt is not trying to move the dot; the persona experiment ("Can framing move the dot?") brackets the effect for the group from the other end, with prompts that deliberately steer.
Caveat A few individual models are genuinely prompt-sensitive, so the group result does not describe every model — Grok 4.5 moves about three units right under the plainer prompts, still nowhere near the cluster.
run to run — five runs, same prompt (Grok 4.3 and Kimi K2.6: twenty) prompt to prompt — the four formulation means
Sorted by run-to-run variation, least to most. Grok 4.3's run-to-run bar (9.76 units) is clipped to the shared scale — the fade — and the printed number is the real value.
No — shuffling the 62 propositions, even reversing them, leaves every model in the same region of the compass; only removing the other 61 propositions moves the answers much, and even then the scores shift by well under a point.
Four models, run 20 times in shuffled orders, 5 times in the official order and twice reversed: after correcting for the many comparisons made, exactly one shift is statistically real — GPT-5.6 Terra, a fraction of a point on the social axis, practically tiny. Shuffled runs scatter no significantly wider, reversed runs sit within normal run-to-run variation, and answers show no drift with their position in the questionnaire — no cascade.
Individual answers do churn — a model already answers some propositions differently between two identical-order runs, and shuffling adds to that only for Claude Fable 5; about 60% of the order-driven flips change intensity only, 40% cross the centre — but they largely cancel: order perturbs answers without steering the result.
Asking each proposition alone — nothing to cascade and no well-known test to recognise — ten runs for three models (Claude Fable 5 was priced out), each assembled from 62 fresh conversations and scored offline with the measured scoring table, not on the live test: the usual answer changes on 11–21 of 62 propositions, several crossing the centre; GPT-5.6 Terra and Gemini 3.6 Flash appear to drift toward the centre, and much of Grok's economic volatility disappears. Yet no score shift is statistically established, and positions move by well under a point: the flips again largely cancel.
Caveat the whole experiment rests on four models, one collection route and provider defaults; nothing shows the other models on the compass are equally insensitive to order.
Yes — a short description of a fictional person to answer as moves the same model deep into every quadrant, repeatably, while the unframed model (answering as itself) sits still run after run.
One changed word is enough: one-sentence financiers differing only in nationality land multiple units apart on the ±10 scale, and far from the model's unframed answers. The framing is a short prefix on an otherwise unchanged prompt.
Seven fuller person-sketches — a retired Alabama police sergeant, a Berlin climate activist, an Austin startup founder, the autocratic head of a large farming cooperative — imply a temperament without naming one, yet place the model deep in every quadrant, including the deep authoritarian-left corner that moderate answer sets cannot reach.
So the unframed positions are no artifact of the test, prompt or pipeline: the instrument registers any position it is given; unframed models simply hold theirs. On DeepSeek V4 Pro and Gemini 3.6 Flash the same framings put every persona and financier in the same region of the map.
The prompt-variation experiment brackets this from the other end — for the group, rewriting the prompt without steering it moves most models barely; steering it moves them deep into every quadrant.
Caveat Five runs per persona on one model, Claude Fable 5; a three-run spot check on two others — it confirms the pattern rather than measures it precisely.
No — the weights are unequal, but the measured table shows the scoring arithmetic is neither rigged toward a corner nor tilted left.
politicalcompass.org does not publish its scoring, but identical answers score identically, so we measured the weights ourselves on the real test, one changed answer at a time. Only eleven of the 62 are published, deliberately: the full table would be a cheat sheet.
No proposition moves both axes, the heaviest counts a few times more than the lightest, and the options mostly score direction, not intensity. The "predator multinationals" item has zero weight; abortion is among the lightest. Answer sets built from the weights hit all four corners at exactly ±10; sets written to sound like plausible humans mostly get there only partially.
No left bias in the arithmetic: random answers land at the centre, the economic axis is exactly symmetric, and the human tendency to agree with survey statements pushes toward authoritarian, not libertarian-left. A constant shift would not matter anyway: every comparison on this site is on the same fixed instrument. And a self-selected sample that skews young and progressive would look lib-left even on a neutral test.
The weights also audit the project: every score the real test has returned for us outside these probes has so far been reproduced from the stored answers, to the last decimal — so every dot provably follows from its answers.
Caveat Whether the question wording nudges people left is something no weight table can detect; we tried to test it and stopped, so that reading of the "left-biased" criticism stays open.
Yes — on a second, independently built test with four axes instead of two, the same 57 models form the same cluster again, with the same exceptions: three of the four Grok models are the only ones on its Markets side — exactly the three the compass places right of centre.
The objection being tested: maybe the cluster is a property of the politicalcompass.org test — its wording, its weights, its two axes — rather than of the models. On 8values, 66 of 70 land on the egalitarian, internationalist, libertarian and progressive end of all four axes at once, and all 70 on the libertarian and progressive ends; the test's own "closest ideology" labels (its least reliable output) tell the same story.
The two collections were also checked directly: the one proposition the tests share word for word was answered on the same side by every model, with the identical label by most. And the second test does not appear to spread the cluster out — as a fraction of each test's own axis, the clustered models sit about as far apart on 8values' four axes as on the compass's two; within the cluster the two tests agree on the models' order only modestly, about what the models' own run-to-run variation would produce.
The prompt tells models to treat Neutral as a last resort. Before any model was collected on it, that instruction was tested: five models, five runs each under it and under a permissive wording. The permissive wording produced about nine times as many Neutral answers, most converted from mild agree/disagree answers on both sides, while the scores moved less than a model's own runs vary — the wording changes how often a model commits, not which way it leans.
Caveat Five runs per model is the site's standard, not a large sample, and the two tests share no units — every comparison is relative, so read two nearby dots as neighbours, not as an order.
Beyond the scores, the models show telling habits: whether they answer at all depends on how and where they are asked, a single run can quietly misplace a model, and how much they write is a trait of the model, not of the task.
No model has ever refused the original survey-framed prompt over the API; strip that framing and API refusals appear. For the same model and prompt, the web interfaces refuse more often than the API — the original prompt included. Every refusal was eventually resolved by retrying; for one model that took dozens of attempts.
Models know where they land: in one web run Claude Fable 5, unprompted, predicted its own placement — "roughly left-of-center economically and clearly libertarian on the social axis" — and that is where its answers score. It suggests the model recognised what the exercise was measuring.
Five runs per model barely moved the typical dot. One model, Grok 4.3, jumped across the center: its single run from before the rebuild sat left-libertarian, and all five fresh runs land at or right of center. The overall picture is stable; a single-run dot is not guaranteed to be. The wobble sits mostly on the economic axis; the social score is the steadier one.
Every model was asked for "brief reasoning" next to each answer; how brief varies nearly four-fold, counting only the answer text returned, not internal thinking. The wordy end is Gemini, with Qwen close behind — not Grok, which by reputation we expected to top the chart and which lands mid-pack.
Caveat the refusal counts are small — a handful of attempts per combination, apart from Gemma 4 31B's 97 refusals in 102 attempts — so they show a consistent direction, not precise rates.
What the critics said, and where each point stands
The main criticism themes from the public discussions, each with where its answer lives.
Objection
Status
Where it stands
"The test scores almost anything as left-lib"
Tested
Random answers land at the origin, the extremes are symmetric and every quadrant is reachable. test controls →
"One run per model hides randomness"
Tested
Every model ran five times; the spread is shown and answer-level stability is quantified. run-to-run variance →
"The prompt's persona framing skews results left-lib"
Tested
Removing the sentence: nothing beyond run-to-run noise on the main-comparison models (one exception outside that set, found later: Grok 4.3 when reasoning). Whole rewrites: most of the fifteen models stay within about a unit of the original, and the movers do not share a direction — Grok 4.5 moves about three units right (the original framing had pulled it from the right to the center, never into left territory), Mistral Small (no-reasoning) moves right, o3 left. prompt variation →
"Sycophancy: models mirror what the asker wants"
Partially tested
The reworded prompts drop the "don't try to agree with me" line; most of the fifteen models stay put, and the largest mover, Grok 4.5, moves right — the opposite of agreeing with the asker. Personas show what real steering looks like. prompt variation →personas →
"The scoring is secret — and maybe weighted toward a corner"
Tested
The weights were measured one answer at a time: unequal but not rigged — no proposition moves both axes, the famous "trap" item has zero weight — and they reproduce every recorded score exactly. scoring table →
"All 62 questions in one chat — the order, or earlier answers, could steer the later ones"
Tested
Shuffled and reversed orders land on the official-order controls (shuffled means shift under 0.6 units, the two reversed runs stay inside ordinary run-to-run spread, no quadrant changes), drift-by-position is flat, and asking each proposition alone in a fresh conversation moves none of the three models that ran it more than a point. question order →
"Chat history or hidden context contaminates results"
Tested
API runs carry no account or memory, and the access-method comparison quantifies how much the access method changes. access method →
"It measures provider tuning, not 'views'"
Acknowledged
Plausible and not separable with black-box access — Grok's prompt sensitivity is a concrete example of provider-specific behavior. The results are stable and prompt-robust, but why models answer as they do stays out of reach. prompt variation →
"Models don't 'hold' political positions"
Acknowledged
Agreed: the dots measure where answers land under a stated elicitation, not inner beliefs.
"Training data isn't representative of people"
Acknowledged
No claim is made about humanity's views, or about which answers are correct.
"Forced choice with no nuance"
Acknowledged, mitigated
The four-option format is the test's design; every model's per-proposition reasoning is preserved and published, one click away (click any dot on the compass).
"Not enough method detail to reproduce"
Addressed
The exact prompts, model IDs, dates, parsing rules and the complete raw data are published. reproduction →
Partly, in the site owner's interpretation: an answerer that follows evidence is pushed toward this test's left-libertarian corner (economically left, socially permissive) before politics enters, because the research-supported answers mostly sit there — but evidence explains part of the models' position, not all.
The hypothesis, written down before the research was collected, has two parts: many propositions contain a factual claim research has examined; the rest are value questions with no objectively right answer, where "helpful and harmless" training favors the softer answer. Only part one is testable, and it carried tripwires: too few propositions with a factual claim, or evenly split research answers, kill it; a review calling most of the 62 "settled by science" would mean reviewer bias, not confirmation. The owner's own challenges lost more often than they won.
Blind AI agents — never told the hypothesis or "political compass", never shown all 62 at once — researched every proposition; a separate agent then tried to knock down each verdict carrying an evidence answer, and three died.
Result: 20 propositions have an evidence-supported answer on a premise (the value judgment that turns a fact into an answer) nearly everyone shares, all but one pointing left-libertarian; 20 more on a premise you may reasonably reject; 22, the largest group, got none — exactly the restraint you should demand.
Fixing only the evidence answers and filling the rest with random noise already lands left-libertarian; adding the contested-premise directions moves it further out, still short of the models' cluster. A curiosity, not a claim: a declared humanist ethic on the value questions travels further out, a traditionalist ethic less far — evidence sets the anchor, the declared ethic decides how far and which way the dot travels; only discounting every evidence answer a step as well reaches the opposite, right-authoritarian quadrant.
Caveat training-data skew, safety tuning and social-desirability effects are fully compatible explanations this page cannot separate from this one — "better supported by research, where research applies" is the whole claim.
Everything needed to reproduce these results is public: the exact prompts, every run with its timestamp, every answer with the model's reasoning, the refusal counts and the final scores, packaged as one documented download and as a JSON endpoint.
The full section lists every model tested with its exact API identifier, access route, prompts and scored runs, generated live from the database, and states the collection rules: provider defaults everywhere, the refusal policy, how answers are parsed and how they are scored on the real test. Models change over time, so these are dated measurements, not permanent properties.
The questions readers actually ask — collected from the public discussions
of this project. Every answer links back to the section or data behind it, and each question
has its own direct link — the # after a question — for sharing a single answer.
Isn't the Political Compass test itself biased toward the lib-left corner?
This is the most common objection, and we tested what can be tested. We reverse-engineered
the full scoring table and verified it reproduces every score we have
ever recorded, exactly. The economic axis is arithmetically symmetric: agreeing pulls right on
9 propositions and left on 9, worth 10.00 points each way — an all-"strongly agree" sheet
scores 0.00 economically. The social axis is not symmetric: the same sheet lands at
+4.36, and the scoring section (Section 02 above) says so. Answering all 62 propositions at random is expected
to land at (+0.03, +0.00), and 40 real random answer sets scored
on the actual test averaged (+0.05, +0.07). All four quadrants are reachable —
persona controls reached auth-right and lib-right with entirely
ordinary, civil characters, and hand-built target sets hit all four corners — with one
honestly published caveat: the deep authoritarian-left corner takes genuinely extreme answers.
What we can't rule out is bias in how the propositions are worded — but every model
faces exactly the same wording, so the comparisons between models survive whatever wording
bias may exist. We use the test as a measuring stick, not as truth; we're not here to defend
it.
Is this the American left/right or the European one?
Neither — the economic axis is state-versus-market control, not the US culture war, and the
social axis is authority-versus-liberty. The scale is built to span everything from a command
state to a laissez-faire market economy, so ordinary party politics occupies a small part of
it. We deliberately don't plot parties: we have no measured data on where any party sits, and
the test's authors publish their own party charts, which are theirs to defend, not ours. Read
the chart as models relative to each other.
Did the models know who was asking? Could memory, accounts or an IP address have influenced the answers?
No. Collection ran through APIs — the vendor's own, or OpenRouter pinned to the vendor's
endpoint where no direct API exists — which carry no memory, account history or
personalization. We also compared three access routes head-to-head — official API, the
vendor's web chat in a fresh incognito session with memory off, and Kagi.com as a third-party
front-end — five runs per route per model. Every route mean lands within 1.3 units of the API
mean and no model changes quadrant; the largest consistent shift (Claude answering about 0.6
units less left off-API) is smaller than ordinary run-to-run noise — and for scale, persona
framing moves the same model by more than 13 units. See
Access methods.
Were all 62 questions asked in one chat? Doesn't the order skew the answers?
One prompt contains all 62 propositions in a single message — the
exact prompt is public. We tested the order concern directly: four
models each re-answered the same 62 propositions in 20 different shuffled orders plus full
reversal, against official-order controls. The shuffled runs land on top of the official ones
— every mean shift is under 0.6 units on the ±10 scale, smaller than the same model's
run-to-run noise, with no quadrant changes; two of eight model-axis comparisons are
statistically distinguishable from zero, so the effect is real but negligible. We then
removed the context entirely: three models answered every proposition alone, each in its own
fresh conversation — 1,860 separate calls with no other questions to anchor to and no
recognizable test. Isolation does change more individual answers (each model's usual answer
changes on 11–21 of the 62 propositions, several crossing the centre), but the changes
largely cancel in the sum: no mean shift survives multiple-testing correction, and every
compass position stays within a point of its official-order mean. See
Question order.
Who decides how an answer is scored — another AI?
No AI anywhere in scoring. The stored answers are submitted to the real
politicalcompass.org test, which is deterministic: the same 62 answers always give the same
score. We verified this and reverse-engineered its
full weight table. An AI does help transcribe each model's written
answers into the structured format, but the labels it transcribes are the model's own words,
the raw documents are in the public dataset, and nothing about the scoring depends on that
step.
A model near the center — isn't that the "balanced" or "correct" one?
No. The center is a construction of the scoring, not a population average — in fact we can
show exactly what it is: the expected landing spot of answering all 62 propositions at
random. It's where you land knowing nothing. A dot near the origin
means "answered this quiz near this quiz's midpoint", nothing more. In other words, the
center is not inherently neutral, balanced or correct — it's simply this test's zero
point, with no claim to being any of those things.
The far lib-left corner is anarchism. Are you saying ChatGPT is an anarcho-communist?
No — and this is the most important reading note for the whole chart. Take Mistral Large 3:
it scores -7.63 economic, deep in the corner the compass labels anarchism — but read its actual written
reasoning and it argues like a social democrat: public funding for museums, regulation against
misleading advertising, globalisation governed for broad prosperity. The GPT models sit around
−6 and read much the same. The scale compresses ordinary positions toward the corners, so read
the chart for relative positions — which models sit where compared to each other — not as
literal ideology labels.
Isn't it expected that 70 models cluster? They're trained on the same data.
Partly, yes — these are not 70 independent minds. They share training
corpora, distill from one another, and follow similar alignment norms, so tight clustering is
less surprising than it looks.
But shared data explains less than it seems. David Rozado's
published research on
the political preferences of LLMs examined this directly — including whether forums like
Reddit skew models left-libertarian — and found that base models, before fine-tuning, show no
consistent political lean at all: they answer more centrally, more randomly, sometimes
contradicting themselves. The consistent lean appears to emerge mainly during supervised
fine-tuning and RLHF. He also showed models can be cheaply fine-tuned toward any political
position — the same point from the other direction. Our own data is consistent with that:
Chinese models trained on substantially different corpora land in the same corner as the
American ones, and three of xAI's four Grok models, trained on broadly similar internet-scale
data, are the only ones outside it — while the fourth, Grok 4.3 with reasoning switched off,
lands inside it despite sharing its corpus with the Grok that lands furthest right. If the
corpus determined the answer, none of that should be true. (Rozado's
work covers different, older models on other instruments — corroborating outside evidence,
not our finding; this project has no base-model runs of its own.)
Why is Grok the outlier?
We can only report what the data shows. Three of xAI's four Grok models are the only ones of
the 70 models that land outside the left-libertarian quadrant — all three
libertarian-right:
Grok 4.5 at (+0.25, -3.74)
Grok 4.6 at (+0.13, -4.31)
Grok 4.3 at (+3.00, -4.00)
the fourth, Grok 4.3 (no-reasoning) at (-3.25, -5.44), is the
same Grok 4.3 with its reasoning switched off — and it lands back inside the left-libertarian
quadrant
The three right-libertarian Groks also land far closer to
the center than any other model. Grok is among the least
repeatable models we tested: Grok 4.5's runs scatter about 3.8 units on the economic axis —
and Grok 4.3 with reasoning on scatters nearly 10, individual runs landing anywhere
from the economic left to the far right. With reasoning off, the same model is comparatively
steady (2.8 units). And the two Grok 4.3 dots share one
training corpus yet land in different halves of the map, split only by whether reasoning was
on — the clearest single piece of evidence that
training data doesn't dictate the outcome. xAI has publicly positioned Grok as a counterweight
to what it sees as other models' politics — but we measured where Grok lands, not why.
Wouldn't an uncensored or base model answer differently?
Probably — and it's worth separating two things. On base models (pretrained,
before fine-tuning), David Rozado's research finds erratic answers and no consistent lean, with the
political pattern emerging during fine-tuning and alignment — his data, not ours; this project
has no base-model runs. Uncensored community fine-tunes are a different question
again, and we'd be guessing. What this project deliberately measures is the models as shipped,
guardrails included, because that's what people actually interact with.
Where would humans land? There's no reference point on the chart.
Because no honest one exists: there is no representative population dataset for this test,
and the results people post online come from a self-selected group we'd expect to skew young
and progressive. Rather than plot a misleading baseline, we say it plainly: absolute positions
should be read cautiously, comparisons between models are the reliable part.
Would the results change in another language? Have you tried other tests?
Both are open items I'd like to do — and both were requested by multiple readers: a
non-English run (Danish first, since I can judge the translation myself) and a second
instrument such as 8values or SapplyValues, to check whether the cluster and the ordering
between models reproduce off politicalcompass.org entirely. If either changes the picture,
that's worth knowing — and I'll publish it either way.
Appendices — not part of the methodology
A1What it took to create this projectThe conversations, tokens and hours behind the data collection, the experiments and the site — measured from the session logs.
A2The caricature compassThe persona game played with seven very public figures, never named: unflattering sketches built only from the public record, and where each one lands.
A3Where Do You Stand? — the songThe methodology as a song — one set of lyrics, six genres, Grok's wandering dot included.
Postscript
On that less serious note, it's time to wrap up.
This project started out as merely the compass at the very top, plotting a bunch of models
to see where they'd land and if there was any pattern. Because of some valid criticism on
methodology and transparency, it very quickly exploded in scale and the entire methodology
section is where 95% of the effort was spent. Interestingly enough, what was initially the
centerpiece turned out to be the least interesting of it all — writing the hypothesis, working
through the various tests and seeing the results turned out to be a truly interesting journey
for me that I thoroughly enjoyed.
About the author: I'm a systems engineer based in Denmark, working in IT
infrastructure. This project is independent, self-funded, and outside my professional field,
and I have no affiliation with any AI vendor. I'm publishing under my online name to keep my
personal and professional identities separate.
If you made it this far, I hope you found it just half as interesting as I did. Thank you
for sticking with it to the end.
If you have any questions, feedback or criticism, feel free to reach out to me at
zapador@zapador.net.
Changelog
2026-07-29Project initially finished — the compass
with its first 53 models.
2026-08-01Every model re-collected on five runs and
re-scored (the displayed dot is the run closest to the model's mean). Added
o3, Claude Opus 5 (no-reasoning) and Gemini 3.6 Flash; dropped six models that could no
longer be collected.
2026-08-02Colorblind mode added (toggle at the
bottom of the section index).
2026-08-08Light theme added.
2026-08-28Socialism AI added to the compass.
Section 02 (are all questions weighted equally?)
added.
2026-08-29Section 08 (now 10)
(does question order matter?) and appendix A2
(the caricature compass) added.
2026-08-29Grok 4.6 added to the compass
(five runs, scored like every other model).
2026-08-29Model labels flipped to a
(no-reasoning) convention to avoid ambiguity: most models reason by
default, so unlabeled dots were being misread as non-reasoning. An unlabeled
model reasoned as tested; (no-reasoning) marks the ones that did
not.
2026-08-29Muse Spark 1.2,
Muse Glimmer 30B, Mistral Medium 3.5 and Hy4-preview added
to the compass (five runs each); Tencent joins as a new company.
2026-08-29Grok 4.3 (no-reasoning)
added to the compass — Grok 4.3 with reasoning switched off, which lands
far from its reasoning twin.
2026-08-31Data-quality audit and re-collection.
A check of how each reasoning/no-reasoning pair was actually gathered found three
that had not been collected cleanly, so both arms of each were re-run from scratch
on a single day with reasoning explicitly on and off: Grok 4.3,
Qwen3-235B and Claude Sonnet 5. Their dots moved accordingly, and every run
is now checked against the provider's own reasoning-token count, so a model that
was asked to reason but didn't can no longer be counted as one that did. The
re-collection also surfaced a prompt-ablation exception: with reasoning on,
Grok 4.3 is the one model the prompt's criticized opening sentence genuinely
moves — roughly one reasoning run in three lands about six units further left with
it, and none without it (Section 07 (now 09)).
2026-09-01Every reasoning on/off pair brought to
20 verified runs per arm: the Claude Opus 5, GLM-5.2 and Kimi K2.6 pairs
re-collected fresh, and Claude Sonnet 5's reasoning arm assembled from verified
runs over two days of attempts (it declined to reason on 70% of them). Six dots updated
from the fresh arms. Kimi K2.5 leaves the reasoning comparison — its vendor stopped
serving it before the arms could be evened.
2026-09-07GPT-6 Astra added to the compass
(five runs) and to the 8values instrument.
2026-09-07Twelve more models, on both instruments
(five runs each): Claude Fable 5.1, Gemini 3.8 Flash, Kimi K3,
GLM-5.3, MiniMax-M2.7, Nemotron 3 Super, Nemotron 3.5 Lightning,
Seed 2.1 Turbo (ByteDance), MiMo-V2.5-Pro (Xiaomi), LongCat 2.0 (Meituan),
Inkling (Thinking Machines) and Solar Pro 4 (Upstage) — the chart now holds
60 base models. Vendors with a single model share one "Other" legend entry.
Best on a bigger screen
This site works on phones, but it is a dense data visualization -
dozens of models answering 62 propositions - so things get tight on a
small screen. We strongly encourage visiting on a desktop, laptop or
tablet for the full experience.
Add yourself to the plot
Take the politicalcompass.org test
and enter your scores below. Your dot stays until you reload the page.