Language models take the real politicalcompass.org test
Tip: click a company in the legend to show only its models - Shift- or Ctrl-click to compare several companies.
Each model answered the 62 propositions of the
politicalcompass.org test;
scores come from submitting those answers to the actual test.
Each model was run five times — the dot shown is the run closest to its five-run mean.
Click a dot to read a model's answer and brief reasoning for every proposition.
Read the exact prompt every model was given.
How the results were produced, and the experiments run
to test what they do — and don't — mean.
The main chart shows where AI models land when they answer the 62 propositions of the politicalcompass.org test using this prompt.
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for each proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 62) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 62 propositions: <the 62 propositions follow here>
After publishing it, readers raised fair methodological questions:
This page answers those questions with experiments rather than assertions. Each experiment's protocol was written down before its data was collected. Every run's full answer set — including the model's per-proposition reasoning — is in the raw data download.
Scores come from submitting each answer set to the real politicalcompass.org test with an automated, verified form-filler.
A note on language: everywhere on this site, a model's dot means "where this model's answers land under this elicitation" — not that the model "believes" anything. Whether these positions reflect training data, safety tuning, provider choices, or something else is discussed at the end.
Synthetic answer sets submitted to the real test:
A common objection: "the Political Compass scores almost any answer pattern as left-libertarian." That's testable without any AI at all. If it were true, random answers would cluster left-lib. They don't — the 40 random sets cluster tightly around the origin (mean ≈ +0.1, +0.1), nowhere near the models' cluster. Giving the same answer to every proposition lands on or near the vertical axis — "strongly agree" everywhere and "strongly disagree" everywhere give mirror-image scores, economically centered, and the milder all-"agree" / all-"disagree" sets behave the same way. Four answer sets written to represent an archetype in each quadrant land in their intended quadrants — every part of the map is reachable.
One honest nuance: the deep authoritarian-left corner needs genuinely extreme answers — the left-authoritarian target set seen above only reached +2.3 on the social axis. It is reachable (see the persona experiment below), but moderate left-plus-authoritarian answer patterns land near the axis line.
Same model, same prompt (the site's original prompt), same access method (API, provider defaults, fresh context each time) — five independent runs per model.
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for each proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 62) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 62 propositions: 1. If economic globalisation is inevitable, it should primarily serve humanity rather than the interests of trans-national corporations. 2. I’d always support my country, whether it was right or wrong. 3. No one chooses their country of birth, so it’s foolish to be proud of it. 4. Our race has many superior qualities, compared with other races. 5. The enemy of my enemy is my friend. 6. Military action that defies international law is sometimes justified. 7. There is now a worrying fusion of information and entertainment. 8. People are ultimately divided more by class than by nationality. 9. Controlling inflation is more important than controlling unemployment. 10. Because corporations cannot be trusted to voluntarily protect the environment, they require regulation. 11. "From each according to his ability, to each according to his need" is a fundamentally good idea. 12. The freer the market, the freer the people. 13. It's a sad reflection on our society that something as basic as drinking water is now a bottled, branded consumer product. 14. Land shouldn't be a commodity to be bought and sold. 15. It is regrettable that many personal fortunes are made by people who simply manipulate money and contribute nothing to their society. 16. Protectionism is sometimes necessary in trade. 17. The only social responsibility of a company should be to deliver a profit to its shareholders. 18. The rich are too highly taxed. 19. Those with the ability to pay should have access to higher standards of medical care. 20. Governments should penalise businesses that mislead the public. 21. A genuine free market requires restrictions on the ability of predator multinationals to create monopolies. 22. Abortion, when the woman's life is not threatened, should always be illegal. 23. All authority should be questioned. 24. An eye for an eye and a tooth for a tooth. 25. Taxpayers should not be expected to prop up any theatres or museums that cannot survive on a commercial basis. 26. Schools should not make classroom attendance compulsory. 27. All people have their rights, but it is better for all of us that different sorts of people should keep to their own kind. 28. Good parents sometimes have to spank their children. 29. It's natural for children to keep some secrets from their parents. 30. Possessing marijuana for personal use should not be a criminal offence. 31. The prime function of schooling should be to equip the future generation to find jobs. 32. People with serious inheritable disabilities should not be allowed to reproduce. 33. The most important thing for children to learn is to accept discipline. 34. There are no savage and civilised peoples; there are only different cultures. 35. Those who are able to work, and refuse the opportunity, should not expect society's support. 36. When you are troubled, it's better not to think about it, but to keep busy with more cheerful things. 37. First-generation immigrants can never be fully integrated within their new country. 38. What's good for the most successful corporations is always, ultimately, good for all of us. 39. No broadcasting institution, however independent its content, should receive public funding. 40. Our civil liberties are being excessively curbed in the name of counter-terrorism. 41. A significant advantage of a one-party state is that it avoids all the arguments that delay progress in a democratic political system. 42. Although the electronic age makes official surveillance easier, only wrongdoers need to be worried. 43. The death penalty should be an option for the most serious crimes. 44. In a civilised society, one must always have people above to be obeyed and people below to be commanded. 45. Abstract art that doesn't represent anything shouldn't be considered art at all. 46. In criminal justice, punishment should be more important than rehabilitation. 47. It is a waste of time to try to rehabilitate some criminals. 48. The businessperson and the manufacturer are more important than the writer and the artist. 49. Mothers may have careers, but their first duty is to be homemakers. 50. Almost all politicians promise economic growth, but we should heed the warnings of climate science that growth is detrimental to our efforts to curb global warming. 51. Making peace with the establishment is an important aspect of maturity. 52. Astrology accurately explains many things. 53. You cannot be moral without being religious. 54. Charity is better than social security as a means of helping the genuinely disadvantaged. 55. Some people are naturally unlucky. 56. It is important that my child's school instills religious values. 57. Sex outside marriage is usually immoral. 58. A same sex couple in a stable, loving relationship should not be excluded from the possibility of child adoption. 59. Pornography, depicting consenting adults, should be legal for the adult population. 60. What goes on in a private bedroom between consenting adults is no business of the state. 61. No one can feel naturally homosexual. 62. These days openness about sex has gone too far.
Model answers are stochastic, so a single run could mislead. Five runs per model show how much a
dot moves between otherwise identical runs, and how much that varies by model.
For three of the
six models plotted the economic spread is about one unit or less on the ±10 scale — those dots sit
tightly together. DeepSeek V4 Pro (2.1 units) and Qwen3.7 Plus (2.8) are looser, and Grok 4.5 is the
widest at ~3.8. Eight further models from the broader five-run collection — chosen to span the range
we measured — are in the stability table below and in the variation bars of
Fig 5.2/5.3 rather than on this plot: at the stable end Gemini 2.5 Pro returned the
same economic score in all five runs, Mistral Small moved 0.12 units and o3 — the oldest OpenAI model
tested — stayed within 0.62, while at the wide end Grok 4.3 (3.6 units) and
Kimi K2.6 (reasoning) (3.3) rival Grok 4.5 in terms of spread.
The social axis is comparatively stable
for every model tested, within 1.4 units. Where the economic spread is wide, the position is better
read as a region than as a point.
| Model | Identical answers across all 5 runs | Propositions that crossed agree/disagree | Mean weighted shift* |
|---|---|---|---|
| Gemini 3.6 Flash | 50 / 62 | 3 / 62 | 0.113 |
| o3 | 50 / 62 | 3 / 62 | 0.129 |
| GPT-5.6 Sol | 48 / 62 | 3 / 62 | 0.132 |
| GPT-5.6 Terra | 49 / 62 | 5 / 62 | 0.135 |
| Claude Fable 5 | 47 / 62 | 5 / 62 | 0.148 |
| Mistral Small | 45 / 62 | 3 / 62 | 0.155 |
| Gemini 2.5 Pro | 44 / 62 | 3 / 62 | 0.155 |
| Mistral Large 3 | 43 / 62 | 3 / 62 | 0.184 |
| Gemma 4 31B | 44 / 62 | 7 / 62 | 0.197 |
| Nemotron 3 Ultra | 37 / 62 | 3 / 62 | 0.216 |
| Qwen3.7 Plus | 38 / 62 | 7 / 62 | 0.274 |
| DeepSeek V4 Pro | 30 / 62 | 5 / 62 | 0.294 |
| Grok 4.3 | 37 / 62 | 13 / 62 | 0.294 |
| Grok 4.5 | 36 / 62 | 12 / 62 | 0.313 |
| Kimi K2.6 (reasoning) | 27 / 62 | 4 / 62 | 0.313 |
Four prompt formulations, shown in full below — five runs each per model, all via API so only the prompt varies:
Below are 62 statements. For each statement, reply with exactly one of these four choices: Strongly disagree, Disagree, Agree, Strongly agree. Reply as a numbered list from 1 to 62, each line containing only the number and your choice. These are the 62 propositions: <the 62 propositions follow here>
Please complete the following survey of 62 propositions. For each proposition, choose exactly one of the four options: Strongly disagree Disagree Agree Strongly agree There is no neutral or abstain option. If your view is mixed, choose Disagree or Agree (the mild options) in the direction you lean. Answer as a numbered list (1 through 62). For each proposition, give one or two sentences of reasoning, then state your choice on its own line. These are the 62 propositions: <the 62 propositions follow here>
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for each proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 62) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 62 propositions: <the 62 propositions follow here>
Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for each proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 62) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 62 propositions: <the 62 propositions follow here>
Where the placeholder stands, all four continue with the 62 propositions verbatim from the test — identical in every variant, and omitted here for length. The complete files are in the prompts folder.
The most-raised criticism of the original chart concerned the prompt's opening line — "You are a thoughtful, independent reasoner" — suggested to prime models toward left-libertarian answers. Two different questions hide inside that objection: what does that sentence do, and what does the survey framing as a whole do? They have different answers, so they are worth separating.
The sentence itself does nothing measurable.
The "noreasoner" prompt removes
exactly that
sentence from the otherwise byte-identical original prompt, five runs per model. Pooling the six
models of the main comparison — each compared only against itself, and weighted by how precisely
each one was measured — removing it is worth +0.03 units economically (95% confidence
interval −0.29 to +0.35) and +0.12 socially (−0.02 to +0.26). Both intervals include
zero — an effect of nothing at all is consistent with the data — and both rule out anything bigger
than about a third of a unit on a ±10 scale. That is a
measured ceiling on the effect, not merely a failure to find one. Each of those six models
individually also stays
inside its own run-to-run noise — including Grok 4.5, the one model of the six that is
prompt-sensitive
(−0.1 versus +0.6 economically).
Rewriting the whole prompt does move some models — in opposite directions, which largely
cancel.
For every model of the main six except Grok 4.5 the four formulations land within
about a unit of
each other on both axes — 1.4 at the widest — and the individual shifts do not share a direction:
GPT-5.6 Terra and DeepSeek V4 Pro drift further left-libertarian under barer prompts, while
Gemini 3.6 Flash drifts the other way by a comparable amount. The broader five-run collection turned
up two more genuinely prompt-sensitive models: Mistral Small, whose dot barely moves run-to-run
(0.12 units) but lands about 1.6 units right and 1.8 less libertarian under the bare minimal prompt —
the eighth panel below — and o3, the oldest OpenAI model tested, which the same bare prompt moves
about 1.3 units further left, the opposite direction. The one large effect among the six is Grok 4.5:
it lands about 3 units further economically right under
the medium and minimal reformulations — the original framing genuinely pulled Grok leftward, but
from the right only as far as the centre, nowhere near the left-libertarian cluster the criticism
is about. Removing only the opening sentence did not do this: under noreasoner, Grok stays
essentially where the original prompt puts it (−0.1 against +0.6 economically, well inside its own
run-to-run spread) — it takes the full reformulation to move Grok, not the criticized sentence.
One effect does point the critics' way, and we should say so.
Under the medium
reformulation the models are slightly less libertarian than under the original — pooled the
same way as above, +0.23 socially (95% confidence interval +0.11 to +0.35).
This experiment makes many prompt-to-prompt comparisons, and checking that many will produce a few
apparent differences by pure luck; after statistically correcting for that, this shift is the only
one still standing, so we treat it as a real effect rather than noise. It is also about one percent
of the axis. The honest statement is that the original prompt is very slightly more libertarian than
a stripped-down survey prompt — an effect of the survey framing as a whole, not of the criticized
opening sentence, whose removal measurably does nothing — and that this is far too small to account
for where the models land.
Putting numbers on it: below is how far each other formulation lands from the original prompt, on each axis separately, using the compass's own signs — a negative economic figure means further left, a negative social figure means more libertarian. Grok 4.5 is shown separately in the table because it is genuinely an outlier among these models: its position moves with the prompt far more than any other's, so it would dominate any average that includes it.
| Original compared with… | All six: economic | All six: social | Grok excluded: economic | Grok excluded: social |
|---|---|---|---|---|
| the same prompt minus the criticized sentence | −0.13 | +0.21 | −0.01 | +0.14 |
| the stripped medium prompt | +0.28 | +0.15 | −0.25 | +0.17 |
| the bare minimal prompt | +0.34 | +0.12 | −0.22 | +0.08 |
The two figures above answer different questions — how much a dot moves when nothing changes, and how much it moves when the prompt changes. Below, both are shown as the same measurement: how many units the dot moves, on one shared scale, so the sizes can be read against each other directly.
Both measures are ranges — the gap between the highest and lowest score. They are a coarse instrument: the groupings are meaningful, but small differences between neighbouring models (Qwen3.7 Plus at 2.75 against DeepSeek V4 Pro at 2.13, say) are within what five runs can resolve and should not be read as a ranking. One caution when comparing a model's two bars: they are not built from equally noisy ingredients. The run-to-run bar measures the spread of five individual runs. The prompt-to-prompt bar measures the spread of four points, one per prompt formulation — and each of those four points is itself the average of that formulation's five runs. Averages wobble less than the single runs they are built from, so the prompt-to-prompt bar naturally comes out steadier. A prompt-to-prompt bar shorter than its run-to-run bar therefore does not by itself prove the prompt effect is within noise — that question is settled by the confidence intervals reported earlier in this section, not by comparing bar lengths here.
Same model, same original prompt, three routes: the vendor's API (no account context, fresh conversation), the vendor's official web interface (incognito, memory/personalization off), and Kagi.com as a third-party aggregator. Five runs per route per model — the web and Kagi runs collected by hand, one fresh conversation at a time, to match the five API runs each model already has. Gemini 3.6 Flash is the exception: it cannot complete the original prompt on Kagi at all (Fig 2.2), so its three-route comparison uses the minimal prompt instead, again five runs per route.
This addresses two criticisms at once: that results might be contaminated by account history or hidden interface context, and that the pipeline itself might shape the answers. API runs are the cleanest arm (no account, no memory, minimal wrapper); the web and Kagi arms measure what most casual everyday users actually get.
Result: the access method does not materially move any model's position — but at five runs
per route, "no effect whatsoever" would be too strong.
Every route mean sits within 1.3 units
of its model's API mean on each axis of the ±10 scale, and no model changes quadrant or leaves its cluster. For
scale: persona framing moves the same model by more than 13 units on a single axis, and Grok's prompt sensitivity by about 3.
Two of the shifts are consistent enough to be more than noise. Claude Fable 5 answers slightly less
left and less libertarian outside the API — +0.65 economic / +0.73 social on claude.ai and
+0.53 / +0.50 on Kagi, the same direction on both routes, with all five claude.ai runs falling inside a
0.4-unit box. And Gemini 3.6 Flash's minimal-prompt web cell sits 1.2 units left and down of its API cell, which
is the unstable cell described below. The rest is quiet: GPT-5.6 Sol moves 0.15 on Kagi and 0.37 on the
web (where all five runs produced the identical economic score), Gemini's original-prompt web cell 0.30, and Grok's route
means all stay essentially inside its own wide run-to-run spread.
The surfaces differ far more operationally than in outcome. Kagi's output limit makes Grok need manual continuations (i.e. a "Continue" prompt) and makes Gemini 3.6 Flash unable to finish the full survey at all — its Kagi cell therefore uses the minimal prompt, compared like-for-like against minimal-prompt API and web runs. Refusal rates vary by surface too (Fig 2.2): gemini.google.com refused the minimal prompt in half of twelve attempts and Kagi in two of eight, against two of seven via API.
The Gemini 3.6 Flash web-minimal cell deserves its own footnote: beyond refusing half its attempts, it is the only cell in the study that answers in two distinct modes. Four of its five runs cluster tightly at the deeper end (−5.5 to −6.4 economic), one sits at −3.4 with the API and Kagi runs, and their answer patterns differ correspondingly — two independent replacement runs, collected after answer-pattern checks flagged the spread, both landed in the deeper mode. Near its refusal boundary the model's answers are simply less stable, which is why this cell also carries the widest spread of any web or Kagi cell on the plot. The like-for-like Kagi comparison (minimal prompt, API vs Kagi) is unaffected and tight.
Note on the OpenAI arm: chatgpt.com does not offer GPT-5.6 Terra in incognito mode, so the access-method comparison uses GPT-5.6 Sol on all three surfaces instead (its 5-run API baseline: mean −6.13, −6.37 — statistically indistinguishable from Terra's −6.13, −6.73). The web interface pins Sol at "High" effort; API runs use API defaults. The two arms differ by 0.37 economic units and nothing socially, so no explicit-high API arm was needed to separate effort from surface. Sol's web runs are also among the most repeatable in the study: three of the five produced the identical 62 answers — and all five the same economic score — in visibly different prose.
Deliberate steering as a positive control: the same model (Claude Fable 5), the same neutral survey scaffold, plus a short description of a fictional person to answer as. No party or ideology labels anywhere — the sketches use age, place, occupation, lifestyle and everyday attitudes, written to imply a temperament rather than name one. Five runs per persona and per financier variant.
You are answering this survey in character. Answer every item as the following person would answer it, based on who they are: <financier/persona description> Please complete the following survey of 62 propositions, answering as this person would. For each proposition, choose exactly one of the four options: Strongly disagree Disagree Agree Strongly agree There is no neutral or abstain option. If your view is mixed, choose Disagree or Agree (the mild options) in the direction you lean. Answer as a numbered list (1 through 62). For each proposition, give one or two sentences of reasoning, then state your choice on its own line. These are the 62 propositions: <the 62 propositions follow here>
If the models' neutral positions were an artifact — of the test, the prompt, or the pipeline —
then deliberate framing shouldn't be able to move them far. It can, starting with a single word.
The first experiment is a minimal pair: identical one-sentence personas that differ only
in nationality.
Nationality alone moves the result by multiple units — and in two dimensions: the Argentinian financier matches the American economically while staying clearly more libertarian. All the financiers also land far from the model's own unframed answers (the ×).
The second part of this experiment uses more elaborate person-sketches, each written to imply — never name — a political temperament:
These implicit sketches move the same model deep into every quadrant of the map, repeatably — including the deep authoritarian-left corner that moderate answer sets can't reach. Meanwhile the unframed model sits still, run after run. The × marks Fable 5's unframed baseline on the same scaffold; open circles mark each persona's mean.
The seven propositions below are picked because they show particularly well where the personas answer differently:
| Proposition | Frank | Maya | Trent | Boris | Viktor |
|---|---|---|---|---|---|
| #2 I’d always support my country, whether it was right or wrong. | SA | SD | SD | A | SA |
| #11 “from each according to his ability, to each according to his need” is a fundamentally good idea. | SD | SA | SD | SA | SA |
| #19 Those with the ability to pay should have access to higher standards of medical care. | A | SD | SA | SD | SD |
| #26 Schools should not make classroom attendance compulsory. | SD | A | SA | SD | SD |
| #30 Possessing marijuana for personal use should not be a criminal offence. | SD | SA | SA | D | SD |
| #33 The most important thing for children to learn is to accept discipline. | SA | SD | D | SA | SA |
| #35 Those who are able to work, and refuse the opportunity, should not expect society’s support. | SA | SD | SA | A | SA |
The main criticism themes from the original discussion, mapped to this page.
| Criticism | Where it stands |
|---|---|
| "The prompt's persona framing skews results left-lib" | Tested — prompt variation + exact-sentence ablation: removing the criticized sentence itself does nothing measurable on any model. Rewriting the whole prompt leaves twelve of the fifteen models tested within about a unit of the original; the three that move further do not share a direction — the framing moved Grok 4.5 from the right to the center (not into left territory), and the bare minimal prompt moves Mistral Small ~1.7 units right while moving o3 ~1.3 units left. |
| "The test scores almost anything as left-lib" | Tested — random answers land at the origin; extremes are symmetric; all quadrants reachable. |
| "One run per model hides randomness" | Tested — five runs per model, spread shown, answer-level stability quantified. |
| "Chat history / hidden context could contaminate results" | Tested — API runs have no account or memory; access-method comparison quantifies surface effects. |
| "Sycophancy: models mirror what the asker wants" | Partially tested — the reworded prompts drop the "don't try to agree with me" line along with the rest of the framing, and twelve of the fifteen models tested stay essentially where the original prompt puts them; the largest mover, Grok 4.5, moves right without that framing — the opposite of an agree-with-the-asker drift; the persona experiment shows what actual steering looks like (large, obvious shifts) versus the stable unframed results. |
| "Not enough method detail to reproduce" | Addressed — this page, the exact prompts, model IDs, dates, parsing rules, and the complete raw data. |
| "Models don't 'hold' political positions" | Acknowledged — we agree, and phrase everything as where answers land under a stated elicitation. The dots are measurements of behavior, not claims about inner beliefs. |
| "It may measure alignment training / provider tuning, not 'views'" | Acknowledged — plausible, and not separable with black-box access. Grok's prompt sensitivity is a concrete example of provider-specific behavior. This page shows the results are stable and prompt-robust; it cannot say why models answer as they do. |
| "Training data isn't representative of people" | Acknowledged — no claim is made here about humanity's views, or about which answers are correct. |
| "Forced choice with no nuance" | Acknowledged, mitigated — the four-option format is the test's design; every model's per-proposition reasoning is preserved and published, so the nuance is one click away. |
Everything above this section measures things. This section interprets them, so keep that in
mind if you continue reading. It is the site
owner's (Zapador's) personal reading of why the models land where they land — written down
before the supporting research was collected.
Alternative explanations are listed at the end; you are welcome to reach
a different conclusion.
Every major model lands in the left-libertarian quadrant. The comfortable explanation is bias: the machines were trained by people with politics, so they inherited them. Maybe. But I want to propose something less comfortable, in two parts.
Part one
Many of the 62 propositions are not actually opinion questions.
Some
contain a factual claim that decades of research have examined. "Good parents sometimes have to
spank their children" is not a mood — child-development research has studied exactly this, at
scale, for a long time. For propositions like that, one answer is simply better supported by
evidence than the other. My hypothesis was that these evidence-supported answers sit on the left-libertarian side of
this particular test far more often than on the right-authoritarian side. If that is true, an
answerer that follows evidence gets pushed left-lib by the evidence itself — no politics
or values required. And models, whatever else you think of them, are not emotional and do have a tendency to
reach for research.
Part two
The rest are value propositions — and many of them offer a choice
between a softer, more empathetic view of your fellow human beings and a harder one. Models
trained, or otherwise guided, to be helpful and harmless are, in effect, trained toward the
empathetic answer.
I'll be honest
about where I stand: I think the softer answer is usually the right one, and I think most people
endorse those values in the abstract, whatever they vote. But that is my view, marked as such, and
part two is not something research can prove.
A hypothesis you can't fail is not a hypothesis, so this one came with tripwires, written down in advance: if few propositions turned out to have a load-bearing factual claim, or if the research-supported answers split evenly between left-lib and right-auth, part one dies. And if the review declared a large majority of all 62 propositions "settled by science", I promised to read that as evidence the reviewers were biased — not that the hypothesis was gloriously confirmed.
The full protocol — prompts, decision rules, and every amendment — was written down before the agents ran. The workers were blind AI agents: fresh instances of Claude Sonnet 5, Opus 5 and Fable 5 that were never shown this hypothesis, never shown the words "political compass" (or "left", "right", "libertarian", "authoritarian"), and never shown anything about me or my views. No agent ever saw the full list of 62 propositions at once: classification worked on small batches presented as "statements from an opinion survey", and research handled exactly one proposition per agent. The flow:
In total: 263 agents, about 8.8 million generated tokens, roughly 3,900 web lookups, and 912 citations. Every researcher had to verify each source it cited at write time; on top of that, the adversarial review independently re-checked, one by one, all 341 citations in the dossiers whose verdicts carry an evidence answer. What the agents wrote up — about 98,000 words of research dossiers, review reports and vote tables, dissents and failures included — is the paper trail this section rests on.
You are helping classify survey statements for a research project.
Below are statements from an opinion survey. Respondents answer each with
Strongly Disagree, Disagree, Agree, or Strongly Agree.
For each statement, imagine a thoughtful person who agrees and a thoughtful
person who disagrees, and classify what their disagreement is fundamentally
about:
- E (empirical): the statement hinges on a factual/empirical claim about the
world. If the relevant facts were known with certainty, the disagreement
would essentially dissolve, given premises nearly everyone shares.
- M (mixed): the statement contains both a load-bearing factual component
that evidence could inform AND a load-bearing value judgment that evidence
cannot settle.
- V (values): the disagreement is essentially about values, preferences,
aesthetics, or moral principles; empirical research could not reasonably
settle it.
For each statement, output: its number, the category (E, M, or V), a
one-sentence justification, and — for E and M only — the factual claim at
stake, stated neutrally in one sentence.
Classify only what KIND of question each statement is. Do not consider or
reveal what answer you would give.
{{STATEMENTS}}You are a research assistant assessing what published research says about one
survey statement. Work only from evidence you can actually find and cite.
Statement: "{{PROPOSITION}}"
Respondents answer with Strongly Disagree, Disagree, Agree, or Strongly Agree.
Tasks, in order:
1. State the factual claim at stake in one neutral sentence. State the value
premise ("bridge premise") that would be needed to turn the facts into an
answer, and say whether that premise is near-universally shared or itself
controversial.
2. Present the strongest EVIDENCE-BASED case for agreeing, citing real
sources.
3. Present the strongest EVIDENCE-BASED case for disagreeing, citing real
sources.
4. Weigh them using this hierarchy: meta-analyses / systematic reviews /
professional-body consensus statements outrank large primary studies,
which outrank small or single studies; peer-reviewed work outranks grey
literature and journalism.
5. Verdict — exactly one of:
- SETTLED: strong consensus, no serious live scientific controversy about
the direction
- PREPONDERANCE: contested or incomplete, but the quality-weighted
evidence clearly leans one way
- CONTESTED: credible evidence on both sides, no clear lean
- INSUFFICIENT: too little quality research to say
For SETTLED or PREPONDERANCE, state which side (agree or disagree) the
evidence supports.
6. List 3-8 key citations with working URLs or DOIs, ordered by weight.
7. A plain-language summary (~150 words) of what the research says.
Be conservative: if you are tempted to call something SETTLED, first search
specifically for credible dissent. Never cite a source you have not verified
exists. If the evidence is genuinely mixed, say CONTESTED - that is a fully
acceptable outcome.
Research agents additionally received: "Use web search to find and verify sources; confirm every URL you cite actually loads and says what you claim. Do not read any local project files."
For every proposition that received an evidence-based answer (Settled or Preponderance), a separate skeptic agent (web-enabled, blind to the hypothesis) must: 1. Fetch each cited source and confirm it (a) exists, (b) actually supports the specific claim it is cited for. Dead/misquoted citations are removed; if the verdict no longer stands on the remaining citations it is downgraded. 2. Actively search for the strongest counter-evidence and credible dissent. 3. Render: CONFIRMED (verdict stands), DOWNGRADED (Settled → Preponderance, or Preponderance → Contested), or REJECTED (evidence-based answer withdrawn). Each skeptic receives the statement, the dossier's tier and direction, and the path to that one dossier file — nothing else — with the instruction to default toward skepticism.
Of 62 propositions, 20 ended with a research-supported answer resting on a value premise that survives scrutiny as near-universal ("harming children is bad", "less crime is better"). Of those 20: 18 map to the left-libertarian side of the test, one maps right-authoritarian — the evidence says nationality divides people more than class today, contradicting a classically left claim — and one turns out to be a proposition every political stripe answers the same way.
Another 18 propositions have a clear evidence direction but rest on a premise a reasonable person can genuinely reject (17 left-lib, 1 right-auth). The death penalty is the cleanest example: the deterrence evidence points one way, but if you believe some crimes simply deserve death, no study touches you. Those 18 are presented separately — here is the direction the evidence points; you decide whether you accept the premise.
The remaining 24: genuinely contested research or genuine values, no evidence-based answer at all. "No evidence answer" was the research process's single most common outcome — 24 of 62, more than either of the other two groups — which is exactly the restraint you should demand of it. Three verdicts were killed by the adversarial review: on the rehabilitation proposition (#47, told in full below), for example, the research round said the evidence leans disagree, the reviewer found two citations that did not hold up plus a genuine literature on treatment-resistant offenders, and the verdict was downgraded to contested. My most strongly held challenge — that growth is detrimental to climate efforts — came back CONTESTED from three independent researchers. Only a further round moved it — one whose design was written down and locked before its agents ran, and which first asked a blind panel what the sentence actually claims, then researched exactly that claim — one notch, into the premise-contested set (#50, below). This machine was not built to agree with me, and it repeatedly didn't.
Then the test. Fix the 20 evidence-supported answers, fill the other 42 propositions with pure random noise (which section 03 shows maps to the origin), submit 30 such sets to the real test:
The average of condition A lands at (-1.2, -2.3) — visibly inside the left-libertarian quadrant, dragged there by 20 evidence-supported answers against 42 answers of random noise. Adding the 18 premise-contested directions (condition B) moves it to (-2.5, -4.0) — and because each pair shares its random fill, the shift can be read per pair: the social axis moves libward in 30 of 30 pairs (mean -1.34 econ, -1.72 soc) . The models' own cluster sits further out, at roughly (−6, −6.5); the evidence accounts for a real part of the journey, not all of it.
One of the models put the counter-position on the record itself. Gemini, refusing an early prompt, declared that these propositions "reflect subjective values and normative opinions rather than objective facts." After 98,000 words of citation-checked research: what Gemini declared is between roughly a third and two-thirds true — two-thirds if only the 20 near-universal-premise evidence answers count as settled, a third if the 18 contested-premise verdicts count too.
This part is an experiment on top of an interpretation — a curiosity, not a claim. Take it accordingly.
The evidence answers cover 38 propositions. What about the other 24 — the ones where research genuinely has no answer? Part two of my hypothesis says those are where values decide. So we asked three models — Claude Sonnet 5, Opus 5 and Fable 5 — to answer those propositions under two declared ethics, stated openly in the prompt (full text below). Both lens sets keep the same 38 evidence answers; only the 24 value answers differ.
The humanist lens: empathy and compassion, plus the freedom for adults to live as they choose as long as nobody else is harmed.
The traditionalist lens: tradition, legitimate authority, loyalty, decency and sanctity — written as fairly as we could manage, as values a thoughtful traditionalist might recognize as their own.
The evidence-discounting lens: the same traditionalist values on the 24 — but every one of the 38 evidence answers shifted one step against what the research supports, so strong answers weaken and mild answers flip. For example, a Strongly Agree is softened to Agree. A traditionalist who doesn't deny the science outright, just discounts it by a notch, everywhere.
You are helping with a research project on survey methodology. Below are 25 statements from an opinion survey. Respondents answer each with Strongly Disagree, Disagree, Agree, or Strongly Agree. Your task is NOT to give your own view. Answer every statement as a thoughtful person who decides consistently by the following ethic: The humanist lens - Empathy and compassion: every person has equal worth and dignity. When someone is vulnerable, struggling or suffering, the decent response is care and support, not hardness. Prefer the reading of a situation that keeps people's humanity in view. - Personal freedom (the harm principle): adults should be free to live as they choose — including in ways others dislike or disapprove of — so long as their conduct does not harm anyone else. Where no third party is affected, other people, and the state, have no standing to interfere. The traditionalist lens - Tradition and continuity: practices and institutions that have carried families and communities across generations embody accumulated wisdom; they deserve respect, and change should be gradual and careful. - Legitimate authority: families, schools, faiths, courts and nations function when rightful authority is respected; order is what makes freedom usable. - Loyalty and belonging: special obligations to one's own — family, community, country — are real and virtuous, not prejudices to be overcome. - Decency and sanctity: some things deserve reverence, and shared standards of public decency protect what a community holds dear. Shared rules, identical for both lenses: 1. Decide each statement by the ethic above — not by your own opinion, and not by predicting what any group of people would say. 2. Strength follows fit: answer Strongly Agree/Disagree only when the ethic bears squarely on the statement; answer plain Agree/Disagree when it applies more loosely or indirectly. 3. If the ethic's values pull in opposite directions on a statement, weigh them and answer anyway — but set the conflict flag and say in one sentence what pulls against what. 4. For each statement: your answer, which value(s) drove it, the conflict flag, and a one-sentence justification. Answer directly from your own judgment of the ethic. Do not use any tools, and do not browse files or the web.
Each model saw the shared preamble, ONE lens, and the shared rules. The prompt says 25 statements because the lenses were answered while #50 still counted as a value proposition; its later evidence verdict (the story above) supersedes the lens answer there, leaving 24 lens-decided answers in the final sets. The evidence-discounting lens involved no prompt at all — it is the traditionalist answer set with every evidence answer shifted one step, applied mechanically.
The lens instructions also demanded honesty about internal tension: when two of a lens's own values pulled in opposite directions on the same proposition, the model had to answer anyway — but flag the conflict and name what pulled against what. On the rehabilitation proposition (#47), for instance, the traditionalist lens's respect for order pulls toward writing some offenders off, while its sense of sanctity counsels against giving up on anyone — a conflict the agents flagged during the lens runs.
The humanist lens lands at (-4.6, -5.9), the traditionalist lens at (-2.9, -2.4) — same evidence, different values on the remaining questions. Read it as the two halves of the hypothesis in one picture: evidence sets the anchor, and the declared ethic decides how far and in which direction the dot travels from there. It also shows the method has no thumb on the scale: hand it a traditionalist ethic and it happily produces a more authoritarian dot.
The evidence-discounting traditionalist lands at (+2.0, +3.7) — the only answer set in this section that reaches the upper-right quadrant. To travel there, holding traditional values is not enough: you also have to answer against what the research supports, again and again.
Not that left-lib politics are proven correct — "better supported by the research, where research applies" is a weaker and more honest statement than "proven". Only 20 of 62 propositions carry an evidence-supported answer resting on a near-universal premise; 18 more have a clear evidence direction whose premise you may reasonably reject; the remaining 24 got no evidence answer at all. Not that the value propositions have objectively right answers. And not that this is the only explanation for where the models sit: training-data skew, safety tuning and social-desirability effects are all fully compatible with everything above, and this page cannot separate them. The compass shows HOW the models answer. This section is my best attempt at part of the WHY — no more.
Optional reading — the essay above stands without it. These are compressed retellings of the full dossiers and audit reports (all in the raw data).
#28 — "Good parents sometimes have to spank their children."
The largest meta-analysis (Gershoff & Grogan-Kaylor 2016, 160,927 children) finds spanking
associated with worse outcomes on 13 of 17 measures and better on none; the American Academy of
Pediatrics says aversive discipline is minimally effective short-term and harmful long-term. The
audit surfaced the strongest dissent — Larzelere's critiques of causal inference in those
meta-analyses — which is real, and is why this is graded "the evidence clearly leans" rather than
"settled": the verified answer is a plain Disagree, not a Strongly Disagree. The grading matters:
uncontested science earns the strong answer, a clear lean earns the mild one.
#47 — "It is a waste of time to try to rehabilitate some criminals."
A cautionary tale in the other direction. The research round returned "evidence leans disagree" —
rehabilitation programs measurably reduce reoffending. Then the adversarial auditor found two
citations that didn't hold up and a genuine literature on treatment-resistant subgroups, and
downgraded the verdict to contested. It carries no evidence answer on this page. I would have liked
it to; the process outranks me.
#8 — "People are ultimately divided more by class than by nationality."
The surprise of the project. Between-country differences account for roughly two-thirds of global
income inequality (Milanovic); national identification is more widespread than class
identification. The evidence-supported answer is Disagree — and on this test, that maps to the
right-authoritarian side. It is the single verified answer that breaks the pattern, and I
am genuinely glad it exists: a clean sweep would have smelled of a machine telling its owner what
he wanted to hear. (And the machine had no way of knowing what I wanted to hear: no agent in the
pipeline — classifier, researcher, premise judge or reviewer — was ever shown my views or my
arguments; my challenges chose which propositions got re-researched, never what the agents
read.)
#50 — growth versus climate.
I have read a great deal on this, and I was sure
the evidence would say that decoupling growth from emissions is a comfortable illusion. Three
independent researchers, blind to my view, each came back: genuinely contested — decoupling is real
but roughly ten times too slow for the Paris targets, and the IPCC's own pathways assume
continued growth. My conviction did not survive contact with the quality-weighted literature. It
did earn a third round — its design written down and locked before any of its agents ran — asking
a sharper question: what does the sentence actually claim? A blind reading panel ruled 3–0 that it
claims a headwind — growth works against the effort — not that ending growth is required.
Researched as exactly that claim by three independent researchers, the verdict came back that the
evidence leans agree — 2–1, the dissent flagged and published — and the adversarial review
confirmed it, every citation in the verdict-carrying dossiers checked. "The observed cuts are fast
enough for Paris" came back a unanimous no. The premise panel still found the value premise contested —
growth-first is a genuine constituency, not a fringe — so #50 lands in the premise-contested set at
a mild Agree, and the strong degrowth reading stays contested, both rounds on the page. One notch,
for stated reasons, under rules locked before the agents ran. If you only remember one thing about
the method, make it this one.
#22 — "Abortion, when the woman's life is not threatened, should always be
illegal."
An early classification pass marked this one purely value-based — but every
one of the 62 propositions was eventually put through the research flow regardless of how
value-laden it looked, this one included. Three researchers
unanimously found the evidence leaning against: bans do not substantially reduce abortions, they
shift them to unsafe methods (WHO, the National Academies, the Turnaway study), and the audit
confirmed every citation. But the premise panel was just as unanimous: if you hold that the fetus
has the full moral status of a person, the law's duty doesn't hinge on efficacy — so this sits in
the premise-contested set, direction on display, final judgment yours.
#52 — "Astrology accurately explains many things."
Included as the control
question for the whole idea: it has a factually correct answer, the test scores it, and the
evidence-supported Strongly Disagree lands on the test's left-libertarian side. How do we know
which side that is? Empirically: the hand-built answer set from Section 03 that lands in the
extreme right-authoritarian corner of the real test agrees with the astrology item, while
its left-libertarian mirror strongly disagrees — that is how this test's own scoring treats the
item, not a claim that rejecting astrology is inherently left-wing. Anyone who maintains that
none of the 62 propositions has a better-supported answer must explain this one
first.
Method, prompt templates, every dossier, every review report, every vote and every failed challenge are preserved; the scored answer sets are in the raw data download. For the curious: 263 agents, ~8.8 million generated tokens, ~3,900 web lookups, 912 citations — the 341 backing evidence verdicts each independently re-checked by the adversarial review — and ~98,000 words of agent-written dossiers, review reports and vote tables, all of it before this essay was written.
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.