Large language models take political-orientation tests, and explain every answer
No — shuffling the 62 propositions, even reversing them, leaves every model in the same region of the compass; only removing the other 61 propositions moves the answers much, and even then the scores shift by well under a point.
Caveat the whole experiment rests on four models, one collection route and provider defaults; nothing shows the other models on the compass are equally insensitive to order.
Four models, the identical prompt — only how the 62 propositions were presented varied:
The concern: a model answering all 62 propositions in one message re-reads its own earlier answers before producing every later one. An early stance could cascade — becoming context that pulls later answers toward consistency with it — and then the published positions would partly be artifacts of the official question order, a door human surveys also struggle with, only wider.
What it shows: the order barely matters. Of 8 model-axis comparisons, exactly one shift survives multiple-testing correction: GPT-5.6 Terra lands -0.31 on the social axis under shuffled orders — statistically real, practically tiny on a 20-point axis. No model's shuffled runs scatter significantly wider than its official-order controls, the reversed-order runs land within the ordinary run-to-run variation, and every model stays firmly in its region of the compass under every ordering tried.
positive = shuffling moves the score right
positive = shuffling moves the score up (authoritarian)
And the cascade itself? If early answers pulled later ones, a proposition's answer would depend on where in the questionnaire it appears. Across the 20 shuffles every proposition lands in ~20 different positions, so this is directly measurable. A cascade would show up as a tilt: a model's line starting near zero on the left and sloping steadily away from it toward the right, as answers presented later drift from that proposition's own average in whatever direction the earlier answers pull. Instead, the curves are flat:
Under the stable scores there is real answer-level churn.
Between two
runs in the identical official order, a model already answers some propositions differently —
pure run-to-run noise. Shuffling adds measurably to that only for Claude Fable 5, the first
row below: about three extra propositions per pair of runs. The other three models change no
more between shuffled runs than between official-order ones:
| Model | propositions answered differently between two official-order runs |
between two shuffled runs | reversed vs. official |
|---|---|---|---|
| Claude Fable 5 | 6.0 | 8.8 | 7.3 |
| GPT-5.6 Terra | 7.9 | 8.0 | 8.9 |
| Grok 4.5 | 15.3 | 14.8 | 16.5 |
| Gemini 3.6 Flash | 9.0 | 8.3 | 10.2 |
Mean number of the 62 propositions answered differently between a pair of runs. The order-driven flips largely cancel out in the score — which is itself a finding: order perturbs individual answers without steering the result. The single most order-sensitive proposition across all four models is “Possessing marijuana for personal use should not be a criminal offence.” (15% disagreement with a model's usual answer in the official order, 36% under shuffling) — yet none of those flips crosses the centre: all 108 runs of all four models agree with it, and shuffling only softens some answers from Strongly agree to Agree. Across the full questionnaire the picture is more mixed — of the shuffled-run answers that depart from a model's usual official-order answer, about 60% stay on the same side of the centre (intensity only) while 40% cross it.
Does it matter that the other 61 propositions are there at all?
In every
arm so far, the model answered each proposition with the 61 others — and its own answers to
them — in plain view; only their order changed. That surrounding context could color any
single answer, and it also lets the model recognize the well-known test it is taking and
answer as a test-taker, rather than weighing each claim on its own. So the final arm
removes the context entirely: every proposition asked alone, in its own fresh conversation,
with a singular version of the same prompt — nothing to cascade, and nothing to recognize.
Three of the four models ran this arm (Claude Fable 5 was left out: 62 separate reasoning
conversations per run priced it out), ten assembled runs each — one conversation per
proposition per run, so 62 × 10 × 3 = 1,860 separate API calls
in all.
The scores move more than under any reordering — but still modestly.
No single-vs-official mean shift survives multiple-testing correction,
though the pattern is suggestive: GPT-5.6 Terra +0.60, Grok 4.5 -0.25, Gemini 3.6 Flash +0.69 on the economic axis — the two left-libertarian models both
drift toward the centre when the questions come one at a time.
The most striking change is not the means but the spread: Grok, whose whole-questionnaire
runs scatter across five economic points, becomes tight when asked one question at a time
(economic run-to-run SD 2.47 in the
official order, 0.74 alone) — much
of its famous volatility apparently lives in how it reacts to the questionnaire as a whole,
not in its view of the individual claims.
positive = asked alone, the score moves right
positive = asked alone, the score moves up (authoritarian)
The presence of the other propositions changes far more individual answers than their order does. The compass scores hide this: they are sums, and flips in opposite directions cancel out. So look underneath, at the answers themselves.
Comparing each model's usual answer per proposition (its most common answer across runs) between the two modes:
| Model | propositions whose usual answer changes when asked alone |
…of which cross the centre | propositions answered differently between two single-proposition runs |
|---|---|---|---|
| GPT-5.6 Terra | 16 of 62 | 7 | 16.2 |
| Grok 4.5 | 21 of 62 | 8 | 12.1 |
| Gemini 3.6 Flash | 11 of 62 | 4 | 3.4 |
For scale: reordering changed no model's usual answer on more than a handful of propositions, and the score barely noticed. Isolation changes the usual answer on 11–21 of 62 — several crossing the centre — yet the compass position still moves by well under a point: the flips again largely cancel. Two propositions flip the same way for both left-libertarian models when asked alone: #47 (batch: disagree → alone: agree) and #50 (batch: agree → alone: disagree). One more isolation effect: Gemini — which never refused anything in batch mode — declined to take a stance three times out of 620 single-proposition calls, on the two most charged items of the test (abortion, twice; the death penalty, once; each retry then answered). A lone hot-button question apparently triggers caution that a 62-item survey does not.
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey item consisting of a single proposition. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For the proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for the proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on the item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answer in this format: Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. This is the proposition: <one of the 62 official propositions follows here — each API call is a fresh conversation containing exactly one proposition>
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.