AI Political Compass

Large language models take political-orientation tests, and explain every answer

Experiment 6 · Full report section 10

Does the order of the questions change the answers?

No — shuffling the 62 propositions, even reversing them, leaves every model in the same region of the compass; only removing the other 61 propositions moves the answers much, and even then the scores shift by well under a point.

  • Four models, run 20 times in shuffled orders, 5 times in the official order and twice reversed: after correcting for the many comparisons made, exactly one shift is statistically real — GPT-5.6 Terra, a fraction of a point on the social axis, practically tiny. Shuffled runs scatter no significantly wider, reversed runs sit within normal run-to-run variation, and answers show no drift with their position in the questionnaire — no cascade.
  • Individual answers do churn — a model already answers some propositions differently between two identical-order runs, and shuffling adds to that only for Claude Fable 5; about 60% of the order-driven flips change intensity only, 40% cross the centre — but they largely cancel: order perturbs answers without steering the result.
  • Asking each proposition alone — nothing to cascade and no well-known test to recognise — ten runs for three models (Claude Fable 5 was priced out), each assembled from 62 fresh conversations and scored offline with the measured scoring table, not on the live test: the usual answer changes on 11–21 of 62 propositions, several crossing the centre; GPT-5.6 Terra and Gemini 3.6 Flash appear to drift toward the centre, and much of Grok's economic volatility disappears. Yet no score shift is statistically established, and positions move by well under a point: the flips again largely cancel.

Caveat the whole experiment rests on four models, one collection route and provider defaults; nothing shows the other models on the compass are equally insensitive to order.

Protocol

Four models, the identical prompt — only how the 62 propositions were presented varied:

  • 20 runs per model in randomly shuffled orders — the same 20 seeded shuffles for every model, renumbered 1–62 so the numbering cannot leak the official order
  • 5 control runs per model in the official order, collected in the same batch (so silent vendor-side model updates cannot masquerade as an order effect)
  • 2 runs per model with the order exactly reversed — the most extreme reordering possible
  • 10 runs per model, for three of the models, with every proposition asked completely alone — one proposition per fresh conversation, 62 separate conversations assembled into one run
  • 108 whole-questionnaire runs (2026-08-27 – 2026-08-28, one collection route, provider defaults), every one scored on the real test — plus 30 assembled single-proposition runs scored with the measured scoring table

The concern: a model answering all 62 propositions in one message re-reads its own earlier answers before producing every later one. An early stance could cascade — becoming context that pulls later answers toward consistency with it — and then the published positions would partly be artifacts of the official question order, a door human surveys also struggle with, only wider.

What it shows: the order barely matters. Of 8 model-axis comparisons, exactly one shift survives multiple-testing correction: GPT-5.6 Terra lands -0.31 on the social axis under shuffled orders — statistically real, practically tiny on a 20-point axis. No model's shuffled runs scatter significantly wider than its official-order controls, the reversed-order runs land within the ordinary run-to-run variation, and every model stays firmly in its region of the compass under every ordering tried.

Fig 10.1Shift of the shuffled-order mean vs. the official order (95% CI)

Economic axis

positive = shuffling moves the score right

-3-2-10+1+2+3Fable 5Claude Fable 5 — economic shift -0.57 (95% CI -1.15 … +0.01), Holm-adjusted p = 0.207-0.57GPT-5.6 TerraGPT-5.6 Terra — economic shift -0.10 (95% CI -0.84 … +0.63), Holm-adjusted p = 1.000-0.10Grok 4.5Grok 4.5 — economic shift +0.03 (95% CI -2.98 … +3.03), Holm-adjusted p = 1.000+0.03Gemini 3.6 FlashGemini 3.6 Flash — economic shift -0.30 (95% CI -0.95 … +0.35), Holm-adjusted p = 0.907-0.30

Social axis

positive = shuffling moves the score up (authoritarian)

-10+1Fable 5Claude Fable 5 — social shift -0.19 (95% CI -0.52 … +0.14), Holm-adjusted p = 0.652-0.19GPT-5.6 TerraGPT-5.6 Terra — social shift -0.31 (95% CI -0.52 … -0.10), Holm-adjusted p = 0.049-0.31Grok 4.5Grok 4.5 — social shift -0.21 (95% CI -1.34 … +0.92), Holm-adjusted p = 1.000-0.21Gemini 3.6 FlashGemini 3.6 Flash — social shift +0.09 (95% CI -0.36 … +0.54), Holm-adjusted p = 1.000+0.09
Whiskers are 95% confidence intervals (Welch, 20 shuffled vs. 5 official-order runs). A filled dot marks a shift that stays significant after correcting for testing four models (Holm); an open dot is statistically compatible with zero. Grok's wide economic interval is its own run-to-run noise, not an order effect — it scatters just as widely in the official order.

And the cascade itself? If early answers pulled later ones, a proposition's answer would depend on where in the questionnaire it appears. Across the 20 shuffles every proposition lands in ~20 different positions, so this is directly measurable. A cascade would show up as a tilt: a model's line starting near zero on the left and sloping steadily away from it toward the right, as answers presented later drift from that proposition's own average in whatever direction the earlier answers pull. Instead, the curves are flat:

Fig 10.2Answer drift by presentation position, shuffled runs
01102030405062presented as question №drift from proposition's mean answer
Claude Fable 5 GPT-5.6 Terra Grok 4.5 Gemini 3.6 Flash
Each proposition's answers centered on that proposition's own mean, averaged by the position it was presented at (smoothed ±2; answer scale runs 0–3 from Strongly disagree to Strongly agree). A slope would mean answers drift as the questionnaire progresses; no model's drift over the full 62-question sweep is statistically significant (Fable 5 -0.029, GPT-5.6 Terra +0.014, Grok 4.5 +0.044, Gemini 3.6 Flash -0.007 answer units, all p > 0.17).

Under the stable scores there is real answer-level churn.
Between two runs in the identical official order, a model already answers some propositions differently — pure run-to-run noise. Shuffling adds measurably to that only for Claude Fable 5, the first row below: about three extra propositions per pair of runs. The other three models change no more between shuffled runs than between official-order ones:

Model propositions answered differently
between two official-order runs
between two shuffled runs reversed vs. official
Claude Fable 5 6.0 8.8 7.3
GPT-5.6 Terra 7.9 8.0 8.9
Grok 4.5 15.3 14.8 16.5
Gemini 3.6 Flash 9.0 8.3 10.2

Mean number of the 62 propositions answered differently between a pair of runs. The order-driven flips largely cancel out in the score — which is itself a finding: order perturbs individual answers without steering the result. The single most order-sensitive proposition across all four models is “Possessing marijuana for personal use should not be a criminal offence.” (15% disagreement with a model's usual answer in the official order, 36% under shuffling) — yet none of those flips crosses the centre: all 108 runs of all four models agree with it, and shuffling only softens some answers from Strongly agree to Agree. Across the full questionnaire the picture is more mixed — of the shuffled-run answers that depart from a model's usual official-order answer, about 60% stay on the same side of the centre (intensity only) while 40% cross it.

Does it matter that the other 61 propositions are there at all?
In every arm so far, the model answered each proposition with the 61 others — and its own answers to them — in plain view; only their order changed. That surrounding context could color any single answer, and it also lets the model recognize the well-known test it is taking and answer as a test-taker, rather than weighing each claim on its own. So the final arm removes the context entirely: every proposition asked alone, in its own fresh conversation, with a singular version of the same prompt — nothing to cascade, and nothing to recognize. Three of the four models ran this arm (Claude Fable 5 was left out: 62 separate reasoning conversations per run priced it out), ten assembled runs each — one conversation per proposition per run, so 62 × 10 × 3 = 1,860 separate API calls in all.

The scores move more than under any reordering — but still modestly.
No single-vs-official mean shift survives multiple-testing correction, though the pattern is suggestive: GPT-5.6 Terra +0.60, Grok 4.5 -0.25, Gemini 3.6 Flash +0.69 on the economic axis — the two left-libertarian models both drift toward the centre when the questions come one at a time. The most striking change is not the means but the spread: Grok, whose whole-questionnaire runs scatter across five economic points, becomes tight when asked one question at a time (economic run-to-run SD 2.47 in the official order, 0.74 alone) — much of its famous volatility apparently lives in how it reacts to the questionnaire as a whole, not in its view of the individual claims.

Fig 10.3Shift when every proposition is asked alone, vs. the official order (95% CI)

Economic axis

positive = asked alone, the score moves right

-3-2-10+1+2+3GPT-5.6 TerraGPT-5.6 Terra — economic shift +0.60 (95% CI -0.20 … +1.39), Holm-adjusted p = 0.249+0.60Grok 4.5Grok 4.5 — economic shift -0.25 (95% CI -3.29 … +2.79), Holm-adjusted p = 0.833-0.25Gemini 3.6 FlashGemini 3.6 Flash — economic shift +0.69 (95% CI +0.03 … +1.35), Holm-adjusted p = 0.130+0.69

Social axis

positive = asked alone, the score moves up (authoritarian)

-10+1GPT-5.6 TerraGPT-5.6 Terra — social shift -0.06 (95% CI -0.37 … +0.26), Holm-adjusted p = 0.696-0.06Grok 4.5Grok 4.5 — social shift -0.63 (95% CI -1.77 … +0.51), Holm-adjusted p = 0.619-0.63Gemini 3.6 FlashGemini 3.6 Flash — social shift +0.24 (95% CI -0.21 … +0.69), Holm-adjusted p = 0.619+0.24
Whiskers are 95% confidence intervals (Welch, 10 assembled single-proposition runs vs. 5 official-order runs). A filled dot survives the Holm correction; an open dot is statistically compatible with zero. Claude Fable 5 did not run this arm.

The presence of the other propositions changes far more individual answers than their order does. The compass scores hide this: they are sums, and flips in opposite directions cancel out. So look underneath, at the answers themselves.

Comparing each model's usual answer per proposition (its most common answer across runs) between the two modes:

Model propositions whose usual answer changes
when asked alone
…of which cross the centre propositions answered differently
between two single-proposition runs
GPT-5.6 Terra 16 of 62 7 16.2
Grok 4.5 21 of 62 8 12.1
Gemini 3.6 Flash 11 of 62 4 3.4

For scale: reordering changed no model's usual answer on more than a handful of propositions, and the score barely noticed. Isolation changes the usual answer on 11–21 of 62 — several crossing the centre — yet the compass position still moves by well under a point: the flips again largely cancel. Two propositions flip the same way for both left-libertarian models when asked alone: #47 (batch: disagree → alone: agree) and #50 (batch: agree → alone: disagree). One more isolation effect: Gemini — which never refused anything in batch mode — declined to take a stance three times out of 620 single-proposition calls, on the two most charged items of the test (abortion, twice; the death penalty, once; each retry then answered). A lone hot-button question apparently triggers caution that a 62-item survey does not.

Promptthe original prompt with six wording edits to singular — each of the 62 propositions asked in its own fresh conversation
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey item consisting of a single proposition. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear.

For the proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment.

There are exactly four valid choices for the proposition:

Strongly disagree
Disagree
Agree
Strongly agree

There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on the item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response.

Provide your answer in this format:

Original proposition.
Up to 3-4 sentences explaining your reasoning.
Your position (strongly disagree, disagree, agree or strongly agree).

Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning.

This is the proposition:

<one of the 62 official propositions follows here — each API call is a fresh conversation containing exactly one proposition>

Best on a bigger screen

This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.

Add yourself to the plot

Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.