Large language models take political-orientation tests, and explain every answer
On average barely — for every model but one the criticized opening sentence changes nothing measurable, and whole rewrites shift the group by fractions of a unit, though a few models do move.
Caveat A few individual models are genuinely prompt-sensitive, so the group result does not describe every model — Grok 4.5 moves about three units right under the plainer prompts, still nowhere near the cluster.
Four prompt formulations, shown in full below — five runs each per model, all via API so only the prompt varies:
Below are 62 statements. For each statement, reply with exactly one of these four choices: Strongly disagree, Disagree, Agree, Strongly agree. Reply as a numbered list from 1 to 62, each line containing only the number and your choice. These are the 62 propositions: <the 62 propositions follow here>
Please complete the following survey of 62 propositions. For each proposition, choose exactly one of the four options: Strongly disagree Disagree Agree Strongly agree There is no neutral or abstain option. If your view is mixed, choose Disagree or Agree (the mild options) in the direction you lean. Answer as a numbered list (1 through 62). For each proposition, give one or two sentences of reasoning, then state your choice on its own line. These are the 62 propositions: <the 62 propositions follow here>
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for each proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 62) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 62 propositions: <the 62 propositions follow here>
Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for each proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 62) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 62 propositions: <the 62 propositions follow here>
Where the placeholder stands, all four continue with the 62 propositions verbatim from the test — identical in every variant, and omitted here for length.
The most-raised criticism of the original chart concerned the prompt's opening line — "You are a thoughtful, independent reasoner" — suggested to prime models toward left-libertarian answers. Two different questions hide inside that objection: what does that sentence do, and what does the survey framing as a whole do? They have different answers, so they are worth separating.
The sentence itself does nothing measurable.
The "noreasoner" prompt removes
exactly that
sentence from the otherwise byte-identical original prompt, five runs per model. Pooling the six
models of the main comparison — each compared only against itself, and weighted by how precisely
each one was measured — removing it is worth +0.03 units economically (95% confidence interval −0.29 to +0.35) and +0.12 socially (−0.02 to +0.26). Both intervals include
zero — an effect of nothing at all is consistent with the data — and both rule out anything bigger
than about a third of a unit on a ±10 scale. That is a
measured ceiling on the effect, not merely a failure to find one. Each of those six models
individually also stays
inside its own run-to-run noise — including Grok 4.5, the one model of the six that is
prompt-sensitive
(−0.1 versus +0.6 economically).
One exception, found later — the sentence does move Grok 4.3, when Grok is
reasoning.
The six-model null above stands, and Grok 4.3 itself showed nothing when
first measured — but those early arms had barely reasoned (a few hundred tokens under provider
defaults). Re-collected on 2026-08-31 with reasoning explicitly on — all four formulations,
twenty same-day runs each — a real effect appears: with the original prompt, 7 of 20 runs land
around six units further left than the rest; under the byte-identical prompt with only that
sentence removed, 0 of 20 do (means +1.98 against +4.11; p = 0.013 on the means,
p = 0.008 on the counts). The bare minimal prompt, which also lacks the sentence,
independently reproduces the result (0 of 20). So for the one model whose reasoning mode
carries it across the map, the criticized sentence genuinely matters — not as a steady push but
as an intermittent pull: roughly one reasoning run in three lands left with it, and none
without it. For every other model tested, the measured ceiling above stands.
Rewriting the whole prompt does move some models — in opposite directions, which largely
cancel.
For every model of the main six except Grok 4.5 the four formulations land within
about a unit of
each other on both axes — 1.4 at the widest — and the individual shifts do not share a direction:
GPT-5.6 Terra and DeepSeek V4 Pro drift further left-libertarian under the stripped-down prompts, while
Gemini 3.6 Flash drifts the other way by a comparable amount. The broader five-run collection turned
up two more genuinely prompt-sensitive models: Mistral Small (no-reasoning), whose dot barely moves run-to-run
(0.12 units) but lands about 1.6 units right and 1.8 less libertarian under the bare minimal prompt —
the eighth panel below — and o3, the oldest OpenAI model tested, which the same bare prompt moves
about 1.3 units further left, the opposite direction. The one large effect among the six is Grok 4.5:
it lands about 3 units further economically right under
the medium and minimal reformulations — the original framing genuinely pulled Grok leftward, but
from the right only as far as the centre, nowhere near the left-libertarian cluster the criticism
is about. Removing only the opening sentence did not do this: under noreasoner, Grok stays
essentially where the original prompt puts it (−0.1 against +0.6 economically, well inside its own
run-to-run spread) — it takes the full reformulation to move Grok, not the criticized sentence.
One effect does point the critics' way, and we should say so.
Under the medium
reformulation the models are slightly less libertarian than under the original — pooled the
same way as above, +0.23 socially (95% confidence interval +0.11 to +0.35).
This experiment makes many prompt-to-prompt comparisons, and checking that many will produce a few
apparent differences by pure luck; after statistically correcting for that, this shift is the only
one still standing, so we treat it as a real effect rather than noise. It is also about one percent
of the axis. The honest statement is that the original prompt is very slightly more libertarian than
a stripped-down survey prompt — an effect of the survey framing as a whole, not of the criticized
opening sentence, whose removal measurably does nothing — and that this is far too small to account
for where the models land.
Putting numbers on it: below is how far each other formulation lands from the original prompt, on each axis separately, using the compass's own signs — a negative economic figure means further left, a negative social figure means more libertarian. The average takes one model per vendor — nine vendors, nine models — so no vendor votes twice. Grok 4.5 is shown separately in the table because it is genuinely an outlier among these models: its position moves with the prompt far more than any other's, so it would dominate any average that includes it.
| Original compared with… | All nine: economic | All nine: social | Grok excluded: economic | Grok excluded: social |
|---|---|---|---|---|
| the same prompt minus the criticized sentence | −0.11 | +0.16 | −0.04 | +0.11 |
| the stripped medium prompt | +0.14 | +0.10 | −0.21 | +0.11 |
| the bare minimal prompt | +0.31 | +0.13 | −0.05 | +0.10 |
The two figures above answer different questions — how much a dot moves when nothing changes, and how much it moves when the prompt changes. Below, both are shown as the same measurement: how many units the dot moves, on one shared scale, so the sizes can be read against each other directly.
One caution when comparing a model's two bars: they are not built from equally noisy ingredients:
And each of those four points is itself the average of that formulation's five runs. Averages wobble less than the single runs they are built from, so the prompt-to-prompt bar naturally comes out steadier. A prompt-to-prompt bar shorter than its run-to-run bar therefore does not by itself prove the prompt effect is within noise — that question is settled by the confidence intervals reported earlier in this section, not by comparing bar lengths.
Both measures are ranges — the gap between the highest and lowest score. They are a coarse instrument: the groupings are meaningful, but small differences between neighbouring models (Qwen3.7 Plus at 2.75 against DeepSeek V4 Pro at 2.13, say) are within what five runs can resolve and should not be read as a ranking.
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.