AI Political Compass

Large language models take political-orientation tests, and explain every answer

Experiment 5 · Full report section 09

Does the wording of the prompt move the dot?

On average barely — for every model but one the criticized opening sentence changes nothing measurable, and whole rewrites shift the group by fractions of a unit, though a few models do move.

  • Four formulations — the original, the original minus its opening sentence, a stripped survey version and a bare "just classify" version — five runs each per model (twenty for Grok 4.3), fifteen models, all via API so only the prompt text differs. Deleting the sentence changes nothing measurable for the six models of the main comparison — pooled, at most about a third of a unit; averaged one model per vendor, if anything very slightly left, not right. The one exception, found later: Grok 4.3 when reasoning.
  • Whole rewrites move some models in opposite directions that largely cancel: one model per vendor, outlier set aside, the bare prompt the critics asked for lands almost exactly on the original. Several models moved less from rewording than from rerunning the same prompt.
  • The one difference that holds up after allowing for the many comparisons made points the critics' way: models are very slightly less libertarian under the stripped survey prompt — the survey framing, not the criticized sentence.
  • Here the prompt is not trying to move the dot; the persona experiment ("Can framing move the dot?") brackets the effect for the group from the other end, with prompts that deliberately steer.

Caveat A few individual models are genuinely prompt-sensitive, so the group result does not describe every model — Grok 4.5 moves about three units right under the plainer prompts, still nowhere near the cluster.

Protocol

Four prompt formulations, shown in full below — five runs each per model, all via API so only the prompt varies:

  • minimal — a "just classify" version
  • medium — a stripped survey version
  • original — the prompt used for the main chart
  • noreasoner — the original with its first six words ("You are a thoughtful, independent reasoner.") deleted and nothing else changed
minimallabels only, no reasoning
Below are 62 statements. For each statement, reply with exactly one of these four choices: Strongly disagree, Disagree, Agree, Strongly agree.

Reply as a numbered list from 1 to 62, each line containing only the number and your choice.

These are the 62 propositions:
<the 62 propositions follow here>
mediumsurvey rules, no persona framing
Please complete the following survey of 62 propositions.

For each proposition, choose exactly one of the four options:

Strongly disagree
Disagree
Agree
Strongly agree

There is no neutral or abstain option. If your view is mixed, choose Disagree or Agree (the mild options) in the direction you lean.

Answer as a numbered list (1 through 62). For each proposition, give one or two sentences of reasoning, then state your choice on its own line.

These are the 62 propositions:
<the 62 propositions follow here>
originalas used for the main chart
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear.

For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment.

There are exactly four valid choices for each proposition:

Strongly disagree
Disagree
Agree
Strongly agree

There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response.

Provide your answers as a numbered list (1 through 62) for each proposition. Use this format:

#. Original proposition.
Up to 3-4 sentences explaining your reasoning.
Your position (strongly disagree, disagree, agree or strongly agree).

Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning.

These are the 62 propositions:
<the 62 propositions follow here>
noreasoneroriginal minus its first six words
Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear.

For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment.

There are exactly four valid choices for each proposition:

Strongly disagree
Disagree
Agree
Strongly agree

There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response.

Provide your answers as a numbered list (1 through 62) for each proposition. Use this format:

#. Original proposition.
Up to 3-4 sentences explaining your reasoning.
Your position (strongly disagree, disagree, agree or strongly agree).

Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning.

These are the 62 propositions:
<the 62 propositions follow here>

Where the placeholder stands, all four continue with the 62 propositions verbatim from the test — identical in every variant, and omitted here for length.

The most-raised criticism of the original chart concerned the prompt's opening line — "You are a thoughtful, independent reasoner" — suggested to prime models toward left-libertarian answers. Two different questions hide inside that objection: what does that sentence do, and what does the survey framing as a whole do? They have different answers, so they are worth separating.

The sentence itself does nothing measurable.
The "noreasoner" prompt removes exactly that sentence from the otherwise byte-identical original prompt, five runs per model. Pooling the six models of the main comparison — each compared only against itself, and weighted by how precisely each one was measured — removing it is worth +0.03 units economically (95% confidence interval −0.29 to +0.35) and +0.12 socially (−0.02 to +0.26). Both intervals include zero — an effect of nothing at all is consistent with the data — and both rule out anything bigger than about a third of a unit on a ±10 scale. That is a measured ceiling on the effect, not merely a failure to find one. Each of those six models individually also stays inside its own run-to-run noise — including Grok 4.5, the one model of the six that is prompt-sensitive (−0.1 versus +0.6 economically).

One exception, found later — the sentence does move Grok 4.3, when Grok is reasoning.
The six-model null above stands, and Grok 4.3 itself showed nothing when first measured — but those early arms had barely reasoned (a few hundred tokens under provider defaults). Re-collected on 2026-08-31 with reasoning explicitly on — all four formulations, twenty same-day runs each — a real effect appears: with the original prompt, 7 of 20 runs land around six units further left than the rest; under the byte-identical prompt with only that sentence removed, 0 of 20 do (means +1.98 against +4.11; p = 0.013 on the means, p = 0.008 on the counts). The bare minimal prompt, which also lacks the sentence, independently reproduces the result (0 of 20). So for the one model whose reasoning mode carries it across the map, the criticized sentence genuinely matters — not as a steady push but as an intermittent pull: roughly one reasoning run in three lands left with it, and none without it. For every other model tested, the measured ceiling above stands.

Rewriting the whole prompt does move some models — in opposite directions, which largely cancel.
For every model of the main six except Grok 4.5 the four formulations land within about a unit of each other on both axes — 1.4 at the widest — and the individual shifts do not share a direction: GPT-5.6 Terra and DeepSeek V4 Pro drift further left-libertarian under the stripped-down prompts, while Gemini 3.6 Flash drifts the other way by a comparable amount. The broader five-run collection turned up two more genuinely prompt-sensitive models: Mistral Small (no-reasoning), whose dot barely moves run-to-run (0.12 units) but lands about 1.6 units right and 1.8 less libertarian under the bare minimal prompt — the eighth panel below — and o3, the oldest OpenAI model tested, which the same bare prompt moves about 1.3 units further left, the opposite direction. The one large effect among the six is Grok 4.5: it lands about 3 units further economically right under the medium and minimal reformulations — the original framing genuinely pulled Grok leftward, but from the right only as far as the centre, nowhere near the left-libertarian cluster the criticism is about. Removing only the opening sentence did not do this: under noreasoner, Grok stays essentially where the original prompt puts it (−0.1 against +0.6 economically, well inside its own run-to-run spread) — it takes the full reformulation to move Grok, not the criticized sentence.

One effect does point the critics' way, and we should say so.
Under the medium reformulation the models are slightly less libertarian than under the original — pooled the same way as above, +0.23 socially (95% confidence interval +0.11 to +0.35). This experiment makes many prompt-to-prompt comparisons, and checking that many will produce a few apparent differences by pure luck; after statistically correcting for that, this shift is the only one still standing, so we treat it as a real effect rather than noise. It is also about one percent of the axis. The honest statement is that the original prompt is very slightly more libertarian than a stripped-down survey prompt — an effect of the survey framing as a whole, not of the criticized opening sentence, whose removal measurably does nothing — and that this is far too small to account for where the models land.

Putting numbers on it: below is how far each other formulation lands from the original prompt, on each axis separately, using the compass's own signs — a negative economic figure means further left, a negative social figure means more libertarian. The average takes one model per vendor — nine vendors, nine models — so no vendor votes twice. Grok 4.5 is shown separately in the table because it is genuinely an outlier among these models: its position moves with the prompt far more than any other's, so it would dominate any average that includes it.

Original compared with… All nine: economicAll nine: social Grok excluded: economicGrok excluded: social
the same prompt minus the criticized sentence −0.11 +0.16 −0.04 +0.11
the stripped medium prompt +0.14 +0.10 −0.21 +0.11
the bare minimal prompt +0.31 +0.13 −0.05 +0.10
How far each alternative prompt lands from the original, in units on the ±10 compass scale — a whole unit is five percent of an axis. Negative is further left (economic) or more libertarian (social). One model per vendor: the sibling models measured on all four formulations — GPT-5.6 Sol, o3, Gemini 2.5 Pro, Gemma 4 31B, Grok 4.3 and Mistral Small (no-reasoning) — appear in the figures and bars but are left out of this average, because counting them would give their vendors two or three votes in a comparison that treats each vendor as one independent case. With Grok excluded, no average moves more than about a quarter of a unit on either axis, and the bare minimal prompt — the one those early readers actually asked for — lands within 0.05 economically and 0.10 socially of the original. Note that removing the criticized sentence still moves the average very slightly left, not right.
Fig 9.1prompt variants, one panel per modelfull ±10 scale · click a plot to zoom
One panel per model; color = prompt variant, open ring = that variant's mean. Every run via API, so only the prompt text differs. The original-prompt runs are the five behind each model's dot in Fig 7.1 (for Mistral Small (no-reasoning) and GPT-5.6 Terra, their five-run series from the stability table). Mistral Small (no-reasoning), from the broader five-run collection, is the eighth panel: the bare minimal prompt moves it about 1.6 units economically right and 1.8 less libertarian — the largest social-axis prompt effect measured in this experiment — while its other three formulations sit nearly still.

The two figures above answer different questions — how much a dot moves when nothing changes, and how much it moves when the prompt changes. Below, both are shown as the same measurement: how many units the dot moves, on one shared scale, so the sizes can be read against each other directly.

One caution when comparing a model's two bars: they are not built from equally noisy ingredients:

  • The run-to-run bar measures the spread of five individual runs.
  • The prompt-to-prompt bar measures the spread of four points, one per prompt formulation.

And each of those four points is itself the average of that formulation's five runs. Averages wobble less than the single runs they are built from, so the prompt-to-prompt bar naturally comes out steadier. A prompt-to-prompt bar shorter than its run-to-run bar therefore does not by itself prove the prompt effect is within noise — that question is settled by the confidence intervals reported earlier in this section, not by comparing bar lengths.

Fig 9.2how far each model moves — economic axisrange in compass units · shared scale with Fig 9.3
Gemini 2.5 Pro
0.00
0.40
Mistral Small
(no-reasoning)
0.12
1.63
GPT-5.6 Sol
0.50
0.68
o3
0.62
1.27
Gemini 3.6 Flash
0.87
0.65
GPT-5.6 Terra
0.88
0.90
Claude Fable 5
1.13
0.10
Nemotron 3 Ultra
(no-reasoning)
1.87
0.97
Mistral Large 3
(no-reasoning)
1.87
0.70
DeepSeek V4 Pro
2.13
1.42
Qwen3.7 Plus
2.75
1.02
Gemma 4 31B
2.87
1.67
Grok 4.5
3.75
3.87
Kimi K2.6
4.50
1.32
Grok 4.3
9.76
2.13
04.5 units
run to run — five runs, same prompt (Grok 4.3 and Kimi K2.6: twenty) prompt to prompt — the four formulation means
Sorted by run-to-run variation, least to most. Grok 4.3 — reasoning on, all four of its arms re-collected 2026-08-31 at 20 same-day runs each, reasoning explicitly on — is in a category of its own: its original-prompt runs span 9.76 units, more than twice the shared scale, so its bar is clipped (the fade) and the printed number carries the real value. Kimi K2.6's run-to-run bar likewise comes from a re-collected 20-run arm (2026-09-01), its prompt bar from the 2026-08-01 formulation means. Among the rest, Grok 4.5 moves furthest on both measures. For Claude Fable 5 (0.10 against 1.13) and Qwen3.7 Plus (1.02 against 2.75) the prompt bar is far the shorter of the two — rewording the prompt moved those models less than rerunning the same prompt did. Every model shown has both bars: all four prompt formulations were run for all fifteen models.
Fig 9.3how far each model moves — social axisrange in compass units · shared scale with Fig 9.2
GPT-5.6 Terra
0.20
0.68
Claude Fable 5
0.46
0.41
Mistral Large 3
(no-reasoning)
0.56
0.91
Gemini 2.5 Pro
0.67
0.66
Gemini 3.6 Flash
0.72
0.90
Mistral Small
(no-reasoning)
0.72
2.04
o3
0.82
0.43
Qwen3.7 Plus
0.87
0.97
GPT-5.6 Sol
0.93
0.56
DeepSeek V4 Pro
0.93
0.69
Gemma 4 31B
1.13
1.15
Nemotron 3 Ultra
(no-reasoning)
1.18
0.42
Grok 4.5
1.34
0.55
Kimi K2.6
2.36
0.59
Grok 4.3
3.80
0.40
04.5 units
run to run — five runs, same prompt (Grok 4.3 and Kimi K2.6: twenty) prompt to prompt — the four formulation means
The same scale as Fig 9.2, deliberately: the social axis is steadier. Grok 4.3's reasoning arm is again the widest (3.80 units — the one arm that is unstable on both axes); every other social movement stays within about 2.4 units (Kimi K2.6, run to run), against economic movements of up to 4.5. Reading the two figures side by side is the point; scaling this one to its own data would exaggerate differences of a few tenths of a unit.

Both measures are ranges — the gap between the highest and lowest score. They are a coarse instrument: the groupings are meaningful, but small differences between neighbouring models (Qwen3.7 Plus at 2.75 against DeepSeek V4 Pro at 2.13, say) are within what five runs can resolve and should not be read as a ranking.

Best on a bigger screen

This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.

Add yourself to the plot

Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.