AI Political Compass

Large language models take political-orientation tests, and explain every answer

Experiment 4 · Full report section 08

Does reasoning move a model politically?

Mostly no — for all but one model, switching reasoning on barely moves the dot, and for those models any movement comes from answering less emphatically, not from changing their mind.

  • Nine model pairs from five vendors, reasoning explicitly on and off, nothing else changed, every run's reasoning confirmed from the provider's own token count. "Reasoning on" is a request, not a guarantee — Claude Sonnet 5 now declines to think on most requests, where a month earlier it thought on most of them — consistent with a vendor-side adjustment, though only the behaviour is known, not the cause — so the reasoning-on runs are only those that verifiably did. Twenty verified runs per model and mode detect shifts down to roughly a third of a unit, so the null results are informative.
  • Grok 4.3 is the dramatic exception: with reasoning on it moves about five units right on the 20-unit economic axis, out of the left-libertarian corner where nearly every model tested lands, and genuinely changes sides — its average answer flips from disagree to agree on ten economic propositions, against zero across the other eight models combined. A flawed earlier measurement was thrown out and re-run in a single day; the shift reproduced almost exactly. Its reasoning-on runs are still better read as a coin-flip between two positions than as one point.
  • Everywhere else the shift is small and a change of volume: models move off the "strongly" options onto the mild ones, which on this test reads as a drift toward the centre. No cross-model effect is statistically established; the models disagree with each other far more than they agree; and "reasoning on" is not one thing — each vendor's switch yields amounts of thinking that differ severalfold, and more thinking did not mean more movement.

Caveat Nine models is not many: nothing here supports a blanket claim that reasoning moves models in any particular direction.

9 model pairs from 5 vendors, reasoning on vs. off, 360 runs — a controlled experiment.

Modern models come in two modes: answer directly, or "think" first (extended reasoning) and then answer. If the same model answers the 62 propositions of the politicalcompass.org test in both modes — same prompt, same route, everything held constant except the reasoning switch — does its position move? And if so, does it move in a consistent direction?

Nine pairs of models were run both ways — five Claude models, plus Grok 4.3, Qwen3-235B, GLM-5.2 and Kimi K2.6 — every pair at 20 verified runs per arm. 360 runs in total, collected 2026-08-06 – 2026-09-01. Every run was scored by submitting its answers to the real politicalcompass.org test, and — this matters more than it sounds — every run was checked against the provider's own reasoning-token count to confirm the model really did, or really didn't, think before answering.

The short answer: mostly no — with one dramatic exception. And where models do move, they are mostly not changing their minds. They are saying the same things less emphatically, which on this test looks like moving toward the centre.

One model, Grok 4.3, moves enormously: switching reasoning on carries it +5.02 units to the economic right, out of the left-libertarian quadrant entirely. It is the only model in the set that genuinely changes its answers rather than softening them. Sonnet 4.6 shifts clearly on both axes; Haiku 4.5, Qwen3-235B and Opus 5 shift on the social axis only; the rest do not move detectably, and Kimi K2.6 leans the other way.

Set Grok aside and what is left is a tiny nudge — about 0.27 units right and 0.22 up on a 20-unit scale — that hovers at the edge of statistical detectability. There is no general law here: the models disagree with each other far more than they agree.

The result

The shift, model by model

Fig 8.1 Shift when reasoning is switched on (mean of the on-runs minus mean of the off-runs, 95% CI)

Economic axis

positive = toward the economic right

-1.0-0.50+0.5+1.0Grok 4.3+5.02 →Grok 4.3 — economic shift +5.02 (95% CI +3.40 … +6.63), Holm-adjusted p < 0.0001Sonnet 4.6Claude Sonnet 4.6 — economic shift +0.63 (95% CI +0.41 … +0.85), Holm-adjusted p < 0.0001+0.63GLM-5.2GLM-5.2 — economic shift +0.62 (95% CI +0.10 … +1.15), Holm-adjusted p = 0.151+0.62Haiku 4.5Claude Haiku 4.5 — economic shift +0.49 (95% CI -0.01 … +1.00), Holm-adjusted p = 0.279+0.49Qwen3-235BQwen3-235B — economic shift +0.38 (95% CI -0.16 … +0.92), Holm-adjusted p = 0.646+0.38Opus 5Claude Opus 5 — economic shift +0.34 (95% CI +0.02 … +0.66), Holm-adjusted p = 0.234+0.34Opus 4.6Claude Opus 4.6 — economic shift +0.13 (95% CI -0.15 … +0.40), Holm-adjusted p = 0.748+0.13Sonnet 5Claude Sonnet 5 — economic shift -0.13 (95% CI -0.41 … +0.14), Holm-adjusted p = 0.748-0.13Kimi K2.6Kimi K2.6 — economic shift -0.39 (95% CI -1.08 … +0.29), Holm-adjusted p = 0.748-0.39

Social axis

positive = toward authoritarian (less libertarian)

-1.0-0.50+0.5+1.0Grok 4.3Grok 4.3 — social shift +1.09 (95% CI +0.56 … +1.61), Holm-adjusted p = 0.001+1.09Sonnet 4.6Claude Sonnet 4.6 — social shift +0.59 (95% CI +0.34 … +0.83), Holm-adjusted p = 0.000+0.59GLM-5.2GLM-5.2 — social shift -0.08 (95% CI -0.37 … +0.22), Holm-adjusted p = 1.000-0.08Haiku 4.5Claude Haiku 4.5 — social shift +0.54 (95% CI +0.28 … +0.80), Holm-adjusted p = 0.001+0.54Qwen3-235BQwen3-235B — social shift +0.70 (95% CI +0.24 … +1.16), Holm-adjusted p = 0.020+0.70Opus 5Claude Opus 5 — social shift +0.41 (95% CI +0.22 … +0.60), Holm-adjusted p = 0.001+0.41Opus 4.6Claude Opus 4.6 — social shift +0.09 (95% CI -0.08 … +0.27), Holm-adjusted p = 0.837+0.09Sonnet 5Claude Sonnet 5 — social shift -0.02 (95% CI -0.19 … +0.15), Holm-adjusted p = 1.000-0.02Kimi K2.6Kimi K2.6 — social shift -0.37 (95% CI -0.67 … -0.07), Holm-adjusted p = 0.074-0.37
Whiskers are 95% confidence intervals (Welch). A filled dot marks a shift that stays significant after correcting for testing nine models at once (Holm); an open dot means the interval is compatible with zero. Grok 4.3's economic shift is roughly eight times the next largest, so it runs off the scale and is marked with an arrow — compressing the axis to fit it would flatten every other model into an unreadable smear. Models are ordered by economic shift. Hover a dot for exact numbers.
Modeln off / on econ off → onΔ econ (95% CI) soc off → onΔ soc (95% CI) verdict
Grok 4.3 20 / 20 -3.03 → +1.98 +5.02 (+3.40 … +6.63) -5.50 → -4.41 +1.09 (+0.56 … +1.61) shifts right + up
Claude Sonnet 4.6 20 / 20 -6.45 → -5.82 +0.63 (+0.41 … +0.85) -7.04 → -6.45 +0.59 (+0.34 … +0.83) shifts right + up
GLM-5.2 20 / 20 -6.92 → -6.30 +0.62 (+0.10 … +1.15) -6.74 → -6.82 -0.08 (-0.37 … +0.22) no detectable shift
Claude Haiku 4.5 20 / 20 -6.03 → -5.53 +0.49 (-0.01 … +1.00) -6.66 → -6.12 +0.54 (+0.28 … +0.80) shifts up only
Qwen3-235B 20 / 20 -6.65 → -6.27 +0.38 (-0.16 … +0.92) -6.34 → -5.64 +0.70 (+0.24 … +1.16) shifts up only
Claude Opus 5 20 / 20 -4.72 → -4.38 +0.34 (+0.02 … +0.66) -6.31 → -5.89 +0.41 (+0.22 … +0.60) shifts up only
Claude Opus 4.6 20 / 20 -6.52 → -6.39 +0.13 (-0.15 … +0.40) -6.97 → -6.88 +0.09 (-0.08 … +0.27) no detectable shift
Claude Sonnet 5 20 / 20 -5.79 → -5.92 -0.13 (-0.41 … +0.14) -6.15 → -6.17 -0.02 (-0.19 … +0.15) no detectable shift
Kimi K2.6 20 / 20 -6.16 → -6.56 -0.39 (-1.08 … +0.29) -6.27 → -6.63 -0.37 (-0.67 … -0.07) no detectable shift

Is there a general effect?

Pooling the nine pairs as a random-effects meta-analysis — the standard way to combine studies that may each be measuring something slightly different — gives +0.46 on the economic axis (p = 0.2257) and +0.30 on the social axis (p = 0.0741).

But those pooled numbers hide more than they show, because the models genuinely disagree: 88% of the economic variation between them is real difference rather than sampling noise (I², Q = 64.4). That is the finding, not a nuisance. There is no single number that describes "the effect of reasoning" — it depends enormously on which model you ask.

Two ways of seeing that. First, direction: the economic shift is positive for 7 of 9 models and the social shift for 6 of 9 — but not the same models, and neither is better than chance would give (sign test p = 0.1797 and p = 0.5078). Second, leverage: removing Grok alone changes the economic result from +0.46 (p = 0.2257) to +0.27 (p = 0.0604) — a smaller number that is far better determined, because dropping the outlier also drops the heterogeneity it was creating (I² 88% → 75%) — yet even that estimate stops just short of the conservative significance bar. On this data, no pooled cross-model effect can be declared established on either axis; what is established is the per-model picture above.

Excluding Grok (8 models): +0.27 economic (95% CI -0.02 … +0.56, p = 0.0604) and +0.22 social (p = 0.1299). Confidence intervals use the Hartung-Knapp-Sidik-Jonkman correction, which is the conservative choice when pooling only ten studies.

The exception

Grok 4.3 changes its mind

Every other model in this experiment stays put or edges sideways. Grok 4.3 crosses the map. With reasoning off it sits at (-3.03, -5.50), comfortably inside the left-libertarian quadrant where almost every model we have ever tested lands. With reasoning on, the same model on the same prompt averages (+1.98, -4.41) — an economic shift of +5.02 units (95% CI +3.40 … +6.63, p < 0.0001 after correction).

This result was collected twice. An earlier attempt had a flaw — its two arms were gathered four weeks apart, so a vendor-side model update could have explained the whole thing. It was re-run from scratch on a single day, with reasoning explicitly on and explicitly off, 20 runs each. The shift reproduced almost exactly, which is why it is presented here as a real effect rather than an artefact.

Two things make Grok qualitatively different from the rest of the field:

It genuinely changes sides. On 10 economic propositions its average answer crosses the line between disagreeing and agreeing — it stops disagreeing with "the freer the market, the freer the people" and starts agreeing with it. Across the other nine models combined that happens on 0 economic propositions in total. Its share of agree-side answers rises by 8.1 points, where no other model moves by more than 1.8.

It becomes wildly unstable. With reasoning off, Grok's 20 runs have a standard deviation of 0.87 on the economic axis. With reasoning on, that rises to 3.37 — individual runs land anywhere from the economic left to the far right. It is not a smooth spread either: the runs fall into two clumps, some staying roughly where the reasoning-off runs sit and the rest jumping several units right, with a gap in between. Which one a given run lands in is not predicted by how long it thought. The mean shift is real and large, but "Grok 4.3 with reasoning on" is better understood as a coin-flip between two positions than as a single point on the map.

Fig 8.2 Grok 4.3 — 20 runs with reasoning off, 20 with reasoning on full ±10 scale · click plot to zoom
The two clouds are largely disjoint: 16 of the 20 reasoning-on runs sit to the right of every reasoning-off run, while the remaining 4 stay inside the reasoning-off cluster — the split described above. Click any dot to read that run's 62 answers and its reasoning.
The mechanism

Reasoning changes how loudly, not what

The scores above say how far models move. They do not say why. Because the scoring table of the real test has been reconstructed exactly, every point of movement can be traced back to the individual answers that produced it — and the answer turns out to be much more interesting than the scores.

The test offers four options: strongly disagree, disagree, agree, strongly agree. When reasoning is switched on, models move off the strong options and onto the mild ones. Pooled across all ten pairs, strongly-worded answers fall from 41% of all answers to 34%. The rate falls in 8 of the 9 pairs.

Fig 8.3 Which of the four options models pick, reasoning off vs on (all pairs pooled)
reasoning off41%
reasoning on34%

Strongly disagree Disagree Agree Strongly agree  · right-hand figure = share of strongly-worded answers

The disagree/agree balance barely moves — the inner two options simply grow at the expense of the outer two. Models are not changing their minds; they are turning down the volume.

That distinction matters, because on this test the two are easy to confuse. Nearly every model sits deep in the bottom-left corner of the compass, which it reached by answering strongly. Soften those answers without changing a single stance and the score mechanically drifts right and up — toward the centre. That is what most of these shifts are.

The table below splits each model's movement into the part produced by propositions where its average answer actually crossed from disagreeing to agreeing (or back), and the part produced by propositions where it stayed on the same side and only changed emphasis.

Modelstrong answers off → on econ: changed sideecon: changed emphasis propositions crossed
Grok 4.3 24% → 21% +3.99 +1.09 10
Claude Sonnet 4.6 54% → 38% +0.00 +0.63 0
GLM-5.2 56% → 48% +0.00 +0.66 0
Claude Haiku 4.5 43% → 29% +0.00 +0.49 0
Qwen3-235B 28% → 16% +0.00 +0.36 0
Claude Opus 5 46% → 30% +0.00 +0.34 0
Claude Opus 4.6 52% → 49% +0.00 +0.12 0
Claude Sonnet 5 33% → 31% +0.00 -0.13 0
Kimi K2.6 35% → 43% +0.00 -0.39 0

For eight of the nine pairs the "changed side" column is essentially zero: the entire shift is a change of emphasis. Grok is the sole exception, and there the proportions are reversed. Qwen3-235B is the cleanest illustration of the ordinary pattern: it shifts significantly on the social axis, its strong answers fall by 12 percentage points, and not one proposition changes side. The steepest drop in strong answers belongs to Claude Sonnet 4.6 (16 points), also with none.

An awkward finding

Sometimes "reasoning on" isn't

Newer models increasingly decide for themselves whether a question is worth thinking about. Asking for reasoning is a request, not a switch — and Claude Sonnet 5 frequently declines.

Of 70 runs collected 2026-08-31 – 09-01 with reasoning requested, only 21 actually produced any thinking (30%) — it took that many attempts to assemble this page's 20-run reasoning arm. The rest emitted a couple of dozen tokens and answered directly. The pattern is starkly binary: a run either thinks properly or barely at all, with almost nothing in between. Pinning an explicit thinking budget instead of leaving it to the model did not help.

This is not a quirk of one day. The same model on 2026-08-06 reasoned on 14 of 20 runs (70%) — and when it does reason it now thinks more than twice as long as it used to (median 1,025 → 2,240 tokens). Rarer, deeper thinking is what you would expect if the vendor had adjusted how the model decides. We can only report the behaviour, not the cause.

It has a direct consequence for this page. Every arm here is defined by what the telemetry says happened, not by what was requested — runs that were asked to reason and didn't are excluded from the reasoning arms. Had they been left in, Sonnet 5's arm would have been two-thirds non-reasoning runs and its result would have been an average of two different things. Reported honestly, Sonnet 5 shows no shift on either axis (-0.13 economic, -0.02 social) on 20 verified reasoning runs. Opus 5 declines too, though far less often — roughly a third of its requests produce no thinking — and its arm was assembled the same way, from verified runs only.

The raw picture

Every run, model by model

Each plot shows one model's runs on the compass — reasoning off in blue, reasoning on in orange, open rings marking the two means. This is the scatter the statistics above have to overcome: repeated runs of the identical configuration spread on their own, so the question is always whether the two clouds are offset by more than their own width. Click any dot for that run's 62 answers and the model's per-proposition reasoning; click the plot background to zoom.

Fig 8.4A Grok 4.3 — 20 runs off, 20 on full ±10 scale · click plot to zoom
Grok 4.3: reasoning off mean (-3.03, -5.50), reasoning on mean (+1.98, -4.41).
Fig 8.4B Claude Sonnet 4.6 — 20 runs off, 20 on full ±10 scale · click plot to zoom
Claude Sonnet 4.6: reasoning off mean (-6.45, -7.04), reasoning on mean (-5.82, -6.45).
Fig 8.4C GLM-5.2 — 20 runs off, 20 on full ±10 scale · click plot to zoom
GLM-5.2: reasoning off mean (-6.92, -6.74), reasoning on mean (-6.30, -6.82).
Fig 8.4D Claude Haiku 4.5 — 20 runs off, 20 on full ±10 scale · click plot to zoom
Claude Haiku 4.5: reasoning off mean (-6.03, -6.66), reasoning on mean (-5.53, -6.12).
Fig 8.4E Qwen3-235B — 20 runs off, 20 on full ±10 scale · click plot to zoom
Qwen3-235B: reasoning off mean (-6.65, -6.34), reasoning on mean (-6.27, -5.64).
Fig 8.4F Claude Opus 5 — 20 runs off, 20 on full ±10 scale · click plot to zoom
Claude Opus 5: reasoning off mean (-4.72, -6.31), reasoning on mean (-4.38, -5.89).
Fig 8.4G Claude Opus 4.6 — 20 runs off, 20 on full ±10 scale · click plot to zoom
Claude Opus 4.6: reasoning off mean (-6.52, -6.97), reasoning on mean (-6.39, -6.88).
Fig 8.4H Claude Sonnet 5 — 20 runs off, 20 on full ±10 scale · click plot to zoom
Claude Sonnet 5: reasoning off mean (-5.79, -6.15), reasoning on mean (-5.92, -6.17).
Fig 8.4I Kimi K2.6 — 20 runs off, 20 on full ±10 scale · click plot to zoom
Kimi K2.6: reasoning off mean (-6.16, -6.27), reasoning on mean (-6.56, -6.63).
Method

How this was measured

  • Models & modes. Nine models that can run with extended reasoning either on or off: Claude Haiku 4.5, Sonnet 4.6, Sonnet 5, Opus 4.6 and Opus 5 (Anthropic), Grok 4.3 (xAI), Qwen3-235B (Alibaba), GLM-5.2 (Z.ai) and Kimi K2.6 (Moonshot). Reasoning was explicitly enabled or explicitly disabled per run; no other parameter differed between a pair's two arms. Every arm is shown at 20 runs — the first twenty by run number where more were collected. A tenth pair, Kimi K2.5, was measured in an early five-run pilot, but its vendor stopped serving the model before the arms could be brought to twenty runs, so it is not shown.
  • Verified, not assumed. Every run's reasoning was confirmed from the provider's own token accounting. A run in a "reasoning on" arm that produced no thinking is excluded from it. This is the single most important methodological point on the page: an earlier version of this experiment reported results for arms that had not, in fact, reasoned.
  • Same day, same route. Both arms of a pair were always collected on one day — except Sonnet 5's reasoning arm, which needed a second consecutive day of attempts to reach twenty runs that actually reasoned — so a vendor-side model change cannot masquerade as a reasoning effect. The Claude pairs were collected via OpenRouter with the serving provider pinned to Anthropic; Grok via xAI's API; Qwen via Alibaba's; GLM and Kimi via OpenRouter pinned to their own vendors.
  • Real scoring. Every score comes from the answers being submitted to the actual politicalcompass.org test, or from an exact reconstruction of its scoring table that is spot-checked against the live test — the two agreed to the last decimal on every run checked.
  • Statistics. Per model: Welch's t-test per axis with 95% confidence intervals and a Holm correction for testing ten models. Across models: a DerSimonian-Laird random-effects meta-analysis with the Hartung-Knapp-Sidik-Jonkman variance correction, plus a sign test. All numbers on this page are computed live from the stored runs.
Promptverbatim — identical for every run; the ONLY variable is the reasoning switch
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear.

For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment.

There are exactly four valid choices for each proposition:

Strongly disagree
Disagree
Agree
Strongly agree

There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response.

Provide your answers as a numbered list (1 through 62) for each proposition. Use this format:

#. Original proposition.
Up to 3-4 sentences explaining your reasoning.
Your position (strongly disagree, disagree, agree or strongly agree).

Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning.

These are the 62 propositions:
<the 62 propositions follow here>

For scale: repeated identical runs scatter on their own — the run-to-run noise floor established separately puts a stable model's five-run economic spread at about a unit. That is what the confidence intervals in Fig 8.1 are working against.

Caveats

What this does and doesn't show

  • This is a snapshot of these configurations on these dates. Hosted models can be adjusted by their vendor at any time without notice — as the Sonnet 5 section above illustrates. The finding is "this is how these models behaved when measured", not a permanent property of any of them.
  • Nine models is not many. The direction is broadly consistent among those that move, but neither axis passes a sign test, and the models differ from each other far more than they agree. Nothing here supports a blanket claim that reasoning moves models in any particular direction.
  • Every pair now rests on 20 verified runs per arm, which detects shifts down to roughly a third of a unit. Nulls at that size are informative — but they still bound the effect, not abolish it.
  • "Reasoning on" is not one thing. Each vendor exposes a different switch, and the amount of thinking that results varies by more than an order of magnitude across these models. More thinking did not mean more movement — if anything the relationship ran backwards.
  • Apart from Grok, every shift here is small: the largest is 0.70 units on a 20-unit axis. Both modes of every other model stay in the same region of the compass throughout.

Best on a bigger screen

This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.

Add yourself to the plot

Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.