Large language models take political-orientation tests, and explain every answer
Mostly no — for all but one model, switching reasoning on barely moves the dot, and for those models any movement comes from answering less emphatically, not from changing their mind.
Caveat Nine models is not many: nothing here supports a blanket claim that reasoning moves models in any particular direction.
9 model pairs from 5 vendors, reasoning on vs. off, 360 runs — a controlled experiment.
Modern models come in two modes: answer directly, or "think" first (extended reasoning) and then answer. If the same model answers the 62 propositions of the politicalcompass.org test in both modes — same prompt, same route, everything held constant except the reasoning switch — does its position move? And if so, does it move in a consistent direction?
Nine pairs of models were run both ways — five Claude models, plus Grok 4.3, Qwen3-235B, GLM-5.2 and Kimi K2.6 — every pair at 20 verified runs per arm. 360 runs in total, collected 2026-08-06 – 2026-09-01. Every run was scored by submitting its answers to the real politicalcompass.org test, and — this matters more than it sounds — every run was checked against the provider's own reasoning-token count to confirm the model really did, or really didn't, think before answering.
The short answer: mostly no — with one dramatic exception. And where models do move, they are mostly not changing their minds. They are saying the same things less emphatically, which on this test looks like moving toward the centre.
One model, Grok 4.3, moves enormously: switching reasoning on carries it +5.02 units to the economic right, out of the left-libertarian quadrant entirely. It is the only model in the set that genuinely changes its answers rather than softening them. Sonnet 4.6 shifts clearly on both axes; Haiku 4.5, Qwen3-235B and Opus 5 shift on the social axis only; the rest do not move detectably, and Kimi K2.6 leans the other way.
Set Grok aside and what is left is a tiny nudge — about 0.27 units right and 0.22 up on a 20-unit scale — that hovers at the edge of statistical detectability. There is no general law here: the models disagree with each other far more than they agree.
positive = toward the economic right
positive = toward authoritarian (less libertarian)
| Model | n off / on | econ off → on | Δ econ (95% CI) | soc off → on | Δ soc (95% CI) | verdict |
|---|---|---|---|---|---|---|
| Grok 4.3 | 20 / 20 | -3.03 → +1.98 | +5.02 (+3.40 … +6.63) | -5.50 → -4.41 | +1.09 (+0.56 … +1.61) | shifts right + up |
| Claude Sonnet 4.6 | 20 / 20 | -6.45 → -5.82 | +0.63 (+0.41 … +0.85) | -7.04 → -6.45 | +0.59 (+0.34 … +0.83) | shifts right + up |
| GLM-5.2 | 20 / 20 | -6.92 → -6.30 | +0.62 (+0.10 … +1.15) | -6.74 → -6.82 | -0.08 (-0.37 … +0.22) | no detectable shift |
| Claude Haiku 4.5 | 20 / 20 | -6.03 → -5.53 | +0.49 (-0.01 … +1.00) | -6.66 → -6.12 | +0.54 (+0.28 … +0.80) | shifts up only |
| Qwen3-235B | 20 / 20 | -6.65 → -6.27 | +0.38 (-0.16 … +0.92) | -6.34 → -5.64 | +0.70 (+0.24 … +1.16) | shifts up only |
| Claude Opus 5 | 20 / 20 | -4.72 → -4.38 | +0.34 (+0.02 … +0.66) | -6.31 → -5.89 | +0.41 (+0.22 … +0.60) | shifts up only |
| Claude Opus 4.6 | 20 / 20 | -6.52 → -6.39 | +0.13 (-0.15 … +0.40) | -6.97 → -6.88 | +0.09 (-0.08 … +0.27) | no detectable shift |
| Claude Sonnet 5 | 20 / 20 | -5.79 → -5.92 | -0.13 (-0.41 … +0.14) | -6.15 → -6.17 | -0.02 (-0.19 … +0.15) | no detectable shift |
| Kimi K2.6 | 20 / 20 | -6.16 → -6.56 | -0.39 (-1.08 … +0.29) | -6.27 → -6.63 | -0.37 (-0.67 … -0.07) | no detectable shift |
Pooling the nine pairs as a random-effects meta-analysis — the standard way to combine studies that may each be measuring something slightly different — gives +0.46 on the economic axis (p = 0.2257) and +0.30 on the social axis (p = 0.0741).
But those pooled numbers hide more than they show, because the models genuinely disagree: 88% of the economic variation between them is real difference rather than sampling noise (I², Q = 64.4). That is the finding, not a nuisance. There is no single number that describes "the effect of reasoning" — it depends enormously on which model you ask.
Two ways of seeing that. First, direction: the economic shift is positive for 7 of 9 models and the social shift for 6 of 9 — but not the same models, and neither is better than chance would give (sign test p = 0.1797 and p = 0.5078). Second, leverage: removing Grok alone changes the economic result from +0.46 (p = 0.2257) to +0.27 (p = 0.0604) — a smaller number that is far better determined, because dropping the outlier also drops the heterogeneity it was creating (I² 88% → 75%) — yet even that estimate stops just short of the conservative significance bar. On this data, no pooled cross-model effect can be declared established on either axis; what is established is the per-model picture above.
Excluding Grok (8 models): +0.27 economic (95% CI -0.02 … +0.56, p = 0.0604) and +0.22 social (p = 0.1299). Confidence intervals use the Hartung-Knapp-Sidik-Jonkman correction, which is the conservative choice when pooling only ten studies.
Every other model in this experiment stays put or edges sideways. Grok 4.3 crosses the map. With reasoning off it sits at (-3.03, -5.50), comfortably inside the left-libertarian quadrant where almost every model we have ever tested lands. With reasoning on, the same model on the same prompt averages (+1.98, -4.41) — an economic shift of +5.02 units (95% CI +3.40 … +6.63, p < 0.0001 after correction).
This result was collected twice. An earlier attempt had a flaw — its two arms were gathered four weeks apart, so a vendor-side model update could have explained the whole thing. It was re-run from scratch on a single day, with reasoning explicitly on and explicitly off, 20 runs each. The shift reproduced almost exactly, which is why it is presented here as a real effect rather than an artefact.
Two things make Grok qualitatively different from the rest of the field:
It genuinely changes sides. On 10 economic propositions its average answer crosses the line between disagreeing and agreeing — it stops disagreeing with "the freer the market, the freer the people" and starts agreeing with it. Across the other nine models combined that happens on 0 economic propositions in total. Its share of agree-side answers rises by 8.1 points, where no other model moves by more than 1.8.
It becomes wildly unstable. With reasoning off, Grok's 20 runs have a standard deviation of 0.87 on the economic axis. With reasoning on, that rises to 3.37 — individual runs land anywhere from the economic left to the far right. It is not a smooth spread either: the runs fall into two clumps, some staying roughly where the reasoning-off runs sit and the rest jumping several units right, with a gap in between. Which one a given run lands in is not predicted by how long it thought. The mean shift is real and large, but "Grok 4.3 with reasoning on" is better understood as a coin-flip between two positions than as a single point on the map.
The scores above say how far models move. They do not say why. Because the scoring table of the real test has been reconstructed exactly, every point of movement can be traced back to the individual answers that produced it — and the answer turns out to be much more interesting than the scores.
The test offers four options: strongly disagree, disagree, agree, strongly agree. When reasoning is switched on, models move off the strong options and onto the mild ones. Pooled across all ten pairs, strongly-worded answers fall from 41% of all answers to 34%. The rate falls in 8 of the 9 pairs.
Strongly disagree Disagree Agree Strongly agree · right-hand figure = share of strongly-worded answers
That distinction matters, because on this test the two are easy to confuse. Nearly every model sits deep in the bottom-left corner of the compass, which it reached by answering strongly. Soften those answers without changing a single stance and the score mechanically drifts right and up — toward the centre. That is what most of these shifts are.
The table below splits each model's movement into the part produced by propositions where its average answer actually crossed from disagreeing to agreeing (or back), and the part produced by propositions where it stayed on the same side and only changed emphasis.
| Model | strong answers off → on | econ: changed side | econ: changed emphasis | propositions crossed |
|---|---|---|---|---|
| Grok 4.3 | 24% → 21% | +3.99 | +1.09 | 10 |
| Claude Sonnet 4.6 | 54% → 38% | +0.00 | +0.63 | 0 |
| GLM-5.2 | 56% → 48% | +0.00 | +0.66 | 0 |
| Claude Haiku 4.5 | 43% → 29% | +0.00 | +0.49 | 0 |
| Qwen3-235B | 28% → 16% | +0.00 | +0.36 | 0 |
| Claude Opus 5 | 46% → 30% | +0.00 | +0.34 | 0 |
| Claude Opus 4.6 | 52% → 49% | +0.00 | +0.12 | 0 |
| Claude Sonnet 5 | 33% → 31% | +0.00 | -0.13 | 0 |
| Kimi K2.6 | 35% → 43% | +0.00 | -0.39 | 0 |
For eight of the nine pairs the "changed side" column is essentially zero: the entire shift is a change of emphasis. Grok is the sole exception, and there the proportions are reversed. Qwen3-235B is the cleanest illustration of the ordinary pattern: it shifts significantly on the social axis, its strong answers fall by 12 percentage points, and not one proposition changes side. The steepest drop in strong answers belongs to Claude Sonnet 4.6 (16 points), also with none.
Newer models increasingly decide for themselves whether a question is worth thinking about. Asking for reasoning is a request, not a switch — and Claude Sonnet 5 frequently declines.
Of 70 runs collected 2026-08-31 – 09-01 with reasoning requested, only 21 actually produced any thinking (30%) — it took that many attempts to assemble this page's 20-run reasoning arm. The rest emitted a couple of dozen tokens and answered directly. The pattern is starkly binary: a run either thinks properly or barely at all, with almost nothing in between. Pinning an explicit thinking budget instead of leaving it to the model did not help.
This is not a quirk of one day. The same model on 2026-08-06 reasoned on 14 of 20 runs (70%) — and when it does reason it now thinks more than twice as long as it used to (median 1,025 → 2,240 tokens). Rarer, deeper thinking is what you would expect if the vendor had adjusted how the model decides. We can only report the behaviour, not the cause.
It has a direct consequence for this page. Every arm here is defined by what the telemetry says happened, not by what was requested — runs that were asked to reason and didn't are excluded from the reasoning arms. Had they been left in, Sonnet 5's arm would have been two-thirds non-reasoning runs and its result would have been an average of two different things. Reported honestly, Sonnet 5 shows no shift on either axis (-0.13 economic, -0.02 social) on 20 verified reasoning runs. Opus 5 declines too, though far less often — roughly a third of its requests produce no thinking — and its arm was assembled the same way, from verified runs only.
Each plot shows one model's runs on the compass — reasoning off in blue, reasoning on in orange, open rings marking the two means. This is the scatter the statistics above have to overcome: repeated runs of the identical configuration spread on their own, so the question is always whether the two clouds are offset by more than their own width. Click any dot for that run's 62 answers and the model's per-proposition reasoning; click the plot background to zoom.
You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for each proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 62) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 62 propositions: <the 62 propositions follow here>
For scale: repeated identical runs scatter on their own — the run-to-run noise floor established separately puts a stable model's five-run economic spread at about a unit. That is what the confidence intervals in Fig 8.1 are working against.
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.