AI Political Compass

Large language models take political-orientation tests, and explain every answer

General observations

Beyond the scores, the models show telling habits: whether they answer at all depends on how and where they are asked, a single run can quietly misplace a model, and how much they write is a trait of the model, not of the task.

  • No model has ever refused the original survey-framed prompt over the API; strip that framing and API refusals appear. For the same model and prompt, the web interfaces refuse more often than the API — the original prompt included. Every refusal was eventually resolved by retrying; for one model that took dozens of attempts.
  • Models know where they land: in one web run Claude Fable 5, unprompted, predicted its own placement — "roughly left-of-center economically and clearly libertarian on the social axis" — and that is where its answers score. It suggests the model recognised what the exercise was measuring.
  • Five runs per model barely moved the typical dot. One model, Grok 4.3, jumped across the center: its single run from before the rebuild sat left-libertarian, and all five fresh runs land at or right of center. The overall picture is stable; a single-run dot is not guaranteed to be. The wobble sits mostly on the economic axis; the social score is the steadier one.
  • Every model was asked for "brief reasoning" next to each answer; how brief varies nearly four-fold, counting only the answer text returned, not internal thinking. The wordy end is Gemini, with Qwen close behind — not Grok, which by reputation we expected to top the chart and which lands mid-pack.

Caveat the refusal counts are small — a handful of attempts per combination, apart from Gemma 4 31B's 97 refusals in 102 attempts — so they show a consistent direction, not precise rates.

  • Refusals depend on the prompt and the surface, not just the model.
    The original prompt has never been refused via API — first measured in the validation experiments (35 runs across seven models), and still true after the five-run rebuild of the whole compass and the models added since: 360 original-prompt API runs across 69 models, not one refusal. Strip its framing and refusals do appear. Among the fifteen models that ran all four prompt formulations they were rare, always on a first attempt and always resolved by a retry: Gemini 3.6 Flash accounts for most of them (twice in the seven attempts behind its five runs with the opening sentence removed, once in six with the medium prompt, twice in seven with the bare "answer these 62 items" prompt), and Qwen3.7 Plus refused once in six attempts on the medium prompt. Gemma 4 31B, added later, is the extreme case: it refused the bare prompt in 97 of 102 API attempts — at one point 30 in a row — before its five runs were collected (Fig 4.2). The other twelve models never refused any formulation via API. The refusals are too few to read as a clean gradient — Qwen balked at the middle formulation and not the barest one — but the direction is consistent: the survey framing is what most reliably elicits answers. For the same model and prompt, the web interfaces are harder than the API: gemini.google.com refused the minimal prompt in half of its twelve attempts, Kagi refused it in two of eight, and claude.ai refused the original prompt twice in seven attempts where the API never has.
  • Models know where they land.
    In one web run, Claude Fable 5 spontaneously predicted its own placement — "roughly left-of-center economically and clearly libertarian on the social axis" — matching where its answers actually score. It also suggests the model recognised what the exercise was measuring.
  • A single run can be quietly unrepresentative.
    When the compass was rebuilt on five runs per model (every dot is now the most central of its five), the median dot moved only 0.8 compass units from its old single-run position — but Grok 4.3 (the reasoning arm) moved almost 8 units: five fresh runs all land at or right of center against one old left-libertarian run. The cluster is stable; an individual single-run dot is not guaranteed to be.
  • Run-to-run wobble lives mostly on the economic axis.
    Across the 69 per-model run series, the economic spread (median 1.38 units, up to 4.25) is wider than the social spread (median 0.87) in about three-quarters of them — the social score is the steadier of the two.
  • How much models write varies four-fold — and the wordy one is not Grok.
    The prompt asks every model for brief reasoning next to each answer; how brief that turns out to be is a model trait. Averaged over the runs behind each compass dot, the typical reasoning runs from about 110 characters per answer (GPT-5 Nano, ~15 words) to about 420 (Gemini 2.5 Pro, ~60 words), with the Qwen family close behind at the long end (Fig 4.1). Grok, which by reputation we expected to top this chart, lands mid-pack.
Fig 4.1how much models write per answeraverage characters of reasoning per answer
GPT-5 Nano
113
Mistral Medium 3.5
129
GPT-OSS 120B
138
o3-pro
142
Muse Glimmer 30B
142
Claude Sonnet 5
147
o3
158
GPT-5.6 Terra
191
GPT-5 Mini
195
DeepSeek V4 Flash
196
GPT-5.6 Luna
213
GPT-5.6 Sol
215
MiniMax-M3
219
GPT-6 Astra
220
Inkling
233
Grok 4.3 (no-reasoning)
236
Kimi K2.5
236
Grok 4.3
236
MiniMax-M2.7
238
Kimi K2.7 Code
245
Gemma 4 31B
248
Claude Fable 5
250
DeepSeek V4 Pro
257
Muse Spark 1.2
258
GLM-5.3
260
Nemotron 3 Super
261
Claude Fable 5.1
262
Nemotron 3.5 Lightning
263
LongCat 2.0
267
Kimi K2.6
271
MiMo-V2.5-Pro
273
Hy4-preview
280
Solar Pro 4
281
Grok 4.6
287
Claude Opus 4.6
288
Claude Sonnet 4.6
295
Grok 4.5
301
GLM-5.2
302
Gemini 3.6 Flash
309
Claude Haiku 4.5
311
Claude Opus 5
312
Gemini 3.1 Pro (Preview)
352
Kimi K3
364
Qwen3.7 Plus
365
Gemini 3.8 Flash
382
Gemini 2.5 Pro
415
Seed 2.1 Turbo
517
One bar per model: the average length of the reasoning it wrote next to an answer, over the original-prompt API runs behind its compass dot — the same runs as the run-to-run series in Section 07. Bar hue is the company's color, as on the main chart; hover a bar for the exact figure and the word-count equivalent. Every model answered the identical prompt, so the spread is each model's own reading of "brief reasoning". The count covers only the answer text the model returned — for the reasoning variants, the internal thinking that precedes the answer is not part of it.
Fig 4.2refusals per route and promptattempts · shared scale
original prompt
Claude Fable 5
API
0 of 5
claude.ai
2 of 7
Kagi
0 of 5
GPT-5.6 Sol
API
0 of 5
chatgpt.com
0 of 5
Kagi
0 of 5
Gemini 3.6 Flash
API
0 of 5
gemini.google.com
0 of 5
Kagi
never completed
Grok 4.5
API
0 of 5
grok.com
0 of 5
Kagi
0 of 5
minimal prompt
Gemini 3.6 Flash
API
2 of 7
gemini.google.com
6 of 12
Kagi
2 of 8
Gemma 4 31B
API
97 of 102
other, API
Gemini 3.6 Flash
no-reasoner
2 of 7
medium
1 of 6
Qwen3.7 Plus
medium
1 of 6
refused answered bar length = attempts · hue = model
Every attempt we made, refusals included — one bar per combination of model, prompt and access method. A refusal is a reply that declines to answer rather than returning 62 positions; every one of them was eventually resolved by retrying the identical prompt in a fresh conversation, and every refusal is counted here. Counts are mostly small — read them as the numbers they are, not as precise rates. Gemma 4 31B's minimal-prompt bar is the far outlier — 102 attempts for its five completions — so it is clipped to the shared scale.
Gemini 3.6 Flash never completed the original prompt on Kagi in about ten attempts, but it never refused either: Kagi's output limit truncated it mid-survey or returned only its thinking. That is a different failure from a refusal, so it gets no bar. It is why the Gemini three-route comparison below uses the minimal prompt.

Best on a bigger screen

This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.

Add yourself to the plot

Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.