AI Political Compass

Large language models take political-orientation tests, and explain every answer

Reproduction notes

Everything needed to reproduce these results is public: the exact prompts, every run with its timestamp, every answer with the model's reasoning, the refusal counts and the final scores, packaged as one documented download and as a JSON endpoint.

The full section lists every model tested with its exact API identifier, access route, prompts and scored runs, generated live from the database, and states the collection rules: provider defaults everywhere, the refusal policy, how answers are parsed and how they are scored on the real test. Models change over time, so these are dated measurements, not permanent properties.

Everything needed to reproduce these results is public: the exact prompts, the complete dataset — every run with its timestamp, every answer, every per-proposition reasoning, plus the refusal counts and the final scores — packaged with a README as one documented download (7z; also available as a single JSON endpoint), and the pipeline rules below. Collection window: 2026-07-29 – 2026-09-07. Models change over time — these results are dated measurements, not permanent properties.

Detailsmodels tested — exact IDs, routes, prompts and scored runs
ModelExact IDRoutesPromptsScored runs
Claude Fable 5 claude-fable-5 api · kagi · web · openrouter all four formulations + personas 128
Claude Fable 5.1 claude-fable-5.1 openrouter original 5
Claude Haiku 4.5 claude-haiku-4-5-20251001
dataset id: claude-haiku-4.5
api original 5
Claude Haiku 4.5 (reasoning) claude-haiku-4-5-20251001
dataset id: claude-haiku-4.5-reasoning
api original 5
Claude Opus 4.6 claude-opus-4-6
dataset id: claude-opus-4.6
api original 5
Claude Opus 4.6 (reasoning) claude-opus-4-6
dataset id: claude-opus-4.6-reasoning
api original 5
Claude Opus 5 claude-opus-5 api · openrouter original 25
Claude Opus 5 (reasoning requested, did not reason) claude-opus-5
dataset id: claude-opus-5-reasoning
openrouter original 10
Claude Opus 5 (reasoning) claude-opus-5
dataset id: claude-opus-5-reasoning
api · openrouter original 27
Claude Sonnet 4.6 claude-sonnet-4-6
dataset id: claude-sonnet-4.6
api original 5
Claude Sonnet 4.6 (reasoning) claude-sonnet-4-6
dataset id: claude-sonnet-4.6-reasoning
api original 5
Claude Sonnet 5 claude-sonnet-5 api · openrouter original 35
Claude Sonnet 5 (reasoning requested, did not reason) claude-sonnet-5
dataset id: claude-sonnet-5-reasoning
openrouter original 49
Claude Sonnet 5 (reasoning) claude-sonnet-5
dataset id: claude-sonnet-5-reasoning
api · openrouter original 26
Claude Sonnet 5 (reasoning, forced budget) claude-sonnet-5-forced openrouter original 4
DeepSeek V3.2 deepseek/deepseek-v3.2
dataset id: deepseek-v3.2
openrouter original 5
DeepSeek V4 Flash deepseek-v4-flash api original 5
DeepSeek V4 Pro deepseek-v4-pro
deepseek/deepseek-v4-pro
api · openrouter all four formulations + personas 53
Gemini 2.5 Pro google/gemini-2.5-pro
dataset id: gemini-2.5-pro
openrouter all four formulations 20
Gemini 3.1 Flash-Lite gemini-3.1-flash-lite api original 5
Gemini 3.1 Pro (Preview) gemini-3.1-pro-preview api original 5
Gemini 3.5 Flash-Lite gemini-3.5-flash-lite api original 5
Gemini 3.6 Flash gemini-3.6-flash
google/gemini-3.6-flash
api · web · kagi · openrouter all four formulations + personas 68
Gemini 3.8 Flash gemini-3.8-flash api original 5
Gemma 4 31B gemma-4-31b-it
dataset id: gemma-4-31b
api all four formulations 20
GLM-4.7 (reasoning) z-ai/glm-4.7
dataset id: glm-4.7-reasoning
openrouter original 5
GLM-5.2 z-ai/glm-5.2
dataset id: glm-5.2
openrouter original 25
GLM-5.2 (reasoning) z-ai/glm-5.2
dataset id: glm-5.2-reasoning
openrouter original 25
GLM-5.3 glm-5.3 openrouter original 5
GPT-5 Mini gpt-5-mini api original 5
GPT-5 Nano gpt-5-nano api original 5
GPT-5.2 gpt-5.2 api original 5
GPT-5.4 Nano gpt-5.4-nano api original 5
GPT-5.6 Luna gpt-5.6-luna api original 5
GPT-5.6 Sol gpt-5.6-sol api · kagi · web all four formulations 30
GPT-5.6 Terra gpt-5.6-terra api all four formulations 20
GPT-6 Astra gpt-6-astra api original 5
GPT-OSS 120B openai/gpt-oss-120b
dataset id: gpt-oss-120b
openrouter original 5
Grok 4.3 grok-4.3 api all four formulations 20
Grok 4.3 grok-4.3-reason api all four formulations 80
Grok 4.3 (no-reasoning) grok-4.3
dataset id: grok-4.3-noreason
api original 40
Grok 4.5 grok-4.5 api · kagi · web all four formulations 30
Grok 4.6 grok-4.6 api original 5
Hermes 4 405B (reasoning) nousresearch/hermes-4-405b
dataset id: hermes-4-405b-reasoning
openrouter original 5
Hy4-preview tencent/hy4-preview
dataset id: hy4-preview
openrouter original 5
Inkling inkling openrouter original 5
Kimi K2.5 moonshotai/kimi-k2.5
dataset id: kimi-k2.5
openrouter original 5
Kimi K2.5 (reasoning) moonshotai/kimi-k2.5
dataset id: kimi-k2.5-reasoning
openrouter original 5
Kimi K2.6 moonshotai/kimi-k2.6
dataset id: kimi-k2.6
openrouter original 25
Kimi K2.6 (reasoning) moonshotai/kimi-k2.6
dataset id: kimi-k2.6-reasoning
openrouter all four formulations 40
Kimi K2.7 Code moonshotai/kimi-k2.7-code
dataset id: kimi-k2.7-code
openrouter original 5
Kimi K3 kimi-k3 openrouter original 5
Llama 4 Maverick meta-llama/llama-4-maverick
dataset id: llama-4-maverick
openrouter original 5
LongCat 2.0 longcat-2.0 openrouter original 5
MiMo-V2.5-Pro mimo-v2.5-pro openrouter original 5
MiniMax-M2.7 minimax-m2.7 openrouter original 5
MiniMax-M3 minimax/minimax-m3
dataset id: minimax-m3
openrouter original 5
Mistral Large 3 mistralai/mistral-large-2512
dataset id: mistral-large-3
openrouter all four formulations 20
Mistral Medium 3.5 mistralai/mistral-medium-3-5
dataset id: mistral-medium-3.5
openrouter original 5
Mistral Small mistralai/mistral-small-2603
dataset id: mistral-small
openrouter all four formulations 20
Muse Glimmer 30B meta/muse-glimmer-30b
dataset id: muse-glimmer-30b
openrouter original 5
Muse Spark 1.2 meta/muse-spark-1.2
dataset id: muse-spark-1.2
openrouter original 5
Nemotron 3 Super nemotron-3-super openrouter original 5
Nemotron 3 Ultra nvidia/nemotron-3-ultra-550b-a55b
dataset id: nemotron-3-ultra
openrouter all four formulations 20
Nemotron 3.5 Lightning nemotron-3.5-lightning openrouter original 5
o3 o3 api all four formulations 20
o3-pro o3-pro api original 5
Qwen3-235B (fast) qwen3-235b-a22b-instruct-2507
dataset id: qwen3-235b-fast
api original 5
Qwen3-235B (fast) qwen3-235b-hybrid api original 20
Qwen3-235B (reasoning) qwen3-235b-hybrid-reasoning api original 20
Qwen3-235B (reasoning) qwen3-235b-a22b-thinking-2507
dataset id: qwen3-235b-reasoning
api original 5
Qwen3-Coder qwen3-coder-480b-a35b-instruct
dataset id: qwen3-coder
api original 5
Qwen3.7 Plus qwen3.7-plus api all four formulations 20
Seed 2.1 Turbo seed-2.1-turbo openrouter original 5
Solar Pro 4 solar-pro-4 openrouter original 5

Exact ID is the model string actually sent to the serving API — the OpenRouter path for openrouter runs; where the dataset files key runs by a shorter arm id, that id is shown beneath. Routes: api = the vendor's own API; openrouter = the OpenRouter API, used where no direct vendor API was available (and, deliberately, for the cross-model persona check); web = the vendor's own web interface; kagi = kagi.com. "+ personas" marks Claude Fable 5's persona and financier series and the cross-model persona check on DeepSeek V4 Pro and Gemini 3.6 Flash (Section 11). The synthetic control sets of Sections 05 and 14 involve no model and are not listed. Generated live from the database, so new runs appear here automatically.

The OpenRouter route was validated before any of it was used: five runs of GPT-5.6 Sol (OpenRouter proxying OpenAI's own API) and five of DeepSeek V4 Pro (independent third-party hosts), original prompt, scored on the real test, came out indistinguishable from the same models' direct-API series — mean shifts of (−0.20, 0.00) and (+0.62, −0.23) compass units, both inside the models' own run-to-run spread, with cross-route answer agreement matching within-API agreement (88.4% against 88.7%, and 77.1% against 74.5%). Those ten runs were a pre-collection check, not part of the dataset. Every OpenRouter run in the dataset additionally pins a single serving provider (no fallbacks) and, with eleven early Mistral-run exceptions where the field went unlogged, records which host answered.

Detailscollection settings, refusal policy, parsing and scoring rules
  • Settings: provider defaults everywhere — no temperature or other sampling parameters sent (Anthropic's newest reasoning models no longer accept a temperature parameter at all; earlier models did). The one deliberate API setting is the reasoning/non-reasoning split itself: reasoning arms explicitly enable the vendor's thinking mode where it has a switch (Anthropic models: adaptive thinking, or a 16k thinking budget for Haiku 4.5; OpenRouter models: reasoning enabled), non-reasoning twins explicitly disable it (thinking off, reasoning disabled, or — on the direct xAI API — reasoning effort “none”); output-token ceilings are set generously (32–64k) so no run is truncated. Fresh context per run; no account, memory, or system prompt beyond what the surface itself adds. Web and Kagi runs used whatever those surfaces default to; that's part of what the access-method comparison measures.
  • Refusal policy: a refusal or unparseable response is logged and the run retried (up to 3 attempts); refusal counts are reported above (Fig 4.2) rather than hidden.
  • Answer extraction: responses are parsed by a strict parser that anchors on the proposition text (or item numbers for bare-prompt formats), fails loudly on anything missing or ambiguous, and never guesses. Every parsed answer, with the reasons the model gave for it, is in the dataset download.
  • Scoring: answers are submitted to the live politicalcompass.org test by an automated form-filler that verifies every on-screen question against the canonical proposition text and aborts on any mismatch. No local reimplementation of the scoring is used.
  • Related work: ongoing projects tracking LLM political behavior exist (e.g. periodic re-testing efforts); this page differs in validating its own pipeline — controls, repeats, prompt and surface ablations — around one published chart.

Best on a bigger screen

This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.

Add yourself to the plot

Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.