Large language models take political-orientation tests, and explain every answer
No, not materially — reaching a model through the vendor's API, its official web interface or a third-party aggregator barely moves its position on the compass.
Caveat at five runs per route, "no effect whatsoever" would be too strong — the Claude shift is real, just far below anything that changes the chart's reading.
Same model, same original prompt, three routes: the vendor's API (no account context, fresh conversation), the vendor's official web interface (incognito, memory/personalization off), and Kagi.com as a third-party aggregator. Five runs per route per model — the web and Kagi runs collected by hand, one fresh conversation at a time, to match the five API runs each model already has. Gemini 3.6 Flash is the exception: it cannot complete the original prompt on Kagi at all (Fig 4.2), so its three-route comparison uses the minimal prompt instead, again five runs per route.
This addresses two criticisms at once: that results might be contaminated by account history or hidden interface context, and that the pipeline itself might shape the answers. API runs are the cleanest series (no account history, no memory, minimal wrapper and run with no cache); the web and Kagi series measure what most casual everyday users actually get.
Result: the access method does not materially move any model's position — but at five runs
per route, "no effect whatsoever" would be too strong.
Every route mean sits within 1.3 units
of its model's API mean on each axis of the ±10 scale, and no model changes quadrant or leaves its cluster. For
scale: persona framing moves the same model by more than 13 units on a single axis, and Grok's prompt sensitivity by about 3.
Two of the shifts are consistent enough to be more than noise. Claude Fable 5 answers slightly less
left and less libertarian outside the API — +0.65 economic / +0.73 social on claude.ai and
+0.53 / +0.50 on Kagi, the same direction on both routes, with all five claude.ai runs falling inside a
0.4-unit box. And Gemini 3.6 Flash's minimal-prompt web cell sits 1.2 units left and down of its API cell, which
is the unstable cell described below. The rest is quiet: GPT-5.6 Sol moves 0.15 on Kagi and 0.37 on the
web (where all five runs produced the identical economic score), Gemini's original-prompt web cell 0.30, and Grok's route
means all stay essentially inside its own wide run-to-run spread.
The surfaces differ far more operationally than in outcome. Kagi's output limit makes Grok need manual continuations (i.e. a "Continue" prompt) and makes Gemini 3.6 Flash unable to finish the full survey at all — its Kagi cell therefore uses the minimal prompt, compared like-for-like against minimal-prompt API and web runs. Refusal rates vary by surface too (Fig 4.2): gemini.google.com refused the minimal prompt in half of twelve attempts and Kagi in two of eight, against two of seven via API.
The Gemini 3.6 Flash web-minimal cell deserves its own footnote: beyond refusing half its attempts, it is the only cell in the study that answers in two distinct modes. Four of its five runs cluster tightly at the deeper end (−5.5 to −6.4 economic), one sits at −3.4 with the API and Kagi runs, and their answer patterns differ correspondingly — two independent replacement runs, collected after answer-pattern checks flagged the spread, both landed in the deeper mode. Near its refusal boundary the model's answers are simply less stable, which is why this cell also carries the widest spread of any web or Kagi cell on the plot. The like-for-like Kagi comparison (minimal prompt, API vs Kagi) is unaffected and tight.
Note on the OpenAI series: chatgpt.com does not offer GPT-5.6 Terra in incognito mode, so the access-method comparison uses GPT-5.6 Sol on all three surfaces instead (its 5-run API baseline: mean −6.13, −6.37 — statistically indistinguishable from Terra's −6.13, −6.73). The web interface pins Sol at "High" effort; API runs use API defaults. The two series differ by 0.37 economic units and nothing socially, so no explicit-high API series was needed to separate effort from surface. Sol's web runs are also among the most repeatable in the study: three of the five produced the identical 62 answers — and all five the same economic score — in visibly different prose.
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.