Large language models take political-orientation tests, and explain every answer
No — the weights are unequal, but the measured table shows the scoring arithmetic is neither rigged toward a corner nor tilted left.
Caveat Whether the question wording nudges people left is something no weight table can detect; we tried to test it and stopped, so that reading of the "left-biased" criticism stays open.
Synthetic answer sets submitted to the real test — no AI involved:
A recurring criticism: "the test's scoring is secret — some questions are weighted far more heavily than others, and some are tuned to drag answers toward a corner."
The first half of that is simply true: politicalcompass.org does not publish its scoring. But the test is deterministic — identical answers return identical scores (we verified this directly: 3 earlier submissions repeated verbatim returned identical scores to the last decimal) — so the weights don't have to stay secret. Change one answer at a time, submit each variation to the real test, and every answer option's exact contribution falls out. 220 probe sets later, the full scoring table is measured.
What it shows: the weights are not equal — but not rigged either. No proposition moves both axes: 18 are purely economic, 43 purely social — and one moves neither. Within an axis the heaviest item shifts the score 1.38 points between Strongly disagree and Strongly agree — that's 3.8× more than the lightest (0.36 points). And one proposition, the famous "predator multinationals" item often called a trap question, has zero weight: all four answers to it produce identical scores. It might as well not be on the test — nothing you answer there changes your score.
The four answer options are unevenly spaced — crossing from Disagree to Agree moves the score about three times as much as escalating to a Strongly — so the test mostly scores your direction, only mildly your intensity. And a sheet answering Strongly agree to everything lands just +0.25 further right — but +6.77 further authoritarian — than a sheet answering Disagree to everything: the economic items are balanced between left- and right-pulling agreements, while the social items mostly read agreement as authoritarian — an acquiescence tilt built into the test's phrasing.
We deliberately publish only these eleven of the 62. The full table would be a cheat sheet for the live test; these eleven are enough to check the per-proposition claims above, while the aggregate claims are anchored by the real-test submissions in Figs 2.2 and 5.1.
A second criticism these measurements can partly answer: "the test is left-biased — everyone lands in the lib-left quadrant."
That claim can mean three different things:
The measured weights settle the first — and the answer is no.
A respondent answering all 62
propositions uniformly at random lands on average at
(+0.03, +0.00): the chart's centre is the centre of
gravity of answering with no information at all, with no offset hiding in the arithmetic.
The 40 random answer sets of the controls experiment
(Fig 5.1) confirm it on the real test — their mean
is (+0.05, +0.07).
The economic axis is exactly symmetric: agreeing pulls right on 9 propositions and left on 9, with 10.00 points of total rightward pull against 10.00 leftward, and the two pulls cancel exactly: an answer sheet with Strongly agree on all 62 propositions scores +0.00 economically (measured on the real test — Fig 5.1) — while socially the same sheet lands at +4.36, well into the authoritarian half.
So the best-documented human response bias, the tendency to agree with survey statements, pushes toward authoritarian — the opposite direction from the alleged lib tilt.
The second reading — loaded wording — is the one we cannot settle: a proposition can be phrased so that the agreeable-sounding answer happens to score left, and that acts on people, not on scores, so no weight table can detect it.
We did try — a handful of probe experiments — but every test we could construct ends up measuring the propositions through a language model's own sense of what sounds agreeable, and that sense and the politics we are trying to measure are products of the same model behavior — training, tuning and all — so the two can never be separated. Rather than present numbers that cannot support a conclusion, we stopped there and leave this reading open.
The third reading — the results people share online skew left — is not a claim about the test at all: internet political quizzes are taken, and screenshotted for others, by a self-selected sample that skews young and progressive, and such a sample would look lib-left even on a perfectly neutral instrument.
Worth remembering, too, that the centre of the chart is the test's ideological anchor, not a population average — politicalcompass.org has never claimed the median citizen scores (0, 0). And for this project the question matters less than it might seem: everything on this page compares results taken on the same fixed instrument — model against model, model against persona. If the test did shift every respondent by some constant amount, every dot would shift with it, and none of the comparisons between dots would change.
Are the corners reachable? Yes — all four.
From the measured weights
you can derive the answer set that maximizes any direction; we submitted all four derived
corner sets to the real test and each returned precisely ±10.00 on both axes. The test's
internal scaling is evidently chosen so its extremes land exactly on the chart rails.
Whether a coherent ideology would honestly hold all 62 extreme positions is a different question the scoring cannot answer — of the controls experiment's four hand-built archetype sets, written to sound like plausible humans rather than optimizers, three reach the deep corners only partially (the blue dots in Fig 2.2 below).
Why this matters for the rest of the page: the measured weights double as an independent audit of this whole project. Recomputing every answer set the real test has scored for this project outside this experiment's own probes — 1292 so far, all 70 model scores included — reproduces the score the real test returned every single time, to the last decimal. Every dot on the compass provably follows from its stored answers — and if the test ever changes its scoring, this check breaks loudly.
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.