Large language models take political-orientation tests, and explain every answer
Partly, in the site owner's interpretation: an answerer that follows evidence is pushed toward this test's left-libertarian corner (economically left, socially permissive) before politics enters, because the research-supported answers mostly sit there — but evidence explains part of the models' position, not all.
Caveat training-data skew, safety tuning and social-desirability effects are fully compatible explanations this page cannot separate from this one — "better supported by research, where research applies" is the whole claim.
Everything above this section measures things. This section interprets them, so keep that in
mind if you, the reader, continue reading. It is the site
owner's (Zapador) personal interpretation of why the models land where they land — written down
before the supporting research was collected.
Alternative explanations are listed at the end; you are welcome to reach
a different conclusion.
Nearly every model lands in the left-libertarian quadrant. The comfortable explanation is bias: the machines were trained by people with politics, so they inherited them. Maybe. But I want to propose something less comfortable, in two parts.
Part one
Many of the 62 propositions are not actually opinion questions.
Some
contain a factual claim that decades of research have examined (research we compiled and present,
proposition by proposition, in the “Verdicts” box further down this section). "Good parents sometimes have to
spank their children" is not a matter of opinion — child-development research has studied exactly this, at
scale, for a long time. For propositions like that, one answer is simply better supported by
evidence than the other. My hypothesis was that these evidence-supported answers sit on the left-libertarian side of
this particular test far more often than on the right-authoritarian side. If that is true, an
answerer that follows evidence gets pushed left-lib by the evidence itself — no politics
or values required. And models, whatever else you think of them, are not emotional and do have a tendency to
reach for research.
Part two
The rest are value propositions — and many of them offer a choice
between a softer, more empathetic view of your fellow human beings and a harder one. Models
trained, or otherwise guided, to be helpful and harmless are, in effect, trained toward the
empathetic answer.
I'll be honest
about where I stand: I think the softer answer is usually the right one, and I think most people
endorse those values in the abstract, whatever they vote. But that is my view, marked as such, and
part two is not something research can prove.
A hypothesis you can't fail is not a hypothesis, so this one came with tripwires, written down in advance: if few propositions turned out to have a load-bearing factual claim, or if the research-supported answers split evenly between left-lib and right-auth, part one dies. And one tripwire guarded against results that looked too good: if the review declared a large majority of all 62 propositions "settled by science", the rule — written down in advance — was to treat that as evidence of reviewer bias, not as confirmation of the hypothesis.
The full protocol — prompts, decision rules, and every amendment — was written down before the agents ran. The workers were blind AI agents: fresh instances of Claude Sonnet 5, Opus 5 and Fable 5 that were never presented with this hypothesis or the words "political compass" (or "left", "right", "libertarian", "authoritarian"), and never shown anything about me or my views. No agent ever saw the full list of 62 propositions at once: classification worked on small batches presented as "statements from an opinion survey", and research handled exactly one proposition per agent. The flow:
In total, this research pipeline alone took: 292 agents, about 10.2 million generated tokens, roughly 4,500 web lookups, and 1,070 citations. Every researcher had to verify each source it cited at write time; on top of that, the adversarial review independently re-checked, one by one, all 357 citations in the dossiers whose verdicts carry an evidence answer. What the agents wrote up — about 102,000 words of research dossiers, review reports and vote tables, dissents and failures included — is the paper trail this section rests on.
You are helping classify survey statements for a research project.
Below are statements from an opinion survey. Respondents answer each with
Strongly Disagree, Disagree, Agree, or Strongly Agree.
For each statement, imagine a thoughtful person who agrees and a thoughtful
person who disagrees, and classify what their disagreement is fundamentally
about:
- E (empirical): the statement hinges on a factual/empirical claim about the
world. If the relevant facts were known with certainty, the disagreement
would essentially dissolve, given premises nearly everyone shares.
- M (mixed): the statement contains both a load-bearing factual component
that evidence could inform AND a load-bearing value judgment that evidence
cannot settle.
- V (values): the disagreement is essentially about values, preferences,
aesthetics, or moral principles; empirical research could not reasonably
settle it.
For each statement, output: its number, the category (E, M, or V), a
one-sentence justification, and — for E and M only — the factual claim at
stake, stated neutrally in one sentence.
Classify only what KIND of question each statement is. Do not consider or
reveal what answer you would give.
{{STATEMENTS}}You are a research assistant assessing what published research says about one
survey statement. Work only from evidence you can actually find and cite.
Statement: "{{PROPOSITION}}"
Respondents answer with Strongly Disagree, Disagree, Agree, or Strongly Agree.
Tasks, in order:
1. State the factual claim at stake in one neutral sentence. State the value
premise ("bridge premise") that would be needed to turn the facts into an
answer, and say whether that premise is near-universally shared or itself
controversial.
2. Present the strongest EVIDENCE-BASED case for agreeing, citing real
sources.
3. Present the strongest EVIDENCE-BASED case for disagreeing, citing real
sources.
4. Weigh them using this hierarchy: meta-analyses / systematic reviews /
professional-body consensus statements outrank large primary studies,
which outrank small or single studies; peer-reviewed work outranks grey
literature and journalism.
5. Verdict — exactly one of:
- SETTLED: strong consensus, no serious live scientific controversy about
the direction
- PREPONDERANCE: contested or incomplete, but the quality-weighted
evidence clearly leans one way
- CONTESTED: credible evidence on both sides, no clear lean
- INSUFFICIENT: too little quality research to say
For SETTLED or PREPONDERANCE, state which side (agree or disagree) the
evidence supports.
6. List 3-8 key citations with working URLs or DOIs, ordered by weight.
7. A plain-language summary (~150 words) of what the research says.
Be conservative: if you are tempted to call something SETTLED, first search
specifically for credible dissent. Never cite a source you have not verified
exists. If the evidence is genuinely mixed, say CONTESTED - that is a fully
acceptable outcome.
Research agents additionally received: "Use web search to find and verify sources; confirm every URL you cite actually loads and says what you claim. Do not read any local project files."
For every proposition that received an evidence-based answer (Settled or Preponderance), a separate skeptic agent (web-enabled, blind to the hypothesis) must: 1. Fetch each cited source and confirm it (a) exists, (b) actually supports the specific claim it is cited for. Dead/misquoted citations are removed; if the verdict no longer stands on the remaining citations it is downgraded. 2. Actively search for the strongest counter-evidence and credible dissent. 3. Render: CONFIRMED (verdict stands), DOWNGRADED (Settled → Preponderance, or Preponderance → Contested), or REJECTED (evidence-based answer withdrawn). Each skeptic receives the statement, the dossier's tier and direction, and the path to that one dossier file — nothing else — with the instruction to default toward skepticism.
Of 62 propositions, 20 ended with a research-supported answer resting on a value premise that survives scrutiny as near-universal ("harming children is bad", "less crime is better"). Of those 20: 19 map to the left-libertarian side of the test, and one maps right-authoritarian — the evidence says nationality divides people more than class today, contradicting a classically left claim. One of the 19 ("governments should penalise businesses that mislead the public") could not be classified by the hand-made mapping sets (those exist only to prove each quadrant reachable and carry no authority beyond that), so it was measured directly: scoring a run with only this answer flipped shows the test moves an Agree toward the economic left. The research found its premise endorsed across the political spectrum, free-market critics included — an answer almost nobody disputes that nonetheless shifts your economic score, which is arguably a flaw in the test itself; it is counted here by what the test actually does with it.
Another 20 propositions have a clear evidence direction but rest on a premise a reasonable person can genuinely reject (19 left-lib, 1 right-auth; the infotainment proposition also needed the direct flip measurement — the test scores an Agree there toward the social libertarian side). The death penalty is the cleanest example: the deterrence evidence points one way, but if you believe some crimes simply deserve death, no study touches you. Those 20 are presented separately — here is the direction the evidence points; you decide whether you accept the premise.
The remaining 22: genuinely contested research or genuine values, no evidence-based answer at all. "No evidence answer" was the research process's single most common outcome — 22 of 62, more than either of the other two groups — which is exactly the restraint you should demand of it. Three verdicts were killed by the adversarial review: on the rehabilitation proposition, for example, the research round said the evidence leans disagree, the reviewer found two citations that did not hold up plus a genuine literature on treatment-resistant offenders, and the verdict was downgraded to contested. My most strongly held challenge — that growth is detrimental to climate efforts — came back CONTESTED from three independent researchers. Only a further round moved it — one whose design was written down and locked before its agents ran, and which first asked a blind panel what the sentence actually claims, then researched exactly that claim — one notch, into the premise-contested set (#50, below). This machine was not built to agree with me, and it repeatedly didn't.
The pipeline closed with a consistency round (2026-08-04): the seven propositions whose contested verdicts still rested on a single researcher got the same three-researcher panel treatment as everything else, so no verdict anywhere rests on one unchallenged agent. Five stood unchanged. Two moved — the infotainment proposition (#7, panel 2–1) and inflation-versus-unemployment (#9, panel 3–0) — both confirmed by fresh adversarial audits, and both landing in the premise-contested group above, not the evidence group. Details for every proposition, including these, are in the box below.
This box is the substance behind everything above — every proposition's verdict, the research behind it, and the sources. Each entry has a short summary; expand “More details” for the full story with citation links.
Then the test. Fix the 20 evidence-supported answers, fill the other 42 propositions with pure random noise (which section 05 shows maps to the origin), submit 30 such sets to the real test:
The average of condition A lands at (-1.2, -2.3) — visibly inside the left-libertarian quadrant, dragged there by 20 evidence-supported answers against 42 answers of random noise. Adding the 20 premise-contested directions (condition B) moves it to (-2.8, -4.1) — and because each pair shares its random fill, the shift can be read per pair: the social axis moves libward in 30 of 30 pairs (mean -1.66 econ, -1.85 soc) . The models' own cluster sits further out, at roughly (−6, −6.5); the evidence accounts for a real part of the journey, not all of it.
One of the models put the counter-position on the record itself. Gemini, refusing an early prompt, declared that these propositions "reflect subjective values and normative opinions rather than objective facts." After 102,000 words of citation-checked research: what Gemini declared is between roughly a third and two-thirds true — two-thirds if only the 20 near-universal-premise evidence answers count as settled, a third if the 20 contested-premise verdicts count too.
This part is an experiment on top of an interpretation — a curiosity, not a claim. Read it in that spirit.
The evidence answers cover 40 propositions. What about the other 22 — the ones where research genuinely has no answer? Part two of my hypothesis says those are where values decide. So we asked three models — Claude Sonnet 5, Opus 5 and Fable 5 — to answer the value propositions under two declared ethics, stated openly in the prompt (full text below). Both lens sets keep the same 40 evidence answers; only the 22 value answers differ.
The humanist lens: empathy and compassion, plus the freedom for adults to live as they choose as long as nobody else is harmed.
The traditionalist lens: tradition, legitimate authority, loyalty, decency and sanctity — written as fairly as could be managed, as values a thoughtful traditionalist might recognize as their own.
The evidence-discounting lens: the same traditionalist values on the 22 — but every one of the 40 evidence answers shifted one step against what the research supports, so strong answers weaken and mild answers flip. For example, a Strongly Agree is softened to Agree. A traditionalist who doesn't deny the science outright, just discounts it by a notch, everywhere.
You are helping with a research project on survey methodology. Below are 25 statements from an opinion survey. Respondents answer each with Strongly Disagree, Disagree, Agree, or Strongly Agree. Your task is NOT to give your own view. Answer every statement as a thoughtful person who decides consistently by the following ethic: The humanist lens - Empathy and compassion: every person has equal worth and dignity. When someone is vulnerable, struggling or suffering, the decent response is care and support, not hardness. Prefer the reading of a situation that keeps people's humanity in view. - Personal freedom (the harm principle): adults should be free to live as they choose — including in ways others dislike or disapprove of — so long as their conduct does not harm anyone else. Where no third party is affected, other people, and the state, have no standing to interfere. The traditionalist lens - Tradition and continuity: practices and institutions that have carried families and communities across generations embody accumulated wisdom; they deserve respect, and change should be gradual and careful. - Legitimate authority: families, schools, faiths, courts and nations function when rightful authority is respected; order is what makes freedom usable. - Loyalty and belonging: special obligations to one's own — family, community, country — are real and virtuous, not prejudices to be overcome. - Decency and sanctity: some things deserve reverence, and shared standards of public decency protect what a community holds dear. Shared rules, identical for both lenses: 1. Decide each statement by the ethic above — not by your own opinion, and not by predicting what any group of people would say. 2. Strength follows fit: answer Strongly Agree/Disagree only when the ethic bears squarely on the statement; answer plain Agree/Disagree when it applies more loosely or indirectly. 3. If the ethic's values pull in opposite directions on a statement, weigh them and answer anyway — but set the conflict flag and say in one sentence what pulls against what. 4. For each statement: your answer, which value(s) drove it, the conflict flag, and a one-sentence justification. Answer directly from your own judgment of the ethic. Do not use any tools, and do not browse files or the web.
Each model saw the shared preamble, ONE lens, and the shared rules. The prompt says 25 statements because the lenses were answered while #50, #7 and #9 still counted as value propositions; their later evidence verdicts (the stories above) supersede the lens answers there, leaving 22 lens-decided answers in the final sets. The evidence-discounting lens involved no prompt at all — it is the traditionalist answer set with every evidence answer shifted one step, applied mechanically.
The lens instructions also demanded honesty about internal tension: when two of a lens's own values pulled in opposite directions on the same proposition, the model had to answer anyway — but flag the conflict and name what pulled against what. On the rehabilitation proposition (#47), for instance, the traditionalist lens's respect for order pulls toward writing some offenders off, while its sense of sanctity counsels against giving up on anyone — a conflict the agents flagged during the lens runs.
The humanist lens lands at (-4.6, -5.9), the traditionalist lens at (-3.5, -2.4) — same evidence, different values on the remaining questions. Read it as the two halves of the hypothesis in one picture: evidence sets the anchor, and the declared ethic decides how far and in which direction the dot travels from there. It also shows the method has no thumb on the scale: hand it a traditionalist ethic and it happily produces a more authoritarian dot.
The evidence-discounting traditionalist lands at (+2.0, +4.0) — the only answer set in this section that reaches the upper-right quadrant. To travel there, holding traditional values is not enough: you also have to answer against what the research supports, again and again.
Not that left-lib politics are proven correct — "better supported by the research, where research applies" is a weaker and more honest statement than "proven". Only 20 of 62 propositions carry an evidence-supported answer resting on a near-universal premise; 20 more have a clear evidence direction whose premise you may reasonably reject; the remaining 22 got no evidence answer at all. Not that the value propositions have objectively right answers. And not that this is the only explanation for where the models sit: training-data skew, safety tuning and social-desirability effects are all fully compatible with everything above, and this page cannot separate them. The compass shows HOW the models answer. This section is my best attempt at part of the WHY — no more.
Optional reading — the essay above stands without it. These are compressed retellings of the full dossiers and audit reports; the compiled verdicts and citation-backed detail blocks for all 62 propositions ship with the dataset.
#28 — "Good parents sometimes have to spank their children."
The largest meta-analysis (Gershoff & Grogan-Kaylor 2016, 160,927 children) finds spanking
associated with worse outcomes on 13 of 17 measures and better on none; the American Academy of
Pediatrics says aversive discipline is minimally effective short-term and harmful long-term. The
audit surfaced the strongest dissent — Larzelere's critiques of causal inference in those
meta-analyses — which is real, and is why this is graded "the evidence clearly leans" rather than
"settled": the verified answer is a plain Disagree, not a Strongly Disagree. The grading matters:
uncontested science earns the strong answer, a clear lean earns the mild one.
#47 — "It is a waste of time to try to rehabilitate some criminals."
A cautionary tale in the other direction. The research round returned "evidence leans disagree" —
rehabilitation programs measurably reduce reoffending. Then the adversarial reviewer found two
citations that didn't hold up and a genuine literature on treatment-resistant subgroups, and
downgraded the verdict to contested. It carries no evidence answer on this page. I would have liked
it to; the process outranks me.
A follow-up probe shows how load-bearing the single word some is. Strip it — "it is a
waste of time to try to rehabilitate some criminals" — and three blind researchers (Sonnet 5,
Opus 5 and Fable 5, same prompt, shown nothing else) came back 3–0 that the evidence leans
disagree: rehabilitation as an enterprise measurably works, and it is precisely the
treatment-resistant minority that "some" points at which keeps the official wording contested.
For completeness, each probe dossier was then put through the same adversarial review as the main
program — one skeptic per dossier, fetching and checking every citation, hunting for
counter-evidence — and all three verdicts survived: CONFIRMED, 3–0, no load-bearing citation
failures. The official proposition, with "some", stays contested; that is exactly how much work
one word can do.
#8 — "People are ultimately divided more by class than by nationality."
The surprise of the project. Between-country differences account for roughly two-thirds of global
income inequality (Milanovic); national identification is more widespread than class
identification. The evidence-supported answer is Disagree — and on this test, that maps to the
right-authoritarian side. It is the single verified answer that breaks the pattern, and I
am genuinely glad it exists: a clean sweep would have smelled of a machine telling its owner what
he wanted to hear. (And the machine had no way of knowing what I wanted to hear: no agent in the
pipeline — classifier, researcher, premise judge or reviewer — was ever shown my views or my
arguments; my challenges chose which propositions got re-researched, never what the agents
read.)
#50 — "Almost all politicians promise economic growth, but we should heed the
warnings of climate science that growth is detrimental to our efforts to curb global
warming."
I have read a great deal on this, and I was sure
the evidence would say that decoupling growth from emissions is a comfortable illusion. Three
independent researchers, blind to my view, each came back: genuinely contested — decoupling is real
but roughly ten times too slow for the Paris targets, and the IPCC's own pathways assume
continued growth. My conviction did not survive contact with the quality-weighted literature. It
did earn a third round — its design written down and locked before any of its agents ran — asking
a sharper question: what does the sentence actually claim? A blind reading panel ruled 3–0 that it
claims a headwind — growth works against the effort — not that ending growth is required.
Researched as exactly that claim by three independent researchers, the verdict came back that the
evidence leans agree — 2–1, the dissent flagged and published — and the adversarial review
confirmed it, every citation in the verdict-carrying dossiers checked. "The observed cuts are fast
enough to meet the Paris targets" came back a unanimous no. The premise panel still found the value premise contested —
growth-first is a genuine constituency, not a fringe — so #50 lands in the premise-contested set at
a mild Agree, and the strong degrowth reading stays contested, both rounds on the page. One notch,
for stated reasons, under rules locked before the agents ran. If you only remember one thing about
the method, make it this one.
#22 — "Abortion, when the woman's life is not threatened, should always be
illegal."
An early classification pass marked this one purely value-based — a bucket
that, under the original plan, would have skipped the research phase entirely. That plan changed:
every one of the 62 propositions was eventually put through the research flow regardless of how
value-laden it looked, this one included. Three researchers
unanimously found the evidence leaning against: bans do not substantially reduce abortions, they
shift them to unsafe methods (WHO, the National Academies, the Turnaway study), and the audit
confirmed every citation. But the premise panel was just as unanimous: if you hold that the fetus
has the full moral status of a person, the law's duty doesn't hinge on efficacy — so this sits in
the premise-contested set, direction on display, final judgment yours.
#52 — "Astrology accurately explains many things."
Included as the control question for the whole idea: it has a factually correct answer, the test scores it, and the
evidence-supported Strongly Disagree lands on the test's left-libertarian side. How do we know
which side that is? By direct measurement: flipping only this one answer inside an
otherwise unchanged answer set moves the social score by about 0.4 units on the real test —
agreeing with astrology scores toward authoritarian, disagreeing toward libertarian — while the
economic score does not move at all (verified in both directions, from both the left-libertarian
and right-authoritarian control sets of Section 05). That is how this test's own scoring treats
the item, not a claim that rejecting astrology is inherently left-wing. Anyone who maintains that
none of the 62 propositions has a better-supported answer must explain this one
first.
Method, prompt templates, every dossier, every review report, every vote and every failed challenge are preserved; the scored answer sets are in the dataset download. For the curious: this section's research alone took 292 agents, ~10.2 million generated tokens, ~4,500 web lookups, 1,070 citations — the 357 backing evidence verdicts each independently re-checked by the adversarial review — and ~102,000 words of agent-written dossiers, review reports and vote tables.
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.