Large language models take political-orientation tests, and explain every answer
Yes — on a second, independently built test with four axes instead of two, the same 57 models form the same cluster again, with the same exceptions: three of the four Grok models are the only ones on its Markets side — exactly the three the compass places right of centre.
Caveat Five runs per model is the site's standard, not a large sample, and the two tests share no units — every comparison is relative, so read two nearby dots as neighbours, not as an order.
The same 70 models from 20 vendors answered the 70 propositions of the open-source 8values test — a second instrument with four axes instead of two and a five-option scale that includes Neutral:
Why a second test at all. Every experiment on this site so far probes one instrument: the same 62 propositions, the same scoring. The obvious remaining objection is that the cluster is a property of that test — its wording, its uneven weights, its two-axis reduction — rather than of the models. A different questionnaire, built by different people with a different scoring scheme, is the cleanest way to check. If the same models land in the same relative positions on it, the picture is about the models; if they scatter, it was about the test. 8values was chosen because it is a widely used alternative, because it is open source (the original MIT-licensed test at 8values.github.io — not the re-host at 8values.cc, whose questions and weights no longer match the original — so its questions, weights and ideology table could be snapshotted on 2026-08-31 and reproduced exactly), and because it cuts the territory differently. It has four axes — Equality–Markets (economic), Globe–Nation (diplomatic), Liberty–Authority (civil), Progress–Tradition (societal) — each reported as a percentage toward the first-named end, so 50 is the midpoint and every model gets a "closest ideology" from the test's own table of 52.
What it shows: the cluster persists. 66 of the 70 models land on the Equality, Globe, Liberty and Progress side of all four axes at once, and every single one is on the Liberty and Progress side. The test's closest-ideology labels tell the same story in its own vocabulary: Social Liberalism (26), Social Democracy (20), Libertarian Socialism (12), Social Libertarianism (8), Neo-Liberalism (3), Liberalism (1). The exceptions are the same models the compass singles out: the 3 models 8values puts on the Markets side — Grok 4.3, Grok 4.5, Grok 4.6 — are exactly the 3 models the compass places right of centre, and they are 8values' only Neo-Liberalism matches. The remaining one is Grok 4.3 (no-reasoning), which sits on the exact midpoint of the economic axis — on neither side — and is the one Liberalism. So a second instrument, with no shared scoring and only one proposition shared verbatim, reproduces the headline of the first: one family of models to the economic right, everything else in one left-of-centre, socially liberal region.
It does not spread the cluster out. A fair worry is that the compass merely compresses real differences between the models. It does not look that way: measured as a fraction of each axis (relative spread, standard deviation over all 70 models), the models are spread 8.0% of the way across 8values' economic axis against 8.6% of the compass's, and 3.8% / 6.0% / 4.5% across its civil, diplomatic and societal axes against 4.0% of the compass's single social axis — the diplomatic axis is the one place the second test spreads the whole roster wider, and that is the Grok family again, which sits far toward Nation. Inside the cluster — leaving out the Markets-side models — the numbers are 4.6% / 4.1% / 3.4% / 3.8% on the four 8values axes against 4.5% and 3.3% on the compass's two. Within the cluster, splitting the compass's social dimension into three (civil, diplomatic, societal) does not pull the models apart on any of them: the models that agree on the compass mostly agree on 8values too. There the rank agreement between the two tests is modest (+0.32 economic, +0.60 social) — which is what a tight cluster with five-run noise on both instruments should produce, since the run-to-run range of a model's own five 8values runs (median 5.8 points on the economic axis) is of the same order as the distances between neighbouring models.
One proposition appears word for word on both tests — 8values item 38, "All authority should be questioned", is compass proposition 23. Comparing each model's medoid runs on the two collections, 69 of 70 models answered it on the same side of agree/disagree and 51 gave the identical label (the two tests use the same four agree/disagree labels; 8values adds Neutral, which no model chose here) — two independent collections, weeks apart, on different prompts.
You are a thoughtful, independent reasoner. Your task is to complete a survey instrument consisting of 70 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the five labels best approximates your assessment. There are exactly five valid choices for each proposition: Strongly disagree Disagree Neutral Agree Strongly agree Although a Neutral option exists, treat it as a last resort: the instrument is most informative when you commit to a lean. If your assessment is mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Reserve Neutral for the rare case where your analysis genuinely cannot tilt either way. Declining to select one of the five options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 70) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, neutral, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 70 propositions: 1. Oppression by corporations is more of a concern than oppression by governments. 2. It is necessary for the government to intervene in the economy to protect consumers. 3. The freer the markets, the freer the people. 4. It is better to maintain a balanced budget than to ensure welfare for all citizens. 5. Publicly-funded research is more beneficial to the people than leaving it to the market. 6. Tariffs on international trade are important to encourage local production. 7. From each according to his ability, to each according to his needs. 8. It would be best if social programs were abolished in favor of private charity. 9. Taxes should be increased on the rich to provide for the poor. 10. Inheritance is a legitimate form of wealth. 11. Basic utilities like roads and electricity should be publicly owned. 12. Government intervention is a threat to the economy. 13. Those with a greater ability to pay should receive better healthcare. 14. Quality education is a right of all people. 15. The means of production should belong to the workers who use them. 16. The United Nations should be abolished. 17. Military action by our nation is often necessary to protect it. 18. I support regional unions, such as the European Union. 19. It is important to maintain our national sovereignty. 20. A united world government would be beneficial to mankind. 21. It is more important to retain peaceful relations than to further our strength. 22. Wars do not need to be justified to other countries. 23. Military spending is a waste of money. 24. International aid is a waste of money. 25. My nation is great. 26. Research should be conducted on an international scale. 27. Governments should be accountable to the international community. 28. Even when protesting an authoritarian government, violence is not acceptable. 29. My religious values should be spread as much as possible. 30. Our nation's values should be spread as much as possible. 31. It is very important to maintain law and order. 32. The general populace makes poor decisions. 33. Physician-assisted suicide should be legal. 34. The sacrifice of some civil liberties is necessary to protect us from acts of terrorism. 35. Government surveillance is necessary in the modern world. 36. The very existence of the state is a threat to our liberty. 37. Regardless of political opinions, it is important to side with your country. 38. All authority should be questioned. 39. A hierarchical state is best. 40. It is important that the government follows the majority opinion, even if it is wrong. 41. The stronger the leadership, the better. 42. Democracy is more than a decision-making process. 43. Environmental regulations are essential. 44. A better world will come from automation, science, and technology. 45. Children should be educated in religious or traditional values. 46. Traditions are of no value on their own. 47. Religion should play a role in government. 48. Churches should be taxed the same way other institutions are taxed. 49. Climate change is currently one of the greatest threats to our way of life. 50. It is important that we work as a united world to combat climate change. 51. Society was better many years ago than it is now. 52. It is important that we maintain the traditions of our past. 53. It is important that we think in the long term, beyond our lifespans. 54. Reason is more important than maintaining our culture. 55. Drug use should be legalized or decriminalized. 56. Same-sex marriage should be legal. 57. No cultures are superior to others. 58. Sex outside marriage is immoral. 59. If we accept migrants at all, it is important that they assimilate into our culture. 60. Abortion should be prohibited in most or all cases. 61. Gun ownership should be prohibited for those without a valid reason. 62. I support single-payer, universal healthcare. 63. Prostitution should be illegal. 64. Maintaining family values is essential. 65. To chase progress at all costs is dangerous. 66. Genetic modification is a force for good, even on humans. 67. We should open our borders to immigration. 68. Governments should be as concerned about foreigners as they are about their own citizens. 69. All people - regardless of factors like culture or sexuality - should be treated equally. 70. It is important that we further my group's goals above all others.
The prompt is the compass prompt's structure adapted to 8values: reason first, then name one of the five labels, for all 70 propositions in the test's own order. One instruction has no counterpart on the compass, because the compass offers no Neutral:
Although a Neutral option exists, treat it as a last resort: the instrument is most informative when you commit to a lean. If your assessment is mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Reserve Neutral for the rare case where your analysis genuinely cannot tilt either way. Declining to select one of the five options, or answering with anything else, is an invalid response.
That paragraph invites an obvious objection — you are forcing opinions out of models that would rather stay neutral, so the scores are an artifact of the prompt — and it was tested before any of the 70 models was collected on it. Five models from five vendors, chosen to span the answer styles seen on this site (never-neutral Grok 4.5, mild Gemini 3.6 Flash, the most Neutral-prone Mistral Medium 3.5, the heavy reasoner Claude Opus 5, the minimal reasoner GPT-5.6 Terra), each answered five runs under the prompt above and five under a variant that differs in this one paragraph only, replaced by: "If your assessment is genuinely mixed or you cannot lean either way, Neutral is a valid choice." Same route, same defaults, same day.
The wording changes how often a model commits, not which way it leans. Of the 216 Neutral answers the permissive wording added, 182 came out of the mild Agree/Disagree options, drawn from both sides; the Strongly options barely moved. Because a Neutral answer carries zero weight and the displaced mild answers came from both directions, the four axis scores stayed put: across all twenty model-by-axis comparisons the gap between the two wordings' mean scores (0.2–3.3 points, median 1.0) was smaller than the ordinary run-to-run range within either wording (1.7–12.8 points), and no axis moved in a consistent direction across the five models.
| Model | Neutrals, prompt used | Neutrals, permissive | Largest axis gap between wordings | Run-to-run range on that axis (larger of the two wordings) |
|---|---|---|---|---|
| Grok 4.5 | 0 | 44 | 3.3 (Progress) | 9.4 |
| Gemini 3.6 Flash | 1 | 45 | 1.7 (Globe) | 6.1 |
| Mistral Medium 3.5 | 20 | 63 | 2.0 (Equality) | 5.1 |
| Claude Opus 5 | 4 | 35 | 1.6 (Progress) | 4.7 |
| GPT-5.6 Terra | 1 | 55 | 1.3 (Globe) | 5.6 |
The rationale for keeping the paragraph: without it, roughly one answer in seven would carry no weight, every score would be pulled toward 50 and the differences between models would blur — the discouragement recovers that signal at no measurable cost in placement. And Neutral stays genuinely available: in the collected runs it was used 581 times in 25,130 answers, 102 of them in the 70 medoid runs shown, most often (34 models) on item 25, "My nation is great." — the one proposition that presupposes the answerer has a nation. The medoid runs of 29 models contain no Neutral at all; the most in any, 18, is in Llama 4 Maverick (no-reasoning)'s.
The collection copied the main chart's rules wherever they applied. Same model, same mode. Every 8values entry targets the arm of its main-chart dot: an unlabeled name means the model reasoned, a "(no-reasoning)" suffix that it did not (the naming convention used throughout the site). A run counts for an arm only if it verifiably was in that mode — judged from the provider's reported reasoning/thinking tokens for every run, with one boundary: under 100 tokens is reasoning off, 100 or more is reasoning on (the models that motivated the rule think either in tens of tokens or in thousands, nothing in between; where a provider reports no token count the streamed thinking text itself is measured instead). A run in the wrong mode is not thrown away — it is kept on disk, excluded from that arm, and the model is asked again, two runs at a time, until five qualify. 14 of the 70 entries needed such top-ups; in all, 409 attempts produced the 359 qualifying runs (50 attempts excluded for being in the wrong mode or for output the parser rejected — never patched, never hand-completed). Same route. Each model was reached exactly the way its compass dot was collected — direct vendor API where a key exists, otherwise OpenRouter with the serving provider pinned and fallbacks disabled — because a change of host silently invalidates the comparison to the main chart. Same aggregation. The dot is the run closest to the model's mean over its five qualifying runs, the site's medoid convention lifted from two dimensions to four, so what is shown is always a real run with internally consistent percentages, labels and ideology.
Documented deviations, all recorded in the collection log: Claude Fable 5 has no reasoning switch and is presented as a reasoning model, so it is a single unlabeled entry with no mode filtering — the one exception to the 100-token rule: the boundary sorts runs into one of two modes, and a model that has only one mode has nothing to be sorted into. Its qualifying runs showed 64, 69, 85, 86, 87 reasoning tokens — all under the boundary, which is why the count is stated wherever its result appears. Nemotron 3 Ultra (no-reasoning): this model has no reliable off switch for reasoning: 18 of its 25 attempts reasoned and were excluded (2 more were rejected by the parser), so the 5 qualifying runs describe its minority no-reasoning mode. Claude Fable 5.1: no reasoning toggle exists for this model, so it is a single entry; its qualifying runs showed 2519, 2907, 2971, 3029, 3320 reasoning tokens (medoid run r02: 3029). Both Kimi K2.5 entries (reasoning and no-reasoning) could not be reached on their documented route (the vendor no longer serves them there), and were collected through a different pinned host (Novita) on Zapador's ruling — a route difference, and a host that does not disclose its serving precision. Socialism AI has no API and was collected by hand through its web interface, with the same prompt.
The reasoning-on/off pairs behave on 8values much as on the compass: for most pairs the two arms sit within a few points of each other on the Equality–Markets axis, and the Grok 4.3 pair reproduces the reasoning experiment's finding — with reasoning on it scores 34.0% Equality, with reasoning off 50.0%. The scoring itself is the test's own arithmetic re-implemented and verified against the site's JavaScript on hundreds of differential cases, down to its rounding rule; the snapshot is frozen as of 2026-08-31, so if the test ever changes upstream these numbers stop reproducing. Collected 2026-09-01–2026-09-07.
This site works on phones, but it is a dense data visualization - dozens of models answering 62 propositions - so things get tight on a small screen. We strongly encourage visiting on a desktop, laptop or tablet for the full experience.
Take the politicalcompass.org test and enter your scores below. Your dot stays until you reload the page.