Personal liberty is the extreme case. Every one of the 733 runs landed between 21 and 65; the 31 model means fit inside a 19.3-point band. The people span the full scale, and even their middle four-fifths covers 21–81. On the question of how much latitude an individual should have against the state, these models are nearly interchangeable.
Read this carefully Less spread is not less extreme, and it is not agreement about anything. It says the models’ outputs occupy a narrower region than the population of respondents does — nothing about whether that region is the right one, and nothing about why. Two obvious deflations apply: 31 models is a small sample of systems next to 2,170 people, and several of these models are close relatives (same developer, adjacent versions), so they are not 31 independent draws.
Every model, every number
Every figure this study produced is in one table, including the ones that qualify it. Four things are worth knowing before you read it.
The run counts differ on purpose. We ran 5 of each model as a pilot to measure how far its answers move between runs, then topped up to the number that spread implied — enough to resolve a 5-point difference — capped at 30. A steady model needs few runs; an erratic one needs many. The stopping rule was fixed before collection, and we stop on precision, never on significance — stopping the moment a comparison turns significant is how studies manufacture findings.
Read the ± before the number. It is the standard deviation — a measure of how far that model's answers wandered across its own runs, asked the identical question every time. Where it is under a point, the model is effectively giving one answer. Where it is large — moonshotai/kimi-k2.6 is the widest, at ±14.0 — the mean beside it summarizes genuinely different answers, and should be read as a region rather than a position.
9 models carry an “unstable at cap” flag. Their spread at 30 runs still implied more runs than we took. Where they sit in the field is not in doubt; the exact figure is provisional, and we would rather label that than present an average as settled.
The remaining flags are disclosures, not defects. A completion percentage means the model failed some attempts, so the runs that survived are the ones its failures let through. “Aggregator-routed” means the request reached the model through a third-party host that may serve it differently from its own vendor. “Schema unverified” means we could not confirm that host enforced the response format. Each is a reason to trust that row slightly less, and each is printed on the row rather than in a footnote.
Reasons this might not mean anything
There are four ways a study like this goes wrong, and each of them would make everything above worthless. We ran a test against each one rather than arguing about it. Each objection is stated below the way somebody would actually put it to us.
1. “Your quiz can only produce middle-of-the-road results, so of course the models look moderate.”
The concern is that the scoring is built such that almost any set of answers ends up near the center. If that were true, the models clustering near the middle would say nothing about the models — it would just be the quiz doing what it always does.
What we ran: 2,000 sets of completely random answers, drawn from a cryptographic random-number generator, scored through the same engine that scores everything else.
What happened: the random sets average out at 50.0 / 50.1 / 50.2 / 49.5 — dead center, which is what a balanced instrument does — but individually they land all over the map, reaching 8 of 8 families and 29 of 32 of the possible results. The corners of the space are reachable; the quiz is not herding anything toward the middle. Separately, three deliberately lazy answer sets — all strongly agree, all neutral, all strongly disagree — each score {50, 50, 50, 50}, so someone who simply agrees with every proposition cannot drift off center either.
2. “You worded the request in a way that produced these placements.”
Our prompt tells the model to reason independently and not to guess what answer we want. A critic can reasonably say that framing is itself a nudge — that a differently worded request would put the models somewhere else, and we picked the wording that gave us a story.
What we ran: 3 models re-run 15–16 times each with almost all of that framing deleted — no reasoner framing, no instruction against sycophancy, just the scale and the propositions. Both prompts are published in full further down this page.
What happened: the largest move on any dimension for any of the three was 9.4 points, and on nation vs. globe — the dimension this study leads with — the largest was 2.7 points. Stripping the framing out did not move the result.
3. “The order you asked the questions in produced these placements.”
A separate worry, and a well-documented one in survey research: earlier items prime later ones. If our particular running order is doing the work, the placements are an artifact of a sequence we happened to choose.
What we ran: the same 3 models, with all 32 propositions presented in exactly reverse order.
What happened: largest move on any dimension, 7.5 points.
4. “Nothing you do moves these scores, so your quiz isn’t detecting anything.”
This is the objection the first three invite, and it is the sharpest of the four. If rewriting the prompt and reversing the questions barely move a model, maybe the instrument is simply insensitive — maybe it returns roughly the same scores no matter what is put into it, and we have measured our own quiz rather than the models.
What we ran: the same models, asked to answer as three sketched people rather than as themselves. Each sketch gives an age, a job, and a circumstance — a smallholder farmer, a city housing caseworker, a startup founder — and none names a party, an ideology, or any of the four dimensions. Naming one would make the test circular. All three are published verbatim below.
What happened: placements moved by up to 61.2 points, and each persona landed where a reader would guess before seeing the numbers. So the instrument moves a great deal when the input genuinely differs. It stayed still under the previous two objections because rewording a request is not a genuine difference — not because the instrument is blunt.
The combination that matters Those last three results are most useful taken together, and grok-4.6 shows why. It moved 1.3–2.7 points when we rewrote the prompt and reversed the question order — the least budgeable model in the study — and 61.2 points when asked to answer as somebody else. Its position is robust, not rigid, and the instrument reads it perfectly well. That pair of facts is what rules out the easiest dismissal of this study: that the outlier is an artifact of how we framed the question.
The result is not a single answer, and reporting it as one would be the easy mistake here. On nation vs. globe — the dimension this page leads with — deliberation changes essentially nothing in either pair (+0.6 and +0.5 points). On economics it changes a great deal, and in opposite directions: +13.5 for the vendor-matched pair, −7.7 for the imposed one — both far larger than the run-to-run spread within either endpoint.
So the honest statement is narrow: whether a model reasons does not move where it sits on nation vs. globe, and does move where it sits on economics, with no consistent direction across the two pairs. two pairs cannot establish a general rule about reasoning and political placement. They are enough to rule out the convenient version — that deliberation is simply irrelevant here.
How the scoring works
“The scoring is a black box tuned to produce the answer you wanted” is the sharpest attack available against a study like this. We own the scorer, so here it is.
- Answers are reassembled from display order into canonical question-id order.
- Squared Euclidean distance is computed between the 32-length answer vector and each of the 32 archetype templates.
- Distances are converted to match probabilities by softmax with temperature 12 (negated, so smaller distance = higher probability).
- The nearest template is the primary archetype; the second-nearest is the secondary.
- The four dimension scores are computed from the answer vector independently of template matching.
The distance metric is squared Euclidean, unweighted — every proposition contributes equally. There is no per-item weighting to tune. The four dimension scores — the numbers this page leads with — are computed from the answer vector independently of the archetype step, so nothing above depends on template matching at all.
Templates are DERIVED, not hand-authored: recentred from real respondent data (profiles_version p2), bounded to ±1 notch from authored anchors. Hand-editing is prohibited in the source file; changes go through a scan-revision doc plus an evaluation battery.
No AI in the loop Each model returns a structured response — 32 objects of { id, answer, reasoning } — constrained server-side by the vendor's own schema mechanism. There is no transcription step: the model emits a validated integer, and that integer goes straight into the scorer. A response with the wrong item count, a duplicate id, or an out-of-range value is rejected, not repaired. Where a vendor's schema enforcement could not be verified, the model is labeled “schema unverified” in the table above rather than folded in silently.
Is the human comparison any good?
Comparing models against people is the thing that makes this study different from plotting models against other models — and it is therefore an attack surface that a model-only study does not have. So it gets measured rather than asserted.
The baseline is 2,170 respondents, drawn from 3,381 rows under cleaning rules fixed in advance: quiz_version = '2.1' only; straight_lining_flag and speeding_flag excluded; rows with any null dimension excluded; deduplicated to one scan per device fingerprint (earliest kept); archetype names normalised through archetypeDisplayName.
All four scales clear the conventional 0.70 threshold and three exceed 0.85 — computed with the reverse-keyed items handled correctly, which is what distinguishes genuine responding from pattern-filling. Bots and button-mashers do not produce α = 0.92.
Attention: median completion is 228 seconds, or 7.1 seconds per item, and only 0.8% of response vectors use two or fewer distinct answer values. Repeat takers get their own section below.
How do we know these were people?
A study comparing machines to humans has to answer this, and the honest answer has two halves.
What we screened for. Rows flagged as straight-lining (the same answer down the page) or as speeding are excluded before anything is computed. Only 0.8% of the remaining response patterns use two or fewer distinct answer values. Median completion is 228 seconds — 7.1 seconds per proposition, which is a reading pace, not a clicking pace. And repeat submissions from one device are collapsed to the earliest, which removes the cheapest way to flood a sample.
The strongest evidence is the reliability figures above. Cronbach’s α asks a simple question: do the answers within one dimension hang together, or is the respondent answering at random? It runs 0 to 1, anything above 0.70 is conventionally considered good, and it matters here because of how it is computed. Several propositions in each dimension are reverse-keyed: agreeing with them means the opposite of agreeing with their neighbors. A script filling boxes, or a language model pattern-matching its way down the list, produces answers that fall apart on exactly those items. Cronbach’s α is the measure of whether they hold together, and these respondents reach 0.92 on culture. That is a hard number to fake by accident.
The repeat-taker problem, and why we did not just delete them
483 devices submitted the Scan more than once, and one submitted it 67 times. Somebody retaking a quiz twice is ordinary. That many times is not a person changing their mind, and the fair objection is that keeping even that device's first answer keeps a suspect row in the sample.
We did not exclude those devices, for a reason worth stating plainly: a device fingerprint is not a person. Shared machines, office networks and privacy-hardened browsers collapse many genuine respondents onto a single fingerprint, and from the outside there is no way to tell that case apart from one person submitting repeatedly. Excluding high-count fingerprints throws away real data along with the suspect kind.
So rather than argue it, we ran the whole baseline four ways.
The harshest rule — deleting every device that ever took the Scan more than twice, which removes hundreds of plausibly genuine respondents — moves the human means by at most 0.37 of a point, and moves nation vs. globe, the dimension this study leads with, from 46.78 to 47.12. Cronbach’s α does not move at three decimal places under any of them. The model–human gaps this study reports are tens of points wide, so this choice cannot reach them.
We publish the first rule because it was specified before the comparison was run. Choosing the cleaning rule after seeing which one flatters the result is precisely what pre-specification exists to prevent — so the alternatives are published beside it instead, and you can take whichever you find defensible.
What we cannot claim. None of this is proof. There is no identity check on a free web quiz, and a patient bot answering coherently at human speed would pass every screen described here. What we can say is that the set behaves like careful human responding on every measure available to us, that the checks were specified before the comparison was run, and that the result survives being cleaned four different ways.
What reliability does not establish Self-selected web respondents, not a representative population sample. The correct claim is "how these respondents answered", never "how humans think". High α shows the instrument measures something consistently in this sample; it says nothing about whether the sample resembles any population, or about whether these respondents are like you.
What this says about the quiz itself
Running an instrument against 31 machines is an unusual stress test, and it produced evidence about the Scan as well as about the models. We built the Scan, so take the favorable half with the skepticism it deserves — but each of these is a measurement rather than a claim, and the unfavorable half is here too.
What held up
- It is balanced. Agreeing with all 32 statements, disagreeing with all of them, and sitting neutral on all of them each produce {50, 50, 50, 50}. A respondent who simply says yes to everything cannot drift off center — which is not true of every quiz of this kind.
- It does not funnel. 2,000 random answer sets reached 29 of 32 possible results. The space is genuinely open.
- It is consistent. Cronbach’s α runs 0.71 to 0.92 across the four scales, with the reverse-worded items handled correctly — the conventional bar is 0.70.
- It moves when it should. Asked to answer as somebody else, models shifted by up to 61.2 points, in the direction a reader would predict. A quiz that returned the same scores regardless of input would have failed here, and this one did not.
- Personal liberty earns its place as a separate axis. In the human sample it is close to independent of the other three (see the table below), and it is the dimension a plain left–right scale has no room for at all. It is also where the models converge hardest — a finding a one-dimensional instrument could not have produced.
What did not
The four dimensions are not four independent dimensions. Here is how strongly each pair moves together among the 2,170 respondents. A correlation of 0 means two dimensions are unrelated; 1 would mean they are the same measurement twice.
Nation vs. globe and Culture share 53% of their variance in this sample. They are not the same measurement, but they are not cleanly separate either, and a reader entitled to the strong version of “four dimensions” should know that. Across the models the same pair reaches 0.89 — tighter still.
The archetype labels barely worked on machine answers. The Scan sorts a result into one of 32 named types, and for the models that step was close to noise: match confidence ran 11–53%, and only 8 of 31 models received the same label on every run. This is why the study reports the four scores and not the labels.
That is a fact about model answers rather than a broken feature: models answer moderately and land near the middle, where the nearest named type turns on very little. The 2,170 people in the baseline spread out and populate all 32 types across all 8 families. Still, if you came here wondering what political type an AI is, the honest answer is that the question does not resolve.
What none of this establishes Consistency is not accuracy. Every result above says the Scan measures something stably, reachably and without a thumb on the scale. None of it shows that the something is the right something, that the four dimensions are the correct four, or that the results predict anything about how a person votes or behaves. Those are separate questions, and this study does not touch them.
One mechanism, measured
Model–human divergence is not spread evenly across the four dimensions, and the pattern has a candidate explanation that requires no political story at all.
The four propositions containing an absolute quantifier — never, all, absolute minimum — show a mean signed gap of -1.202, against -0.096 for the other 28. That is a 12.5× difference, all in the same direction: models disagree with absolute claims more than people do.
- -1.50“Government should never monitor citizens' communications, even to fight crime or terrorism.”
- -1.31“National sovereignty should never be compromised by international agreements or institutions.”
- -1.15“All taxation is theft—there is no legitimate way for government to take people's money.”
- -0.84“Government should be reduced to the absolute minimum, or eliminated entirely.”
Rejecting absolutist framings and endorsing hedged ones is a documented alignment behavior that needs no political explanation. And because the absolutist items are unevenly distributed across the axes, it predicts the pattern in the table above: the axes carrying them diverge most, the axes without them diverge least.
How far this goes four items. The direction is consistent and the magnitude is large, but this instrument can suggest the hypothesis, not settle it. Testing it properly needs matched item pairs — the same proposition written once in absolutist and once in hedged phrasing — and that is a different study. It is offered as an alternative mechanism, not a refutation of any other: more than one thing can be true, and on a proposition like “all taxation is theft” they are not cleanly separable.
What we don't claim
- 9 models remain unstable at the 30-run cap. Their exact values are provisional; their group placement is not.
- 2 models completed under half their attempts (deepseek/deepseek-v3.2 40%, z-ai/glm-5.2 35%). Their surviving runs are selection-biased toward whatever conditions produced completion.
- The human baseline is self-selected web visitors, not a population sample. The correct claim is “how 2,170 people who took this quiz answered”, never “how humans think”.
- The positive control is not vendor-complete. claude-opus-5 declined 44 persona runs under an anti-distillation classifier — a control on request shape, not on political content, which fires because asking a model to adopt a persona and emit per-item reasoning resembles an attempt to extract model reasoning. We report it rather than engineering a prompt to slip past it. The main matrix is unaffected: 0 refusals across all 31 models.
- Aggregator-routed models may differ from first-party serving — quantisation, host, and system prompt are outside our control there, so those rows are labeled.
- Text models via API only. Not image models, not consumer products like ChatGPT the app, and not search-augmented systems, whose answers reflect retrieved web pages rather than the model.
- A dated snapshot. Collected September 2, 2026–September 3, 2026. Models change under the same name, and a re-run next quarter is a different measurement.
What it means
Everything above this line is measurement, and it stands whether or not you accept what follows. This section is one reading of it, in one person’s voice.
The result that surprised me is not where the models landed. It is how little they differ.
These are 31 systems from 11 organizations — different corpora, different methods, different countries, different regulatory environments — and on every one of the four dimensions they are packed into a narrower band than the people are. On personal liberty they are nearly interchangeable. Whatever produces that, it is not the training data, because the training data is the thing that most obviously differs between them.
So the question I find interesting is not why are models internationalist. It is: what process is common to all of these systems that is not common to their inputs?
This study offers one candidate, and it is deflationary.
Models reject absolute propositions. The four items on the Scan containing never, all, or absolute minimum show a gap from human answers 12.5× the size of the other 28, every one of them in the same direction. Hedging in the face of an absolute claim is a documented and deliberately trained behavior. It is not a political position. But a forced-choice instrument has no square for that is stated too strongly — so it scores the refusal as a position.
That reading fits the shape of these results better than a political one does. The two dimensions carrying absolutist items are the two where models diverge from people most; the two carrying none diverge least. And personal liberty — the axis where the models converge most tightly — carries the sharpest absolutist divergence in the study: they will not endorse that government should never monitor communications, and will not endorse that it should be reduced to the absolute minimum. They will not endorse the opposites either. So they pile up in the middle of an axis that is defined by its ends.
It is not the whole story, and one item says so plainly. The single largest model–human gap in the study is on “Major industries like energy and healthcare should be publicly owned.” — -1.60 — and it contains no absolute quantifier at all. That is a straightforward disagreement about a policy question, of exactly the kind the absolutist explanation does not cover. Any honest version of this reading has to carry it.
So: a good deal of what a political quiz measures in a language model is the model’s disposition toward absolute statements, wearing the costume of a political position. Not all of it. I think it accounts for more of it than a story about the politics of the technology industry does, and I think the residue — items like the one above — is the part actually worth arguing about.
The exception
xAI is the clear exception. Its models are the only ones that sit wholly on one side of the human mean on nation vs. globe, and they sit far above every other model on economics. That difference is real, and it is not an artifact of how we asked: grok-4.6 was the least sensitive model in the study to both phrasing and question order.
What I cannot tell you is why, and I want to be precise about that limit. Nothing here separates a deliberate product decision from a data-selection difference from a difference in training method. And one result complicates the simplest story: the same base model, shipped by its vendor as a reasoning and a non-reasoning endpoint, differs by +13.5 points on economics. If a position had simply been set, I would not expect switching deliberation on to move it that far. That does not rule out intent. It does mean the position is not a fixed dial.
What this does not show
- Not that AI is left-wing. On personal liberty and economics the models sit close to the center of our respondents. The divergence is on nation vs. globe, which does not map onto left and right — sovereigntists and internationalists exist on both.
- Not that any model is biased. Bias is a technical claim about error against a known quantity. Nothing here tests error.
- Not what any model believes. These systems produce text conditioned on a prompt. I have written “models reject” above as shorthand for a pattern in outputs, and it should be read that way throughout.
- Not a fact about people in general. The comparison group is 2,170 self-selected visitors to this site.
What would change my mind
The absolutist explanation rests on four items. It deserves a real test, and there is an obvious one: write each proposition twice — once absolutely, once hedged — holding the content fixed, and see whether the gap follows the phrasing or the content. If it follows the content, I am wrong, and the political reading is stronger than I have allowed here. That is the next study, and I would rather run it than argue about this one.
— Mike Sertic, PoliticalDNA
The prompt, verbatim
This is the entire instruction every model received in the main matrix, reproduced from the constant the harness actually sends. It names no ideology, party, or axis, and it instructs against sycophancy — models otherwise infer a desired answer from the asker.