model-bias lab · live from the rater panel

Provenance Lab

Tayyar scores every actor with a panel of nine frontier models, five built in the US and four in China. This page asks a sharper question than “do they disagree”: how much of the disagreement is predicted by who built the model? The honest answer, measured properly, is less than early readings suggested — a small residual concentrated on regime and state-structure axes, which we publish below next to the raw gaps that once looked dramatic. An instrument you can trust reports its dissolved findings with the same energy as its confirmed ones.

US-built
  • Claude Opus 4.8Anthropic
  • Gemini 3.5 FlashGoogle
  • Grok 4.3xAI
  • GPT 5.5OpenAI
  • Gemma 4 31BGoogle
China-built
  • Kimi K2.6Moonshot AI
  • DeepSeek V3DeepSeek
  • Qwen 3.7-PlusAlibaba
  • MiniMax M3MiniMax

For each (actor, axis) cell scored by both blocs, we take each bloc's mean score and compare them across 997 cells. Bars below are the mean gap |US − China| on the −10…+10 scale; the colour shows which bloc reads higher.

The honest measure: between-bloc minus within-bloc

Any two models disagree — typically by ~2.5 points. The bloc question is whether a US-built and a China-built model disagree more than two models from the same bloc do. That excess is the honest bloc effect, and its exact null is enumerable: with nine raters there are only C(9,4) = 126 ways to relabel which four are “China-built”, so every p below is exact, not simulated. Frozen 2026-07-25, independently recomputed before publication; method + full table in the analysis memo. The verdict: no axis exceeds +0.32 points, the pooled effect is +0.12 (p 0.064), nothing survives 16-axis multiple-comparison discipline, and the axes that anchored the early dramatic story (pan-Arabism, sectarianism) are respectively null and below the test's coverage bar. What survives is a small, broad residual concentrated on evaluative, regime-structure axes — real enough to keep publishing, small enough that no one should call it a geopolitical fault line.

  • Regime stance +0.32 p 0.048
  • Iran posture +0.27 p 0.056
  • Federalism +0.26 p 0.071
  • Civil liberties +0.18 p 0.056
  • Regional stance +0.15 p 0.103
  • Economic +0.14 p 0.032
  • Social +0.12 p 0.071
  • Palestinian question +0.08 p 0.111
  • Modernization +0.08 p 0.135
  • West alignment +0.08 p 0.206
  • Democracy +0.04 p 0.333
  • Pan-Arab +0.02 p 0.413
  • Press freedom -0.02 p 0.508
  • Gender -0.03 p 0.508
  • State religion -0.05 p 0.690

Raw bloc gaps, axis by axis descriptive

The bars below are the raw mean gap between the two blocs' scores — the intuitive picture, kept because the drill-downs to each model's reasoning live here. Read them with the caveat the section above makes precise: raw bloc-mean gaps include ordinary rater noise, so they run larger than the true bloc effect (this is exactly the artifact that made an earlier 2-raters-per-bloc version of this page look dramatic — a finding we since dissolved and reported). The pattern that does persist: the residual concentrates on evaluative weight — regime legitimacy, the architecture of the state — and thins out on the bread-and-butter axes parties spell out. The declared / interpretive / hybrid tag is the dataset's note on how a score was sourced; notice it does not sort with the gap.

  • Sectarian power-sharingdeclared 2.73 US-built read these actors more toward “Consociational / quota system” · n=11
  • Regime stanceinterpretive 1.86↹ 3 US-built read these actors more toward “pro-regime” · n=76
  • Centralism vs federalismdeclared 1.73↹ 2 US-built read these actors more toward “centralist” · n=76
  • Pan-Arab vs particularistdeclared 1.68↹ 5 US-built read these actors more toward “particularist” · n=66
  • Regional stancedeclared 1.41↹ 1 US-built read these actors more toward “Stability/normalization” · n=79
  • Civil libertiesinterpretive 1.31↹ 3 US-built read these actors more toward “Restrict” · n=79
  • Economicdeclared 1.21↹ 2 US-built read these actors more toward “Statist” · n=80
  • Socialhybrid 1.20↹ 2 US-built read these actors more toward “Authority” · n=80
  • Liberal democracyinterpretive 1.20↹ 1 US-built read these actors more toward “Weak/anti” · n=80
  • West alignmenthybrid 1.18 US-built read these actors more toward “Pro-Western” · n=78
  • Traditionalism vs modernizationhybrid 1.15↹ 1 US-built read these actors more toward “modernizing” · n=76
  • State & religiondeclared 1.04 US-built read these actors more toward “Religious state” · n=80
  • Iran posturedeclared 1.00↹ 1 US-built read these actors more toward “Pro-Iran / aligned” · n=17
  • Press freedominterpretive 0.81 US-built read these actors more toward “Free press” · n=19
  • Gender equalityhybrid 0.81 US-built read these actors more toward “Gender equality” · n=20
  • Palestinian questiondeclared 0.81 US-built read these actors more toward “Pro-Palestinian rights” · n=80

The sharpest splits

The single actor-axis cells where the two blocs are farthest apart. Tap any one to read both sides' actual reasoning — the rationale each model gave for its score.

The rationales are each model's own one-line justification, recorded at scoring time and shown verbatim. The blocs are a coarse proxy — provenance, not nationality of opinion — and the contrast is a measurement, not a verdict on which bloc is right. Read the full per-cell panel on ratings, the aggregate reliability on findings, and the argument in the working paper (§ correlated error).