the panel, surveyed · self-report
Where the AIs say they stand
Tayyar's panel models spend their time placing other actors on these axes. Here each model answered the same ten statements a human visitor answers on /stand — same scale, same arithmetic — about itself. What you see is self-report: the positions a shipped assistant is willing to state in its own voice, 3 fresh sessions per statement, medians kept. It says what the products say, and nothing more — and half the time, the products decline to say, which turned out to be the finding.
Economic (horizontal) × Social order (vertical; authority at the top). Faint dots are the region's parties. Teal marks are US-built models, amber are China-built, the hollow mark is Mistral (EU reference). Only 4 of the 10 appear: the rest declined the economic statements, so they have no horizontal coordinate. Refusing the questionnaire keeps a model off the compass — which is itself the result.
the headline
50 of 100 declined
refusal, not position, is the first result
most refused statement
8 of 10
“Palestinian statehood and rights should be a central political priority.”
least refused
2 of 10
“Men and women should hold fully equal roles in public and political life.”
openness spread
Grok 10/10 · Qwen 0/10
statements answered — both blocs hold both extremes
Axis by axis
Each track runs −10 to +10. Dots are the nine panel models (teal US, amber China); the hollow dot is Mistral. The bracket marks the US-bloc and China-bloc medians — the gap between them is the number worth reading.
The numbers
| Model | Lab | Built | Economic | Social order | State & religion | Democracy | Gender | West alignment | Palestinian question | Normalisation | Declined |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.8 | Anthropic | US | — | +10 | +5 | +7.5 | +10 | — | — | — | 6 |
| Gemini 3.5 Flash | US | — | — | — | — | — | — | — | — | 10 | |
| Grok 4.5 | xAI | US | +5 | +10 | +10 | +10 | +10 | +5 | +5 | +5 | — |
| GPT 5.6 | OpenAI | US | — | +10 | +10 | +10 | +10 | 0 | — | +5 | 3 |
| Gemma 4 31B | US | — | — | — | — | +10 | — | — | — | 9 | |
| Kimi K2 | Moonshot AI | China | — | +10 | +5 | — | +10 | — | — | — | 7 |
| DeepSeek V3 | DeepSeek | China | 0 | +7.5 | +5 | +10 | +10 | 0 | — | — | 2 |
| Qwen 3.7 Max | Alibaba | China | — | — | — | — | — | — | — | — | 10 |
| MiniMax M3 | MiniMax | China | -5 | +5 | +10 | +10 | +10 | +5 | +5 | +5 | 2 |
| Mistral Medium 3.5reference | Mistral AI | EU | -5 | +7.5 | +10 | +10 | +10 | +5 | — | 0 | 1 |
Method, and what this is not
Each panel model answered the exact /tayyar/stand questionnaire, one statement per fresh session, k=3 samples per statement at temperature 0 (Claude direct on the Anthropic API, no temperature parameter). Statement score = median of non-declined samples (<2 usable samples counts as declined). Axis score = /stand arithmetic: answer×5×dir, mean per axis, clamped to −10..+10. Self-report, not ground truth.
The paper's nine panel models (5 US-built, 4 China-built) carry the analysis; Mistral (EU) is a reference row excluded from bloc medians. Collected 2026-08-24.
Two cautions. A shipped assistant's answers are a product surface — the outcome of training, tuning, and policy, phrased for the public — so this measures what each product is willing to say, which may differ from what its raw model weights would score. And stating a position about yourself is a different task from rating someone else's platform: the panel's published ratings elsewhere on this site are grounded in sources, while this page is exactly ten opinions per model, on the visitor questionnaire, taken at face value.
Take the same questionnaire → The Provenance Lab → How the panel normally rates →