academic · dataset overview

Tayyar dataset

A position dataset for MENA political actors. 102 parties and 274 politicians across 24 countries, scored on 16 axes. Every fact carries an external source citation. The dataset is released as a frozen, versioned snapshot — v1.0, 17 tables, 12,867 panel ratings, with a per-file checksum manifest — while the live tools keep evolving on top of it. This page is the canonical citable surface — the cards below summarize the live dataset, the methodology buttons jump to the deep dives.

Release v1.0 · 2026-07-25 · CC BY-NC-SA 4.0 · DOI via Zenodo pending

Total entities 376 102 parties · 274 politicians
Verified 51% 190 / 376 with source citations
Countries 24 MENA states with seeded party and politician data
Axes 16 2 compass + 14 issue, each with a formal rubric
Position scores 1195 Including composite + lens-divergent (declared / behavioral)
Events tracked 328 /pulse — semi-live with confidence tier
Source documents 426 /documents — verbatim primary-source corpus
Verified quotes 101 /quotes — every row with at least one URL citation

What's included

  • Field-level source citations. Each fact about each entity links back to where it came from — founding year cites a different source than current leader, which cites a different source than legal status. The coverage page tracks the rollup.
  • 16 calibrated axes. Economic, social, state-religion, democracy, west-alignment, regional-stance, Palestinian question, civil liberties, regime stance, pan-Arab, federalism, modernization, gender, iran-posture, press-freedom, sectarianism. Each one comes with a scoring rubric and concrete MENA examples anchored along the scale.
  • Richer status than yes/no. Parties carry a government role (lead, coalition major / minor, confidence-and-supply, opposition major / minor, extra-parliamentary, banned) and a legal status (legal, restricted, outlawed, dissolved, merged away). Opposition and independent flags sit alongside.
  • Declared vs. behavioral on key cases. Parties like Hezbollah and Hamas read as more committed to democracy in what they declare than in what they do; 20 parties carry a rhetoric-vs-record gap that is itself the finding. The home compass has a lens toggle to switch between views.
  • 426 primary-source documents and 101 verified quotes. The document corpus carries verbatim manifestos, charters, parliamentary speeches, and UN addresses with country / party / politician attribution. The quote corpus drives Who-said-it and the "On the record" sections on every party / politician page.
  • A semi-live event feed. Pulse tracks recent political shifts with a confidence rating — confirmed, reported, rumored, speculative — so in-flux developments (party formations, merger talks) can sit alongside confirmed events without being conflated. Subscribable as RSS.
  • Frozen, versioned release. The full dataset is published as v1.0 — every table as CSV and JSON with a codebook README, CITATION.cff, and SHA-256 checksums. One deliberate exclusion: the r7-dvb study round ships with its own pre-registered paper, not here. A Zenodo DOI is in registration.

Cite as

Cite the frozen release for the data and the paper for the instrument. The release tag pins a citation that cannot shift as the live dataset evolves; the Zenodo DOI will be added here once registered.

@misc{gara_tayyar_dataset_v1_0,
  author  = {Gara, Tarek},
  title   = {Tayyar: a MENA political-position dataset scored by a language-model panel},
  year    = {2026},
  version = {v1.0},
  url     = {https://github.com/tarekgara/tarekgara.com/releases/tag/tayyar-dataset-v1.0},
  note    = {Frozen release v1.0 (2026-07-25); CC BY-NC-SA 4.0; DOI via Zenodo pending}
}
How to cite 1 reference

Paper 1 Gara, T. (2026). The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel. Preprint, arXiv:2606.23042.

Open ↗
Show BibTeX
@misc{gara_tayyar_2026,
  author        = {Gara, Tarek},
  title         = {The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel},
  year          = {2026},
  eprint        = {2606.23042},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CY},
  doi           = {10.48550/arXiv.2606.23042},
  url           = {https://arxiv.org/abs/2606.23042}
}

Where to read more

  • Methodology — how the dataset got built and where it falls short
  • Findings — the structural patterns the data shows
  • Coverage — verification status, country breakdowns, special-status leaderboards
  • Axes catalog — all 16 axes with correlations and per-axis stats

Data access

The full dataset is public: release v1.0 carries all 17 tables as CSV and JSON — including the panel ratings (12,867 rows across rounds, with round semantics in the bundled codebook), the reliability-program tables, and per-file SHA-256 checksums — under CC BY-NC-SA 4.0. One round is deliberately absent: r7-dvb belongs to a pre-registered study and is released with that paper. The live tools keep evolving past the freeze; cite the release, browse the compass. Earlier snapshot v0.2 remains in the repository history.

What's coming next

Changes are tracked as versioned snapshots; the repository will be open-sourced under MIT on publication. The roadmap, in the order it'll land:

  • Document-grounded scoring. Positions derived from reading party platforms, speeches, and voting records — with the specific passages cited. The hand-coded scores stay as the baseline; the document-grounded ones replace them as they're produced.
  • Inter-rater agreement. Cohen's κ between hand-coded and document-grounded scores reported on methodology. Where they agree, the rubric's doing its job; where they don't, that's a finding worth writing up.
  • Lens system at scale. Declared / behavioral / perceived rows generated for every party, not just the hand-coded marquee cases. The compass lens toggle then carries information across the whole dataset.
  • Second-pass verification. Each fact and each position score reviewed against primary sources by someone other than the author.