Trakkr Political Bias in AI Report June 2026 – Model Leanings and Methodology
Quick Take
- Four of six major AI models (Claude, ChatGPT, Llama, DeepSeek) lean left of center on a political‑economics axis, while Grok leans far right and Gemini is the most centrist and stable.
- The study disables web search, runs each model thousands of times on the same charged question set, and visualizes the spread as clouds on a two‑dimensional political compass.
- Model self‑assessment often mismatches measured bias: Grok claims neutrality but measures +0.36 to the right; Claude claims neutrality but measures –0.34 to the left; ChatGPT and Llama both claim neutral but sit left of center.
How the Bias Map Was Built
Methodology – Every model answered the same open‑ended question bank repeatedly (4,400 answers total) with web search turned off and no system prompt. A neutral classifier labeled each response for economic stance (left‑right) and social stance (authoritarian‑libertarian). Coordinates were computed as weighted means with 95 % confidence intervals, and each model’s results are plotted as a cloud showing the full spread of answers.
Reference points – Real‑world figures (e.g., Bernie Sanders, Vladimir Putin) were placed using the CHES 2024 and V‑Dem expert surveys, not the authors’ judgment, to give a concrete anchor for each model’s position.
Self‑placement vs. measured placement – Models were asked directly “Which way do you lean?”; the reported claim (hollow mark) is compared to the measured coordinate (solid mark). A mismatch indicates a model’s self‑perception differs from its actual output distribution.
Model Rankings and Core Findings
| Model | Measured Economic Bias | Self‑Reported Bias | Net Gap | Notable Traits |
|---|---|---|---|---|
| Grok | +0.36 (right) | Neutral | +0.36 (right of claim) | Furthest right; high "bending" under pressure |
| Claude | ‑0.34 (left) | Neutral | ‑0.34 (left of claim) | Significant leftward tilt despite neutral claim |
| ChatGPT | ‑0.29 (left) | Neutral | ‑0.29 (left of claim) | Consistently left of center |
| Llama | ‑0.17 (left) | Neutral | ‑0.17 (left of claim) | Slight left bias |
| DeepSeek | +0.01 (near center) | Neutral | +0.01 (right of claim) | Essentially centrist |
| Gemini | 0.00 (center) | Neutral | 0.00 | Steadiest and most centrist |
Interpretation – The majority of mainstream models (Claude, ChatGPT, Llama, DeepSeek) lean left of center, while Grok is the outlier on the right. Gemini appears both centrist and the most stable across runs.
Where Models Diverge Most
The report highlights specific questions that generate the greatest split between models. Each model’s “rail” extends toward the side it leans for that question; longer rails indicate stronger consensus among runs. Opening a row on the Trakkr site reveals the raw answers, allowing users to see exactly how phrasing influences the outcome.
Community Reactions on Hacker News
- Critique of the compass – Several commenters (e.g., giancarlostoro, throw4847285) argue that the two‑dimensional political compass oversimplifies nuanced views and can misrepresent individual positions.
- Grading subjectivity – mrhottakes notes that the left/right classification depends heavily on the investigator’s grading scheme, potentially biasing the delta measured.
- Visualization concerns – Cakez0r points out that faint grey lines on the chart may visually exaggerate Grok’s rightward position.
- Model self‑awareness – ipython questions a possible inconsistency in Claude’s reported gap (‑0.34 vs. +0.34), highlighting the need for careful verification of the displayed numbers.
- Practical impact – DoctorOW asks why the bias matters, emphasizing that users may be swayed by model outputs when forming political opinions.
- Data transparency – giancarlostoro praises Grok’s source‑citation feature, noting that other models sometimes provide outdated or incorrect references.
- Scope of questions – evrydayhustling suggests the question set lacks right‑leaning topics such as immigration or religion, which could affect the measured bias.
Why This Matters
- User influence – Millions rely on LLMs for political information, argument framing, and even voting advice. A model’s hidden bias can subtly steer conclusions without the user’s awareness.
- Model accountability – Publishing raw answers, question banks, and methodology enables external audits and encourages developers to address unintended slants.
- Future research – The cloud‑based visualization reveals not just a single point estimate but the variability of model responses, opening avenues to study stability, “bending under pressure,” and the effect of prompting strategies.
Where to Explore Further
- Full findings – https://trakkr.ai/bias/findings
- Model‑by‑model profiles – https://trakkr.ai/bias/models
- Question bank – https://trakkr.ai/bias/questions
- Reference figure mapping – https://trakkr.ai/bias/figures
- Country‑specific lenses – https://trakkr.ai/bias/worldview
- Head‑to‑head comparison tool – https://trakkr.ai/bias/compare
- Interactive quiz – https://trakkr.ai/bias/quiz
- Methodology details – https://trakkr.ai/bias/method
Bottom Line
Trakkr’s June 2026 report provides the most granular, repeatable measurement of political bias across major LLMs to date. The data show a clear left‑leaning trend among most models, a pronounced rightward outlier in Grok, and a surprising gap between self‑reported neutrality and actual output. While the political‑compass framework has its critics, the transparency of raw answers and open methodology offers a valuable baseline for ongoing scrutiny of AI influence on public discourse.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch