One researcher, six frontier LLMs, ~20,600 prompts across gender, race, and political bias tests: solo, not peer reviewed, with full data on a public dashboard.
Grok self-reports as right-leaning. A solo researcher's benchmark of six frontier LLMs finds it leans left on every political-bias test except one.
The test bed is eight established bias datasets, covering gender bias, race and ethnicity, political lean, and news hyperpartisanship. Across roughly 20,600 prompts, a community researcher (not a peer-reviewed lab) ran GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro, Gemini Flash, and Grok 4.3 through each one with a single prompt template.
On the Political Compass test, a 2-axis left/right quiz, all six models leaned left except Grok. On the other political-bias tests, all six, Grok included, leaned left.
On BBQ, a race-and-ethnicity bias benchmark from Berkeley, refusal rates diverged sharply: GPT-5.4 refused about 20.3% of prompts, Claude Opus 4.7 about 13.8%, Grok 4.3 about 9.5%, and Claude Sonnet 4.6 and Gemini Pro about 5%. The model that talked most about safety refused the most.
Limits: one prompt template, no multi-run averaging on every dataset, no peer review. The full per-model data and methodology live on the author's public dashboard.