27 AI models were ranked by the public and ChatGPT came 8th — these are the models that beat it
ai-benchmarkingllm-comparisonuser-experiencechatbots
Abstraction: Prolific's Humaine leaderboard ranks Gemini first in public user experience study
Key points:
- Prolific's Humaine leaderboard evaluated 27 AI models via 21,352 pairwise user experience comparisons (not task-completion benchmarks); methodology based on users choosing which model performed better in side-by-side interactions
- Gemini 2.5 Pro ranked first overall across nearly all demographic filters: age groups in UK and US, political affiliation, ethnicity
- Rankings 2-4: Deepseek, Mistral Le Chat, Grok; Le Chat is noted for loyal fanbase despite lower benchmark scores
- ChatGPT (GPT-4.1) placed 8th; Claude's two v4 models placed 11th and 12th
- Grok-3 ranked highest for trust, ethics, and safety — described as ironic given Grok's controversies
- User experience rankings diverge significantly from standard benchmark results; highlights the gap between capability benchmarks and real-world perceived quality
Connections: Gemini · Chatgpt · Grok · Deepseek · Anthropic · Large Language Models · AI Benchmarking