AI Benchmarking
concepts · 2 notes linked
Related: Large Language Models · Metr · Task Completion Horizon · Gemini · Chatgpt · Grok · Deepseek · Anthropic
Notes
- 27 AI models were ranked by the public and ChatGPT came 8th — these are the models that beat it — Prolific's Humaine leaderboard ranks Gemini first in public user experience study
- Large Language Model Performance Doubles Every 7 Months — LLM task-completion capability doubles every seven months exponentially