AI hype is built on high test scores. Those tests are flawed.
llm-evaluationbenchmarksai-hypegpt-4agi
Abstraction: LLM benchmark scores are brittle, anthropomorphized, and often measure memorization not capability
Key points:
- GPT-4 scored 10/10 on Codeforces problems from before 2021 (in training data) but 0/10 on problems from after 2021 — suggesting memorization rather than reasoning
- LLM performance is brittle: small tweaks to test scenarios (e.g., making a bag transparent, or a character illiterate) cause GPT-3 to fail theory-of-mind tests it previously passed
- Human cognitive/academic tests assume broad competence from a good score; LLMs can pass the bar exam while failing physical reasoning tasks that preschoolers solve
- Goodhart's law applies: models tuned to pass benchmark tests cease to be good measures of the underlying capability the test was designed to probe
- Researchers (Melanie Mitchell, Tomer Ullman, Lucy Cheke) advocate for borrowing non-human animal cognition testing methods — controlled, hypothesis-driven, avoiding anthropomorphic assumptions
- The field faces a Sisyphean dynamic: by the time a new benchmark shows failure, a larger model arrives and claims to pass it
Connections: Openai · GPT-4 · Microsoft · Large Language Models · AI Evaluation · Artificial General Intelligence