AI hype is built on high test scores. Those tests are flawed.

llm-evaluationbenchmarksai-hypegpt-4agi

Abstraction: LLM benchmark scores are brittle, anthropomorphized, and often measure memorization not capability

Key points:

Connections: Openai · GPT-4 · Microsoft · Large Language Models · AI Evaluation · Artificial General Intelligence

Source: https://www.technologyreview.com/2023/08/30/1078670/large-language-models-arent-people-lets-stop-testing-them-like-they-were/