AI Evaluation
concepts · 1 notes linked
Related: Openai · GPT-4 · Microsoft · Large Language Models · Artificial General Intelligence
Notes
- AI hype is built on high test scores. Those tests are flawed. — LLM benchmark scores are brittle, anthropomorphized, and often measure memorization not capability