Metr
entities · 3 notes linked
Related: Zvi Mowshowitz · Anthropic · Model Evaluation · AI Safety · Large Language Models · AI Benchmarking · Task Completion Horizon · Claude
Notes
- Claude Fable 5 and Mythos 5: The System Card — Reading the Fable/Mythos 319-page system card
- GPT-5.6: The System Card — Review of OpenAI GPT-5.6 Sol/Terra/Luna system card
- Large Language Model Performance Doubles Every 7 Months — LLM task-completion capability doubles every seven months exponentially