Model Evaluation
concepts · 15 notes linked
Related: Zvi Mowshowitz · Openai · Anthropic · Claude · Large Language Models · AI Safety · AI Policy · AI Agents
Notes
- AI #172: The First Fable — Weekly roundup excluding the Fable model itself
- AI #174: You're It — Weekly AI roundup, Claude Tag, medical scanners, agent security
- AI #175: The Fable Continues — Weekly AI roundup covering Fable's return aftermath
- Claude Fable 5 and Mythos 5: Capabilities — Fable 5 capability review, benchmarks, classifiers, reception
- Claude Fable 5 and Mythos 5: The System Card — Reading the Fable/Mythos 319-page system card
- Claude Sonnet 5 Is Not Frontier But Has Its Uses — Sonnet 5 system card review, cheaper faster non-frontier model
- GLM-5.2 Is The New Best Open Model — GLM-5.2 open model capabilities review, benchmarks
- GPT-5.6: The System Card — Review of OpenAI GPT-5.6 Sol/Terra/Luna system card
- How Smart is ChatGPT? — GPT-4 vs GPT-3.5 exam-percentile benchmark comparison
- LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others — Artificial Analysis leaderboard ranking 100+ LLMs
- OpenAI Offers A New Policy Blueprint — OpenAI's federal frontier-AI safety framework blueprint
- The False Promise of Imitating Proprietary LLMs — Finetuning open models on ChatGPT outputs mimics style but not factuality or capability
- Trump Signs Executive Order For AI Testing Prior To Frontier Model Releases — Trump EO mandating pre-release frontier cyber testing
- WSJ Article Claiming China Has Matched Anthropic Is Obvious Nonsense — Debunking WSJ headline that China matched Mythos on cyber
- What you wanted to know about AUC — AUC-ROC metric explained as threshold-invariant ranking score