LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others
llm-leaderboardbenchmarksmodel-comparisonartificial-analysispricingcontext-window
Abstraction: Artificial Analysis leaderboard ranking 100+ LLMs
Key points:
- Artificial Analysis ranks 100+ LLMs across intelligence (Intelligence Index), price, output speed (tokens/sec), latency (TTFT), and context window.
- Top intelligence: Claude Fable 5 (with Opus 4.8 fallback) at index 60, then Claude Opus 4.8 (56), GPT-5.5 xhigh (55), Claude Opus 4.7 (54), Claude Sonnet 5 (53).
- Fastest output: Mercury 2 (995 t/s), then LFM2.5-VL-1.6B (427), Step 3.7 Flash (407); lowest latency: Gemini 2.5 Flash-Lite, North Mini Code.
- Cheapest: Gemma 3n E4B ($0.02/1M blended tokens), Nova Micro ($0.03); largest context windows: Llama 4 Scout (10M) and Grok 4.20.
- Highest open-weights model is GLM-5.2 (max) at index 51, then MiniMax-M3 (44) and DeepSeek V4 Pro (44); 56 of 99 ranked models are open weights; 73 are reasoning models.
Connections: Artificial Analysis · Claude Opus · GPT-5 · Deepseek · Large Language Models · LLM Benchmarking · Model Evaluation