How Smart is ChatGPT?
gpt-4chatgptbenchmarksexam-scoresopenaidata-visualization
Abstraction: GPT-4 vs GPT-3.5 exam-percentile benchmark comparison
Key points:
- Visual Capitalist (Marcus Lu, 2023-04-26) visualizes exam results from OpenAI's GPT-4 technical report (released March 27, 2023), scored as percentiles.
- GPT-4 dramatically beats GPT-3.5: Uniform Bar Exam 90th vs 10th percentile; GRE Verbal 99th vs 63rd; GRE Quant 80th vs 25th; SAT Math 89th vs 70th.
- No improvement on AP English Language (14th), AP English Literature (8th), or competitive programming (Codeforces <5th percentile both models; GPT-4 avg rating 392).
- GPT-4 improvements over GPT-3.5: internet access via plugins, visual/image inputs (multimodal), and larger context—8,192 tokens (~6,000 words) or 32,768 tokens (~24,000 words) vs 4,096.
- GPT-3.5 was trained on data only up to June 2021 with no internet access.
Connections: Openai · GPT-4 · Chatgpt · Large Language Models · Model Evaluation · Multimodality
Source: https://www.visualcapitalist.com/how-smart-is-chatgpt/