Benchmark Evaluation
concepts · 2 notes linked
Related: Kevin Musgrave · Metric Learning · Bayesian Optimization · GPT-4 · Mit · Hugging Face · Large Language Models · Prompt Engineering
Notes
- Paper page - Exploring the MIT Mathematics and EECS Curriculum Using Large Language Models — GPT-4 achieves near-perfect scores on MIT Math and EECS curriculum
- Updates to "A Metric Learning Reality Check — Updates to arXiv paper exposing unfair comparisons in metric learning benchmarks