Nick Haber
- Fantastic Bugs and Where to Find Them in AI Benchmarks
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
- ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
- When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration