Humanity’s Last Exam
Summary
This partial ingest is based on the HLE website clip and the citation embedded there. Humanity’s Last Exam is presented as a 2,500-question multimodal benchmark designed to remain difficult for frontier models after older academic benchmarks have saturated, with public questions, a private held-out set, and benchmark reporting that includes both accuracy and calibration error.
Key Claims
- Popular academic benchmarks saturate too quickly to remain informative about frontier model capabilities.
- A broad, expert-contributed benchmark can better track advanced closed-ended academic reasoning.
- Reporting calibration error alongside accuracy is important because benchmark usefulness depends on more than answer correctness alone.
- Strong performance on HLE would indicate high technical competence on closed-ended questions, but not by itself autonomous research ability or AGI.
Methods / Formalism
- Assemble 2,500 difficult questions across more than a hundred subjects from a large contributor base.
- Keep some test material private to reduce direct overfitting pressure from public release.
- Evaluate models on both accuracy and a confidence-linked calibration error measure.
Evidence / Experiments
- The clipped page reports low absolute accuracy for several frontier models relative to older benchmark saturation patterns.
- It also highlights large calibration errors, suggesting frontier systems remain poorly matched between confidence and correctness on difficult questions.
- Because this ingest comes from the website rather than the full paper, details about sampling, judging, and metric definitions should still be checked against the Nature article when needed.
Connections
- Anchors a new AI Evaluation and Benchmarking area in the wiki.
- Motivates the Calibration concept operationally by showing that benchmark dashboards already report confidence quality next to accuracy.
- Adjacent to Wikipedia2026 - Brier Score and PredictAddict2026 - Calibration vs Refinement Thread as part of a broader probabilistic-evaluation cluster.
Open Questions
- How stable is HLE against prompt-overfitting and benchmark contamination over time?
- Which calibration metric is used on the site, and how does it compare to proper scoring-rule based reporting?
- What should replace HLE once closed-ended frontier exams also begin to saturate?
Citation
Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., and HLE Contributors Consortium. (2026). Humanity’s Last Exam. Nature / HLE website. The clip includes the Nature citation “A benchmark of expert-level academic questions to assess AI capabilities.”