Humanity’s Last Exam

Summary

This partial ingest is based on the HLE website clip and the citation embedded there. Humanity’s Last Exam is presented as a 2,500-question multimodal benchmark designed to remain difficult for frontier models after older academic benchmarks have saturated, with public questions, a private held-out set, and benchmark reporting that includes both accuracy and calibration error.

Key Claims

  • Popular academic benchmarks saturate too quickly to remain informative about frontier model capabilities.
  • A broad, expert-contributed benchmark can better track advanced closed-ended academic reasoning.
  • Reporting calibration error alongside accuracy is important because benchmark usefulness depends on more than answer correctness alone.
  • Strong performance on HLE would indicate high technical competence on closed-ended questions, but not by itself autonomous research ability or AGI.

Methods / Formalism

  • Assemble 2,500 difficult questions across more than a hundred subjects from a large contributor base.
  • Keep some test material private to reduce direct overfitting pressure from public release.
  • Evaluate models on both accuracy and a confidence-linked calibration error measure.

Evidence / Experiments

  • The clipped page reports low absolute accuracy for several frontier models relative to older benchmark saturation patterns.
  • It also highlights large calibration errors, suggesting frontier systems remain poorly matched between confidence and correctness on difficult questions.
  • Because this ingest comes from the website rather than the full paper, details about sampling, judging, and metric definitions should still be checked against the Nature article when needed.

Connections

Open Questions

  • How stable is HLE against prompt-overfitting and benchmark contamination over time?
  • Which calibration metric is used on the site, and how does it compare to proper scoring-rule based reporting?
  • What should replace HLE once closed-ended frontier exams also begin to saturate?

Citation

Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., and HLE Contributors Consortium. (2026). Humanity’s Last Exam. Nature / HLE website. The clip includes the Nature citation “A benchmark of expert-level academic questions to assess AI capabilities.”