
An exam hall. Image: Homayoon soleimani / Wikimedia Commons, CC0, cropped
Humanity’s Last Exam, one of the hardest tests used to measure how smart AI models really are, has a new, stricter edition. The Center for AI Safety and Scale AI have released HLE-Diamond, a 1,000-question version cleaned up over the past year, and the first results put OpenAI’s GPT-6 Astra in front, Anthropic close behind, and xAI’s Grok 4.7 at the bottom of the table.
What is HLE-Diamond?
The original Humanity’s Last Exam launched in January 2025 with 2,500 questions written by nearly 1,000 experts across more than 100 subjects, from advanced math and physics to classics and medicine. It was designed to be the test AI couldn’t easily beat, after models had started acing older benchmarks.
It also drew criticism: researchers argued that some of its answers were wrong or disputed, which makes scores harder to trust. HLE-Diamond is the answer to that. It’s a refined subset built after “a year-long process of cleaning and refinement with input from research communities,” according to the organizers, and it’s split evenly into two halves:
- 500 reasoning questions, which test whether a model can work through hard problems
- 500 knowledge questions, which test expert-level knowledge
Models take it “closed book”: no web search and no code tools.
The scores
Here’s how the leading models did without tools:
- GPT-6 Astra (OpenAI): 60.6%
- Claude Opus 5.5 (Anthropic): 55.0%
- Claude Fable 5.1 (Anthropic): 51.3%
- Claude Opus 5 (Anthropic): 38.6%
- Gemini 3.8 Flash (Google): 34.3%
- GPT-6 Sol (OpenAI): 33.8%
- GPT-5.6 Sol (OpenAI): 31.2%
- Muse Spark 1.3 (Meta): 25.4%
- Grok 4.7 (xAI): 23.4%
Astra’s lead comes almost entirely from reasoning, where it scored 75.6% against 63.2% for Opus 5.5. On the knowledge questions, the order flips: Opus 5.5 edges ahead with 46.8% to Astra’s 45.6%. In other words, today’s best models are far better at thinking through a hard problem than at simply knowing obscure expert facts.
Give them tools and the gap shrinks
The organizers also tested the top models with web search and code tools switched on, and scores jumped:
- GPT-6 Astra: 82.9%
- Claude Opus 5.5: 73.9%
- Claude Fable 5.1: 72.4%
That’s a big leap for a test named to sound unbeatable, and a sign it may not stay hard for long once models can look things up.
Bad news for Grok
The results are awkward for Elon Musk’s xAI. Grok 4.7, which launched this week with 2 trillion parameters, finished last, behind Meta’s Muse Spark and well behind even OpenAI’s cheaper GPT-6 Sol. Musk has said Grok 4.9 is meant to match “Astra/Fable class” models. On this test, it has a long way to go.
For Anthropic, it’s a solid showing: Opus 5.5 is second overall and top for expert knowledge, and Anthropic’s models take second, third and fourth place.
Why it matters
Benchmarks shape which AI models companies and developers choose, and a test is only useful if its answers can be trusted. By cleaning up its questions, HLE-Diamond gives a clearer picture of where each model actually stands. The dataset is publicly available on Hugging Face, so anyone can check the results for themselves.


