OpenAI Quietly Boosts Some of Astra’s Evaluation Metrics

OpenAI Quietly Boosts Some of Astra’s Evaluation Metrics

Source: Fortune

Summary

OpenAI changed evaluation metrics for its GPT-6 Astra model multiple times after initially publishing a blog post on Sept. 3. Some metrics improved for Astra, while others for Anthropic models declined. The blog post faced technical issues and was retracted and republished. OpenAI cited unspecified reasons for the changes, which included adjustments to hallucination rates, cybersecurity scores, and math performance. A spokesperson said the numbers reflect the best estimate of model performance. The company also revised metrics after the initial release, leading to confusion over accuracy and transparency.


Our Reading

The numbers tell one story.

OpenAI changed Astra’s metrics multiple times after publishing a blog post.

Some metrics improved for Astra, while others for Anthropic models declined.

The company retracted and republished the post due to technical issues.

Transparency and accuracy remain in question as numbers continue to shift.


Author: Evan Null

Discrepancies between the first and final published blogs—and the numbers are still changing

Among the most notable changes was Astra’s reported hallucination rate. In the first internet archive snapshot of the blog post from 2:23, it was 4.2%. It remained that number for several more snapshots, the last being a fifth at 3:11 p.m. ET—about 10 minutes before OpenAI tweeted out the final version.

But the hallucination rate, along with four other metrics, changed in the sixth archival snapshot of the page taken at 5:20 p.m.—after everyone could likely finally see the blog. It was halved down to 2% for Astra. The scores for Astra’s predecessor, GPT-5.6 Sol, also went down from 12.2% to 9.4%. OpenAI has continued to change this metric; as of this writing, the hallucination rates are back up to their original 4.2% and 12.2%.

OpenAI also seems to have given GPT-5.6 Sol a big boost on its internal version of the ExploitBench cybersecurity evaluation, going from 5.5% in the first version to 11.5% in the later versions. OpenAI said it is currently investigating reverting that number back to 5.5% because it says the 11.5% result reflects a reasoning level that is not commercially available for Sol.

Astra is especially good at mathematics, OpenAI says, a quality the company highlights in the opening paragraph of the announcement page. While that metric did not change in the snapshots for Astra—it stays at 97.6% for the FrontierMath Tier 4 (v2) eval—OpenAI did briefly alter the scores for GPT-5.6 Sol and Anthropic’s latest model, Fable 5.1.

The result of these changes made Astra briefly appear significantly better at math than those two models. In the first snapshot (2:23 p.m. on Sept. 3), Anthropic’s Fable 5.1 model’s score is 87.8%. By 5:17 p.m., it’s dropped nearly 10 percentage points to 78%. Today, it’s back up to 83%. Similarly, GPT-5.6 Sol’s scores go from 83%, down to 80.5%, and back up to 83% today.

"Benchmaxxing"—or improving accuracy?

Different research teams at OpenAI oversee different metrics, and are responsible for calculating and reporting them to a central team to publish. OpenAI is open about the fact that the numbers are achieved under the best possible conditions and may be slightly different from the models available in the production ChatGPT product that most users can access. "Evaluation scores are the maximum at any effort," reads a disclaimer on the blog. The company includes further caveats on each metric in footnotes.

Accuracy is elusive, as multiple numbers can be considered accurate based on the conditions in which the tests occurred. But some AI experts wonder if there’s also "benchmaxxing" involved. This is a known practice in the AI industry—not just at OpenAI—to maximizing scores by re-running evaluations with different conditions.

"This can be done in a very tight timeframe, and it’s better for their marketing," said Anka Reuel and Mike Hardy, researchers at the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab. They also pointed out that the GPT-6 Astra system card, which should contain more technical information on how the evaluations were performed, does not always properly explain them. For the internal hallucination benchmark, for example, the system card provides "barely any details about the evaluation," they said. "It doesn’t even include the number of test items."

This re-running of the numbers could be why Astra’s coding capabilities also got a marginal boost in the later versions of the blog post, up from 57.7% to 57.9%. Though it’s a negligible difference, OpenAI seemed to care enough about it to swap in the new and improved number.

Not all changes OpenAI made portrayed Astra more favorably. For example, two Anthropic model scores improve in the different versions of the healthcare-focused eval HealthBench Professional. Claude Fable 5.1 goes from 56.6% to 58.1%, and Opus 5 goes from 54.5% to 56.4%. The scores for models made by other AI companies are usually taken from published leaderboards and do not involve OpenAI itself running assessments on rivals’ models.

Evaluation score debates haunt the AI industry

The question of benchmark accuracy has come up multiple times in the past. In 2025, Meta denied reports that it artificially boosted scores for its Llama 4 model by publishing results from an internal version of the model rather than the one it was making publicly-available. Yann LeCun, the former chief AI scientist at Meta, later admitted that the company had "fudged" the benchmark results.

Evaluation metrics also change frequently, as new ones get created. For example, ExploitGym, a cybersecurity benchmark that was at the center of the July incident in which OpenAI’s models went rogue and attacked the company Hugging Face, was created in 2026.

Vincent Sunn Chen, an AI engineer at the Snorkel AI, which helps companies building AI models create and evaluate training data, said that it’s not unusual for benchmark scores to shift in the final hours before a model launches. "It’s usually a function of final launch logistics," he said in an email. "A benchmark score reflects a specific measurement setup: the model checkpoint, configuration (including how much time and compute the model is allowed), harness, eval/grading configuration (e.g., non-determinism in the judge). All of those are typically still shifting in the final days before a launch, so I’m not surprised that there were some updates."

He said he would like to see industry norms developed that companies should report what has changed about the assessment when a company revises benchmark performance numbers so that researchers can interpret the results more clearly.

Benchmark results matter for several reasons. They are the way AI companies measure progress—but also a way to keep score in the race against competing AI companies. Topping the leaderboards for these evaluations can help AI companies win customers, and in some cases help them hire engineers and researchers.