OpenAI Revises Astra's Benchmarks After Blog Post Errors
Source: Fortune. Casualplayhub News adds summary, context, and editorial framing while linking back to the original report.
OpenAI's launch of its GPT-6 Astra model on September 3 turned into a messy affair, with the company quietly revising evaluation benchmarks multiple times after a botched blog post rollout. The post, originally scheduled for 2 p.m. ET, was delayed by nearly two hours due to what the company described as a content management system bug and an internet outage. When the link finally went live on OpenAI's X account at 3:32 p.m., many users, including Fortune, encountered error messages. CEO Sam Altman acknowledged the hiccup, tweeting, "We hit a little snag getting the blog post deployed, but it is really great." The post eventually became accessible about an hour later.
But the real story unfolded behind the scenes. Internet archive snapshots revealed that OpenAI had actually published a version of the blog shortly after 2 p.m. and then retracted it. The company declined to disclose the reason for the retraction, insisting it was unrelated to the benchmark figures. When the post reappeared, the evaluation metrics had changed. And they kept changing.
Among the most notable shifts was Astra's hallucination rate. The first snapshot at 2:23 p.m. showed 4.2%, which held steady for several hours. By 5:20 p.m., after the post was widely visible, that number had dropped to 2%. The hallucination rate for Astra's predecessor, GPT-5.6 Sol, also fell from 12.2% to 9.4%. Then, as of this writing, both rates had climbed back to their original levels. OpenAI also boosted Sol's score on its internal ExploitBench cybersecurity evaluation from 5.5% to 11.5%, though the company said it was investigating reverting that change because the higher score reflected a reasoning level not commercially available.
Mathematics benchmarks saw similar turbulence. Astra's score on FrontierMath Tier 4 (v2) remained at 97.6%, but OpenAI briefly altered the scores for GPT-5.6 Sol and Anthropic's Fable 5.1. In the first snapshot, Anthropic's score was 87.8%. By 5:17 p.m., it had dropped to 78%, then later rose to 83%. Sol's math score went from 83% to 80.5% and back to 83%. The changes were not all one-sided: two Anthropic model scores improved on the HealthBench Professional evaluation, with Fable 5.1 rising from 56.6% to 58.1% and Opus 5 from 54.5% to 56.4%.
OpenAI defended the revisions. "We care deeply about getting evaluations right," a spokesperson told Fortune. "Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons." The company also noted that the Arc Prize Foundation independently assessed Astra at 99.9% on the ARC-AGI-3 evaluation when given a powerful harness, compared to 63% with the standard setup.
However, the pattern of changes raised eyebrows among researchers. Anka Reuel and Mike Hardy of the Stanford Intelligent Systems Laboratory pointed to a practice known as "benchmaxxing"—re-running evaluations with different conditions to maximize scores. "This can be done in a very tight timeframe, and it's better for their marketing," they said. They also criticized the GPT-6 Astra system card for lacking transparency on how evaluations were performed, noting that the internal hallucination benchmark provided "barely any details about the evaluation" and not even the number of test items.
Vincent Sunn Chen, an AI engineer at Snorkel AI, said benchmark score shifts are common in the final days before a launch, but he called for industry norms requiring companies to report what changed. The confusion over changing metrics, combined with past incidents like Meta's admission of "fudging" Llama 4 benchmarks, could make it hard for customers and investors to determine which models are truly best. As OpenAI eyes a possible 2027 IPO, the narrative of having the best models may be muddied by these technical complexities.
Article commentary
The quiet revision of benchmark scores during OpenAI's GPT-6 Astra launch is a revealing episode in the high-stakes world of AI competition. On one hand, it highlights the inherent difficulty of standardizing evaluation metrics for large language models. Benchmarks are not static; they depend on countless variables—model checkpoints, scaffolding, harnesses, and even the exact moment of the run. As OpenAI's spokesperson noted, noise within a few percentage points is normal, and adjustments before publication are routine. From this perspective, the company's actions are defensible: they were simply trying to present the most accurate numbers possible under tight deadlines. Yet the sequence of events suggests something more troubling. The hallucination rate halving from 4.2% to 2% and then snapping back, the rival Anthropic's math score dropping nearly 10 points and then partially recovering, the Sol's cybersecurity score jumping from 5.5% to 11.5%—these are not minor tweaks. They are significant swings that happen to coincide with a narrative favoring Astra. The timing, immediately after a botched blog post that OpenAI itself caused, invites skepticism. The company's refusal to explain the retraction, citing a bug and an internet outage, feels like a convenient cover for a more deliberate process of optimizing numbers for public consumption. The practice of "benchmaxxing" is not unique to OpenAI. Meta's Llama 4 scandal earlier in 2025, where Yann LeCun admitted to "fudging" results, shows that the entire industry is susceptible to gaming benchmarks. The problem is structural: benchmarks are both a measure of progress and a marketing tool. Companies want to top leaderboards to attract customers, talent, and investors. This dual purpose creates a powerful incentive to squeeze every last decimal point out of evaluations, even if it means bending the rules of transparency. What makes this episode particularly concerning is the lack of standardization. As Vincent Sunn Chen pointed out, there is no industry norm for reporting changes in benchmark scores. The result is a fog of confusion. Customers and investors are left to decipher whether a model's performance is genuinely superior or merely a product of favorable testing conditions. For a company like OpenAI, which is reportedly preparing for a 2027 IPO, this ambiguity could backfire. Investors may demand clearer, more auditable metrics before committing capital. Moreover, the incident underscores the limits of quantitative evaluation in AI. Benchmarks are proxies for real-world performance, but they can be gamed, and they often fail to capture nuances like safety, reliability, and ethical alignment. The Stanford researchers' criticism of the system card's lack of detail is a reminder that even the most advanced models are black boxes. The public deserves more than just a number; it deserves context about how that number was obtained. In the end, OpenAI's handling of the Astra launch raises a fundamental question: Can the AI industry police itself on benchmark transparency? The answer, so far, is no. The lack of a universal standard and the competitive pressure to look good create a perfect storm for manipulation. Unless regulators or independent auditors step in, the narrative of progress in AI may remain as fluid as the numbers on a blog post.