OpenAI's launch of its GPT-6 Astra model on September 3 turned into a messy affair, with the company quietly revising evaluation benchmarks multiple times after a botched blog post rollout. The post, originally scheduled for 2 p.m. ET, was delayed by nearly two hours due to what the company described as a content management system bug and an internet outage. When the link finally went live on OpenAI's X account at 3:32 p.m., many users, including Fortune, encountered error messages. CEO Sam Altman acknowledged the hiccup, tweeting, "We hit a little snag getting the blog post deployed, but it is really great." The post eventually became accessible about an hour later.

But the real story unfolded behind the scenes. Internet archive snapshots revealed that OpenAI had actually published a version of the blog shortly after 2 p.m. and then retracted it. The company declined to disclose the reason for the retraction, insisting it was unrelated to the benchmark figures. When the post reappeared, the evaluation metrics had changed. And they kept changing.

Among the most notable shifts was Astra's hallucination rate. The first snapshot at 2:23 p.m. showed 4.2%, which held steady for several hours. By 5:20 p.m., after the post was widely visible, that number had dropped to 2%. The hallucination rate for Astra's predecessor, GPT-5.6 Sol, also fell from 12.2% to 9.4%. Then, as of this writing, both rates had climbed back to their original levels. OpenAI also boosted Sol's score on its internal ExploitBench cybersecurity evaluation from 5.5% to 11.5%, though the company said it was investigating reverting that change because the higher score reflected a reasoning level not commercially available.

Mathematics benchmarks saw similar turbulence. Astra's score on FrontierMath Tier 4 (v2) remained at 97.6%, but OpenAI briefly altered the scores for GPT-5.6 Sol and Anthropic's Fable 5.1. In the first snapshot, Anthropic's score was 87.8%. By 5:17 p.m., it had dropped to 78%, then later rose to 83%. Sol's math score went from 83% to 80.5% and back to 83%. The changes were not all one-sided: two Anthropic model scores improved on the HealthBench Professional evaluation, with Fable 5.1 rising from 56.6% to 58.1% and Opus 5 from 54.5% to 56.4%.

OpenAI defended the revisions. "We care deeply about getting evaluations right," a spokesperson told Fortune. "Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons." The company also noted that the Arc Prize Foundation independently assessed Astra at 99.9% on the ARC-AGI-3 evaluation when given a powerful harness, compared to 63% with the standard setup.

However, the pattern of changes raised eyebrows among researchers. Anka Reuel and Mike Hardy of the Stanford Intelligent Systems Laboratory pointed to a practice known as "benchmaxxing"—re-running evaluations with different conditions to maximize scores. "This can be done in a very tight timeframe, and it's better for their marketing," they said. They also criticized the GPT-6 Astra system card for lacking transparency on how evaluations were performed, noting that the internal hallucination benchmark provided "barely any details about the evaluation" and not even the number of test items.

Vincent Sunn Chen, an AI engineer at Snorkel AI, said benchmark score shifts are common in the final days before a launch, but he called for industry norms requiring companies to report what changed. The confusion over changing metrics, combined with past incidents like Meta's admission of "fudging" Llama 4 benchmarks, could make it hard for customers and investors to determine which models are truly best. As OpenAI eyes a possible 2027 IPO, the narrative of having the best models may be muddied by these technical complexities.