To evaluate AGI claims, ask for a definition, identify the exact system tested, inspect the evaluation conditions and separate demonstrated ability from the conclusion attached to it. Then ask what evidence would disprove the claim. A strong benchmark result can establish meaningful progress while leaving broader questions about general intelligence unanswered.
My September 4 LinkedIn post asked whether GPT-6 Astra counted as AGI or the strongest model of the moment. It reached 35,401 impressions, 276 reactions and 149 comments. A week later, the useful follow-up is to examine the evidence behind that question, including where the numbers in my original post need correction.
in
“Is this AGI, or just the best model of the month?”
How to Evaluate AGI Claims Without Moving the Goalposts
Start by writing down what AGI means in the claim you are assessing. Does it mean learning unfamiliar tasks efficiently, performing most economically valuable work, matching people across cognitive abilities or operating independently over long periods? These definitions overlap, but they require different evidence. Without a stated definition, a discussion can change its finish line every time a new result appears.
Make the definition testable. Specify the range of tasks, the comparison group, the resources allowed and the acceptable failure rate. A claim about human-level performance needs a description of the humans and their working conditions. A claim about autonomy needs a description of the assistance provided. Those details determine what the result can support.
Stanford HAI's framework for validating AI claims asks readers to connect the claim, the actual test and the evidence supporting the interpretation.
What the Astra Launch Actually Establishes
OpenAI's current launch page reports 99.9% on ARC-AGI-3, rather than the 98.6% cited in my post. It also reports strong results across mathematics, computer use and other professional tasks. These are specific capability claims from the developer. The page provides evidence of substantial progress; it does not independently settle every possible definition of AGI.
The current GPT-6 Astra announcement is the source for those reported results. Check its evaluation conditions before comparing the headline percentages.
I am leaving the original post's GPU count, exact browser-vulnerability count, price comparison and attributed executive statements out of this analysis because I have not independently verified all of them. A social post is the starting point for reporting. Repeating an unsupported detail in a longer article would only give it a second place to appear.
Name the System Behind the Score
A model, an agent and a deployed product are different evaluation objects. An agent may combine a model with tools, memory, retrieval, retries and a verifier. A product adds permissions, interface behavior and operational controls. If the measured system includes those components, credit the result to that configuration. Do not silently transfer its score to every use of the underlying model.
Ask for the model version, reasoning settings, tool access, attempt budget and selection procedure. If a system generates many answers and an external process chooses the winner, that is part of the method. The result can still be valuable, but the resources and supervision must remain visible when someone compares it with another system or a human baseline.
Read the Human Baseline Carefully
ARC Prize describes its ARC-AGI-3 human dataset as a controlled study of 458 participants. Its interactive environments examine how people learn and solve unfamiliar problems. That is useful context for interpreting a score: the evaluation has a defined population, task environment and measurement procedure, rather than an undefined comparison with all human intelligence.
Read ARC Prize's human performance study for the experimental setup behind that baseline.
A near-perfect result on a particular suite should prompt the next question: which important abilities remain outside it? Consider ambiguous goals, incomplete evidence, social context and learning from a changed environment. You do not need to dismiss the benchmark to ask that question. You need to keep its measurement boundary attached to its result.
Look for Breadth and Transfer
Google DeepMind proposes assessing AGI through a broader cognitive framework and a staged evaluation process using held-out tasks and human comparisons. Its proposal illustrates why a collection of complementary measurements can be more informative than one leaderboard position. A framework is still a proposal for measurement, rather than a certificate that any particular system has passed.
DeepMind's cognitive framework for measuring AGI describes that approach.
For a claim review, look for transfer to tasks that differ meaningfully from the examples used to develop the system. Ask whether changed instructions, unfamiliar interfaces or missing information alter the result. Repeated success on one narrow task family and dependable adaptation across different families are separate findings. Both deserve to be reported accurately.
Separate Capability from Permission
OpenAI's safety overview says Astra reaches its Critical cybersecurity capability threshold and describes additional deployment protections. It also reports challenges in monitoring written reasoning under adversarial conditions. These statements concern capability, safety and oversight. They should not be collapsed into either a claim that the system is harmless or a conclusion that broad intelligence has been proven.
The Astra safety overview explains the developer's risk assessment and safeguards.
For an operating team, permission is a separate decision. A system can be capable of completing a task without being authorized to perform every action that helps it finish. A credible evaluation should identify boundary violations as failures even when the requested artifact appears. Otherwise, the score rewards behavior the production environment cannot accept.
Build a One-Page Claim Review
Use six fields: claim, definition, tested system, evidence, limits and next decision. Copy the claim precisely and attach the dated primary source. Describe the test conditions in plain language. Record what the evidence supports, what it leaves unknown and which additional finding would change your conclusion. The review should make disagreement possible without forcing another person to reconstruct your research.
Here is an illustrative entry: a system solves nearly every task in a named interactive benchmark under a specified action budget. The supported conclusion concerns that benchmark and configuration. The unresolved questions concern transfer, reliability outside the suite and the broader definition of intelligence. The next decision might be a bounded pilot on a relevant workflow. This entry illustrates the review method; it contains no new experimental result.
Where This Method Stops
A claim review cannot inspect private training data, reproduce a closed evaluation without access or resolve philosophical disagreements by itself. It can make those limits explicit. Missing evidence should remain unknown. It should neither become proof of a hidden breakthrough nor an accusation that the developer fabricated a result. Ask for the evidence needed to narrow the uncertainty.
This article also does not rank models for your company. That requires representative work, acceptance criteria and a cost model specific to your operation. The AGI debate concerns the interpretation of broad capability claims. A purchasing decision concerns whether an available system can complete a defined job at an acceptable quality, cost and level of oversight.
When you are ready for that narrower decision, use the guide to evaluating AI models on real work to design the pilot.
Take the next major AGI announcement and complete the six-field review before changing your operating plan. Keep the impressive result, its conditions and its limits together. That produces a more useful answer than choosing a side in the headline: a clear account of what has been demonstrated and what your team should test next.

