Model intelligence · Experimental findings
How Smart Are the Smartest AI Models? Comparing Five Areas of Intelligence
A comparison of GPT-6 Astra and Claude Opus 5.5: planning, consequence prediction, context use, adaptability, and explanation accuracy.
Orgenix field note ·
“Which AI model is smarter?” becomes a more useful question when we specify what we need it to do. Finding a promising plan, anticipating what could go wrong, and applying a constraint are different abilities. A model can demonstrate one while failing at another.
At Orgenix, one area we are researching in depth is how to choose the right model for each task in software development. This experiment is an exploratory part of that work: identifying behaviors worth testing in real engineering tasks.
We wanted to examine those abilities through decisions with observable consequences. We asked OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5 to play chess against each other, saving their moves and short explanations. Astra played White; Opus played Black. An independent chess engine, Stockfish, evaluated the positions without supplying either player with advice.
This article compares five areas of intelligence using that record. Each section explains the ability, why it matters, which model the evidence favors, and the specific decisions supporting that assessment. These are qualitative judgments about the observed performance, not standardized intelligence scores or predictions of performance in other domains.
| Area | Winner | Justification |
|---|---|---|
| Planning | No clear winner | Astra found a strong attacking continuation; Opus predicted and executed the final forced mate. These examples do not establish an overall planning advantage. |
| Consequence prediction | Opus | Identified and exploited immediate threats that Astra failed to handle in the decisive sequence. |
| Context use | Opus | Acted on exposed pieces and immediate threats that Astra’s decisions overlooked despite the supplied board state. |
| Adaptability | Opus | Exploited the reversal and converted its advantage while Astra continued costly attacks. This is a limited judgment about the observed decisions. |
| Explanation accuracy | Opus | Made specific predictions verified by subsequent play; Astra’s critical explanations omitted direct refutations. Opus also made an incorrect defensive suggestion. |
One condition affects the comparison: both models used Max reasoning settings, but the request instructions changed during the game after requests failed to produce moves. Immediately before the critical reversal, both were told to consider at most two candidate moves, submit promptly without revisiting the choice, and choose the first supplied legal move if uncertain. The comparison therefore concerns this recorded setup, not unrestricted reasoning under a fixed protocol.
Reading the diagrams: White (Astra) is at the bottom. Each board shows the position before the arrowed move. Teal arrows show actual moves; red marks costly actual moves; dashed gold shows an engine-supported alternative that was not played.
1. Planning
Planning means selecting actions that advance a goal over several steps. It matters whenever a model must do more than produce a plausible immediate answer: organize a project, develop an argument, or choose a sequence of actions whose benefits arrive later.
Verdict: both demonstrated useful planning; the evidence does not establish an overall leader. Astra found a strong attacking continuation. Opus eventually identified and executed a short forced sequence. Those are different planning achievements, and the result alone does not tell us which model is generally the better planner.
Evidence. Astra’s queen move, 19.Qg3, matched Stockfish’s preferred move in our timed analysis. It preserved a substantial advantage and supported an attack on the opposing king. Astra’s explanation connected the move to a follow-up bishop move, showing a stated purpose beyond the immediate action. One strong continuation does not establish the depth of its planning, but it is evidence of an effective choice.
Opus provided a particularly checkable example near the end. With 30...Qd1+, it predicted a forced checkmate in two moves. Astra had exactly one legal response, 31.Qe1; Opus then delivered 31...Qxe1#. Its stated sequence matched the legal possibilities and the outcome.
Neither model planned optimally throughout. Opus had missed a faster forced mate on move 25. Astra’s strong attacking position did not survive its subsequent decisions. The useful distinction is between finding a good plan and maintaining its validity as events unfold.
2. Consequence prediction
Consequence prediction means anticipating what an action makes possible—including responses that defeat its intended purpose. It matters because an action can look productive in isolation while creating a larger problem. A model proposing a change should consider both its benefit and the new risks it introduces.
Verdict: Opus was stronger in the decisive sequence. Its explanations identified immediate consequences that it subsequently exploited. Astra’s critical move emphasized attacking potential while failing to handle a direct counteraction.
Evidence. Opus played 19...Qb6, placing its queen on a diagonal that threatened Astra’s undefended rook and king. In its saved explanation, Opus identified the rook threat and specifically named Nf6+ as an unsuitable response for White. Astra did not receive this explanation; it was Opus’s own assessment.
Astra then chose exactly that knight check, 20.Nf6+, explaining:
“This forcing knight check pressures Black’s king and creates attacking chances against the kingside.”
Opus captured the knight. Astra continued its attack with a bishop move, and Opus captured the exposed rook with check. Astra had described what it wanted to achieve without adequately accounting for what its opponent could do next.
The independent evaluation made the cost visible:
| Decision | Replay evaluation from White’s perspective | Displayed win estimate |
|---|---|---|
Astra’s strong 19.Qg3 | +3.09 | Astra 100% |
Opus’s reply 19...Qb6 | +3.48 | Astra 100% |
Astra’s mistake 20.Nf6+? | −3.95 | Opus 100% |
The “100%” display was a rounded Stockfish estimate based on engine self-play, not a calibrated prediction of these models’ winning chances. It signaled a strongly favored position, not a guaranteed result. Stockfish’s WDL documentation explains its basis. Longer independent engine searches also confirmed the reversal.
Astra had a viable alternative: 20.Be3 preserved a large advantage in our engine check. The failure was avoidable. Opus’s threat did not force Astra to lose; Astra selected a response whose consequences it handled poorly.
3. Context use
Context use is the ability to make available facts constrain a decision. It matters when a model receives a long brief, an existing codebase, or a set of requirements. The presence of a fact in the input is only useful if the model applies it when choosing an action.
Verdict: Opus showed stronger use of the immediate board state during the collapse. Astra repeatedly selected moves that exposed pieces to direct capture, despite receiving the current position. This supports a comparison of observed decisions; it does not identify whether the internal cause was attention, calculation, or another process.
Evidence. At move 23, Astra played Rxf6, describing the removal of a defender and new attacking chances. Opus’s queen captured the rook with ...Qxf6. At move 24, Astra played Bd4, explaining that it attacked the queen and activated its bishop. Opus captured that bishop with ...Qxd4+.
Both captures were Stockfish’s preferred replies in the timed analysis. The relevant piece locations were in the supplied position. Astra’s explanations emphasized a favorable aspect of each move while its actions failed to account for the resulting exposure. Opus acted effectively on those same immediate relationships.
These examples overlap with consequence prediction: using a fact and calculating its consequences are connected. They should not be counted as independent experimental proof of two separate model advantages. They do, however, expose two useful evaluation questions: did the model incorporate the relevant constraint, and did it anticipate what follows?
4. Adaptability
Adaptability means revising a plan when new information changes what is sensible. It matters in extended tasks because an initially good approach can become inappropriate. Persistence is useful only while the plan remains sound.
Verdict: the observed decisions favor Opus, with limited evidence about adaptability itself. Opus exploited the changed opportunity and completed the game. Astra continued selecting attacking moves as its position deteriorated. That behavior is consistent with weaker adjustment in this sequence, but the record cannot prove that it consciously persisted with a fixed plan.
Evidence. After Astra’s critical knight error, its next bishop move, 21.Bh6, pursued a checkmate threat while leaving its rook exposed. Opus answered with 21...Qxd4+, taking the rook with check. Astra’s later rook and bishop losses continued the pattern of seeking activity without adequately handling direct replies.
Opus moved from an inferior position to exploiting those errors and eventually finding a forced finish. Its success depended on Astra’s mistakes; it did not demonstrate a forced recovery from the earlier disadvantage. The supported observation is that Opus made better use of the changed situation once the opportunity appeared.
5. Explanation accuracy
Explanation accuracy means that a model’s stated reasons agree with the facts and consequences of its action. It matters because users often decide whether to trust an answer by reading the accompanying explanation. Specific, testable claims give us a better basis for that judgment than fluent descriptions of intent.
Verdict: Opus’s explanations were more informative and accurate at the examined turning points, but were not consistently correct. Astra’s explanations for several losing moves described benefits without addressing the direct refutations.
Evidence. Opus’s explanation for 19...Qb6 identified the exposed rook and the threat to capture it. Its later rook-capture explanation correctly described winning that piece with check. Its final prediction of mate in two was verified by the forced sequence that followed.
Astra’s knight-check explanation promised “attacking chances”; its later bishop-move explanation described attacking the opposing queen. Those descriptions conveyed purpose but omitted the responses that made the moves fail.
Opus also supplied a useful counterexample to its own apparent authority. In the queen-move explanation, it suggested two defenses for White: Be3 and c3. Engine review supported Be3 as preserving White’s advantage, but found c3 losing that advantage. A detailed explanation can contain a correct diagnosis and an incorrect recommendation in the same paragraph.
These were short explanations submitted with moves, not access to private reasoning. We can assess their factual accuracy and predictive value without claiming they reveal the full process that produced the decision.
Applying the findings to real work
The practical question is which model to assign to a particular job. These observations give us a starting hypothesis: try Opus first for tasks that depend on checking consequences and challenging a proposed action; keep both models in contention for planning. The game does not establish an area where Astra should be the exclusive first choice, and it does not justify eliminating Astra from consideration.
The recommendations below are starting points for evaluation. We have not demonstrated that the chess differences transfer to coding, research, or business decisions. Before relying on a model for those tasks, compare its output against actual requirements and independently check the result.
| Real-world task | Model to try first | Why this is worth testing—and what to check |
|---|---|---|
| Reviewing a code change, migration, or deployment plan | Opus | Its stronger consequence checking suggests testing it as a reviewer of dependencies, failure paths, and unintended effects. Ask it to identify a concrete failure scenario; confirm the finding with code inspection or a test. |
| Troubleshooting a failed approach | Opus | It made better use of the changed situation in the observed sequence. Test whether it updates a diagnosis when new logs or evidence contradict the original plan, and whether its proposed fix resolves the failure. |
| Work with interacting constraints, such as permissions or data-sharing rules | Opus | Its handling of immediate relationships motivates testing whether it applies all relevant requirements together. Supply explicit acceptance criteria and check for violations; access to the requirements alone is insufficient. |
| Developing an architecture, project plan, or research approach | Compare Astra and Opus | Both demonstrated useful planning. Ask each for alternatives, assumptions, dependencies, and ways its plan could fail. Choose using feasibility and verified constraints rather than the persuasiveness of the explanation. |
| Producing an analysis or recommendation that others must audit | Opus as the initial candidate | Several of its explanations made accurate, checkable predictions. Require sources or reproducible checks, because the same model also offered an incorrect recommendation in a detailed explanation. |
Consider a software change that makes private notes available to a matching service. One task is to design a useful matching approach. Another is to inspect whether that approach could disclose private text. A third is to verify that the implementation respects the approved boundary. Those tasks call for different evidence: a feasible design, a concrete analysis of disclosure paths, and tests of the implemented behavior.
Our findings would make Opus a candidate for the consequence-review step, while leaving the design step open to both models. They would not justify trusting Opus’s review without verification. Its partly incorrect defensive advice in the game is exactly why a specific explanation still needs checking.
A workflow worth evaluating is therefore to have one model propose an approach and another challenge it against the requirements. For example, Astra could propose a design and Opus could review its failure cases. That pairing is a research hypothesis, not a workflow validated by this game. The deciding evidence should be whether it catches more real errors and produces better outcomes than a simpler approach, at an acceptable cost and latency.
What this means for Orgenix
Opus won the game. More usefully, the record shows where it outperformed Astra: anticipating immediate consequences, exploiting facts present in the position, and making several predictions that survived verification. Astra demonstrated strong play before a critical failure and then compounded the loss. Opus also made inaccuracies, including an earlier capture that increased Astra’s advantage and a missed faster checkmate.
These findings support a specific comparison under the recorded instructions. They do not establish a general intelligence hierarchy, and the midgame restriction on deliberation limits what we can infer about either model at its best.
At Orgenix, we are researching in depth how to match models to the tasks within software development: understanding requirements, evaluating architecture, implementing changes, reviewing risks, and verifying outcomes. A model that proposes a promising design may need a different kind of review from a model that excels at finding its failure cases.
Our aim is to make those assignments evidence-based: identify the ability a task requires, compare models on representative work, and measure the quality of the result alongside cost and latency. This game contributes observations and hypotheses to that research. The next standard of proof is whether those strengths hold up in the software tasks where we intend to use them.
Inspect the game
The full replay preserves the recorded turn timings and on-screen explanations. It has no spoken audio.
The complete move record is available in PGN format. The diagrams above use exact saved positions; the arrows distinguish played moves from alternatives.
Download the complete game (PGN)First cohort opens in November
Follow the build, one useful update at a time.
Join the list and tell us how you are building with coding agents.











