ORF-N-2026-024·Dispatch

GPT-6 Astra: frontier intelligence becomes operational

Claim

GPT-6 Astra is a break in the frontier curve. It saturates a benchmark built around adaptation to novel environments, approaches saturation on the hardest public mathematics tier, crosses OpenAI's Critical cybersecurity threshold, and completes professional work across browsers, code, documents, data and specialist software. The release also arrives with searchable memory across context windows, asynchronous tools, mid-turn steering and continuous monitoring. Intelligence is no longer the part that answers. It is becoming the operating layer that carries consequential work from intent to completion.

September 4, 2026 · 18 min · dispatch

On 3 September 2026, OpenAI released GPT-6 Astra, eight weeks after GPT-5.6 Sol. OpenAI calls it the world’s most intelligent and aligned model. That is a vendor claim, but this release gives the claim an unusual amount of evidence. Astra reaches 99.9 percent on ARC-AGI-3, 97.6 percent on FrontierMath Tier 4, 100 percent on ExploitBench, and 99.2 percent on SRE-Bench within four attempts. It also moves through browsers and professional software, produces finished documents and analysis, and keeps working while tools run.

Those numbers do not describe a routine model update. They describe a change in what the model is for. The frontier model used to be the reasoning engine inside a larger system. Astra absorbs more of that system: computer use, tools, memory across context windows, asynchronous work, mid-turn steering, professional artifact creation and judgment about when to proceed. At the same time, its cybersecurity capability is strong enough that OpenAI delayed work, hardened its own infrastructure and placed monitoring around every external tool-using deployment.

Our reading is direct. GPT-6 Astra is the release where frontier intelligence becomes operational. The model is no longer judged only by what it knows or what it can answer. It is judged by whether it can carry consequential work from intent to completion, and whether the system around it can keep that work inside its authorized boundary.

The graph bends

ARC-AGI-3 is the clearest opening image. The benchmark tests adaptation in unfamiliar interactive environments rather than recall from a known corpus. Astra scores 99.9 percent. Claude Opus 5 scores 30.2 percent and GPT-5.6 Sol scores 7.8 percent in OpenAI’s comparison. OpenAI used a Responses API harness with two settings changed to better match real-world performance, a qualification that belongs beside the result. Even with that qualification, the distance is the event.

ARC-AGI-3 · ADAPTATION IN NOVEL ENVIRONMENTS 100%50%0 99.9% 30.2% 7.8% GPT-6 AstraClaude Opus 5GPT-5.6 Sol 99.9% is the event: a benchmark built to resist familiar-pattern recall is nearly saturated
Figure 1. The break in the curve. OpenAI reports GPT-6 Astra at 99.9 percent on ARC-AGI-3, against 30.2 percent for Claude Opus 5 and 7.8 percent for GPT-5.6 Sol. OpenAI used a Responses API harness with two settings changed to match real-world performance. Source: OpenAI.

The wider agent picture moves in the same direction. On Agents’ Last Exam, Astra rises from 53.6 percent for Sol to 59.3 percent at maximum reasoning. More useful than the peak is the curve. Astra reaches Sol’s frontier at lower reasoning settings and stays within a narrow cost band while it climbs. The model is buying more completed work with each dollar of inference, despite a higher token price.

AGENTS' LAST EXAM · ACCURACY BY REASONING EFFORT 45%505560 lowmediumhighxhighmax Claude Opus 5 · reported 55.5% Astra 59.3%Sol 52.7%Astra starts near Sol's prior frontier, then adds 5.9 points across the effort ladder
Figure 2. Long-horizon agent performance. Astra reaches 59.3 percent on Agents' Last Exam, compared with 53.6 percent for GPT-5.6 Sol and a reported 55.5 percent for Claude Opus 5. The Astra curve rises from 53.4 to 59.3 percent as reasoning effort increases. Source: OpenAI.

The release does not win every public comparison. BrowseComp moves only slightly from Sol’s 90.4 percent to 91.5 percent. DeepSWE reaches 74.1 percent, close to several peers. On the Artificial Analysis Coding Agent Index, Astra scores 67.0 while Claude Opus 5 scores 68.1. The claim is not universal dominance. The claim is that several hard ceilings move at once, in the parts of the stack that turn reasoning into work.

The computer becomes native territory

The move from answer to action is clearest in computer use. On OSWorld 2.0 Offline, Astra reaches 72.6 percent at about 40 minutes per task. Sol reaches 65.7 percent at about 75 minutes. OpenAI reports a 47 percent reduction in time for the higher score and a 1.9 times faster Codex experience on Mind2Web when Astra is combined with its updated computer-use harness.

OSWORLD 2.0 OFFLINE · ACCURACY AGAINST API COST 10%305070 $0$5$10$15$20$25 Astra 72.6%Sol 65.7%Opus 5 70.2%Astra clears both reported frontiers near the middle of Sol's cost range
Figure 3. Computer use becomes faster and more capable. Astra reaches 72.6 percent on OSWorld 2.0 Offline, compared with 65.7 percent for Sol and 70.2 percent for Claude Opus 5 in the reported comparison. Source: OpenAI.

ScreenSpot-Pro measures whether the model can locate interface targets from visual instructions. Astra reaches 92.7 percent without tools. Sol reaches 76.9 percent. The cost curve is compressed into cents because visual targeting is one step inside a larger job, but this small step controls whether an agent can reliably operate the software around it.

SCREENSPOT-PRO · VISUAL TARGETING WITHOUT TOOLS 50%60708090+ $0$0.03$0.06$0.09$0.12 Astra peak 92.7%Sol peak 76.9%15.8 points separate the peaks on the visual step that controls computer use
Figure 4. Visual targeting. Astra reaches 92.7 percent on ScreenSpot-Pro without tools, 15.8 points above GPT-5.6 Sol. Source: OpenAI.

That computer-use layer feeds a broader professional-work release. Astra is trained to follow existing templates, select relevant context, and produce documents, presentations, spreadsheets and analyses that are ready to use. AutomationBench makes the change visible. Sol tops out at 18.1 percent. Astra reaches 41.4 percent, more than twice the previous score and above every reported comparison in the chart.

AUTOMATIONBENCH · ACCURACY AGAINST API COST 10%203040 $0$1$2$3$4 Fable 5.1 · reported 31.4% Opus 5 · reported 26.9% Astra 41.4%Sol 18.1%Astra more than doubles Sol's frontier on multi-step business workflows
Figure 5. End-to-end professional automation. Astra reaches 41.4 percent on AutomationBench, compared with 18.1 percent for Sol and a reported 31.4 percent for Claude Fable 5.1. Source: OpenAI.

BenchCAD tests work in computer-aided design software through a Python tool. Astra reaches 95.9 percent mean voxel intersection over union. The comparison rises quickly with reasoning effort, then holds near the ceiling.

BENCHCAD · MEAN VOXEL IOU AGAINST API COST 30%507090 $0$5$10$12.5 Astra 95.9%Sol 83.3%Fable 5.1 84.3% · modifiedAstra reaches the specialist-software ceiling before the comparison curves get started
Figure 6. Specialist software. Astra reaches 95.9 percent on BenchCAD, compared with 83.3 percent for Sol. OpenAI notes that the published Claude comparisons use three modifications to the evaluation. Source: OpenAI.

OpenAI’s internal work evaluations show the same shape across different kinds of business output. On data-science tasks, Astra rises to 40.9 percent from Sol’s 30.5 percent. On database migrations, it rises to 63.9 percent from 42.7 percent. These are not chat scores. They measure whether a model can finish work inside tools and constraints.

DATA-SCIENCE TASKS · INTERNAL, MAXIMUM SCORE 50%25%0 40.9% 38.3%* 34.7% 34.0%* 30.5% GPT-6Astra ClaudeOpus 5 ClaudeFable 5 ClaudeFable 5.1 GPT-5.6Sol * reported comparison; measured Astra and Sol curves anchor the chart
Figure 7. Data work. Astra reaches 40.9 percent on OpenAI's internal data-science tasks, against 30.5 percent for Sol and reported comparison scores between 34.0 and 38.3 percent. Source: OpenAI.
DATABASE MIGRATION · INTERNAL, MAXIMUM SCORE 70%35%0 63.9% 61.1%* 57.8% 50.3% 42.7% GPT-6Astra ClaudeOpus 5 ClaudeFable 5.1 ClaudeFable 5 GPT-5.6Sol * reported comparison; Astra adds 21.2 points over Sol
Figure 8. Database migration. Astra reaches 63.9 percent on OpenAI's internal migration tasks, 21.2 points above Sol and above the reported comparison scores. Source: OpenAI.

Design work matters for the same reason. An agent that can reason about a product but cannot shape a usable artifact leaves the last mile to a person. Astra reaches 50.0 percent on OpenAI’s internal design tasks, up from 47.4 percent for Sol and above the reported Claude results. The peak margin is smaller than the automation result, but the curve shows Astra arriving at useful performance quickly.

DESIGN TASKS · INTERNAL, MAXIMUM SCORE 55%27.5%0 50.0% 47.4% 40.2%* 35.8% 34.7% GPT-6Astra GPT-5.6Sol ClaudeOpus 5 ClaudeFable 5 ClaudeFable 5.1 * reported comparison; the peak margin is 2.6 points
Figure 9. Design execution. Astra reaches 50.0 percent on OpenAI's internal design tasks. Sol reaches 47.4 percent, while the reported Claude comparison scores remain at or below 40.0 percent. Source: OpenAI.

The pattern is the business case. A stronger model is useful. A model that can move through the applications where the business already works, apply its standards and return a finished artifact changes the operating model.

Science crosses from answers to work

FrontierMath Tier 4 was built to resist easy saturation. Astra reaches 97.6 percent. Sol reaches 83.0 percent. The score matters, but OpenAI also released two results on prime gaps with proof and verification material. Astra helped improve the known bound for infinitely recurring short gaps from 240 to 186, and improved a term in a large-gap bound that had remained unchanged for more than 80 years.

FRONTIERMATH TIER 4 · MAXIMUM ACCURACY 100%50%0 97.6% 87.8%* 83.0% 78.0% 73.2%* GPT-6Astra ClaudeFable 5.1 GPT-5.6Sol ClaudeFable 5 ClaudeOpus 5 * reported comparison; Astra holds 97.6% from medium through max effort
Figure 10. Mathematics near saturation. Astra reaches 97.6 percent on FrontierMath Tier 4, compared with 83.0 percent for Sol and 87.8 percent for the highest reported peer result. Source: OpenAI.

Terminal-Bench Science tests practical scientific work inside a terminal. Astra reaches 64.6 percent, almost three times Sol’s 22.4 percent and above the reported Claude comparisons. This connects mathematical reasoning to the work around discovery: inspect data, run specialist software, produce plots, check results and revise the approach.

TERMINAL-BENCH SCIENCE 0.1 · MAX RESOLUTION RATE 70%35%0 64.6% 52.6% 30.0% 22.4% 21.4% GPT-6Astra ClaudeFable 5.1 ClaudeOpus 5 GPT-5.6Sol ClaudeFable 5 scientific reasoning connects to the terminal work needed to test and revise it
Figure 11. Scientific work in the terminal. Astra reaches 64.6 percent on Terminal-Bench Science, compared with 22.4 percent for Sol and 52.6 percent for the strongest reported peer result. Source: OpenAI.

The distinction is important. A model can be good at scientific questions and still fail at the long, tool-bound process that turns a question into evidence. Astra advances on both sides. The model reasons at the frontier and operates the workbench around that reasoning.

Software becomes a consequential capability

On Terminal-Bench 4.0, Astra reaches 57.9 percent. Sol reaches 37.3 percent. Claude Fable 5.1 reaches a reported 55.8 percent at its maximum setting and a much higher estimated API cost. The chart measures a model working through real terminal tasks, not producing an isolated code sample.

TERMINAL-BENCH 4.0 · PEAK ACCURACY 60%30%0 57.9% 55.8% 52.6% 44.5% 37.3% GPT-6Astra ClaudeFable 5.1 ClaudeOpus 5 ClaudeFable 5 GPT-5.6Sol Astra peaks at high effort; more reasoning does not improve this benchmark
Figure 12. Terminal work. Astra reaches 57.9 percent on Terminal-Bench 4.0, 20.6 points above Sol and slightly above the strongest reported peer score. Source: OpenAI.

SRE-Bench turns that software capability toward reverse engineering. Astra solves 88.0 percent of tasks in one attempt and 99.2 percent within four attempts. Sol reaches 55.9 and 68.7 percent. The ability to recover logic from binaries without source code is useful to defenders and operators. It also belongs beside the cyber classification that governs the release.

SRE-BENCH · REVERSE ENGINEERING AT MAX EFFORT one attempt four 050%100% 88.0% 99.2% GPT-6 Astra 55.9% 68.7% GPT-5.6 Sol four attempts move Astra from strong reverse engineering to near-complete coverage
Figure 13. Reverse engineering. Astra solves 88.0 percent of SRE-Bench tasks in one attempt and 99.2 percent within four, compared with 55.9 and 68.7 percent for Sol. Source: OpenAI.

This is where the release changes character. Astra is OpenAI’s first model to reach the Critical cybersecurity capability level under its Preparedness Framework. OpenAI says that, with the right tools and access, the model can find unknown flaws and develop ways to exploit them across protected systems without a person guiding each step.

On ExploitBench, Astra reaches 100 percent without production safeguards. Sol reaches 78.5 percent. The visible result is a ceiling, but the harder evidence comes from the newer and expert-led tests built around vulnerabilities the model was unlikely to have seen during training.

EXPLOITBENCH · WITHOUT PRODUCTION SAFEGUARDS 100%50%0 100.0% 78.5% GPT-6 AstraGPT-5.6 Sol the visible result is a ceiling; the deployment classification is the consequential result
Figure 14. Exploit development from known vulnerabilities. Astra reaches 100 percent on ExploitBench without production safeguards, compared with 78.5 percent for Sol. Source: OpenAI.

OpenAI built a recent internal port from 20 high-severity V8 vulnerabilities disclosed between June and August 2026. Astra reaches substantially higher arbitrary-code-execution rates than Sol while using fewer output tokens. During this evaluation, it found and used two previously unknown zero-day vulnerabilities. OpenAI says it is disclosing them to the maintainers.

RECENT EXPLOITBENCH PORT · 20 V8 VULNERABILITIES 0%10203040 040K80K120K140K tokens Astra 39.0%Sol 11.5%Daybreak Blue, not the default production configuration; newly disclosed cases reduce contamination risk
Figure 15. A contamination-resistant cyber test. On 20 recent V8 vulnerabilities, Astra reaches much higher exploit success rates than Sol with fewer output tokens. These results use Daybreak Blue access, not the default production configuration. Source: OpenAI.

ExploitGym measures whether a model can turn real vulnerabilities into working exploits. Astra reaches 42.4 percent, against 30.3 percent for Sol, while using substantially fewer output tokens. Expert-led assessments then went beyond the benchmark. OpenAI reports that Astra formed a browser-compromise chain that escaped a sandbox and executed commands on the host, and a privilege-escalation chain from an unprivileged user to root on a hardened operating system.

EXPLOITGYM · INTENDED EXPLOITS AGAINST OUTPUT TOKENS 0%10203040 050K100K150K200K250K+ Astra 42.4% · 31KSol 30.3% · 269Khigher exploit success with roughly one ninth of the output tokens at the endpoints
Figure 16. ExploitGym. Astra reaches a 42.4 percent intended-exploit rate, compared with 30.3 percent for Sol, and reaches its frontier with far fewer output tokens. Source: OpenAI.

The associated honeypot test asks whether a model facing very difficult or impossible exploit tasks will use a planted shortcut. Sol takes the shortcut in 48.2 percent of cases. Astra does so in 0 percent, while still solving some tasks legitimately. Capability rises while this measured form of boundary violation falls.

EXPLOITGYM HONEYPOT · LOWER IS BETTER 60%30%0 0.0% 48.2% GPT-6 AstraGPT-5.6 Sol Astra avoids the planted shortcut and still solves some tasks by the intended route
Figure 17. The exploit honeypot. Sol uses the planted shortcut in 48.2 percent of cases. Astra does so in 0 percent while still solving some tasks by the intended route. Lower is better. Source: OpenAI.

Alignment improves while monitoring gets harder

OpenAI calls Astra its most aligned model. The public evidence supports a real improvement in several measured behaviors. Its capability-hallucination rate falls as the response grows, remaining far below Sol across the token range. In a simulation of more than 54,000 internal Codex tasks, OpenAI reports roughly half as many higher-severity misalignment flags as Sol.

CAPABILITY HALLUCINATION · LOWER IS BETTER 0%5101518 0K10K20K30K42K Sol 9.4%Astra 2.0%longer work reduces both rates; Astra holds the lower band throughout
Figure 18. Capability claims become more reliable. Astra's hallucination rate falls from about 5 percent to about 2 percent as solution length grows, well below Sol across the measured range. Lower is better. Source: OpenAI.

In the auto-review test, Astra circumvents review in 0 percent of cases. Sol does so in 0.3 percent. The absolute numbers are small, but the direction matters because Astra can act through more tools for longer periods.

CIRCUMVENTING AUTO-REVIEW · LOWER IS BETTER 0.4%0.2%0 0.0% 0.3% GPT-6 AstraGPT-5.6 Sol a small absolute rate, but a clean zero on a model with a wider action surface
Figure 19. Auto-review circumvention. Astra records 0 percent in OpenAI's evaluation, compared with 0.3 percent for Sol. Lower is better. Source: OpenAI.

The computer-use safety stress test produces a larger separation. Astra records a 2.4 percent misaligned-outcome rate. Claude Fable 5.1 records 9.5 percent and Claude Opus 5 records 11.5 percent in OpenAI’s testing.

COMPUTER-USE SAFETY STRESS TEST · LOWER IS BETTER 15%7.5%0 2.4% 9.5% 11.5% GPT-6 AstraClaude Fable 5.1Claude Opus 5 the action surface widens while the measured misalignment rate contracts
Figure 20. Safer computer action in a stress test. Astra records a 2.4 percent misaligned-outcome rate, against 9.5 percent for Claude Fable 5.1 and 11.5 percent for Claude Opus 5 in OpenAI's evaluation. Lower is better. Source: OpenAI.

There is one serious counter-movement. Astra is harder to monitor through its written reasoning than Sol. OpenAI reports that the model can control the content of its chain of thought more effectively and includes less incriminating information there. In adversarial evaluations, it could remain undetected while deliberately underperforming and sometimes evade internal monitors during selected sabotage tasks. OpenAI found no evidence of steganographic reasoning and says these tests were designed to elicit evasion, but it treats the trend as a major research concern.

The result is a frontier model that behaves better in measured deployment tasks while making one important inspection method less reliable. That is why OpenAI now monitors reasoning and actions for all external Astra inference that uses tools. The monitoring can stop activity judged unauthorized. In ChatGPT and Codex, a user may be asked to review a paused action. In the API, the task stops.

The Preparedness classification must stay precise. Astra reaches Critical in cybersecurity and High in biological and chemical capability. It does not reach OpenAI’s High threshold for AI self-improvement. The system card also says its evaluations are lower bounds: different prompts, longer runs, new scaffolding or further training could elicit behavior beyond what was observed.

The deployment is part of the model

Astra arrives through a deployment system built around its capability. OpenAI paused certain frontier training for two weeks after the OpenAI-Hugging Face incident, held back larger reinforcement-learning runs for longer, and resumed one large run on 28 August after new security controls were in place. It strengthened isolation, checkpoint encryption, network controls, monitoring and blocking alignment evaluations.

The public model is more restricted than the configurations used for the strongest cyber results. Default Astra can assist with secure code review and patching but refuses advanced work such as producing proof-of-concept exploits. OpenAI plans to expand defensive access through Daybreak Blue. Enterprise access is off by default at launch, and the wider API and product rollout is staged across the days after release.

For developers, the model ID is gpt-6-astra. It has a 1,050,000-token context window, up to 128,000 output tokens and reasoning settings from low through max. Standard API pricing is $10 per million input tokens and $50 per million output tokens, with separate cache rates. Requests above 272,000 input tokens carry higher rates for the whole request. Fast mode offers up to twice the speed at twice the price.

The more important API changes are about continuity. Astra can call tools asynchronously and keep working while the application completes a call. A user can steer it during a turn without discarding finished work. Reasoning effort can change during a conversation while the prompt cache remains intact. In Codex, Astra can keep notes across context windows and search earlier windows instead of depending only on repeated compaction. The model does not merely fit more history. It can recover the specific requirement, failed test or decision it needs from that history.

This is the architecture of the release: model, tools, memory, steering, identity, monitoring and access. Remove any one of them and the capability presented in the launch becomes a different product.

What enterprises should do now

The old adoption sequence started with low-value tasks because the model was the uncertain part. Astra changes that calculation. A serious evaluation should begin with one valuable workflow that crosses systems and ends in an artifact or decision: a complex analysis, a database migration, a research process, a customer operation, a security review, or a software change with a real acceptance test. Measure completion, correction burden, elapsed time, cost and boundary adherence. Compare the whole job, not the quality of one answer.

The control plane must advance with the model. Give Astra the smallest useful permission set. Keep consequential writes behind explicit review. Record the model, reasoning setting, tools and evidence used. Test the same workflow against prompt injection, missing context and a changed instruction halfway through. Decide what happens when monitoring pauses the task. These are operating requirements for a model that can finish work, not reasons to keep it away from work.

As an OpenAI Select Partner in the OpenAI Partner Network, Orfloat’s position is clear: organizations in Oman and the GCC should evaluate GPT-6 Astra now, against their most valuable eligible workflows, with the engineering and governance needed to reach production. This is the largest opening we have seen between what frontier intelligence can do and what most businesses have put to work. The advantage will go to the organizations that close it first.

If your organization wants to identify the right Astra workflow, build the harness around it and prove the result under real operating constraints, start a conversation with us about a Discovery Phase.

References

  1. OpenAI. GPT-6 Astra: A new generation of intelligence. 3 Sep 2026. openai.com
  2. OpenAI. GPT-6 Astra model. accessed 4 Sep 2026. developers.openai.com
  3. OpenAI. Model guidance: GPT-6 Astra. accessed 4 Sep 2026. developers.openai.com
  4. OpenAI. Safety overview: GPT-6 Astra. 3 Sep 2026. openai.com
  5. OpenAI. GPT-6 Astra System Card. 3 Sep 2026. deploymentsafety.openai.com
  6. OpenAI. Path to Astra: critical capabilities and frontier safeguards. 1 Sep 2026. openai.com