OpenAI Releases GPT-6 Astra: A 1.05M-Context Computer-Use Model Gated Behind a ‘Critical’ Cyber Threshold
Today, OpenAI launched GPT-6 Astra. The firm calls it its most clever and aligned mannequin, and positions it primarily as a computer-use system relatively than a chat mannequin. The pitch is that Astra operates software program the best way a individual does, throughout browsers, spreadsheets, desktop functions and terminals, and finishes multi-step jobs as an alternative of describing methods to do them.
Is it deployable? Partly, and never by yourself {hardware}. Astra is a closed, hosted mannequin with no launched weights, so self-hosting shouldn’t be an choice. It is stay immediately just for organizations in OpenAI’s Trusted Access and Daybreak packages.
What is definitely new
The primary change for devs is context dealing with. Codex beforehand used compaction, summarizing earlier turns as soon as context stuffed up. That course of discards the element an agent later wants: why a repair failed, which exams ran, which requirement was added early. Astra as an alternative retains notes throughout context home windows and searches again into earlier messages and power output. The characteristic ships experimental behind a config.toml setting and turns into the Codex default within the coming weeks.
Astra can even ask the person a query whereas persevering with work that doesn’t depend upon the reply. That removes a frequent agent failure the place one unresolved choice stalls a complete job.
On the model page, Astra lists a 1,050,000-token context window, 128,000 max output tokens and an April 30, 2026 information cutoff. Input is textual content and picture, output is textual content solely. reasoning.effort provides two new ranges above excessive: xhigh and max. Tool assist covers computer use, hosted shell, apply patch, abilities, MCP and power search. Fine-tuning shouldn’t be supported.
The benchmark image
OpenAI stories 72.6% on OSWorld V2-Offline towards 65.7% for GPT-5.6 Sol, with common process time falling from roughly 75 minutes to 40. Anthropic stories 77.9% for Claude Fable 5.1 however says it used a completely different OSWorld launch and shouldn’t be in contrast immediately.
Astra scores 98.6% on ARC-AGI-3. That quantity was produced with a Responses API harness that retains reasoning between turns and makes use of compaction for lengthy contexts, and OpenAI has previously shown these settings transfer ARC-AGI-3 scores considerably with out altering the mannequin. The outcome measures the mannequin plus the agent system.
Other reported figures: 97.6% on FrontierMath Tier 4, 95.9% on BenchCAD Vision2Code towards 84.3% for Fable 5.1, and 64.6% on Terminal-Bench Science towards Anthropic’s reported 52.6%. Epoch AI notes OpenAI funded FrontierMath and has unique entry to a part of it.
Coding is the weak spot within the story. Astra scores 74.1% on DeepSWE v1.1 versus 70.8% for Sol. Meta reported 75.4% for Muse Spark 1.3 at most reasoning, and the public leaderboard places Gemini 3.8 Flash and Claude Opus 5 close to 74%. On a 113-task benchmark, these gaps are one or two duties.
Cyber functionality drives the entry mannequin
Astra is the primary mannequin OpenAI has designated as reaching the Critical cybersecurity threshold in its Preparedness Framework. In testing it developed exploits for hardened browsers and working techniques, and located two beforehand unknown V8 vulnerabilities that OpenAI says it’s disclosing to maintainers.
The penalties are sensible. Standard entry refuses superior cybersecurity work together with exploit discovery. For API builders, a cybersecurity security test stops a process outright relatively than pausing for approval. OpenAI’s Mia Glaese warned that customers outdoors trusted-access packages could hit slowdowns, pauses or blocks, generally throughout unrelated work.
OpenAI stories 100% on ExploitBench, an combination capability-coverage rating relatively than a go price, and 42.4% on ExploitGym towards 30.3% for Sol, with the standard six-hour time restrict eliminated for each.
Pricing
Astra prices $10 per million enter tokens and $50 per million output, with cached enter at $1.00. Requests above 272K enter tokens invoice at 2x enter and 1.5x output for the complete request. Batch and Flex run at 50%, Fast mode at 2x. Pro, Business and Enterprise customers additionally get Astra Pro.
Key Takeaways
- Astra is a computer-use mannequin first: 72.6% OSWorld V2-Offline, process time down from ~75 to ~40 minutes.
- Notes change compaction in Codex, so lengthy agent runs cease shedding failure element.
- Coding beneficial properties are marginal: 74.1% DeepSWE v1.1 sits contained in the leaderboard pack.
- First mannequin at OpenAI’s Critical cyber threshold; commonplace entry refuses exploit work.
- No open weights, $10/$50 per million tokens, 1.05M context, API and AWS in coming days.
Check out the OpenAI announcement, OpenAI on X and GPT-6 Astra model page. Also, be happy to comply with us on Twitter and don’t overlook to affix our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to accomplice with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and many others.? Connect with us
The put up OpenAI Releases GPT-6 Astra: A 1.05M-Context Computer-Use Model Gated Behind a ‘Critical’ Cyber Threshold appeared first on MarkTechPost.
