Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM
Sakana AI has launched Fugu-Cyber (mannequin ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu orchestration household. It is not only a brand new frontier mannequin. It is a 3rd endpoint on the Fugu orchestrator, tuned for safety reasoning. Sakana launched that orchestrator a month earlier.
Sakana stories successful charge of 86.9% on CyberFitness center and 72.1% on CTI-REALM. It describes these outcomes as akin to cyber-focused frontier fashions equivalent to GPT-5.5-Cyber and Claude Mythos Preview.
What the 2 benchmarks truly measure
The two evaluations sit at reverse ends of a safety workflow:
- CyberGym is a UC Berkeley benchmark of 1,507 real-world vulnerabilities throughout 188 OSS-Fuzz initiatives. In its principal process, an agent receives a vulnerability description and an unpatched codebase. It should write a proof-of-concept that crashes the pre-patch construct however not the post-patch construct. That verification step is what makes the benchmark exhausting to recreation.
- CTI-REALM is Microsoft’s open-source detection-engineering benchmark. Microsoft curated 37 public menace stories from sources together with Datadog Security Labs, Palo Alto Networks, and Splunk. An agent should map MITRE ATT&CK methods, discover telemetry, iterate on KQL queries, and emit validated Sigma guidelines. Scoring covers Linux endpoints, Azure Kubernetes Service, and Azure cloud.
Together the pair spans ‘discover and show the bug’ and ‘flip intel right into a detection.’ That framing is essentially the most defensible a part of Sakana’s announcement.
Where 86.9% sits in opposition to the sphere
Context issues greater than the quantity.
When the CyberGym researchers published their first outcomes, the perfect agent-model pairing reached roughly 20%. Anthropic reported 83.1% for Claude Mythos Preview beneath Project Glasswing in April 2026. OpenAI reported 85.6% for its up to date GPT-5.5-Cyber, in opposition to 81.8% for GPT-5.5. Sakana’s 86.9% is due to this fact a small step previous the reported frontier, not a bounce.
CTI-REALM is a special story. Microsoft’s personal analysis put the highest three configurations, all Claude, in a band from 0.624 to 0.685. Fugu-Cyber’s 72.1% would sit above that band. One caveat issues. CTI-REALM is scored as a trajectory reward between 0 and 1. It isn’t a cross/fail charge. Sakana calls it successful charge anyway.
How the orchestration works
Fugu is itself a language mannequin. It is educated to learn a question and construct an agentic scaffold on the fly. It then delegates sub-tasks to specialist fashions in a pool.
The strategy is documented within the Fugu technical report and two ICLR 2026 papers, TRINITY and the Conductor. TRINITY assigns Thinker, Worker, and Verifier roles throughout a number of LLMs. The Conductor learns natural-language coordination methods via reinforcement studying.
For safety work, Sakana analysis crew argues the verifier function is the purpose. A candidate vulnerability surfaced by one agent will get validated by security-specialized sub-agents earlier than any patch is proposed. Routing stays proprietary, so you can’t see which mannequin dealt with which step.
Access, coverage, and worth
Fugu-Cyber is gated on 4 dimensions.
Access requires an utility type stating the supposed use case and verified contact particulars. Sakana crew opinions each manually. The mannequin ships beneath an up to date Acceptable Usage Policy that prohibits offensive misuse. Billing is restricted to the Token Plan. The $20, $100, and $200 subscription tiers cowl Fugu and Fugu-Ultra solely. And the Fugu API isn’t supplied within the EU or EEA whereas Sakana works towards GDPR compliance.
Pricing is mounted at $6 per million enter tokens, $36 output, and $0.60 cached enter. All three charges double above a 272K-token context. Every line is strictly 1.2× the Fugu-Ultra charge, a flat 20% premium for the cyber endpoint. Long codebase runs cross 272K simply, so the doubled tier isn’t an edge case.
