buildfastwithaibuildfastwithai
AI WorkshopsAll blogsAgentic AI Launchpad
Agentic AI Launchpad
Download Unrot App
Free AI Workshop
Mentorship

Agentic AI Launchpad

Go from user to builder in 6 weeks.

Explore Program
Claude Mastery Course
Share
Back to blogs
Analysis
Reviews
LLMs

Sakana Fugu-Cyber Review: Benchmarks & Access (2026)

July 21, 2026
13 min read
Share:
Sakana Fugu-Cyber Review: Benchmarks & Access (2026)
Share:

Sakana AI Fugu-Cyber: 86.9% on CyberGym, Reviewed

A model that scores 86.9% at verifying real vulnerabilities in real codebases should be the easiest product launch of the year. Sakana AI instead put it behind a manual approval form. That decision tells you more about the state of AI security in 2026 than the benchmark does.

Fugu-Cyber launched on July 21, 2026 as Sakana AI's cybersecurity-specialised orchestration model, hitting 86.9% on CyberGym and 72.1% on CTI-REALM, which the company describes as state of the art on the industry's hardest security benchmarks and comparable to frontier security models including GPT-5.5-Cyber and Mythos-Preview. It is not a new frontier model. It is a multi-agent system that behaves like one.

This review covers what Fugu-Cyber actually is, what those two benchmarks measure and why they were chosen, how orchestration differs from a single model, how the gated access process works, and the caveat Sakana itself put in the announcement: the model alone is not enough for enterprise security

What Is Fugu-Cyber?

Fugu-Cyber is a cybersecurity-specialised orchestration model from Sakana AI, released July 21, 2026, that dynamically coordinates multiple specialised agents to handle multi-step security tasks while presenting itself as a single model through one API. It targets two defensive workflows: verifying real-world vulnerabilities in complex codebases, and turning threat intelligence reports into detection rules.

The important architectural point is that Sakana did not train a bigger security model. Fugu-Cyber extends the orchestration approach behind the original Fugu, released June 22, 2026, which routes tasks across a pool of strong underlying models rather than relying on one. For security work, that pool gets specialised and the coordination gets tuned for depth and precision instead of speed.

Table 1: Fugu-Cyber at a glance

FUGU-CYBER AT A GLANCE

All figures from Sakana AI's official release announcement, July 21, 2026.

If you have not read our earlier coverage, our Sakana AI Fugu review explains the base orchestration model and the export-control strategy behind it, which is the necessary background for this release.

Benchmarks: 86.9% CyberGym and 72.1% CTI-REALM

Fugu-Cyber posts 86.9% on CyberGym and 72.1% on CTI-REALM, which Sakana positions as state of the art on the industry's most challenging security benchmarks and level with dedicated frontier security models. Those two numbers are the entire quantitative case for the release, so they deserve scrutiny rather than a headline.

Table 2: Fugu-Cyber benchmark results

FUGU-CYBER BENCHMARK RESULTS

Comparison models cited by Sakana: GPT-5.5-Cyber and Mythos-Preview. Independent third-party reproduction was not available at launch.

The gap between the two scores is the more interesting signal. Nearly 87% on vulnerability verification against 72% on detection engineering suggests the system is considerably stronger at analysing code it can read than at the more interpretive work of converting a written threat report into a working rule. That matches my general experience of AI security tooling: reading code is tractable, reading intent is not.

The caveat that matters: these are Sakana's own reported figures, and the comparison to GPT-5.5-Cyber and Mythos-Preview is provider-reported rather than a head-to-head run by a neutral party. That is normal for a launch, and it is still a claim rather than a verified result. Treat 86.9% as a strong signal, not a settled fact.

What CyberGym and CTI-REALM Actually Measure

These two benchmarks map onto the two pillars of defensive security work: finding what is broken, and knowing what to watch for. Sakana chose them deliberately, and understanding what they test tells you whether the scores apply to your job.

CyberGym: vulnerability verification

CyberGym evaluates whether a system can analyse complex real-world codebases and verify genuine vulnerabilities. The critical word is verify. Flagging a suspicious pattern is easy and produces mountains of noise, while confirming that a specific weakness is genuinely exploitable in a specific codebase is the hard, expensive work that occupies real security engineers. An 86.9% success rate on that task is a meaningful result if it holds up outside the benchmark.

CTI-REALM: detection engineering

CTI-REALM tests translating cyber threat intelligence reports into working detection rules. In practice, a human analyst reads a write-up of a new attack technique and converts it into something a monitoring system can act on. It is skilled, repetitive, chronically under-resourced work, which makes it an obvious automation target. The 72.1% score says the system is useful here but clearly still needs a reviewer.

Quotable version: one benchmark measures whether the AI can find the hole, the other measures whether it can describe the burglar. Fugu-Cyber is notably better at the first job.

How Orchestration Works

An orchestration model is a system that routes a task across multiple specialised agents and synthesises their outputs, while exposing a single model interface to the user. You send one API call; internally, Fugu decides whether to answer directly or assemble a team of expert agents, then handles delegation, verification and synthesis without you writing any coordination logic.

Sakana's distinguishing claim is that this coordination is learned rather than hardcoded. Traditional routers use if-else rules: send code questions here, send maths there. Sakana's approach trains the orchestration itself, so the system learns when to delegate, how agents should communicate, and how to combine results. Their published research on TRINITY and Conductor underpins that mechanism.

Table 3: Single model vs orchestration

SINGLE MODEL VS ORCHESTRATION

The swappable pool is the strategic point: orchestration routes around export controls and provider outages.

For security work specifically, the multi-step decomposition is the genuine fit. Verifying a vulnerability is not one question. It is reading a codebase, forming a hypothesis, tracing data flow, checking exploitability, and writing up the finding. Splitting that across agents with a coordinator is a more natural shape than asking one model to hold the entire chain in a single context window.

🚀 Cohort Waitlist Open
Go From AI User to AI Builder

Don't just use ChatGPT. Learn to build custom LLM agents, RAG pipelines, and full-stack Agentic AI apps in our intensive 6-week program.

6 Weeks Live Mentorship
Deploy 5+ Real-world Apps
Weekly App Templates & Code
No Coding Experience Required
Explore Program
Join 1,000+ graduates•Free Registration

Fugu vs Fugu-Cyber: What Changed

Fugu-Cyber is the same orchestration concept narrowed to security reasoning, with different access rules. The base Fugu launched June 22, 2026 as a general-purpose orchestrator with Fugu and Fugu Ultra tiers, available through an OpenAI-compatible API with subscription pricing. Fugu-Cyber arrived a month later, gated.

Table 4: Fugu vs Fugu-Cyber

FUGU VS FUGU-CYBER

Both are API-only products. Neither publishes open weights or discloses the full underlying model pool.

My read: the gating is the actual product decision here. Sakana could have shipped this as another tier on the existing console and captured more signups. Choosing a manual review queue instead costs them growth, which is a reasonable proxy for taking the dual-use problem seriously.

How to Get Access

Fugu-Cyber is available only as an API endpoint through Sakana's Token Plan, and every application is manually reviewed before approval. There is no open signup, no downloadable weights, and no self-serve path.

The application process

  1. Visit the Fugu access page on sakana.ai and open the access request form.
  2. Provide verified contact information. Anonymous or throwaway details will not clear review.
  3. Describe your intended use case specifically. Vague answers give the reviewer no basis to approve.
  4. Wait for manual review. Sakana's team assesses each application before granting endpoint access.
  5. Review the Acceptable Usage Policy, which was updated for this release and prohibits offensive misuse.

WHAT THIS MEANS FOR EVALUATION TIMELINES

If you are planning a security tooling bake-off, factor the approval queue into your schedule. Unlike a normal API you can test the same afternoon, Fugu-Cyber requires a human decision first. Apply before you need it, and expect to justify your use case in terms a reviewer can verify.

The Responsible Release Angle

Sakana paired the launch with an updated Acceptable Usage Policy prohibiting offensive misuse, manual vetting of every applicant, and stated alignment with industry safety standards. For a model this capable at finding vulnerabilities, that combination is the minimum defensible posture, and it is worth crediting because plenty of releases skip it.

The dual-use tension is unavoidable and worth naming plainly. A system that verifies exploitable vulnerabilities in real codebases at 86.9% is enormously valuable to a defensive security team, and the same capability is valuable to an attacker. The difference between the two is authorisation, not technique. Gating access behind identity verification and a stated use case is how you keep the capability pointed at defence.

Whether the gate holds is a separate question. Manual review filters out casual misuse and creates an accountability trail, and it will not stop a determined, well-resourced adversary who can present a plausible corporate identity. Sakana almost certainly knows this. The policy is a speed bump plus attribution, not a wall, and speed bumps plus attribution are genuinely how most security controls work.

The industry pattern worth noticing: GPT-5.6 shipped after a government-gated preview, and Fugu-Cyber ships behind manual approval. Frontier security capability increasingly arrives with a queue in front of it, and that is now normal rather than exceptional.

We covered the government-gated rollout precedent in our GPT-5.6 Sol, Terra and Luna review, which is the closest comparable case this year.

Sakana's Own Caveat: The Model Is Not Enough

The most credible part of this announcement is that Sakana argues against over-relying on its own product. The company states directly that frontier cyber models alone are insufficient for enterprise security, and lists why.

Table 5: Why the raw model is not enough

WHY THE RAW MODEL IS NOT ENOUGH

Sakana's Applied Enterprise team builds specialised harnesses and workflows around the model for production use.

A vendor telling you their model needs a harness, a workflow, and an expert reviewer before you act on its output is a vendor describing reality accurately. It is also, conveniently, a services pitch, since Sakana's Applied Enterprise team sells exactly that wrapper. Both things are true simultaneously, and I would rather have the honest framing plus the upsell than a claim of full autonomy.

My strong opinion: this caveat is the single most useful paragraph in the announcement. A security team that deploys Fugu-Cyber expecting autonomous vulnerability management will drown in false positives. A team that deploys it as a fast first-pass analyst whose output a human verifies will get real value from it.

🚀 Cohort Program Open
Claude Mastery: Cowork & Code

The only comprehensive program designed to take you from basic prompting to building interactive Artifacts, custom integrations, and deploying production-ready code with Claude Code.

No coding experience needed
Build interactive Artifacts & Agents
Deploy apps with Claude Code
Cohort-based learning & mentorship
Explore Program
Cohort-based training•Register Now

Who Should Actually Use This

Fugu-Cyber fits organisations with an existing security function that want to accelerate specific workflows, not teams hoping to substitute for one. Here is my honest read on fit by organisation type.

Table 6: Fit by organisation type

TABLE 6: FIT BY ORGANISATION TYPE

Access requires a verifiable identity and a stated use case, which structurally favours organisations over individuals.

The unifying rule: value scales with the quality of your verification layer. Fugu-Cyber makes a good security engineer meaningfully faster. It does not make an absent security engineer appear, and any team treating it as a replacement will discover that the hard way through a backlog of unvalidated findings.

Honest Limitations

Five constraints are worth weighing before you apply for access, and none of them are disqualifying on their own.

  • Benchmarks are self-reported. The 86.9% and 72.1% figures come from Sakana, and comparisons to GPT-5.5-Cyber and Mythos-Preview are provider-reported rather than independently run.
  • Access is gated and slow. Manual review before approval means you cannot evaluate it on the same day you decide to.
  • No open weights and no self-hosting. For security teams that cannot send code to a third-party API, that is a hard blocker regardless of the scores.
  • Pool composition is undisclosed. Orchestration routes across underlying models Sakana does not fully enumerate, which complicates compliance review and vendor risk assessment.
  • Detection engineering is the weaker half. At 72.1%, CTI-REALM output needs consistent human review before rules reach production.

Overall verdict, 8/10: a genuinely strong security model with an unusually honest launch. The score reflects real capability on vulnerability verification, docked for self-reported benchmarks, undisclosed pool composition, and API-only access that rules out the most security-sensitive environments. The orchestration bet continues to look smart, and the gated release is the right call.

For how Sakana's approach compares against the frontier models it orchestrates, see our July 2026 model ranking.

Frequently Asked Questions

Q: What is Sakana AI Fugu-Cyber?

Fugu-Cyber is a cybersecurity-specialised orchestration model released by Sakana AI on July 21, 2026. It is a multi-agent system that behaves like a single model, coordinating specialised agents to verify vulnerabilities in codebases and convert threat intelligence into detection rules. It scores 86.9% on CyberGym and 72.1% on CTI-REALM.

Q: What is the CyberGym benchmark?

CyberGym evaluates whether an AI system can analyse complex real-world codebases and verify genuine vulnerabilities, rather than simply flagging suspicious patterns. Verification is the harder and more valuable task because it filters out the false positives that make automated scanning noisy. Fugu-Cyber reports an 86.9% success rate.

Q: How good is Fugu-Cyber compared to GPT-5.5-Cyber?

Sakana describes Fugu-Cyber's performance as comparable to leading cybersecurity frontier models including GPT-5.5-Cyber and Mythos-Preview. That comparison is provider-reported rather than an independent head-to-head evaluation, so treat it as a credible claim awaiting third-party verification.

Q: How do I get access to Fugu-Cyber?

Submit an access request form on sakana.ai with verified contact information and a specific description of your intended use case. Sakana's team manually reviews and approves each application before granting access to the API endpoint under the Token Plan. There is no self-serve signup.

Q: Is Fugu-Cyber open source?

No. Fugu-Cyber is available only as a gated API endpoint with no open weights, no self-hosting option, and no public disclosure of the full underlying model pool. Organisations that cannot send code to a third-party API will not be able to use it.

Q: What is an orchestration model in AI?

An orchestration model routes a task across multiple specialised agents and synthesises their outputs while presenting a single model interface. Sakana's approach learns the coordination rather than hardcoding if-else routing rules, so the system decides when to delegate, how agents communicate, and how results combine.

Q: Can AI replace a security team?

No, and Sakana says so directly. The company states that frontier cyber models alone are insufficient for enterprise security, citing false positives, lack of production context, and the need for localised human expertise and verification workflows. Human review remains necessary before patches are proposed.

Q: What is CTI-REALM?

CTI-REALM is a benchmark that measures translating cyber threat intelligence reports into working detection rules, a core detection-engineering task normally done by human analysts. Fugu-Cyber scores 72.1%, notably lower than its 86.9% on CyberGym, which suggests interpretive work remains harder than code analysis.

Recommended Reads

  • Sakana AI Fugu review
  • Best AI models July 2026 ranking
  • GPT-5.6 Sol Terra Luna review
  • GLM-5.2 vs Claude vs Kimi coding
  • Every major LLM ranked in 2026

AI is getting genuinely good at security work, and the teams that win will be the ones who pair it with real human verification. Follow Build Fast with AI for hands-on coverage of every major model release.

References

  • Sakana AI, Introducing Fugu-Cyber
  • Sakana AI, Fugu release announcement
  • Sakana AI, model console
  • VentureBeat, Sakana frontier orchestration
  • DataCamp, Sakana Fugu explained

ExplainX, Fugu benchmarks vs real-world testing

Enjoyed this article? Share it →
Share:
    You Might Also Like
    Latest AI Models April 2026: Rankings & Features
    Benchmarks
    Latest AI Models April 2026: Rankings & Features

    Meta Description GPT-5.4, Gemini 3.1 Ultra, Gemma 4, Muse Spark, GLM-5.1: every major AI model released March-April 2026, compared by benchmark, price, and use case.

    Qwen3.6-27B: 27B Model Beats 397B on Coding (2026)
    Reviews
    Qwen3.6-27B: 27B Model Beats 397B on Coding (2026)

    Qwen3.6-27B scores 77.2% on SWE-bench Verified, beats a 397B MoE, runs on 18GB VRAM, and matches Claude 4.5 Opus on Terminal-Bench. Full review inside.