惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Hacker News: Ask HN
Hacker News: Ask HN
D
DataBreaches.Net
Microsoft Security Blog
Microsoft Security Blog
U
Unit 42
V
Visual Studio Blog
GbyAI
GbyAI
云风的 BLOG
云风的 BLOG
博客园 - Franky
C
CXSECURITY Database RSS Feed - CXSecurity.com
大猫的无限游戏
大猫的无限游戏
P
Privacy & Cybersecurity Law Blog
T
The Exploit Database - CXSecurity.com
Simon Willison's Weblog
Simon Willison's Weblog
L
LangChain Blog
I
Intezer
V2EX - 技术
V2EX - 技术
Google DeepMind News
Google DeepMind News
T
Threat Research - Cisco Blogs
Apple Machine Learning Research
Apple Machine Learning Research
V
V2EX
腾讯CDC
博客园 - 【当耐特】
Know Your Adversary
Know Your Adversary
TaoSecurity Blog
TaoSecurity Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Fortinet All Blogs
Project Zero
Project Zero
Blog — PlanetScale
Blog — PlanetScale
S
Security @ Cisco Blogs
量子位
M
MIT News - Artificial intelligence
美团技术团队
C
Cisco Blogs
S
Schneier on Security
Recent Commits to openclaw:main
Recent Commits to openclaw:main
G
Google Developers Blog
N
News and Events Feed by Topic
MongoDB | Blog
MongoDB | Blog
The Hacker News
The Hacker News
H
Help Net Security
S
Secure Thoughts
Scott Helme
Scott Helme
SecWiki News
SecWiki News
T
Troy Hunt's Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
博客园 - 叶小钗
O
OpenAI News
Application and Cybersecurity Blog
Application and Cybersecurity Blog
博客园 - 司徒正美
T
Tenable Blog

Analytics Vidhya

Handling Imbalanced Classification: What Works Better Than SMOTE GPT-5.6 Is Here: Sol, Terra, and Luna Loop Engineering for AI Agents: How /loop is Changing AI Workflows DeepSeek DSpark: The Speculative Decoding Trick Behind 400% Faster LLM OKF: Redefining Knowledge Bases for AI Agents Modern VLMs Explained: How GPT-4o, Gemini, Claude Vision, and Qwen-VL Work YOLO26 Tutorial: Object Detection, Pose Estimation & More Large Action Models (LAMs) vs Agentic LLMs: What's the Real Difference? Claude Sonnet 5: The Fable 5 at Home The Best $20 AI Plan: ChatGPT Plus vs Claude Pro vs Gemini Pro GraphRAG vs Vector RAG: Which Retrieval Method is Best? Using AI When You Don’t Trust AI The Self-Improving Loop in AI Agents: Architecture, Benefits, and How it Outperforms Traditional Agent Workflows Harness-1: The 20B Retrieval Subagent That Beats GPT-5.4 at Search Claude's Hidden Art Skill: Making Illustrations With Code System Design for ML Interviews: 10 Real Problems Walked Through Most People Use ChatGPT Wrong: 10 Features and Tips That Changed How I Work OpenAI Just Launched 3 Free AI Courses with Certificates Autoregressive Models: Predicting the Future Using the Past Gemini Omni: AI Video Generation Inside Gemini DiffusionGemma: Google’s Diffusion-Based Open Model for Faster Text Generation Top 10 AI Engineering Tools Everyone is Using in 2026 I Tested Claude Fable 5: Can Anthropic’s Newest AI Deliver on the Hype? Prophet vs NeuralProphet vs TimeGPT vs Chronos: A Practical Comparison Build an Emergency Helpline Voice Agent with LangChain Choosing the Right Vector Database for RAG and AI Applications Google Gemma 4 12B: Architecture, Benchmarks, Access, and Hands-on Guide for Developers How to Choose the Right AI Model for Your Needs Agent Observability with LangSmith, Langfuse, and Arize: A Hands-On Comparison How to Use Claude Managed Agents? Google AI Studio vs Gemini App: What’s the Difference? AI Workflows for Sales Teams: Prospect Research, Lead Qualification, and CRM Updates on Autopilot Using LangGraph 25 Most Influential AI Pioneers to Meet at DataHack Summit 2026 Claude Opus 4.8: A Smarter Model in the Right Direction PySpark Optimization: 12 Proven Techniques to Speed Up Your Spark Jobs 10 Everyday Tasks You Can Automate with AI Today (With n8n Templates) Google Antigravity 2.0: The Full Developer Guide (I/O 2026) Build a Claude Cowork-Like Browser Agent Using Playwright MCP and Claude Desktop Pandas vs Polars vs DuckDB: Which Library Should You Choose? Qwen3.7-Max: Alibaba’s New Agent-First LLM for Coding, Reasoning, and Long-Horizon AI Workflows The Biggest Announcements from Google I/O 2026 Top 9 AI Events and Conferences in 2026 that you Must Attend Gemini 3.5 Flash: Frontier Intelligence with Speed Kimi WebBridge: Hands-on Guide to Kimi’s Browser Extension for AI Agents 40 Advanced SQL Window Functions Every Data Scientist Must Know(with examples) Top 10 AI Research Papers of 2025 6 Steps to Crack GenAI Case Study Interviews (With Real Examples) OpenAI Omni Moderation: How to Filter Text & Images for Free DataHack Summit 2026: You Just Cannot Skip This AI Event of the Year OpenAI’s New API Voice Models Will Change the Way You Use AI Hermes Agent Guide: What is it and How to Use it? Top 10 LLM Research Papers of 2026 Agent Memory Patterns in Cognitive Science and AI Systems 10 AI Agents Every AI Engineer Must Build (with GitHub Samples) 23 Tips for Smart Claude Code Token Saving and Workflow Optimization Feature Engineering with LLMs: Techniques & Python Examples ChatGPT is Now Inside Excel and Google Sheets: Here is How to Use it Gemini API File Search: The Easy Way to Build RAG Top 10 Open-Source Libraries to Fine-Tune LLMs Locally ML Intern in Practice: From Prompt to a Shipped Hugging Face Model 15+ Solved Agentic AI Projects with Github Links How People are Figuring Out Life With Claude MemPalace Explained: Building Long-Term Memory for AI Agents Beyond RAG Grok Voice Think Fast 1.0: Build Voice AI Agents That Actually Think Compressing LSTM Models for Retail Edge Deployment: A Practical Comparison MCP vs Agent Skills: Different Altogether GPT 5.5 vs Opus 4.7: Which is the Best AI Model Today? What is Agentic AI? Claude Code vs Codex: A Detailed Terminal Agent Comparison Google Deep Research Max: Build Autonomous AI Research Agents in Minutes Meta Muse Spark Review: Is It Worth the Hype? ChatGPT Images 2.0 vs Nano Banana 2: Which is Better? Cursor V3 Explained: The AI Coding Agent That’s Replacing Traditional IDEs in 2026 DeepSeek-V4: The Most Powerful Open-Source Model Ever Is GPT Image 2 the Best Image Generation Model? Token Economics: Why AI is Getting “Cheaper” From Idea to Output: Claude Does the Design Work Opus 4.7 vs Opus 4.6: Should You Switch? Build Human-Like AI Voice App with Gemini 3.1 Flash TTS How to Structure a Claude Code Project that Thinks Like an Engineer Gemma 4 Tool Calling Explained: Build AI Agents with Function Calling (Step-by-Step Guide) Anthropic Launches Claude Opus 4.7 For “Most Difficult Tasks” Top 28 Claude Shortcuts that will 10X your Speed GPT-5.4-Cyber: Why OpenAI is Keeping its Most Powerful Model Under Lock and Key Google AI Studio Guide: Every Feature Explained Mastering Deep Agents: Context Engineering that Actually Works 21 Computer Vision Projects from Beginner to Advanced (2026 Guide) Excel 101: Excel Agent Mode Explained MiniMax M2.7 Goes Open-Weight to Let You Run Agents Locally Top 10 Gemma 4 Projects That Will Blow Your Mind GLM-5.1: Architecture, Benchmarks, Capabilities & How to Use It Understanding BERTopic: From Raw Text to Interpretable Topics From Karpathy’s LLM Wiki to Graphify: AI Memory Layers are Here 10 Most Important AI Concepts Explained Simply Project Glasswing is World’s Most Powerful AI in Action How to Run Gemma 4 on Your Phone Without Internet: A Hands-On Guide Running Claude Code for Free with Gemma 4 and Ollama LLM Wiki Revolution: How Andrej Karpathy’s Idea is Changing AI Rethinking Enterprise Search: How Cortex Search Turns Data into Business Impact Google’s Gemma 4: Is it the Best Open-Source Model of 2026?
Sakana Fugu: Multi-Agent System as a Model
Harsh Mishra · 2026-06-23 · via Analytics Vidhya

For years, AI progress has centered on scaling individual foundation models: larger parameters, longer context windows, stronger reasoning, and better tool use. Sakana AI’s Fugu points elsewhere, behaving like one model from the outside while coordinating multiple expert agents internally.

A single API call can trigger direct answering, specialist delegation, intermediate verification, and final synthesis, hiding orchestration complexity behind a normal LLM interface. In this article, a practical guide to Fugu’s architecture, variants, pricing, benchmarks, access, code, tests, enterprise fit, trade-offs, and use cases.

Table of contents

  • What is Sakana Fugu? 
    • Why the naming matters 
  • Why Multi-Agent System as a Model Matters 
  • Sakana Fugu Release Overview 
  • Fugu vs Fugu Ultra 
    • Fugu 
    • Fugu Ultra 
    • Comparison table 
  • Architecture: How Fugu Works Internally 
  • Core architecture components 
  • Pricing  
  • Benchmark Results 
  • Technical Hands-on: Using Sakana Fugu API 
  • Conclusion 

What is Sakana Fugu? 

Sakana Fugu is an OpenAI-compatible managed model API that looks like a single LLM but works as a multi-agent system internally. Developers send a prompt to one model ID, such as fugu or fugu-ultra, while Fugu handles agent selection, role assignment, coordination, verification, and final response.

Instead of manually building planner, coder, reviewer, researcher, or supervisor agents with frameworks like LangGraph, AutoGen, or CrewAI, teams get orchestration packaged into the model itself. This reduces the need to manage prompts, routing, retries, memory, state, monitoring, and failure recovery.

Why the naming matters 

The name “Sakana” means fish in Japanese. The company often frames its research around collective intelligence, similar to how a school of fish can behave as one coordinated system. Fugu follows that idea. Many agents coordinate behind one interface. 

Why Multi-Agent System as a Model Matters 

Most production AI systems today fall into one of three patterns: 

  1. Single-model prompting 
  2. Tool-augmented LLM applications 
  3. Manually designed multi-agent workflows 

Single-model prompting is simple, but it can fail on complex tasks that require planning, execution, verification, and iteration. 

Tool-augmented LLMs improve usefulness by connecting models to search, databases, code execution, APIs, or business systems. But the model still usually acts as the central reasoning engine. 

Multi-agent workflows go further. They divide work across specialized agents. For example: 

  • A planner breaks down the task. 
  • A researcher gathers context. 
  • A coder writes code. 
  • A reviewer checks for correctness. 
  • A verifier tests the answer. 
  • A supervisor coordinates the process. 

This can improve reliability on difficult tasks, but building it well is hard. Teams must answer many system design questions: 

  • Which agent should handle which task? 
  • How should agents communicate? 
  • When should the system stop? 
  • How should intermediate outputs be verified? 
  • How should cost and latency be controlled? 
  • How should failures be recovered? 
  • How should compliance restrictions be applied? 

Fugu attempts to make this easier by turning multi-agent orchestration into a model-level capability. The developer does not need to design every agent interaction manually. 

Sakana Fugu Release Overview 

Sakana Fugu was introduced as Sakana AI’s commercial multi-agent orchestration product. The initial beta positioned it as a system that coordinates pools of frontier foundation models for coding, mathematics, scientific reasoning, research, and complex analysis. 

The latest Fugu release makes the product easier to access through Sakana’s console and an OpenAI-compatible API. The core release message is simple: developers can plug multi-agent intelligence into existing workflows without rewriting their application around a new SDK or orchestration framework. 

Fugu vs Fugu Ultra 

Sakana Fugu comes in two main model options: Fugu and Fugu Ultra. 

Fugu 

Fugu is the default model for everyday work. It balances performance and latency. It is suitable for coding support, code review, chatbots, internal assistants, document analysis, and interactive workflows where response time matters. 

A key point is that Fugu can route to the best model based on the task. It also allows users to opt specific agents out of the model pool, which can help with data, privacy, compliance, or organizational requirements. 

Fugu Ultra 

Fugu Ultra is optimized for maximum answer quality. It coordinates a deeper pool of expert agents and is intended for hard, high-stakes, multi-step problems. According to the Sakana, Fugu Ultra can route between one to three agents depending on the problem. 

Fugu Ultra is better suited for workloads where accuracy, depth, and persistence matter more than latency. Examples include: 

  • Paper reproduction 
  • Kaggle-style data science workflows 
  • Cybersecurity analysis 
  • Literature review 
  • Patent investigation 
  • Deep technical research 
  • Complex code review 
  • Scientific reasoning 

Comparison table 

Feature  Fugu  Fugu Ultra 
Best for  Everyday coding, chat, review, interactive workflows  Hard reasoning, research, high-stakes analysis 
Design goal  Balance quality and latency  Maximize quality 
Agent pool  Flexible, with opt-out support  Fixed full pool 
Latency  Lower  Higher 
Cost  Depends on active underlying agent tier  Fixed token pricing 
Recommended users  Developers, product teams, internal tools  Researchers, advanced developers, enterprise analysis teams 
Main trade-off  Less depth than Ultra  Higher cost and response time 

Architecture: How Fugu Works Internally 

Fugu’s architecture can be understood as a managed orchestration layer wrapped inside a model API. 

From the outside, the flow looks like this: 

flowchart

Internally, the system is closer to this: 

Internal orchestrator model

Sakana Fugu exposes a single API while internally coordinating a pool of specialized models. The user sends one request, and Fugu handles routing, delegation, verification, and synthesis.  

Core architecture components 

1. API gateway 

The developer interacts with a standard API surface. This matters because Fugu supports OpenAI-compatible endpoints, so teams can reuse existing OpenAI SDK clients with a different base URL and API key. 

2. Orchestrator model 

The orchestrator is the core intelligence layer. It decides how the task should be handled. For simpler tasks, it may answer with minimal orchestration. For complex tasks, it can coordinate multiple expert agents. 

3. Agent pool 

Fugu has access to a pool of underlying models or agents. These agents may have different strengths across coding, reasoning, research, long-context analysis, or other specialized tasks. 

4. Dynamic routing 

Instead of hardcoding a workflow, Fugu dynamically selects which agent or agents to use. This is important because model strengths are often task-specific. One model may perform better at code generation, another at mathematical reasoning, another at long-context synthesis. 

5. Delegation and communication 

The orchestrator can break down a complex task into subtasks. It can send focused instructions to different agents and control what context each agent receives. 

6. Verification 

For difficult tasks, the system can use verification-style behavior. One agent may solve, another may critique or validate, and the orchestrator may combine the results. 

7. Synthesis 

The final answer is returned as a single response. The user does not see the full internal agent graph. . 

Pricing  

Fugu has two pricing modes: pay-as-you-go and subscription plans. 

Pay-as-you-go 

Pay-as-you-go is designed for heavier production workloads. Sakana says consumption-based tokens are served at higher priority than monthly-plan tokens. 

Fugu pricing 

Fugu pricing depends on the active agent setup. 

Active agents  Billing rule 
1 agent  Pay the standard rate for the specific underlying model 
Multiple agents  Fees are not stacked. You are charged one rate based on the top-tier model involved 

This is important because many multi-agent systems become expensive when each model call is billed separately. Fugu’s pricing model tries to avoid stacking model fees across agents. 

Fugu Ultra pricing 

Fugu Ultra has fixed pricing for fugu-ultra-20260615 per 1M tokens. 

Token type  Standard price  Context greater than 272K 
Input  $5 per 1M tokens  $10 per 1M tokens 
Output  $30 per 1M tokens  $45 per 1M tokens 
Cached input  $0.50 per 1M tokens  $1.00 per 1M tokens 

Subscription plans 

Subscription plans are designed for individuals and everyday hands-on use. Every tier includes both Fugu and Fugu Ultra. 

Plan  Price  Best for  Usage 
Standard  $20/month  Lightweight daily usage, occasional API calls, small experiments  Baseline allowance 
Pro  $100/month  Regular coding, review, research, and analysis sessions  10x Standard usage 
Max  $200/month  Heavy long-running workloads  20x Standard usage 

Benchmark Results 

Sakana reports Fugu and Fugu Ultra benchmark scores across coding, reasoning, science, agentic tasks, long-context reasoning, and cybersecurity-style evaluation. 

Sakana Fugu and Fugu Ultra compared with frontier baseline models across coding, reasoning, science, long-context, and agentic benchmarks.  

Benchmarks are useful, but they should not be treated as direct production guarantees. Fugu’s benchmark profile suggests three practical insights. 

1. Fugu is strongest when tasks require orchestration 

The strongest use case is not a simple one-shot answer. The model is designed for tasks that benefit from decomposition, expert selection, verification, and synthesis. 

Examples: 

  • Debug this repository. 
  • Review this pull request. 
  • Reproduce this research paper. 
  • Investigate this patent landscape. 
  • Analyze a possible security vulnerability. 
  • Compare multiple technical approaches and recommend one. 

2. Ultra is not always automatically better 

Fugu Ultra is optimized for answer quality, but Fugu can outperform it on some benchmarks. Developers should benchmark both models on their own workload before standardizing. 

A practical routing strategy could be: 

Use fugu for interactive work.
Use fugu-ultra for complex, high-value tasks.
Fallback to fugu when latency or cost matters.  

3. Multi-agent performance comes with hidden complexity 

Even though Fugu hides orchestration complexity from the developer, the underlying system still performs additional work. This can affect latency, cost, and observability. 

Teams should monitor: 

  • Total tokens 
  • Orchestration tokens 
  • Latency by task type 
  • Quality by workload category 
  • Failure cases 
  • Model version behavior 
  • Cost per successful outcome 

Technical Hands-on: Using Sakana Fugu API 

Sakana fugu documentation: https://console.sakana.ai/get-started

1: Create an API key 

Go to the Sakana console API key page login and create API: https://console.sakana.ai/api-keys

Create an API key and store it securely. The key is shown only once. 

2: Set environment variables 

export FUGU_API_KEY="your_api_key_here"
export FUGU_BASE_URL="https://api.sakana.ai/v1"  

3: Install the OpenAI Python SDK 

pip install openai  

4: Basic Responses API call 

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["FUGU_API_KEY"],
    base_url=os.environ.get("FUGU_BASE_URL", "https://api.sakana.ai/v1"),
)

response = client.responses.create(
    model="fugu",
    input="Explain Sakana Fugu in simple terms for a software engineer.",
)

print(response.output_text)

Step 5: Use Fugu Ultra for harder reasoning 

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["FUGU_API_KEY"],
    base_url=os.environ.get("FUGU_BASE_URL", "https://api.sakana.ai/v1"),
)

response = client.responses.create(
    model="fugu-ultra",
    instructions="You are a senior AI architect. Be precise and technical.",
    input="""
Compare single-agent LLM systems, manually designed multi-agent workflows,
and Sakana Fugu-style multi-agent systems as a model.
Focus on architecture, cost, latency, observability, and governance.
""",
)

print(response.output_text)

Conclusion 

Sakana Fugu stands out because it shifts the abstraction layer. Instead of offering just another large model, it packages multi-agent orchestration behind a model API.

For developers, this means easier access to agentic workflows without building complex orchestration systems from scratch. For technical leaders, it offers a managed way to improve reasoning, coding, research, and analysis while reducing dependence on a single model provider.

Fugu is best suited for complex, ambiguous, high-value tasks rather than simple chatbot prompts. Still, teams should adopt it carefully, given its limited routing transparency, possible latency, unclear token accounting, and regional constraints.

The simplest way to think about Fugu is this: it is not just a model you prompt. It is a model that manages other models. That makes it an important step toward the next generation of AI applications.

Frequently Asked Questions

Q1. Is Sakana Fugu a single model or a multi-agent system? 

A. It is exposed as a single model API, but internally it behaves as a multi-agent orchestration system. 

Q2. What model IDs should I use? 

A. Use fugu for standard work and fugu-ultra for complex, high-value tasks. Use fugu-ultra-20260615 if you want to pin a specific Ultra version. 

Q3. Is Fugu OpenAI-compatible?

A. Yes. It supports OpenAI-compatible Responses, Chat Completions, and Models APIs. 

Harsh Mishra is an AI/ML Engineer who spends more time talking to Large Language Models than actual humans. Passionate about GenAI, NLP, and making machines smarter (so they don’t replace him just yet). When not optimizing models, he’s probably optimizing his coffee intake. 🚀☕