惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Privacy & Cybersecurity Law Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
D
Docker
V
V2EX
GbyAI
GbyAI
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - Franky
Jina AI
Jina AI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
I
InfoQ
博客园 - 司徒正美
雷峰网
雷峰网
F
Full Disclosure
S
SegmentFault 最新的问题
大猫的无限游戏
大猫的无限游戏
博客园 - 叶小钗
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
D
DataBreaches.Net
M
MIT News - Artificial intelligence
V
Visual Studio Blog
H
Help Net Security
月光博客
月光博客
博客园 - 三生石上(FineUI控件)
博客园_首页
O
OpenAI News
人人都是产品经理
人人都是产品经理
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Attack and Defense Labs
Attack and Defense Labs
Blog — PlanetScale
Blog — PlanetScale
爱范儿
爱范儿
罗磊的独立博客
P
Palo Alto Networks Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
博客园 - 聂微东
Last Week in AI
Last Week in AI
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
Lohrmann on Cybersecurity
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
The Register - Security
The Register - Security
S
Security @ Cisco Blogs
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
酷 壳 – CoolShell
酷 壳 – CoolShell
AWS News Blog
AWS News Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
LINUX DO - 最新话题
Hacker News - Newest:
Hacker News - Newest: "LLM"
T
Threat Research - Cisco Blogs

Crazyrouter Blog (English)

Ideogram AI Guide 2026: Product Mockups, Text Rendering, and API Automation Akool AI Voice Generator Review 2026: API Alternatives for Developers GLM 4.6 API Guide 2026: Build Chinese-English Agents with Tool Calling Google Veo3 API Guide 2026: Batch Video Generation, QA, and Fallbacks AI Lip Sync Tools Comparison 2026: Developer Guide for Localization Pipelines Claude Opus 4.8 vs Opus 4.7: Real API Benchmark Results for Developers Opus 4.8 vs Opus 4.7 Coding Test: What Changed for Developers? Opus 4.8 vs Opus 4.7 for Agents: JSON, Tool Use, and Structured Output Gemini 2.5 Flash-Lite for RAG, Agent Routing, and Cost per Successful Task Gemini 2.5 Flash-Lite for Support Automation and Ticket Triage Gemini 2.5 Flash-Lite Use Cases: The Practical Automation Tier for Developers Claude Jupiter v1-p vs GPT-5.5 Benchmark: Real API Test on Reasoning and Coding Claude Jupiter v1-p vs Claude Opus 4.7 vs Sonnet 4.6: Live API Test Claude Jupiter v1-p vs Claude Opus 4.7 vs Sonnet 4.6: Live API Test Claude Code Pricing 2026: Pro vs Max vs Team vs API Costs Claude Opus 4.7 vs DeepSeek V4 Pro: Real API Compatibility and Coding Benchmark Gemini CLI Complete Guide 2026: Repo Automation, CI Agents, and Multi-Model Routing Ideogram AI Guide 2026: Brand Design Automation, API Workflows, and Alternatives GLM 4.6 API Guide 2026: Agents, RAG, Tool Calling, and Bilingual Apps WAN 2.2 Animate Tutorial 2026: Character Consistency, Shot Control, and API Workflows Google Veo3 API Guide 2026: Production Video Pipelines, Prompts, Pricing, and Fallbacks AI API Pricing Comparison 2026: Text, Image, Video, Caching, and Router Costs Codex CLI Installation Guide 2026: Windows, macOS, Linux, Proxies, and CI Setup How to Get a Claude API Key in 2026: Secure Setup for Teams, CI, and Alternatives Gemini Advanced Review 2026: Is It Worth It for Coding, Research, and API Teams? Claude Code Pricing Guide 2026: Team Agent Budgets, API Fallbacks, and Cost Control Seedance 2.0 Pricing: Convert 46 CNY per Million Tokens to Cost per Second Qwen2.5-Omni Guide 2026: Real-Time Voice, Vision, and Multimodal Agents Kimi K2 Thinking Guide 2026: Reasoning Workflows, Evals, and Cost Control Google Veo3 API Guide 2026: Batch Video Pipelines, Pricing, and Fallbacks Codex CLI Installation Guide 2026: macOS, Linux, WSL, Proxies, and Dev Containers How to Get a Claude API Key in 2026: Safe Production Setup and Alternatives AI API Pricing Comparison 2026: GPT, Claude, Gemini, Video, and Agent Workloads Gemini Advanced Review 2026: Is It Worth It for Developer Teams? Claude Code Pricing Guide 2026: API Fallbacks, Team Seats, and Budget Control Seedream 4.0 API Tutorial 2026: Batch Image Generation, Product Creative, and Pricing Qwen2.5-Omni Guide 2026: Real-Time Voice, Vision, Text Agents, and API Integration Kimi K2 Thinking Guide 2026: Reasoning Agents, Evaluation Workflows, and API Cost Control WAN 2.2 Animate Tutorial 2026: Character Motion, Shot Control, API Pipelines, and Pricing Google Veo3 API Guide 2026: Production Video Workflows, Prompts, Pricing, and Fallbacks AI API Pricing Comparison 2026: OpenAI, Claude, Gemini, DeepSeek, and Router Costs How to Get a Claude API Key in 2026: Setup, Security, Rotation, and Alternatives Codex CLI Installation Guide 2026: macOS, Linux, WSL, Proxies, and Devcontainers Gemini Advanced Review 2026: Is It Worth It for Developers and API Builders? Claude Code Pricing Guide 2026: CI Agents, Team Seats, and API Budget Planning AI API Gateway for Singapore and Malaysia Developers: One Endpoint for GPT, Claude and Gemini AI API Gateway for Thai Developers: Use GPT, Claude and Gemini with One Key One API Key for GPT, Claude and Gemini: A Practical Setup for Central Asia Developers Gemini 3.5 Flash vs Claude Response-Tier Models: Which One Should Developers Use? Gemini 3.5 Flash vs Gemini 3 Flash vs Gemini 2.5 Flash: Real API Benchmark "How to Test Multiple AI Image Models with One API Key" Codex CLI Installation Guide: Setup on macOS, Linux, Windows WSL and CI/CD Seedream 4.0 API Tutorial: ByteDance Image Generation for Production Pipelines Kimi K2 Thinking Model: Complete Developer Guide for Reasoning Workflows Luma Ray 2 Review: AI Video Generation Quality, Speed, and API Guide Pika 2.2 New Features Review: Scene Director, Sound Design, and API Updates Google Veo 3 API Guide: Video Generation with Audio for Developers AI Lip Sync Tools Comparison 2026: Best APIs for Talking Avatars and Video Dubbing Gemini Advanced Review May 2026: Is It Worth $20/Month for AI Power Users? Claude Code Pricing in May 2026: Max Plan, Opus 4, and Real Cost Breakdown Hermes Agent + Crazyrouter: One-Click Setup for 627+ AI Models Text-Embedding-3-Small: Complete Guide to OpenAI's Most Popular Embedding Model (2026) AI Meme Generator & Coloring Book Creator with GPT-image-2 — Fun Projects That Actually Make Money AI Future Baby Prediction with GPT-image-2 — See What Your Child Might Look Like Ghibli Style Photo Transformation with GPT-image-2 — Turn Any Photo Into Anime Art AI Action Figure Generator with GPT-image-2 — Turn Anyone Into a Boxed Toy AI Face Reading & Personal Color Analysis with GPT-image-2 — Two Viral Use Cases in One Guide AI Palm Reading with GPT-image-2 — Generate Professional Palmistry Analysis from a Single Photo Gemini 2.5 Flash-Lite Pricing Explained — The Cheapest Gemini Model for High-Volume Workloads Claude Sonnet 4.6 Pricing Explained — Caching, Tiers, and How to Save 45% with Crazyrouter Gemini Free vs Gemini Advanced: Pricing, Limits, Features, and Is It Worth Paying For? AI Context Window Comparison (2026): GPT, Claude, Gemini Token Limits by Model Claude Sonnet 4.5 Pricing Explained — Caching, Batch API, and How to Save 45% with Crazyrouter Claude Opus 4.7 Pricing Explained — New Tokenizer, Caching, and How to Save 45% with Crazyrouter Claude Opus 4.6 Pricing Explained — Caching, Tiers, and How to Save 45% with Crazyrouter Best AI Models for RAG Applications 2026: Embeddings, Retrieval, and Generation Seedance 2.0 vs Kling 2.1 vs Runway Gen 4 Turbo: Video AI API Comparison 2026 AI Video Generation API Pricing May 2026: Veo3 vs Kling vs Runway vs Sora How to Get Claude API Key in China 2026: Complete Setup Guide AI Coding Tools ROI Calculator: Claude Code vs Codex CLI vs Gemini CLI Cost Analysis 2026 AI API Pricing Comparison May 2026 - Complete Developer Guide Grok 4 API Pricing Complete Guide 2026 DeepSeek R2: The 32B Reasoning Model That Runs on a Single GPU — Complete Guide for Developers "GPT-5.1 Codex Max Pricing Explained — The Code-Specialized Model and How to Save with Crazyrouter" GLM-5 Pricing Explained — Zhipu AI's Flagship Model and How to Access via Crazyrouter Gemini 3 Flash Pricing Explained — Balanced Speed and Cost with Crazyrouter Savings "Gemini 3.1 Pro Pricing Explained — Context Tiers, Caching, and How to Save with Crazyrouter" GPT-5.5 Pricing Explained — OpenAI's Latest Flagship, Reasoning Tokens, and How to Save with Crazyrouter AI Model Pricing Guide 2026: What Every Model Costs on Crazyrouter (and How Much You Save) MiniMax M2 Pricing Explained — China's Competitive AI Model and How to Access via Crazyrouter Grok 4.1 Thinking Pricing Explained — Reasoning Tokens, Caching, and How to Save with Crazyrouter Grok 4.1 Pricing Explained — 2M Context, Caching, Tool Costs, and How to Save with Crazyrouter GPT-5 Pricing Explained — Reasoning Tokens, Caching, Batch API, and How to Save with Crazyrouter GPT-5-nano Pricing Explained — The Cheapest GPT Model for High-Throughput Workloads GPT-5-mini Pricing Explained — Ultra-Low Cost AI with Caching and Batch Discounts GPT-5.4 Pricing Explained — Cached Input, Context Tiers, Batch API, and How to Save with Crazyrouter GPT-5.2 Pricing Explained — Caching, Batch API, and How to Save with Crazyrouter OpenRouter vs Crazyrouter (2026): Pricing, Models, and Which API Gateway Fits Developers Better Suno v4 vs v5 vs v4.5: Which Version Sounds Better and Is Worth Using in 2026? How to Use Claude Code with Crazyrouter: Base URL Setup, Model Routing, and Cost Savings
GPT-4o Pricing Explained — The Legacy Flagship That's Still Worth Using
Crazyrouter · 2026-04-27 · via Crazyrouter Blog (English)

GPT-4o Pricing Explained — The Legacy Flagship That's Still Worth Using#

GPT-4o launched in May 2024 as OpenAI's flagship multimodal model — fast, capable, and significantly cheaper than GPT-4 Turbo. It dominated the API landscape for the better part of a year. Now, with the GPT-5 family taking center stage, GPT-4o has settled into a different role: the reliable, battle-tested workhorse that millions of developers still depend on every day.

And honestly? For a lot of use cases, it's still the smart choice.

This guide breaks down everything you need to know about GPT-4o API pricing in 2026 — base rates, caching discounts, Batch API savings, and how Crazyrouter can cut your costs even further.

Last updated: April 27, 2026

Base Pricing#

GPT-4o uses a straightforward per-token pricing model. Here's what you're looking at on the standard OpenAI API:

ComponentPrice
Input tokens$2.50 / 1M tokens
Cached input tokens$1.25 / 1M tokens
Output tokens$10.00 / 1M tokens

For context, when GPT-4o first launched, these prices represented a massive drop from GPT-4 Turbo (10/10/30 per MTok). Even now, they remain competitive — especially when you factor in caching and batch discounts.

What does this look like in practice?#

A typical API call with ~1,000 input tokens and ~500 output tokens costs roughly:

  • Input cost: 1,000 tokens × (2.50/1,000,000)=2.50 / 1,000,000) = 0.0025
  • Output cost: 500 tokens × (10.00/1,000,000)=10.00 / 1,000,000) = 0.005
  • Total per call: ~$0.0075

That's less than a cent per request for most conversational interactions. For a chatbot handling 10,000 conversations per day with similar token counts, you're looking at about 75/dayor 75/day or ~2,250/month at list price.

128K Context Window#

GPT-4o supports a 128K token context window — that's roughly 96,000 words or about 300 pages of text in a single prompt. This is the same context length as GPT-4 Turbo, and it remains one of the largest context windows available at this price point.

What can you fit in 128K tokens?

  • An entire novel (most novels are 60K–100K words)
  • A full codebase for a medium-sized project
  • Hundreds of pages of documentation
  • Long conversation histories without truncation

The key advantage: GPT-4o charges the same per-token rate regardless of how much of that 128K window you use. There's no "long context surcharge" like some newer models implement. Whether you send 1K tokens or 120K tokens, the input rate stays at $2.50/MTok.

This makes GPT-4o particularly cost-effective for tasks that require large context — document analysis, code review, long-form summarization — where newer models with tiered pricing might actually cost more.

Automatic Caching (Prompt Caching)#

One of the most impactful cost-saving features for GPT-4o is OpenAI's automatic prompt caching. Introduced in late 2024, this feature requires zero code changes — it just works.

How it works#

When you send a request to the API, OpenAI automatically caches the prefix of your prompt. If a subsequent request shares the same prefix (at least 1,024 tokens), the cached portion is served at a 50% discount:

  • Regular input: $2.50 / MTok
  • Cached input: $1.25 / MTok

Caching is automatic. You don't need to set any flags, manage cache keys, or change your API calls. OpenAI handles it server-side.

When caching kicks in#

Caching is most effective when you have:

  • System prompts that stay consistent across requests
  • Few-shot examples that you include in every call
  • Document context that multiple queries reference
  • Conversation history where earlier messages remain unchanged

Real savings example#

Imagine you're building a customer support bot with a 2,000-token system prompt and 3,000 tokens of product documentation included in every request. For 10,000 daily requests:

Without caching:

  • 5,000 tokens × 10,000 requests = 50M input tokens
  • Cost: 50 × 2.50=2.50 = 125/day

With caching (system prompt + docs cached):

  • 5,000 cached tokens × 10,000 requests = 50M cached tokens
  • Cost: 50 × 1.25=1.25 = 62.50/day
  • Savings: 62.50/day( 62.50/day (~1,875/month)

The cache has a lifetime of 5–10 minutes of inactivity, so it works best for applications with steady traffic. For bursty workloads, you'll see partial caching benefits.

Batch API — 50% Off#

For workloads that don't need real-time responses, OpenAI's Batch API is a game-changer. It offers a flat 50% discount on all token costs:

ComponentStandardBatch API
Input tokens$2.50 / MTok$1.25 / MTok
Output tokens$10.00 / MTok$5.00 / MTok

How Batch API works#

Instead of sending individual requests, you upload a JSONL file containing multiple requests. OpenAI processes them asynchronously and returns results within 24 hours (usually much faster).

Best use cases for Batch API#

  • Content generation — blog posts, product descriptions, translations
  • Data processing — classification, extraction, summarization of large datasets
  • Evaluation pipelines — running test suites against your prompts
  • Embedding generation — processing large document collections
  • Offline analytics — sentiment analysis, categorization

Combining Batch + Caching#

Here's where it gets interesting. Batch API and prompt caching can stack. If your batch requests share common prefixes, you get:

  • 50% off from Batch API
  • Additional 50% off cached input tokens

That means cached input in a batch costs just 0.625/MToka750.625/MTok — a 75% reduction from the standard 2.50 rate.

Crazyrouter Pricing — 45% Off Official Rates#

If you're already optimizing with caching and batching, there's one more lever to pull: routing your API calls through Crazyrouter.

Crazyrouter offers GPT-4o at 55% of OpenAI's official pricing — that's a 45% discount on every token:

ComponentOpenAI OfficialCrazyrouterSavings
Input tokens$2.50 / MTok$1.375 / MTok45% off
Cached input$1.25 / MTok$0.6875 / MTok45% off
Output tokens$10.00 / MTok$5.50 / MTok45% off

How to use Crazyrouter#

Switching is dead simple. You just change the base_url — your existing code works as-is.

Python (OpenAI SDK):

curl:

Node.js:

That's it. Same SDK, same parameters, same response format. Just a different URL and API key.

Why Crazyrouter?#

  • Full OpenAI API compatibility — drop-in replacement, no code changes beyond the base URL
  • All models available — GPT-4o, GPT-5.4, o3, o4-mini, and more
  • Pay-as-you-go — no minimums, no commitments
  • Transparent pricing — what you see is what you pay

Real-World Cost Comparison#

Let's put all the savings together with a realistic scenario. Imagine a SaaS product using GPT-4o for customer support, processing 50,000 requests per day with an average of 2,000 input tokens (1,500 cached) and 800 output tokens per request.

Monthly token volumes:

  • Total input: 50,000 × 2,000 × 30 = 3B tokens (3,000 MTok)
  • Cached input: 50,000 × 1,500 × 30 = 2.25B tokens (2,250 MTok)
  • Non-cached input: 750 MTok
  • Output: 50,000 × 800 × 30 = 1.2B tokens (1,200 MTok)
ScenarioInput CostOutput CostMonthly Total
OpenAI standard (no caching)3,000 × 2.50=2.50 = 7,5001,200 × 10.00=10.00 = 12,000$19,500
OpenAI with caching(750 × 2.50)+(2,250×2.50) + (2,250 × 1.25) = $4,687.501,200 × 10.00=10.00 = 12,000$16,687.50
Crazyrouter with caching(750 × 1.375)+(2,250×1.375) + (2,250 × 0.6875) = $2,578.131,200 × 5.50=5.50 = 6,600$9,178.13
Crazyrouter + Batch API(750 × 0.6875)+(2,250×0.6875) + (2,250 × 0.34375) = $1,289.061,200 × 2.75=2.75 = 3,300$4,589.06

From 19,500downto19,500 down to 4,589 — that's a 76% reduction by stacking caching, Batch API, and Crazyrouter together.

Even without Batch API (which requires async processing), Crazyrouter with caching saves you 53% compared to standard OpenAI pricing.

Should You Upgrade to GPT-5.4?#

GPT-5.4 is OpenAI's current flagship, and it's undeniably more capable than GPT-4o. But "more capable" doesn't always mean "better value." Here's how they compare:

FeatureGPT-4oGPT-5.4
Input price$2.50 / MTok$2.50 / MTok
Output price$10.00 / MTok$10.00 / MTok
Context window128K1M
Max output16,384 tokens64K tokens
ReasoningGoodExcellent
CodingStrongStronger
MultimodalText + VisionText + Vision + Audio
SpeedFastComparable
ReliabilityBattle-testedNewer, still stabilizing
Crazyrouter price (input)$1.375 / MTok$1.375 / MTok
Crazyrouter price (output)$5.50 / MTok$5.50 / MTok

When to stick with GPT-4o#

  • Your prompts work well already. If GPT-4o handles your use case reliably, switching introduces risk with minimal upside.
  • You need predictability. GPT-4o has been in production for nearly two years. Its behavior is well-understood and stable.
  • 128K context is enough. Most applications don't need a 1M context window.
  • You're cost-sensitive on output. At the same price point, GPT-4o's shorter, more concise outputs can actually save money if you don't need GPT-5.4's verbosity.

When to upgrade to GPT-5.4#

  • Complex reasoning tasks where GPT-4o falls short
  • Very long documents that exceed 128K tokens
  • Audio processing requirements
  • Coding tasks where the quality difference matters
  • You need the latest capabilities and can absorb the migration effort

The honest take: if GPT-4o is working for you, there's no rush to migrate. It's not going anywhere, and the price-to-performance ratio is still excellent.

Key Takeaways#

  1. GPT-4o remains a strong value proposition at 2.50/2.50/10 per MTok — especially for applications that don't need cutting-edge reasoning.

  2. Automatic caching is free money. Design your prompts with consistent prefixes and you'll save 50% on cached input tokens with zero code changes.

  3. Batch API halves everything for async workloads. If you can tolerate up to 24-hour turnaround, there's no reason not to use it.

  4. Crazyrouter stacks on top with 45% savings across the board. Combined with caching and batching, you can reduce costs by up to 76%.

  5. Don't upgrade just because something newer exists. GPT-4o is battle-tested, fast, and reliable. Upgrade when your use case demands it, not because of FOMO.

Get Started with Crazyrouter#

Ready to cut your GPT-4o costs by 45%? Getting started takes about 30 seconds:

  1. Sign up at crazyrouter.com
  2. Get your API key from the dashboard
  3. Change your base URL to https://crazyrouter.com/v1
  4. That's it. Your existing code works immediately.

No contracts. No minimums. Pay only for what you use.

Get Your API Key →


Pricing information is accurate as of April 27, 2026. OpenAI may adjust pricing at any time. Crazyrouter pricing is subject to change — check crazyrouter.com for the latest rates. This article is for informational purposes only and does not constitute financial advice. Always verify current pricing on the official provider websites before making purchasing decisions.