惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Secure Thoughts
Apple Machine Learning Research
Apple Machine Learning Research
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 司徒正美
博客园_首页
博客园 - 叶小钗
Blog — PlanetScale
Blog — PlanetScale
The Cloudflare Blog
量子位
人人都是产品经理
人人都是产品经理
小众软件
小众软件
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
IT之家
IT之家
Attack and Defense Labs
Attack and Defense Labs
Hacker News: Ask HN
Hacker News: Ask HN
TaoSecurity Blog
TaoSecurity Blog
Forbes - Security
Forbes - Security
罗磊的独立博客
Webroot Blog
Webroot Blog
美团技术团队
Help Net Security
Help Net Security
Google DeepMind News
Google DeepMind News
博客园 - 三生石上(FineUI控件)
The GitHub Blog
The GitHub Blog
Microsoft Security Blog
Microsoft Security Blog
H
Heimdal Security Blog
C
Check Point Blog
V2EX - 技术
V2EX - 技术
The Last Watchdog
The Last Watchdog
Microsoft Azure Blog
Microsoft Azure Blog
AI
AI
Cloudbric
Cloudbric
Application and Cybersecurity Blog
Application and Cybersecurity Blog
W
WeLiveSecurity
P
Privacy International News Feed
T
The Exploit Database - CXSecurity.com
P
Privacy & Cybersecurity Law Blog
B
Blog RSS Feed
Hacker News - Newest:
Hacker News - Newest: "LLM"
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
L
LINUX DO - 热门话题
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
A
Arctic Wolf
O
OpenAI News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 【当耐特】
C
Cyber Attacks, Cyber Crime and Cyber Security
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Y
Y Combinator Blog
Know Your Adversary
Know Your Adversary

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Stop Burning Through Claude Tokens: Practical Tips Every Developer Should Know
Lakshmi Bhargavi Gajula · 2026-06-14 · via DEV Community

Stop Burning Through Claude Tokens: Practical Tips Every Developer Should Know

If you've integrated Claude into your workflow or your app, you've probably had that moment — the bill arrives (or the rate limit hits), and you think: where did all those tokens go? I've been there. And after spending an embarrassing amount of time optimizing prompts, restructuring conversations, and reading through Anthropic's documentation, I have some hard-won opinions to share.

This isn't a dry technical rundown. These are real lessons from real (sometimes painful) experience. Let's get into it.


🧠 First, Understand What You're Actually Paying For

Before you can reduce token usage, you need to understand the two-sided nature of token billing:

  • Input tokens: Everything you send to Claude — your system prompt, the conversation history, the user message.
  • Output tokens: Everything Claude generates back.

Here's the thing most developers underestimate: your system prompt and conversation history are re-sent on every single API call. That silent, invisible re-transmission is often the biggest culprit behind bloated token usage, especially in multi-turn applications.


✂️ Tip #1: Trim Your System Prompt Ruthlessly

System prompts are powerful, but they're also easy to let balloon out of control. I've seen system prompts that read like legal contracts — paragraphs of edge-case handling, personality descriptions, and disclaimers that Claude probably doesn't need.

Be surgical about it:

  • Remove redundant instructions. If you say "be concise" three times, once is enough.
  • Cut examples unless they're truly necessary for format enforcement.
  • Avoid explaining why Claude should do something unless it genuinely affects the output quality.

A well-written system prompt should be the minimum viable instruction set. Treat it like code — refactor it regularly.

# Before (verbose)
You are a helpful assistant. You should always be polite and professional.
Never be rude to users. Make sure your responses are accurate and helpful.
Always try to answer the user's question. Be concise when possible.

# After (lean)
You are a concise, professional assistant. Answer questions accurately and briefly.

That tiny cleanup can save hundreds of tokens per conversation when multiplied across thousands of calls.


🗂️ Tip #2: Manage Conversation History Strategically

This one trips up a lot of developers building chat-based applications. By default, you're probably passing the entire conversation history to Claude on every turn. For long conversations, this gets expensive fast.

Strategies to handle this:

  1. Sliding window: Only pass the last N messages instead of the full history. Works well for most conversational use cases.
  2. Summarization: Periodically ask Claude to summarize the conversation so far, then replace the raw history with that summary. You keep context without the token overhead.
  3. Relevant retrieval: For knowledge-heavy apps, use vector search to pull only the relevant past messages rather than the whole thread.

There's no one-size-fits-all answer here — it depends on your use case. But ignoring history management entirely is almost always a mistake.


📏 Tip #3: Ask for Shorter Outputs Explicitly

Claude is trained to be thorough and helpful, which is great — until you're paying for 800 tokens when 200 would have done the job. Claude won't automatically be terse unless you tell it to be.

Try being explicit:

  • "Answer in 2-3 sentences."
  • "Respond with only the code, no explanation."
  • "Give me a bullet list, max 5 items."

This feels obvious, but it's surprising how many prompts just ask a question and hope for a brief answer. Claude interprets open-ended questions as an invitation to be comprehensive. Set expectations clearly.


⚠️ Tip #4: Watch Out for the Context Window Trap

Here's something to be careful about: just because Claude has a large context window doesn't mean you should use all of it carelessly. Stuffing the context with huge documents, long code files, or massive conversation histories might work, but it's often wasteful.

Before dumping a 10,000-token document into your prompt, ask yourself:

  • Does Claude actually need all of this information?
  • Can I extract the relevant section instead?
  • Could I use a retrieval strategy instead of brute-force context stuffing?

Big context windows are a capability, not an excuse to skip thoughtful prompt design.


🔄 Tip #5: Batch When You Can

If you're running repetitive tasks — classifying a list of items, summarizing multiple documents, extracting structured data from records — consider batching them into a single prompt rather than making individual API calls.

# Instead of this (3 separate calls)
Classify the sentiment of: "I love this product!"
Classify the sentiment of: "This is terrible."
Classify the sentiment of: "It's okay I guess."

# Do this (1 call)
Classify the sentiment of each sentence below. Reply with a JSON array.
1. "I love this product!"
2. "This is terrible."
3. "It's okay I guess."

Batching reduces per-call overhead from system prompts and saves on round-trip latency too. Win-win.


🚩 Things to Be Careful About

Beyond token usage, there are a few broader habits worth watching:

Don't Over-rely on Claude for Simple Tasks

If you're using Claude to do string formatting, basic math, or simple lookups — you're probably over-engineering it. Use Claude where reasoning, language understanding, or generation actually matters.

Prompt Injection is a Real Risk

If you're building an app that feeds user-provided content directly into Claude's context, be aware of prompt injection attacks. Users can craft inputs designed to override your system prompt or manipulate Claude's behavior. Always sanitize and think carefully about what user content ends up in your prompts.

Don't Assume Consistency Without Testing

Claude is not a deterministic function. The same prompt can produce slightly different results, especially at higher temperature settings. If consistency matters in your app (structured data extraction, for example), test extensively and consider using structured output formats or validation layers.

Caching is Your Friend

If you're using the API and making repeated calls with identical or near-identical system prompts and context, look into prompt caching (available via Anthropic's API). It can dramatically reduce costs for repetitive workloads.


🎯 Final Thoughts

Using Claude effectively isn't just about getting good outputs — it's about being intentional with every token you send and receive. The developers who get the most out of Claude (at the lowest cost) are the ones who treat prompting like engineering: iterative, measured, and always looking for inefficiencies to cut.

Here's my honest take: most token waste is preventable. Bloated system prompts, unmanaged conversation history, and vague output instructions account for the majority of unnecessary spending I've seen. Fix those three things first, and you'll probably see a significant drop in usage before you need to do anything more sophisticated.

Start small, measure your token counts, and iterate. Claude is a genuinely powerful tool — treat it like one.