惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
SegmentFault 最新的问题
月光博客
月光博客
大猫的无限游戏
大猫的无限游戏
GbyAI
GbyAI
博客园 - 叶小钗
小众软件
小众软件
WordPress大学
WordPress大学
I
InfoQ
Last Week in AI
Last Week in AI
Vercel News
Vercel News
博客园 - Franky
Stack Overflow Blog
Stack Overflow Blog
P
Proofpoint News Feed
A
About on SuperTechFans
Engineering at Meta
Engineering at Meta
腾讯CDC
D
DataBreaches.Net
有赞技术团队
有赞技术团队
宝玉的分享
宝玉的分享
Jina AI
Jina AI
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
G
Google Developers Blog
V
Visual Studio Blog
酷 壳 – CoolShell
酷 壳 – CoolShell

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Your Session History Is Bleeding Tokens Every Time You Paste
Shouvik Palit · 2026-06-14 · via DEV Community

Shouvik Palit

Everyone optimizes what they type into Claude.
Nobody optimizes what they paste.

But developers paste constantly. GitHub READMEs. Research papers. API docs. Jira tickets. Confluence pages. Slack threads.

Every time you copy a webpage, your clipboard picks up everything. The navigation. The footer. The boilerplate. The cookie banners. The share buttons rendered as plain text.

You wanted the content. Claude got the garbage too. And you paid for every token of it.

The numbers

Real token counts from live Claude Code sessions:

Content Before After Saved
GitHub README (trooper) 4,600 1,200 74%
Research paper (arXiv) 4,800 1,900 60%
GitHub README (caveman) 1,800 600 67%
API documentation 105 35 67%
tokslayer's own README 800 170 79%
Average 67%

Where the savings actually land

First turn input tokens are paid in full. The saving is on two things:

Output tokens. Claude responds to the compressed version, so answers are shorter and more focused.

Session history. The compressed version stays in context, not the bloated original. Every subsequent turn in that session carries 3,400 fewer tokens of history. Long sessions with multiple pastes, this compounds hard.

Note: write-path interception for true input token saving on turn one is on the roadmap.

What's actually happening when you paste

When you copy a GitHub page you get the content plus:

  • "Skip to content" navigation
  • Repo tabs (Code, Issues, Pull requests, Actions...)
  • Breadcrumbs and branch selectors
  • Footer links
  • Share buttons rendered as text
  • License metadata

None of that is the README. All of it hits Claude's context window. All of it costs tokens.

The fix

A Claude Code skill that sits between your clipboard and Claude. Detects pasted content. Strips the noise. Sends only the signal.

No proxy. No server. No MCP. No configuration. One file. Drop it in. Restart Claude Code. Done.

Receipt on every paste:

ORIGINAL:  "Skip to content shouvik12 trooper Repository..."  (~4,600 tokens)
OPTIMIZED: "Trooper: LLM proxy. Local-first. Ollama default..."  (~1,200 tokens)
SAVED:     ~3,400 tokens (74%)

What gets stripped: nav chrome, footers, filler phrases, redundant sentences, marketing boilerplate.

What stays: headings, code blocks, API signatures, URLs, numbers, technical terms, proper nouns.

Install

curl -fsSL https://raw.githubusercontent.com/shouvik12/tokslayer/main/install.sh | bash

Restart Claude Code. Works on every paste automatically from that point on.

The meta test

Ran tokslayer's own README through itself and typed "summarize":

ORIGINAL:  "tokslayer, Slays tokens before they reach Claude..."  (~800 tokens)
OPTIMIZED: "Tokslayer: Claude Code skill. Compresses pasted content..."  (~170 tokens)
SAVED:     ~630 tokens (79%)

A tool that eats its own cooking.

Where it fits

This covers the input side. Pair it with caveman for output compression (65% reduction on Claude responses) and trooper for routing and fallback when Claude quota runs out.

Together: lean input, lean output, resilient routing.


https://github.com/shouvik12/tokslayer