惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 三生石上(FineUI控件)
Blog — PlanetScale
Blog — PlanetScale
B
Blog
GbyAI
GbyAI
爱范儿
爱范儿
月光博客
月光博客
N
Netflix TechBlog - Medium
T
Tailwind CSS Blog
G
Google Developers Blog
大猫的无限游戏
大猫的无限游戏
Vercel News
Vercel News
H
Hackread – Cybersecurity News, Data Breaches, AI and More
WordPress大学
WordPress大学
The GitHub Blog
The GitHub Blog
Recent Announcements
Recent Announcements
腾讯CDC
MyScale Blog
MyScale Blog
V
Visual Studio Blog
The Cloudflare Blog
Microsoft Security Blog
Microsoft Security Blog
A
About on SuperTechFans
Google DeepMind News
Google DeepMind News
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Large Context Windows Are Not a Solved Problem
Lavkesh Dwivedi · 2026-06-19 · via DEV Community

Originally published on lavkesh.com


I recall the excitement a year ago when models could hold a million tokens in context. That's about 750,000 words or ten average novels sitting in a single prompt. The demos were impressive, and researchers posted benchmarks, but soon teams realized that having a massive context window and knowing what to do with it are two different problems.

I'm not dismissing the capability; a million tokens in context is a real technical achievement. However, I think there's a version of the conversation happening right now that treats window size as the finish line, and that's worth pushing back on.

The pattern I've seen play out is that a team gets access to a long-context model, loads in a large document or codebase, sends a query, and gets back results that are okay, sometimes good, but often frustratingly hard to diagnose. The model technically saw everything in the prompt, but whether it used the right parts is a different question entirely.

Researchers have identified a phenomenon called 'lost in the middle,' where models tend to pay disproportionate attention to content at the beginning and end of a context window, underweighting material in the middle. So if you're feeding in a 200-page document and the critical detail is on page 94, you might not get the answer you're looking for.

This is why retrieval-augmented generation hasn't gone away, even as context windows have grown. Targeted retrieval gives you more control over what the model works with, producing more consistent results. However, RAG introduces its own set of problems, such as maintaining a chunking strategy, embedding model, vector store, and retrieval pipeline.

Long context does handle well a specific class of tasks where the signal isn't concentrated in one place and relationships between parts of the document matter. Examples include reviewing code across an entire repository, analyzing a contract, or summarizing a long research transcript.

The cost dimension is also worth discussing honestly. Long-context inference is expensive, with input tokens adding up fast. A lot of teams have gone through a phase of enthusiasm about long context, done the cost modeling, and quietly landed back on retrieval-based approaches as more economical and predictable.

There's also a latency component; long prompts take longer to process, adding friction for interactive applications. For batch workflows, it matters less, but it's another variable that doesn't appear in benchmark numbers.

The capability is moving forward, with context windows growing and models getting better at using what's in them. There's active research on improving mid-context attention and helping models navigate long inputs more reliably.

Right now, there's a meaningful gap between the spec sheet and what you can depend on in production. Teams doing the most interesting work aren't treating large context windows as a solved input problem; they're being deliberate about what goes in and building evaluation pipelines for long-context failure modes.