惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Microsoft Security Blog
Microsoft Security Blog
量子位
大猫的无限游戏
大猫的无限游戏
酷 壳 – CoolShell
酷 壳 – CoolShell
IT之家
IT之家
博客园 - 三生石上(FineUI控件)
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - Franky
美团技术团队
Last Week in AI
Last Week in AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
人人都是产品经理
人人都是产品经理
罗磊的独立博客
Jina AI
Jina AI
小众软件
小众软件
S
SegmentFault 最新的问题
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
雷峰网
雷峰网
博客园 - 聂微东
博客园_首页
The Cloudflare Blog
WordPress大学
WordPress大学
Apple Machine Learning Research
Apple Machine Learning Research
有赞技术团队
有赞技术团队

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
More Context Does Not Mean More Trust
Kair Akhmett · 2026-05-21 · via DEV Community

After publishing my previous post about engineering trust in AI coding, someone asked me a question that cannot be answered honestly in one sentence:

Why would a model start hallucinating more right after increasing context
length and reducing batch size, and then become normal again after 20-30
minutes?

The first thing to say is: without inference-stack logs, we cannot prove the exact cause.

But I would not explain it as "the model needed thirty minutes to learn". During ordinary inference, the model is not training. If the weights did not change, the model itself did not become smarter after half an hour.

The more likely explanation is that the serving mode changed. That is still a hypothesis, not a proven conclusion.

In other words, the problem may not have been only "model quality". It may have been how the production inference system behaved under a new workload.

More Context Is Not a Free Upgrade

It is tempting to think:

Give the model more context and it will make fewer mistakes.

Sometimes that is true. But not always.

If the system is actually sending more relevant input tokens to the model, a longer context window increases the chance that the necessary information reaches the model at all. But it does not guarantee that the model will use the entire context equally well.

There is a well-known effect often called "lost in the middle". In "Lost in the Middle: How Language Models Use Long Contexts", the authors evaluated multi-document QA and key-value retrieval. In those tasks and on the models studied, performance often dropped when the relevant information was placed in the middle of a long context, and was better when it was closer to the beginning or the end of the input.

So "the context got longer" does not automatically mean "the answer became more trustworthy".

Sometimes a long prompt simply gives the important fact a larger place to hide.

Context Length Changes More Than the Prompt

In production inference, a request is not only a semantic object. There is also a runtime profile:

  • prefill;
  • decode;
  • KV cache;
  • dynamic batching;
  • GPU memory pressure;
  • truncation policy;
  • timeout policy;
  • fallback paths;
  • scheduler behavior.

When requests actually start carrying more input tokens, prefill becomes more expensive. The system has to process more prompt tokens before generation begins. Memory pressure and KV-cache pressure increase. The set of requests that can fit into a batch may change.

If batch size is reduced at the same time, the shape of computation changes as well.

With greedy decoding or a fixed seed, the same runtime, and the same backend version, we usually expect reproducibility. But production inference does not always guarantee bitwise-identical behavior even for mathematically equivalent operations. PyTorch documents this class of issue in its Numerical Accuracy notes.

Serving frameworks also treat this as a real concern. At the time of writing, vLLM has a beta batch-invariance mode: under supported conditions, it is meant to make output deterministic and independent of batch size or request order in the batch. The existence of such a mode is a useful reminder that batching is not always just a performance detail.

Why the Effect Might Disappear After 20-30 Minutes

We should not pretend to know the cause without telemetry. But there are several plausible hypotheses.

1. Cache or Runtime Warm-Up

One possible hypothesis: if the stack uses prefix or prompt caching, and the traffic contains repeated prefixes, the first requests may have a worse latency or cache-hit profile. That alone does not prove an increase in hallucination rate. But it can indirectly lead to truncation, timeout, or fallback if there is a wrapper above the model with timeout-based degradation or a similar fallback policy.

A similar idea applies to runtime caches, memory pools, and allocator behavior. After a configuration change, early requests may run in a less stable profile, while later requests see smoother latency. That is a hypothesis to test with metrics, not an explanation to accept by default.

From the outside, this can look like:

The model was confused for the first half hour, then became normal.

But in that version of the story, the model did not stabilize. The system around the model did.

2. Dynamic Batching May Have Reached a More Stable Workload

After a batch-size change, the system may spend some time processing a mixed stream of requests: short, long, old, new, and differently structured contexts.

Dynamic batching builds batches from live traffic. If batch composition changes, latency and memory pressure can change as well. In some stacks, the shape of the batch may also affect bit-level or numeric reproducibility. That does not mean every such difference becomes a semantically different answer.

Again, this does not mean "the model got worse". It may be a transitional
serving-state issue.

3. Truncation, Timeout, or Fallback May Have Fired

In practice, this is a class of causes worth checking before concluding "the model hallucinated".

The model may look as if it invented a fact. But sometimes the simpler root cause is that the fact never reached the prompt.

If we are not talking about a bare model server, but about a RAG, tool-use, or agent pipeline, this can happen at several levels. Especially if there is a wrapper above the model server that shortens evidence, skips a tool step, or uses a fallback path under latency or resource limits.

For example:

  • the retriever did not return part of the evidence;
  • the wrapper truncated the context;
  • a tool returned an incomplete result;
  • part of the system block was shortened;
  • the request went through a fallback path;
  • the response was cut by an output limit.

From the outside, it looks like a semantic failure. Internally, it may be a delivery failure.

How to Test This Instead of Guessing

To prove the cause, you need to log not only model answers, but also the mode in which those answers were produced.

At minimum:

  • input tokens;
  • output tokens;
  • prefill latency;
  • decode latency;
  • time to first token;
  • effective batch size;
  • prefix or prompt cache hit rate;
  • KV-cache pressure or eviction rate;
  • context truncation events;
  • timeout events;
  • fallback events;
  • model version;
  • precision or quantization;
  • temperature, top_p, seed;
  • the first divergent token for the same prompt.

The experiment also needs controls: the same prompt, the same retrieval
snapshot, the same model and backend version, the same sampling parameters, and a fixed seed where the backend supports it.

A simple matrix:

  1. short context + batch size 1;
  2. long context + batch size 1;
  3. long context + batch size N;
  4. long context + dynamic batching;
  5. long context right after a cold restart;
  6. long context after warm-up.

If divergences appear mostly under long context + cold runtime + dynamic
batching, that is an argument for an infrastructure hypothesis worth checking with repeated runs and metrics. It is not proof by itself.

The cause still has to be confirmed with latency, cache, truncation, fallback, and first-divergent-token data.

If divergences remain after warm-up and under batch size 1, then you need to look at the prompt itself, retrieval, placement of evidence, and the model.

Where Undes Fits

This leads to a broader engineering point: a final answer is not enough. We need to understand how the answer was produced.

Undes does not control the provider's inference stack directly. It cannot look inside another provider's KV cache or scheduler. But it can help at a different layer:

  • show which files and snippets were read, sent into context, and recorded as evidence;
  • record which claims are supported by evidence;
  • surface places where the model referenced an unconfirmed method or file;
  • separate grounded findings from assumed implementation;
  • preserve rejected hypotheses;
  • keep open checks visible instead of hiding them inside confident prose;
  • mark an answer as not patch-safe when the verification is incomplete.

This is not magical hallucination suppression. A more precise statement is:

Undes makes unchecked hallucinations harder to hide.

If a model invents a method, a properly instrumented run can raise that as a warning. If a dependency was not read, it should become an open check. If an important objection disappeared between phases, it can become a transition signal.

Some unsupported claims stop being just polished sentences in a final answer. They become visible as warnings, open checks, or diagnostic status that an engineer can inspect and close.

Why "More Context" Is Not the Same as "More Trust"

You can give a model more text. But trust does not come from the amount of text.

Trust comes from a verifiable link between:

  • the request;
  • the context that was read;
  • evidence;
  • claims;
  • critique;
  • unresolved risks;
  • the final answer.

Long context helps only when the system can show what from that context was used and what remained unverified.

Otherwise, a longer prompt may simply become a larger place for mistakes to hide.

Closing

AI coding is not only about model quality. It is about the quality of the entire engineering loop around the model: context delivery, retrieval, batching, evidence, checks, and trust status.

The next layer of AI tooling is not just "give the model more context".

The next layer is reducing the chance that a model can present unverified output as verified output without a warning or diagnostic status.

AI agents can generate code now. Undes helps you decide whether to trust it.

Further Reading


I am building Undes, a local-first AI engineering CLI that generates and verifies engineering answers. If your team is already using AI coding agents, I would like feedback on which trust signals you still need before acting on generated code in a real workflow.