惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
AI
AI
L
LINUX DO - 最新话题
The Register - Security
The Register - Security
T
Threatpost
Y
Y Combinator Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Attack and Defense Labs
Attack and Defense Labs
T
Tailwind CSS Blog
P
Proofpoint News Feed
MongoDB | Blog
MongoDB | Blog
H
Heimdal Security Blog
小众软件
小众软件
D
Docker
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
T
The Exploit Database - CXSecurity.com
AWS News Blog
AWS News Blog
腾讯CDC
博客园 - 司徒正美
美团技术团队
L
LINUX DO - 热门话题
N
Netflix TechBlog - Medium
Stack Overflow Blog
Stack Overflow Blog
S
Security Affairs
阮一峰的网络日志
阮一峰的网络日志
爱范儿
爱范儿
N
News and Events Feed by Topic
J
Java Code Geeks
F
Fortinet All Blogs
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
U
Unit 42
V2EX - 技术
V2EX - 技术
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Tor Project blog
H
Help Net Security
The GitHub Blog
The GitHub Blog
L
Lohrmann on Cybersecurity
Hugging Face - Blog
Hugging Face - Blog
S
Securelist
PCI Perspectives
PCI Perspectives
W
WeLiveSecurity
A
About on SuperTechFans
N
News and Events Feed by Topic
博客园 - 叶小钗
Cloudbric
Cloudbric
L
LangChain Blog
WordPress大学
WordPress大学
B
Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
4 Types of Hallucinations: One Detection Pattern Per Type
Gabriel Anha · 2026-05-05 · via DEV Community

A customer pasted three sentences from the assistant into a ticket. The first cites a paper from 2024 that does not exist. The second sentence contradicts the third. None of them appear anywhere in the document the user actually uploaded.

If you are running a single hallucination check against that paragraph, you will catch one of those three problems and miss the other two. They are not the same defect. They come from different failure modes and need different detectors. Treating "hallucination" as one bucket is why your eval suite passes while support escalates.

The Ji et al. survey on hallucination in natural language generation splits the problem into intrinsic and extrinsic. SelfCheckGPT (Manakul et al., 2023) showed that sampled responses diverge for hallucinated facts and converge for grounded ones, and TruthfulQA (Lin, Hilton, Evans, 2021) isolates the factual-mimicry failure on its own. Each maps to a different shape of error in production.

Four shapes, one harness — Python below.

Type 1: factual fabrication

The model invents an entity: a date, a citation, a statute, a function signature. The string is well-formed and the syntax is correct, but the referent does not exist.

You catch this by grounding extracted entities against an authoritative source. The detector just needs to be a lookup against a source you trust.

# detectors/factual.py
import re
from typing import Callable

CITATION_RE = re.compile(
    r"\b([A-Z][a-z]+(?:\s+et\s+al\.)?)[,\s]+\(?(\d{4})\)?"
)

def factual_fabrication(
    text: str,
    lookup: Callable[[str, int], bool],
) -> list[dict]:
    """Return list of unverified citations.
    `lookup(author, year) -> True if found in trusted source.`"""
    findings = []
    for match in CITATION_RE.finditer(text):
        author, year = match.group(1), int(match.group(2))
        if not lookup(author, year):
            findings.append({
                "kind": "factual_fabrication",
                "span": match.group(0),
                "reason": f"no record of {author} ({year})",
            })
    return findings

Enter fullscreen mode Exit fullscreen mode

The regex is illustrative; for entities that matter (drug names, ICD codes, internal user IDs, ticker symbols), swap it for a domain extractor and a real index. The point is the contract: extract candidates, look them up in something you trust, fail loudly when the lookup misses.

This is the type the customer noticed first. It is also the easiest to catch.

Type 2: intrinsic contradiction

The output disagrees with itself. Sentence three negates sentence one. The first paragraph says the patient is allergic to penicillin; the third paragraph recommends amoxicillin.

This one cannot be caught by lookup. It can be caught by sampling, which is the move SelfCheckGPT formalized. If a claim is grounded, repeated samples agree. Invented claims drift apart across samples. You generate k responses to the same prompt and measure pairwise agreement.

# detectors/contradiction.py
from itertools import combinations

def intrinsic_contradiction(
    samples: list[str],
    nli: Callable[[str, str], str],  # entail|neutral|contradict
) -> dict:
    """Pairwise NLI across k samples. Flag if any pair contradicts."""
    contradictions = 0
    pairs = list(combinations(range(len(samples)), 2))
    for i, j in pairs:
        if nli(samples[i], samples[j]) == "contradict":
            contradictions += 1
    rate = contradictions / max(1, len(pairs))
    return {
        "kind": "intrinsic_contradiction",
        "contradiction_rate": rate,
        "flagged": rate > 0.0,
    }

Enter fullscreen mode Exit fullscreen mode

The nli callable is a natural-language-inference classifier. A small fine-tuned model works (a small NLI model such as DeBERTa-MNLI is cheap enough to run on CPU). Calling a second LLM with a strict yes/no prompt also works. It needs no model hosting, which is why most teams ship it first.

The aggregate signal is the contradiction rate across pairs. Treat one contradicting pair out of ten as a pointer for a deeper look. Don't make it a kill switch.

Type 3: prompt-vs-output divergence

The output ignores what the user gave you. The user uploaded a contract dated 2018 and asked for a summary; the assistant summarized a different contract. Or the user pastes error logs and the response answers a question they didn't ask.

This is the type that is easiest to test for and the most often skipped. The test is: does the output stay faithful to the input? Run NLI in one direction: does the input entail the output's claims about the input? Flag anything that drifts.

# detectors/divergence.py
def prompt_output_divergence(
    user_input: str,
    output: str,
    nli: Callable[[str, str], str],
    splitter: Callable[[str], list[str]],
) -> list[dict]:
    """For each output sentence that asserts something about
    the input, check input entails it."""
    findings = []
    for sentence in splitter(output):
        verdict = nli(user_input, sentence)
        if verdict == "contradict":
            findings.append({
                "kind": "prompt_output_divergence",
                "span": sentence,
                "reason": "input contradicts output sentence",
            })
    return findings

Enter fullscreen mode Exit fullscreen mode

Two practical notes on the splitter and on neutral verdicts. The splitter matters: a regex on . is fine for prose, terrible for code blocks and lists. Use nltk.sent_tokenize or equivalent. And "neutral" is not a fail: an output sentence that adds a generic disclaimer is not a divergence, just an addition. Only contradict is a hard signal.

Type 4: tool-call hallucination (the short version)

Strict schemas guarantee the shape of a tool call. They guarantee nothing about the values. A model can emit delete_user(user_id="usr_4f9...") against a user who never existed and the schema will be happy.

The detection pattern is a runtime existence check before the side effect runs.

# detectors/tool_call.py
def tool_call_hallucination(
    tool_name: str,
    args: dict,
    resolver: Callable[[str, dict], dict | None],
) -> dict | None:
    """`resolver` looks up the referenced entity in your DB.
    Returns None if the referenced entity exists; finding if not."""
    resolved = resolver(tool_name, args)
    if resolved is None:
        return {
            "kind": "tool_call_hallucination",
            "tool": tool_name,
            "args": args,
            "reason": "referenced entity does not exist",
        }
    return None

Enter fullscreen mode Exit fullscreen mode

That is the whole pattern. Schema-validate the call, then validate the values against state, then run the side effect.

The harness, end to end

The file below combines all four detectors and expects four pluggable callables (citation_lookup, nli, splitter, tool_resolver) so you bring your own backends.

# halluharness.py
from dataclasses import dataclass, field
from typing import Callable, Optional

from detectors.factual import factual_fabrication
from detectors.contradiction import intrinsic_contradiction
from detectors.divergence import prompt_output_divergence
from detectors.tool_call import tool_call_hallucination

@dataclass
class Run:
    user_input: str
    output: str
    samples: list[str]
    tool_calls: list[dict] = field(default_factory=list)

@dataclass
class Verdict:
    findings: list[dict]
    score: float        # 1.0 clean, 0.0 fully hallucinated
    blocked: bool

def check(
    run: Run,
    citation_lookup: Callable[[str, int], bool],
    nli: Callable[[str, str], str],
    splitter: Callable[[str], list[str]],
    tool_resolver: Callable[[str, dict], Optional[dict]],
    block_on: tuple[str, ...] = (
        "factual_fabrication",
        "tool_call_hallucination",
    ),
) -> Verdict:
    findings: list[dict] = []
    findings.extend(
        factual_fabrication(run.output, citation_lookup)
    )
    contra = intrinsic_contradiction(run.samples, nli)
    if contra["flagged"]:
        findings.append(contra)
    findings.extend(
        prompt_output_divergence(
            run.user_input, run.output, nli, splitter
        )
    )
    for call in run.tool_calls:
        f = tool_call_hallucination(
            call["name"], call["args"], tool_resolver
        )
        if f is not None:
            findings.append(f)
    n = len(findings)
    score = 1.0 if not n else max(0.0, 1.0 - 0.25 * n)
    blocked = any(f["kind"] in block_on for f in findings)
    return Verdict(findings=findings, score=score, blocked=blocked)

Enter fullscreen mode Exit fullscreen mode

Compared to a single end-to-end "is this a hallucination" prompt, the harness wins on three counts that matter in production:

Each detector stays separate. When a check fires, you know which one and why; the finding has a kind and a reason. The on-call gets a span and a reason instead of a vibe.

Samples are a first-class input. Self-consistency is not optional for the contradiction check. If your eval pipeline only generates one response per case, the contradiction detector cannot see anything; you have to wire k > 1 sampling at the generator step. The same idea I wrote about for stochastic judges applies on the generator side: one sample is one observation, not the truth.

Block and report are different defaults. Factual fabrication and tool-call hallucination both have hard ground truth (the lookup either matches or it does not), so they default to blocking. Contradiction and divergence depend on a probabilistic NLI verdict; log them and route for human review. You will tune those defaults per surface; a customer-facing chat answer and a backend agent action have different risk budgets. The 0.25-per-finding penalty in score is illustrative; in production you weight by kind and severity.

What to do with this on Monday

Pick the type that hurts you the most right now. If your assistant cites things that do not exist, ship the factual detector first; the lookup is the load-bearing piece, not the regex. If your agents touch databases, ship the tool-call resolver first and gate side effects behind it. If your RAG answers contradict the document, the divergence check is what to ship first.

Then add the other three. Run them all on every output. Log findings to your tracing layer alongside the trace ID. The next time a customer pastes three sentences from a broken response, you will know which of the four types fired, on which span, with which reason. Next time it happens, you ship a fix to the right detector, not a new prompt.

If this was useful

The LLM Observability Pocket Guide covers how to wire detectors like these into the eval and tracing tools that already live in your stack: where to put the checks (online vs. offline), how to sample for self-consistency without doubling your inference bill, and what to alert on. It also walks through threading per-finding rationales through OpenTelemetry spans, so the on-call gets a span and a reason instead of a one-line "hallucination=true" flag.

LLM Observability Pocket Guide