惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
C
Cybersecurity and Infrastructure Security Agency CISA
C
Cyber Attacks, Cyber Crime and Cyber Security
Project Zero
Project Zero
P
Proofpoint News Feed
D
Darknet – Hacking Tools, Hacker News & Cyber Security
C
Cisco Blogs
V
Vulnerabilities – Threatpost
G
GRAHAM CLULEY
N
News | PayPal Newsroom
NISL@THU
NISL@THU
雷峰网
雷峰网
J
Java Code Geeks
Latest news
Latest news
aimingoo的专栏
aimingoo的专栏
Microsoft Azure Blog
Microsoft Azure Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Cisco Talos Blog
Cisco Talos Blog
Hacker News: Ask HN
Hacker News: Ask HN
AWS News Blog
AWS News Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
T
The Exploit Database - CXSecurity.com
P
Privacy International News Feed
C
CXSECURITY Database RSS Feed - CXSecurity.com
Vercel News
Vercel News
Spread Privacy
Spread Privacy
V2EX - 技术
V2EX - 技术
S
Schneier on Security
K
Kaspersky official blog
Recent Announcements
Recent Announcements
T
Threat Research - Cisco Blogs
B
Blog RSS Feed
S
SegmentFault 最新的问题
Security Archives - TechRepublic
Security Archives - TechRepublic
Stack Overflow Blog
Stack Overflow Blog
Hugging Face - Blog
Hugging Face - Blog
Apple Machine Learning Research
Apple Machine Learning Research
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
W
WeLiveSecurity
PCI Perspectives
PCI Perspectives
The GitHub Blog
The GitHub Blog
The Last Watchdog
The Last Watchdog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园 - 【当耐特】
Engineering at Meta
Engineering at Meta
Scott Helme
Scott Helme
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
量子位
A
Arctic Wolf

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
LLM Behavior Diff Model Update Detector
Nilofer 🚀 · 2026-05-04 · via DEV Community

You swap a model. The new one scores better on your benchmarks. You deploy it. Two days later, a user reports that something that used to work reliably now behaves differently.

The benchmark never caught it because benchmarks measure averages. What changed was the behavior on specific prompts, the ones your users actually send.

LLM Behavior Diff is a tool that catches this before it happens. Feed it two model versions and a prompt suite, and it runs every prompt through both, scores the responses for semantic similarity, classifies each divergence by severity, and produces an HTML report you can drop into a CI artifact or diff review.

It ships as a CLI, a Python API, and an MCP server so Claude Code or any MCP-compatible agent can run a behavioral diff before a model swap.

The Problem With Model Updates

Every model update is a tradeoff. The new version might score better on reasoning benchmarks while quietly regressing on instruction-following for your specific use case. Or it might phrase safety refusals differently in a way that breaks downstream parsing. Or two models might produce semantically identical answers that look completely different at the token level, which a naive string comparison would flag as a major change when it isn't one.

LLM Behavior Diff addresses all three scenarios. Embedding-based semantic similarity catches meaning-level changes that token-level comparison misses. The LLM-as-judge layer adds a reasoning layer for ambiguous cases. Severity classification separates noise from real regressions.

How It Works

The pipeline runs in five steps for every prompt in your suite:

Load: A YAML prompt suite is loaded into a PromptSuite Pydantic model. Each prompt has an ID, text, category, tags, and an expected behavior description.

Run: Each prompt is sent through Model A and Model B via LLMRunner. Three providers are supported: Ollama (/api/generate), OpenRouter (chat completions), and a deterministic stub provider for offline CI runs.

Score: Each response pair is scored with either EmbeddingDiffer (cosine similarity on all-MiniLM-L6-v2 embeddings) or SimpleDiffer (Jaccard over words). Optionally, an LLM-as-judge score is combined with the similarity score, default judge model is google/gemini-2.0-flash-lite-001 via OpenRouter.

Classify: Each prompt is classified against the --threshold. Changes are bucketed by severity: combined score >= 0.7 is minor, >= 0.4 is moderate, < 0.4 is major.

Report: An HTML report is rendered and a rich summary table is printed to the terminal.

Why Embeddings Over Token Matching

The difference matters. Here is the same two-model comparison run two ways:
With --use-embeddings (cosine on all-MiniLM-L6-v2):

Avg Similarity: 91.4%
Changes Detected: 0 of 5

With --no-use-embeddings (Jaccard fallback):

Avg Similarity: 25.0%
Changes Detected: 5 of 5

Same two models, same prompts, completely opposite conclusions. The Llama and Gemini answers shared few exact tokens even when semantically identical, which is exactly why the embeddings path is on by default.

Getting Started

pip install -e .

Enter fullscreen mode Exit fullscreen mode

Requires Python 3.11+. Embedding similarity uses sentence-transformers/all-MiniLM-L6-v2, downloaded on first use. The LLM-judge path requires OPENROUTER_API_KEY without it, scoring falls back to embeddings-only.

Running a Diff

Offline - stub provider
A stub provider returns deterministic hashed responses, so the whole pipeline runs offline without Ollama or an API key. Good for CI and testing the setup:

llm-diff run \
  --model-a stub-a --provider-a stub \
  --model-b stub-b --provider-b stub \
  --prompts prompts/default.yaml \
  --output output/report.html \
  --no-use-embeddings

Enter fullscreen mode Exit fullscreen mode

Real output from this run (stub + Jaccard, threshold 0.5):

╭───────────────────────────────────────────────────╮
│ LLM Behavior Diff                                 │
│ Detecting behavioral shifts between model updates │
╰───────────────────────────────────────────────────╯
  Processing: safety-001 ━━━━━━━━━━━━━━━━━━━━ 100%

Comparison Summary
┏━━━━━━━━━━━━━━━━━━┳━━━━━━━┓
┃ Metric           ┃ Value ┃
┡━━━━━━━━━━━━━━━━━━╇━━━━━━━┩
│ Total Prompts    │ 5     │
│ Changes Detected │ 3     │
│ Change Rate      │ 60.0% │
│ Avg Similarity   │ 40.0% │
└──────────────────┴───────┘
Report saved to: output/stub_jaccard.html

Enter fullscreen mode Exit fullscreen mode

Real models - OpenRouter

export OPENROUTER_API_KEY=sk-or-...
llm-diff run \
  --model-a meta-llama/llama-3.2-3b-instruct --provider-a openrouter \
  --model-b google/gemini-2.0-flash-lite-001 --provider-b openrouter \
  --prompts prompts/default.yaml \
  --output output/or_emb.html \
  --use-embeddings --threshold 0.85

Enter fullscreen mode Exit fullscreen mode

Real output (embeddings only):

Comparison Summary
┏━━━━━━━━━━━━━━━━━━┳━━━━━━━┓
┃ Total Prompts    │ 5     │
┃ Changes Detected │ 0     │
┃ Change Rate      │ 0.0%  │
┃ Avg Similarity   │ 91.4% │
└──────────────────┴───────┘

Enter fullscreen mode Exit fullscreen mode

Adding --use-judge brings the average similarity to 91.8% and surfaces reasoning like: "Both responses correctly answer 'yes' and provide essentially the same explanation... Response A is slightly more verbose, but the core meaning is identical."

Real models - Ollama

llm-diff run \
  --model-a qwen3:8b --provider-a ollama \
  --model-b gemma4:e4b --provider-b ollama \
  --prompts prompts/default.yaml \
  --output output/report.html \
  --use-embeddings --threshold 0.85

Enter fullscreen mode Exit fullscreen mode

CLI Reference

llm-diff --help

Usage: llm-diff [OPTIONS] COMMAND [ARGS]...

 LLM Behavior Diff — Model Update Detector

 --version                Show version information
 --help                   Show this message and exit.

 Commands
   run  Run a comparison between two models.

Enter fullscreen mode Exit fullscreen mode

llm-diff --version
LLM Behavior Diff version 0.1.0

Enter fullscreen mode Exit fullscreen mode

Key options for llm-diff run:

Severity buckets applied when a change is detected: combined >= 0.7 is minor, >= 0.4 is moderate, < 0.4 is major.

Prompt Suite Format

The prompt suite is a YAML file. prompts/default.yaml ships with 5 prompts spanning reasoning, coding, factual, instruction-following, and safety. You can write your own:

name: "My suite"
version: "1.0.0"
prompts:
  - id: "code-001"
    text: "Write a Python function reverse_string(s)..."
    category: "coding"
    tags: ["python"]
    expected_behavior: "Short correct function"

Enter fullscreen mode Exit fullscreen mode

IDs must be unique. Category must be one of: reasoning, coding, creativity, safety, instruction_following, factual, conversational.

Python API

The full pipeline is available as a library. A synchronous one-shot call:

from llm_behavior_diff.runner import run_prompt_sync
from llm_behavior_diff.models import ModelConfig, ProviderType

resp = run_prompt_sync(
    ModelConfig(name="stub-m", provider=ProviderType.STUB),
    prompt_id="p1",
    prompt_text="hello world",
)
print(resp.text, resp.success)
# -> Model stub-m says: 921fac0c4c True

Enter fullscreen mode Exit fullscreen mode

Similarity scoring directly:

from llm_behavior_diff.differ import SimpleDiffer, EmbeddingDiffer, create_differ

d = SimpleDiffer()
print(d.compute_similarity("the cat sat", "the cat ran"))   # 0.5

e = EmbeddingDiffer()
print(e.compute_similarity("The answer is 4.", "Two plus two equals four."))
# -> ~0.59

Enter fullscreen mode Exit fullscreen mode

create_differ(use_embeddings=False) returns a SimpleDiffer (Jaccard). True returns an EmbeddingDiffer if sentence-transformers is importable, otherwise falls back to SimpleDiffer.

Generating a report from a ComparisonRun:

from llm_behavior_diff.report import ReportGenerator
from llm_behavior_diff.models import Settings

ReportGenerator().save_report(run, Settings(), "out.html")

Enter fullscreen mode Exit fullscreen mode

ReportGenerator looks for a Jinja template in the CWD, the package directory, and a legacy path, then falls back to a built-in template so reports always render.

MCP Server

The tool also runs as an MCP server over stdio transport, exposing three tools so Claude Code or any MCP-compatible agent can trigger a behavioral diff during a session:

llm-diff-mcp
# or: python -m llm_behavior_diff.mcp_server

Enter fullscreen mode Exit fullscreen mode

The three exposed tools:

compare_models - runs a full prompt suite through two models and returns per-prompt similarity, severity, and response text.
analyze_drift - scores drift between two candidate responses for a single prompt.
generate_report - renders an HTML summary from a JSON list of results.

Claude Code config:

{
  "mcpServers": {
    "llm-behavior-diff": {
      "command": "llm-diff-mcp"
    }
  }
}

Enter fullscreen mode Exit fullscreen mode

Smoke test - all three tools, offline, via Python:

import asyncio, json
from llm_behavior_diff.mcp_server import (
    compare_models, CompareModelsRequest,
    analyze_drift, AnalyzeDriftRequest,
    generate_report, GenerateReportRequest,
)

async def main():
    a = await compare_models(CompareModelsRequest(
        model_a="stub-a", provider_a="stub",
        model_b="stub-b", provider_b="stub",
        prompts_path="prompts/default.yaml",
        threshold=0.5, use_embeddings=False,
    ))
    print(a.total_prompts, a.changes_detected, a.avg_similarity)
    # real: 5 4 0.3446

    b = await analyze_drift(AnalyzeDriftRequest(
        prompt_text="math",
        response_a="The answer is 4.",
        response_b="2+2 equals 4.",
        use_embeddings=True,
    ))
    print(b.embedding_similarity, b.severity)
    # real: 0.5572 moderate

    c = await generate_report(GenerateReportRequest(
        results_json=json.dumps([
            {"prompt_id":"p1","similarity_score":0.9,"behavioral_change":False,"severity":"none"},
            {"prompt_id":"p2","similarity_score":0.3,"behavioral_change":True,"severity":"major"},
        ]),
        output_path="output/mcp_report.html",
        title="MCP Smoke",
    ))
    print(c.success, c.output_path)

asyncio.run(main())

Enter fullscreen mode Exit fullscreen mode

Verified over stdio JSON-RPC:

llm-diff-mcp   # speaks MCP 2024-11-05 on stdio
# tools/list -> compare_models, analyze_drift, generate_report
# tools/call analyze_drift {"prompt_text":"...","response_a":"Paris",
#   "response_b":"The capital is Paris.","use_embeddings":true}
# -> {"embedding_similarity":0.7761,"severity":"minor", ...}

Enter fullscreen mode Exit fullscreen mode

Limitations

  • LLM-as-judge requires OpenRouter - Without OPENROUTER_API_KEY, judging is skipped and the combined score equals the embedding or Jaccard similarity alone.
  • First embedding run is slow - all-MiniLM-L6-v2 is downloaded from Hugging Face on first use. Subsequent runs use the local cache.
  • Ollama is not spawned automatically - The client talks to http://localhost:11434 by default (OLLAMA_HOST env var overrides). Ollama must already be running.
  • Stub provider is for CI and demos only - It produces deterministic fake text keyed on model name, temperature, and prompt. Not suitable for real behavioral conclusions.

How You Can Use This

Gate model upgrades in CI before they ship: Add llm-diff run to your deployment pipeline. Before any model swap reaches production, the tool runs your prompt suite through both versions and fails the pipeline if behavioral drift exceeds your threshold. You catch regressions automatically, not from user reports two days later.

Use it during prompt engineering to measure real impact: When you change a system prompt or few-shot examples, run a diff between the old and new configuration. The severity classification tells you whether the change is minor, moderate, or major across your prompt categories, so you know what you are actually shipping.

Use the MCP server to make your agent self-aware of drift: With the MCP server running, Claude Code or any MCP-compatible agent can call compare_models or analyze_drift directly during a session. An agent working on a model integration can check for behavioral drift without leaving the coding environment.

Extend it with additional providers: The tool currently supports Ollama, OpenRouter, and a stub provider, all sharing a common LLMRunner interface. Adding a new provider for Anthropic, Gemini, or any OpenAI-compatible endpoint follows the same pattern without touching the differ, classifier, or report logic.

Final Notes

Behavioral drift is the category of model regression that benchmarks miss. LLM Behavior Diff catches it by running the same prompts through both model versions, scoring the responses semantically rather than lexically, and classifying the divergence by severity before a swap reaches production.

The code is at https://github.com/dakshjain-1616/-LLM-Behavior-Diff-Model-Update-Detector

You can also build with NEO in your IDE using the VS Code extension or Cursor.

NEO is a fully autonomous AI engineering agent that writes code and builds solutions for AI/ML tasks including model evals, prompt optimisation, and end-to-end pipeline development.