惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

大猫的无限游戏
大猫的无限游戏
MyScale Blog
MyScale Blog
雷峰网
雷峰网
量子位
小众软件
小众软件
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 叶小钗
T
Tailwind CSS Blog
月光博客
月光博客
博客园 - 【当耐特】
博客园_首页
罗磊的独立博客
博客园 - 三生石上(FineUI控件)
IT之家
IT之家
爱范儿
爱范儿
阮一峰的网络日志
阮一峰的网络日志
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
WordPress大学
WordPress大学
The Cloudflare Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
S
SegmentFault 最新的问题
人人都是产品经理
人人都是产品经理
V
V2EX
酷 壳 – CoolShell
酷 壳 – CoolShell

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
I Thought My AI Code Reviewer Was Finished. Then a Single...
shy The · 2026-06-03 · via DEV Community

shy The

GitHub “Finish-Up-A-Thon” Challenge Submission

What I Built
ReviewFlow is an automated Code Review pipeline driven by a custom Multi-Agent architecture written in native TypeScript. Instead of relying on a single monolithic prompt, it structures and routes incoming Pull Requests to specialized agents (Security, Logic, and Style) to perform parallel reviews. The workflow is orchestrated through GitHub Actions and posts comments via the Octokit API.

The Blind Spot: When Brilliant AI Meets Silent Failures
Getting the pipeline to trigger GitHub Actions felt like a huge win. But then, I noticed a fatal issue: the comments were disappearing.

The logs showed the LLM was generating brilliant security and logic feedback, but on the actual Pull Request page, nothing was posted. After digging into the Octokit error logs, I found a subtle reliability bug: Coordinate Hallucination.

Even when the model correctly identified a vulnerability, it couldn't reliably anchor that feedback to a valid line in the Git Diff. My system had a strict ⁠VerifierNode⁠ designed to block invalid API requests. When it saw these hallucinated out-of-bound coordinates, it silently dropped the comments. A single hallucinated number was wiping out the entire review pipeline. During initial testing across 5 dummy Pull Requests, the LLM generated 14 highly valuable security and architectural suggestions. However, because of coordinate hallucination, 9 of them were completely dropped by the verifier node due to invalid line ranges. That's a 64% silent failure rate—brilliant engineering ideas lost in the ether simply because the AI couldn't read the Git Diff index accurately.

The Fix: Deterministic Guardrails
I needed an engineering solution. I built a Deterministic Validation Layer (parseDiffToValidLines) to pre-calculate an exact "Valid Diff Index" map before the LLM even sees the prompt.

At the boundary of this layer, we enforce a hard check to strip away any out-of-bounds coordinates or malformed responses before they hit the GitHub API:


typescript
// 3. 执行确定性验证 (Validation Layer)
const validatedComments = rawComments.filter(comment => {
    return typeof comment.path === "string" &&
           Number.isInteger(comment.line) && comment.line >= 1 &&
           typeof comment.body === "string" && comment.body.length > 0;
});

// 4. 透明度反馈:告知用户过滤掉了多少无效内容
if (rawComments.length !== validatedComments.length) {
    console.warn(`⚠️ Filtered ${rawComments.length - validatedComments.length} invalid comments from AI output.`);
}



1. Engineering with Copilot: I outlined the logic in comments, and Copilot rapidly generated the core parsing function and comprehensive unit tests to prove this layer was bulletproof.

2. The Result: The agents now have a hard constraint: Choose a line number strictly from this pre-verified list, or drop the comment.

By using Copilot to offload the heavy lifting of structural verification, I was able to inject deterministic reliability into a non-deterministic LLM pipeline. The silent failures stopped, and the agents finally pinned their feedback to the exact lines of code.

The Takeaway: In an LLM pipeline, non-deterministic components need deterministic guardrails.

**Demo**

1. Automated PR Feedback (GitHub Action):
![ ](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/fe8rj7c6d8770up5top0.png)

2. Engineering Implementation (VS Code):
![ ](https://dev-to-uploads.s3.amazonaws.com/uploads/articles/vl9reugi2nbkghk4jftl.png)

Explore the Code
https://github.com/ywu593412-afk/diffens

Enter fullscreen mode Exit fullscreen mode