惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Troy Hunt's Blog
P
Palo Alto Networks Blog
N
News and Events Feed by Topic
Hacker News - Newest:
Hacker News - Newest: "LLM"
T
Threatpost
C
Cyber Attacks, Cyber Crime and Cyber Security
S
Schneier on Security
Google Online Security Blog
Google Online Security Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Spread Privacy
Spread Privacy
NISL@THU
NISL@THU
Cisco Talos Blog
Cisco Talos Blog
The GitHub Blog
The GitHub Blog
S
SegmentFault 最新的问题
量子位
L
Lohrmann on Cybersecurity
酷 壳 – CoolShell
酷 壳 – CoolShell
Attack and Defense Labs
Attack and Defense Labs
Y
Y Combinator Blog
Project Zero
Project Zero
AWS News Blog
AWS News Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Last Week in AI
Last Week in AI
博客园 - 聂微东
MyScale Blog
MyScale Blog
aimingoo的专栏
aimingoo的专栏
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
S
Securelist
Latest news
Latest news
C
CXSECURITY Database RSS Feed - CXSecurity.com
B
Blog RSS Feed
Webroot Blog
Webroot Blog
Blog — PlanetScale
Blog — PlanetScale
Recent Announcements
Recent Announcements
V2EX - 技术
V2EX - 技术
Schneier on Security
Schneier on Security
F
Full Disclosure
Apple Machine Learning Research
Apple Machine Learning Research
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
P
Proofpoint News Feed
Recent Commits to openclaw:main
Recent Commits to openclaw:main
月光博客
月光博客
L
LINUX DO - 最新话题
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
H
Heimdal Security Blog
F
Fortinet All Blogs
博客园_首页
N
News | PayPal Newsroom
P
Proofpoint News Feed

BlogFinder

日常漫步 Vol.24 之漫步前山河 - 雅余 周报 #1-聊聊本周的收获 - Edwin's Blog 我的OpenCode必装插件与Skill Write Something 掌中之物未必在掌握之中 · CRIVU PiliNara,一个更顺手的 PiliPlus 分支 「NekoEcho」:做一个必有回响的猫娘主题博客 2026-05 书影音总结 简化博客主题 - 安迪 我第一次发布 npm 包 拾花小记#45:中考前的二三事 – 小改学习志 黛西花园5月游 #18 枇杷又熟了的五月月报 一些奇奇怪怪的需求?word仿方正书版的几个小操作 - Xiobb's Blog 0419 御温泉之旅 修复了一些bug,网站基本上趋于稳定了 - 新锐博客 又回到四十年前 如何定义成功 迷鹿屋2026已重新上线 科技冰火两重天+一周回顾 ${title} 热度退了,我反而用得更深了-咕咚同学 我到底该不该换个域名? 随身WIFI折腾记 - 安迪 博客撰写体验提升——hexo pro插件 为什么不用相机把屏幕上的接关密码拍下来? 国清寺与天台山 – Ouroboros ★★★★☆《挽救计划》——久违的经济上行感 - Davidの3号基地 删除右键“打开方式”里多余选项 第三周刊_No.53|一切都会被支付两次 安卓APP通话记录与录音上传踩坑记录 - 子舒的博客 天量下跌 inBox 笔记 2.3.8,把工具栏交给了你-咕咚同学 我把小龙虾搬到了微信-咕咚同学 安好 - 响石潭 Compound Engineering Plugin:让每个工程单元都比上一个更容易 MOSS-TTS Family:开源高质量语音与声音生成模型家族深度解析 Crawl4AI:专为 LLM 设计的开源 Web 爬虫与数据抓取工具 Build Your Own X:从零实现你最喜欢的技术——程序员进阶的终极资源清单 Anthropic Skills:用文件夹教 Claude 专业技能的开源框架 1年的去月球(下) - 梅之夏 欢迎回来。 简单讲讲 ASN.1 与 OID DTV - 直播聚合客户端 5.22-5.27 – 不兴江 还没去过鸭川 – 不兴江 张晶晶同学三刷林志颖 关于我 – 不兴江 爱与嫉妒 – 不兴江 港股被持续做空 备案码花了四百块-咕咚同学 一句话生成封面:我给公众号做了4种风格的AI封面生成技能 「官」方認證 再谈费曼学习法 2026-05-28T00:34:11+08:00 2026-05-28T00:28:45+08:00 离谱的英语学习指南:基于AI的英语进阶系统方法论 iii:零集成架构的后端统一运行时 Claude Code Harness:让 Claude Code 工作有迹可循的工程化框架 Heretic:全自动移除大语言模型审查机制的开源工具 MarkItDown:微软开源的万能文档转 Markdown 利器 Harness:让 Claude Code 秒变多智能体协作工厂 这段时间尽折腾AI Agent了,确实极大地提高了效率 近期动态:两个新站点正式上线啦 误判解除!zhouayuan.com 腾讯安全申诉成功 - 周阿源|玩具设计・插画日常・生活随笔 Ralph:让 AI 编码工具自主循环跑完所有 PRD 任务的量产神器 全都违法 – 个人工作记录 关于zhouayuan.com被误判 “含违规信息” 的说明与申诉记录 - 周阿源|玩具设计・插画日常・生活随笔 小米 MiMo v2.5 Pro 白嫖 最大的人间清醒,兜里有钱,但是不花。 夜晚靓歌(12):于文文现场solo - 王志勇的Blog 今日插画:风扬起的倔强 - 周阿源|玩具设计・插画日常・生活随笔 回门习俗 独立网卡 - 忘记了回忆 500亿入股人工智能企业 从命令行到桌面智能体-咕咚同学 第一性原理读书笔记 行者微评论223-加班の守株待兔-博客|政治与时事-风雨行者 ZOZO开源物理接触求解器:GPU加速的可扩展仿真引擎 OpenStock:开源股票市场交易平台技术深度解析 MoneyPrinterTurbo:基于AI的全自动短视频生成工具深度解析 Claude-Mem:为 Claude Code 构建的持久化记忆压缩系统 Twenty:可代码化定制的企业级开源 CRM 平台技术深度解析 2026-05-26T22:59:17+08:00 企业级开源大模型部署平台 GPUStack 实战教程 1年的去月球(上) - 梅之夏 Sevalla - 静态网站托管服务 不用翻墙、不用注册、不用月费,普通人也能用上 Claude Code 装修灯具要注意⚠️ 黄梅天先锋 - 游子微博 公安备案顺利办结,站点备案全部完成 - 周阿源|玩具设计・插画日常・生活随笔 第三次兑换天猫超市卡了宗宗酱-三维狐少儿编程 Don't think, feel. - Rolen's Blog 人这一辈子,到底图个什么 博客迁移 - Edwin's Blog 情感赛道写作模板 再现本轮行情的典型特征 裁员与平常心-咕咚同学 别让“偷懒”,成为隐私泄露的破绽 片刻 - Jdeal | Life is like a Design.
How I Use AI for Code Review | Viking
Viking Zhang · 2026-05-31 · via BlogFinder

Over the past year, I’ve been writing nearly all my code with AI. Meanwhile, TinyShip has been live for six months now — it keeps gaining features, the codebase keeps growing, and since it’s a template product with real users, quality can’t be an afterthought.

The thing is, AI can spit out hundreds of lines across seven or eight files in a minute. There’s no way I can read every line. What makes it trickier is that TinyShip is a monorepo — three apps (Next.js, Nuxt.js, TanStack Start) that need to change in sync, plus support for both PostgreSQL and SQLite. A change in one place can ripple into the other two.

When a feature commit clocks in at over a thousand lines, my stomach tightens. How do I actually review all that code properly?

I recently came across a great article — Using AI to write better code more slowly. The core idea is that writing code with AI shouldn’t just be about speed. It should be about using AI to produce more robust, higher-quality code. That means slowing down, reading the code carefully. That resonates deeply with me.

So for my own product’s quality, I take a two-pronged approach:

  • Testing — especially for TinyShip, E2E tests are essential. With six different configurations in play, there’s no getting around it. I’ll cover my E2E strategy in a follow-up post.
  • Code Review — I built my own review workflow. I call it Review Forge. It’s not a fancy tool — just a set of rules plus a file structure that lets AI handle the review legwork while I make the final calls. That’s what this post is about.

Overview

I’ve packaged my code review process as a skill. You can find it here: https://github.com/vikingmute/review-forge

Here’s how it works at a high level:

  • Review — Different models each generate a bug report against the current diff or branch, one report per model.
  • Synthesize — Take the bug reports from each model and merge them into a summary.md.
  • Manual Review — Go through each finding and decide which issues to actually fix.
  • Fix — Hand off the selected issues to one model to fix, then run the tests.
  • Verify — Have a different model verify the fixes. Loop until every issue you’ve flagged is resolved.

Let me walk through each step in detail, tied to how the skill actually works.

Review

# Once review-forge is installed and I'm done developing, I kick off a review by sending this prompt to an agent:
/review-forge review feature: checkout-refactor

This creates a folder named checkout-refactor inside /code_review (all review-related content for this feature lives there). The model then writes a review document named after itself — if you’re using Codex, you get codex.md. So the file structure looks like this:

code_review
├── checkout-refactor/
   ├── codex.md
# Then I fire up a different agent with the same prompt to get a second report:
/review-forge review feature: checkout-refactor
code_review
├── checkout-refactor/
   ├── codex.md
   ├── composer.md

I typically use three models: GPT-5.5, Composer 2.5, and DeepSeek Pro V4.

Why Multi-Agent Review

At first, I only used a single service for reviews — I tried Bugbot, CodeRabbit, and Codex’s review mode.

The results were decent. They could catch obvious bugs and logic errors. But when a single model dumped a long list of issues on me, I’d get overwhelmed and then overly cautious. It’s not that I didn’t trust it — it’s that I’d end up spending an extra step wondering: is this really a problem, or is the model misunderstanding something? I’d even go ask other models: “Hey, do these issues actually exist? If so, fix them.”

And then I noticed something else: every model has its blind spots.

For the same feature, each model focuses on whether the code is correct. But edge cases, error paths, cross-file side effects — a single model won’t necessarily cover all of them. It’s not a capability problem; it’s an attention allocation problem. One model doing a single review pass is like one person reviewing code — there are always things they’ll miss.

So I tried having two models review the same code independently — GPT and Claude each produced their own report — and then I compared them.

The results were interesting:

  • About 60% of findings overlapped
  • The remaining 40% each had its own emphasis
  • Crucially, several issues were caught by only one model; the other never mentioned them at all

But the bigger insight here is this: when two models both flag the same issue, it’s almost certainly a real problem.

For overlapping findings, I can act on them directly without spending much time asking “is this a real bug?” Two models independently pointing at the same thing means the issue is obvious enough. High-priority overlaps especially — I don’t second-guess those. I just fix them.

So multi-model review gives me two layers of value:

  1. Cross-validation — issues flagged by multiple models have high confidence; fix them decisively.
  2. Wider coverage — single-model findings fill gaps the other reviewer might have missed, though those I verify myself.

Of course, I don’t run multiple models every time. I decide based on the size of the change: features that touch a dozen or more files get the multi-model treatment; small tweaks get one model and skip the heavy process. An extra review pass isn’t free, after all.

Synthesize

Given all the above, the Synthesize step felt like a natural abstraction. It’s really just handing the different reports to an AI to analyze, rank, and produce a checklist.

# Any agent, any model can run this subcommand. The analysis is straightforward.
/review-forge synthesize feature: checkout-refactor
code_review
├── checkout-refactor/
   ├── codex.md
   ├── composer.md
   ├── summary.md   # <-- the new summary document

The summary document sorts overlapping issues by priority at the top, with each one prefixed by a checkbox, like this:

- [ ] `RF-002`
- `severity`: high
- `reviewer_agreement`: 2/3 (composer, codex)

I manually review every issue and decide which ones to fix, checking the box on the ones I want addressed.

Why I Manually Decide Which Bugs to Fix

This might be the most important step in the entire workflow.

After an AI review, you get a long list of items, each tagged with a severity level. By the AI’s logic, anything marked severe should be fixed. But that’s not how it works in reality. AI lacks project context and the judgment to make trade-offs. This gets harder the larger the project gets, and AI tends to treat every finding as worth fixing — over-optimizing.

So a human has to make the call. You have to know your system well.

My approach is simple: after the AI review, I go through each item one by one. For every finding, I answer two questions:

  1. Would this actually cause a bug in a real-world scenario?
  2. What’s the cost of fixing it?

For most issues the AI flags, the answer to the first question is “no.” It’s just theoretically imperfect. The ones I act on are the ones that would actually break things — users hitting error pages, data getting corrupted, payments failing. Those get fixed, no hesitation.

I skip the high/medium severity items with a poor ROI — like needing to change 100 lines of code to handle an extremely narrow edge case.

After a full review cycle, the AI might list a dozen or more issues. I usually end up fixing three or four. That’s not a knock on the AI — it’s that reviewing code and fixing code are fundamentally different things. AI is good at finding problems, not at making decisions. The latter requires a human’s understanding of the project as a whole, the ability to weigh risk against reward, and a sense of when things are good enough.

Fix / Verify

From here, the workflow is straightforward — a fix-then-verify cycle. I use two separate agents: one for fixing, one for reviewing.

# I send a prompt like this to an agent running Claude:
/review-forge fix feature: checkout-refactor

This creates a fix-plan.md and a status.md to track progress, then carries out the fixes and runs the relevant tests.

code_review
├── checkout-refactor/
   ├── codex.md
   ├── composer.md
   ├── summary.md
   ├── fix-plan.md    # fix plan
   ├── status.md      # status tracker
# In a different agent, I send this prompt — I use Codex with GPT-5.5 here:
/review-forge verify feature: checkout-refactor

This agent reviews the other model’s fixes, checks whether each issue is properly resolved, creates a verify.md document, and updates status.md with feedback. The two agents go back and forth until everything I’ve flagged is confirmed resolved.

code_review
├── checkout-refactor/
   ├── codex.md
   ├── composer.md
   ├── summary.md
   ├── fix-plan.md
   ├── status.md      # updated status
   ├── verify.md      # verification results and details

Why You Can’t Use the Same Model to Fix and Review

Once you’ve decided what to fix, there’s one hard rule: the model that fixes must not be the model that verifies.

The reasoning is the same as with human code review — you can’t have the same person write the code and review it. They won’t catch their own bugs, not because they lack ability, but because of cognitive inertia. AI is the same way.

Plus, using a different model for verification has an extra benefit: if the second model also signs off on the fixes, I can feel reasonably confident. It’s like having two lines of defense.

Closing Thoughts

After using this workflow for a while, here are my honest impressions:

The biggest value is peace of mind. Before, I was hesitant to merge AI-written code because I couldn’t tell if there were hidden traps. Now, after running through this process, I know which issues were evaluated and which ones I consciously chose not to fix. I merge code with my eyes open.

Multi-model review works, but don’t overuse it. Save it for big features. For small changes, it’s not worth it — one model is enough.

Review Forge solves a simple problem: use multiple AIs to surface as many issues as possible, then let a human decide which ones are worth fixing.

If you’re writing most of your code with AI and your review process is struggling to keep up, give a similar workflow a try. You don’t need my exact setup — the core is just three things:

  1. Use multiple AIs to review your code and reduce blind spots
  2. Manually go through every finding and decide what to fix
  3. Fix, run tests, and cross-verify with a different model

Here’s what the final file structure looks like:

TinyShip