惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
量子位
A
Arctic Wolf
L
Lohrmann on Cybersecurity
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
WordPress大学
WordPress大学
V
Vulnerabilities – Threatpost
博客园 - Franky
C
Cyber Attacks, Cyber Crime and Cyber Security
The Cloudflare Blog
Last Week in AI
Last Week in AI
The Hacker News
The Hacker News
I
Intezer
J
Java Code Geeks
P
Privacy International News Feed
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
Secure Thoughts
Cisco Talos Blog
Cisco Talos Blog
阮一峰的网络日志
阮一峰的网络日志
S
Securelist
Security Latest
Security Latest
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
小众软件
小众软件
Jina AI
Jina AI
有赞技术团队
有赞技术团队
人人都是产品经理
人人都是产品经理
博客园_首页
酷 壳 – CoolShell
酷 壳 – CoolShell
T
The Exploit Database - CXSecurity.com
雷峰网
雷峰网
T
Tenable Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
P
Privacy & Cybersecurity Law Blog
Simon Willison's Weblog
Simon Willison's Weblog
博客园 - 【当耐特】
T
Threat Research - Cisco Blogs
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
MongoDB | Blog
MongoDB | Blog
D
DataBreaches.Net
N
News | PayPal Newsroom
Google Online Security Blog
Google Online Security Blog
K
Kaspersky official blog
H
Help Net Security
宝玉的分享
宝玉的分享
罗磊的独立博客
Webroot Blog
Webroot Blog
月光博客
月光博客
B
Blog RSS Feed
Recorded Future
Recorded Future

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
AI coding made us faster. Why did incidents increase?
theanonymous · 2026-05-19 · via Hacker News - Newest: "AI"

You have 1 article left to read this month before you need to register a free LeadDev.com account.

Estimated reading time: 7 minutes

Key takeaways:

  • AI amplifies whatever your delivery process already is. Strong practices get faster. Weak ones get paged at midnight.
  • The self-grading loop is your biggest hidden risk: when the model writes code and tests, it tests its own assumptions.
  • Observability is a release gate, not an afterthought.

Across engineering organizations, a pattern has become consistent. After rolling out an AI-coding assistant, velocity metrics improve quickly: pull requests (PRs) get bigger, cycle times shorten, and sprint records fall. Then, a few months in, the on-call rotation gets brutal. More tooling, faster shipping, and worse reliability. These are connected, not coincidental.

The DevOps Research and Assessment (DORA) 2024 report confirmed the pattern across the industry. Teams with significantly higher AI adoption also showed a higher change failure rate, meaning more deployments requiring hotfixes or causing production failures.

However, that finding cuts both ways. One pattern worth noting: teams that had already adopted contract testing and canary deployments before introducing AI tended to see their change failure rate fall rather than rise, because the AI-accelerated changes were already well-constrained.

Contract testing verifies that services interact correctly at their boundaries, while canary deployments release changes to a small percentage of traffic before rolling out fully. The tools themselves are not the problem. The delivery system around them is.

Three failure patterns explain most of the gap between those two outcomes. These are not new problems. They are old ones running faster.

Your inbox, upgraded.

Receive weekly engineering insights to level up your leadership approach.

3 failure patterns

1. Polished code fools reviewers

AI-generated code looks right. It follows naming conventions, respects linting rules, and reads like something a senior engineer would produce. PRs that would previously have prompted 20 are getting approved in 15. Often, reviewers pattern-match to familiar structure and skip the hard reasoning: side effects, edge cases, business logic that only makes sense if you understand the domain.

In practice, a model can produce a wrong implementation with the same fluency as a correct one. A reviewer under time pressure may not catch the difference.

2. When AI marks its own homework

When the same model writes code and then writes tests, but the tests cover what the model produced, not what the requirements are. Coverage turns green. Edge cases nobody described remain untested. Worse, AI-generated tests frequently verify implementation details rather than observable behavior. Refactor the code later and the tests need rewriting too. The suite slows every future change while providing none of the safety it implied.

3. AI cannot see the whole system

Every model works on the code it is shown. It has no awareness of the broader system: the shared retry queue, the upstream producer that sends late-arriving events, the implicit reliability guarantee held together by a design decision from three years ago. A change that looks like a clean refactor can quietly remove something critical. Those failures do not appear in unit tests. They appear on the on-call rotation.

None of these patterns were introduced by AI. Overconfident review, shallow test coverage, and implicit system assumptions predate coding assistants by years. What AI does is amplify whatever is already there, for good and bad. The question is whether your delivery process is worth amplifying.

A quality-first operating model

The answer is not to slow AI adoption. It is to redesign the delivery process so that speed and reliability reinforce each other. Three principles make the biggest practical difference.

1. Write the spec before you write the prompt

Before prompting the model to write any code, document the expected behavior in plain English: what the code should do, what inputs it handles, and what happens when things go wrong. Two or three sentences are enough. The AI then writes tests against that spec, and the implementation to satisfy those tests, breaking the self-grading loop described above.

Many teams resist this as an extra step. Writing a formal spec used to mean tickets, acceptance criteria, and a refinement session. Writing two sentences before pasting a prompt takes four minutes. The cost of documenting intent has dropped far enough that engineers actually do it, which is the part that made test-driven development (TDD) hard to adopt for 20 years.

Figure 1 (below) contrasts the two workflows: the old pattern where the model writes code and then tests against what it just produced, and the spec-first approach where intent precedes every line of generated code.

Left: the self-grading loop: AI writes code and tests, testing its own assumptions.
Right: spec-first flow: intent is documented before the model is invoked, breaking the loop.

2. Tier changes by risk and enforce contract test

Before AI touches authentication, payment logic, or data model changes, require a human to write the spec and a second reviewer to validate the business logic explicitly. Use feature flags to decouple deployment from release.

Feature flags are toggles that let you ship code without activating it for users yet. For anything touching an external integration, require a contract test that validates actual application programming interface (API) behavior against the live endpoint, not a mock. 

Many teams also find mutation testing scores useful as a merge gate. Mutation testing deliberately introduces small faults into code to confirm that tests actually catch them. If they do not, they are testing implementation rather than behavior.

3. Treat observability as a release gate

For medium and high-risk changes, define which metrics should be affected before deploying. Use canary rollout with automated rollback if error rate or latency crosses a threshold. Require a linked monitoring dashboard before any production-path PR can merge. If the author cannot point to a dashboard, the change is not done.

How one wrong field took down a transaction service at peak load

The following scenario illustrates a class of failure that recurs across engineering teams when AI-assisted development outpaces the controls around it.

The call came at 11.30 pm during a peak transaction window. A FinTech team’s transaction processing service had gone quiet in the worst possible way: not crashing, not throwing visible errors, just silently dropping outbound disbursements. Recipients were seeing no confirmation; funds were not arriving. This was the kind of night when silence from a financial service means something has gone badly wrong.

Earlier that week, an AI assistant had helped refactor the service. The model correctly identified that adding an idempotency key, which is a unique token sent with each request to prevent duplicate processing if the same request is retried, would improve reliability. It generated the change and placed the key in the request header. The downstream service provider only accepted it in the request body. Every unit test passed because the external API was mocked.

The engineer who traced it described a specific feeling: the code looked completely right. Clean, well-commented, the kind you would point to in a review as an example of doing it properly. The bug was not in the logic. It was in an assumption about a third-party API the model had never actually called.

Three gaps made it possible: the refactor had not been risk-tiered, no contract test had validated the integration against the live API, and no canary had monitored error rates during rollout. Any one of those controls would have caught it within minutes. None existed.

The model wrote correct code for the problem as it understood it. The problem was incompletely specified and nobody had built the system to check.

LDX3 London 2026 agenda is live - See who is in the lineup

LondonJune 2 & 3, 2026

Last few weeks left. Secure your spot soon!

What changed, and what to do next

Across teams I have seen adopt these controls, incident frequency has typically dropped by roughly a third within two quarters. Recovery time (mean time to recovery, or MTTR) tends to improve more than raw incident count, because defining metrics before deployment builds the dashboards needed for fast diagnosis.

On-call fatigue also reduced, though that is harder to measure than incident counts. Engineers tend to dread release days less. The tooling has not changed. The process has.

AI will amplify whatever your delivery system already is. In practice, and consistent with the DORA amplifier finding, teams that had strong practices before adopting AI tended to get faster. Teams that did not got a wake-up call late at night, at precisely the moment their system could least afford it. The good news is the gaps are fixable and the fixes do not require slowing down.

3 things to start this week

  1. Adopt spec-first prompting for critical changes. Require two or three sentences of intended behavior before any AI-assisted PR is opened on a production-path service.
  2. Classify AI-generated PRs by risk. Changes touching payment, authentication, or data models require human business-logic review, a contract test against the live API, and a documented rollback path.
  3. Treat observability as a release gate. No production-path merge without a linked monitoring dashboard. Measure success with change failure rate and MTTR – both should fall within two quarters.