惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Vulnerabilities – Threatpost
Know Your Adversary
Know Your Adversary
C
Cyber Attacks, Cyber Crime and Cyber Security
S
Secure Thoughts
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
宝玉的分享
宝玉的分享
Spread Privacy
Spread Privacy
AWS News Blog
AWS News Blog
D
Docker
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
TaoSecurity Blog
TaoSecurity Blog
博客园 - 三生石上(FineUI控件)
Apple Machine Learning Research
Apple Machine Learning Research
Cyberwarzone
Cyberwarzone
V
V2EX
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
WordPress大学
WordPress大学
P
Palo Alto Networks Blog
H
Heimdal Security Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 叶小钗
N
News and Events Feed by Topic
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Simon Willison's Weblog
Simon Willison's Weblog
Project Zero
Project Zero
Martin Fowler
Martin Fowler
大猫的无限游戏
大猫的无限游戏
D
DataBreaches.Net
Engineering at Meta
Engineering at Meta
S
Schneier on Security
Google DeepMind News
Google DeepMind News
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Hugging Face - Blog
Hugging Face - Blog
P
Proofpoint News Feed
S
SegmentFault 最新的问题
Hacker News: Ask HN
Hacker News: Ask HN
小众软件
小众软件
博客园 - 聂微东
S
Security Affairs
T
Tor Project blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
T
Threat Research - Cisco Blogs
T
Threatpost
博客园 - 【当耐特】
L
LINUX DO - 热门话题
G
Google Developers Blog
P
Privacy & Cybersecurity Law Blog
A
About on SuperTechFans
F
Fortinet All Blogs

Hacker News: Ask HN

The New Window Delete ChatGPT Atlas Spyware Tell HN: Qwen Free Tier Is Discontinued Ask HN: SeedLegals Partnerships in London, worth it? Ask HN: How to highlight talent from untraditional backgrounds? Ask HN: We dont need a programming language now? Durable Object alarm loop: $34k in 8 days, zero users, no platform warning What if Time at the subatomic level has multiple arrows? How to add MidnightBSD Key to UEFI Secure Boot DBX? (Revoked and Forbidden Keys) Ask HN: What's your experience working at xAI as an AI tutor? Any engineers here with experience of clinical data standards? Ask HN: Who is using OpenClaw? Agent Skills for Software Test Automation Ask HN: Who needs contributors? Claude Code is thinking too much Ask HN: What Is the Big-O Order of a Jigsaw Puzzle? Ask HN: Stepping into a new role as a Senior, mentoring dos and dont's? Founder from Zurich heading to SF and Austin for the first time Hacker News No Manual Screenshots: I Built a Scalable Screenshot API Using Cloud Playwright Ask HN: Thought experiment: AGI giving us answers we don't like? Ask HN: I quit my job over weaponized robots to start my own venture 1% Vacancy, 81% Preleased: Where Midmarket Compute Deploys in 2026 Ask HN: Preferred pricing model for sound effects libraries? Copy of the email I sent to my undergraduate professors on Nov 30, 2025 Model API Performance | Hacker News Ask HN: Are open-weight LLMs the new offline encyclopedias? Valgrind 3.27 RC1 is out Claude Code OAuth down for >12 hours Ask HN: What's Better?–Tauri or Electron? Technical SEO vs. content optimization: which one moves rankings? Hacker News Ask HN: Can you cut off AI usage immediately? Nvidia's moat is not what it used to be Ask HN: What's your experience with PoW captchas against form spam? Ask HN: What are all the bad things that AI companies have done which we forgot Ask HN: Is Zero Trust Architecture Overkill? Ask HN: What is the best way to get your first users? Tell HN: OpenAI silently removed Study Mode from ChatGPT Ask HN: How to build an "AI native" company? Tell HN: docker pull fails in spain due to football cloudflare block Launchfolio – Create a portfolio in minutes for free, no account needed 120k USD compute credits from various providers Ask HN: How do you retain what you learn from podcasts? Ask HN: How is everyone dealing with the increase of code reviews? Ask HN: Agentic AI just makes me sad Ask HN: Do you trust AI agents with API keys / private keys? Ask HN: Anyone using Nostr as a lightweight back end/DB for rapid prototyping? Ask HN: What should I do with my app? 130 downloads 3 real subscribers Strong feeling: we are in a folded AI reality Hacker News What comes after Open Source? Ask HN: Former grok-code-fast-1 users, what coding model are you using now? When career anxiety becomes gameplay: lessons in China 'young-faculty simulator' I propose a new programming language, CPC Ask HN: Do you remux WebM to MP4 without re-encoding? Ask HN: How to have a macOS devcontainer in VS Code? I built a free 30-day habit tracker in Google Sheets Ask HN: What is the most annoying part of scheduling meetings? Ask HN: Has anyone reconsidered Antivirus software after recent security news? Tell HN: See the AI Doc Ask HN: Why have we not stepped back on the moon again? Ask HN: How did you specialize as a software engineer? Ask HN: Agentic Permutation of Testing Paths In A System Ask HN: Will AI Redefine Programming? Ask HN: How do you stop playing 20 questions with your AI coding tools Ask HN: Is the telehealth consulting for psychiatry even works? Ask HN: Im back end engineer, not front end – is this just excuse? What tools do you use to visualize algorithms? Tor Browser on Android leaks IP in desktop mode Published on Rapid API | Hacker News Persistent vs. Stubborn / Genius vs. Intelligent Is the pitch deck culture making founders worse at building businesses? Do founders' political views affect how you see a product? Ask HN: Easiest UX for Seniors My app hit 1,152 first-time downloads in a single day Claude API Error: 529 | Hacker News My AI workflow evolved from prompts to a near-autonomous workflow Hacker News Ask HN: Best books on building a programming language I collected startup ideas. It changed how I think about ideas completely Is algorithm still relevant in 2026 Is VC the new PMF strategy? Ask HN: Would you take your engineering team to Buenos Aires for an offsite? Hacker News Artemis 2 Coming Home | Hacker News Ask HN: Recommendations on which models to pay how much for? Open Source card game cuttle.cards has its world championship Saturday at 1pm ET Scanners are too late for AI-driven actions Ask HN: Negotiating Intern Pay Ask HN: Hiring in the age of AI-assisted coding: what works? Amazon Luna Shuts Down without refunds? The Weather Channel RetroCast Now Behind the Scenes and Technical / Design V1.21 Update for Gpumkat | Hacker News Valence and HYVE, RT Physics Attention and a "Synthetic Organism" Ask HN: Its either I or Agent code. Both of us on same codebase is a disaster Ask HN: Does Sam Altman know how to code? Ask HN: Is a purely Markdown-based CRM a terrible idea? Optimized for LLM agents Ask HN: Improving as mid-level dev with forced use of LLMs I built ClawIDE: A web-based IDE for managing multiple Claude Code sessions
Ask HN: What does your agentic software dark factory look like?
ElFitz · 2026-04-27 · via Hacker News: Ask HN

In some of the comment threads around here a few of you shared interesting ideas and patterns, enough that I believe everyone interesting in harness engineering is working on some sort of software dark factory or another.

We have OpenAI’s Symphony[1], StrongDM’s Factory[2], Yegge’s GasTown[3], and probably a few others I’ve missed.

So I’m curious. What have you been working on? What have learned? What has worked and what has failed? And what do you think comes after?

I’ll go first. The first thing I tried that yielded interesting results was, when possible, providing a ground truth or reference for the model to iterate against: screenshots or mockups for UI work, API contracts and unit / integration tests for logic. That’s the Ralph Loop we all know and love. A feedback loop.

The second (obvious, I know) was splitting planning and implementation.

Reviews by other models and iterative loops came next, with appreciable results. However the implementing agent would often wiggle out by deferring things into oblivion or saying things that were actually important feedback were out of scope. Another feedback loop. I’ve found turning those reviews into "hard gates" has its own set of issue, as reviewing agents will always find something to nitpick about, turning this iterative implementation approaches into near infinite loops.

Combining these reviews and committing plans alongside the code led to an interesting accident: reviewing agents spontaneously and unexpectedly picked up on those and drastically improved their feedbacks by comparing plan and implementation (should have been obvious, and you’ll imagine my surprise the first time GitHub Copilot actually provided useful feedbacks instead of the usual typo nitpicks).

Then a comment here led me to an adversarial green team / red team process.

A first agent creates a spec (based on StrongDM’s NLSpec) from my initial plan and gets it reviewed, including a detailed API.

A red team agent writes unit and integration test based on these specs, and gets them reviewed.

Then a green team agent is given those same specs and API, and implements the actual feature or fix, and iterates against the tests, without any access to the tests themselves, only which tests failed and what they were testing. This prevents it from gaming the tests.

Finally, once tests pass, a reviewing agent reviews the implementation against the specs.

This was nice. And it allows mixing and matching models, thinking levels, and providers. But both green and red team would sometimes diverge from the initial specs and API, sometimes with good reasons.

So another agent was brought in to evaluate those divergences when they occur and, if they are valid improvements, restart the process from the spec generation point, with the new insights. Yet another feedback loop.

And finally, integrating logs, OTel traces, and stack traces into the process. These agents seem remarkably capable at sifting through these, and end-to-end observability drastically improved results. Again, a feedback loop.

That’s all for me so far. Curious to see what other insights, findings, lessons or learnings everyone else has to share on this!

It’s a fun ride.