惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

MongoDB | Blog
MongoDB | Blog
博客园 - 聂微东
Attack and Defense Labs
Attack and Defense Labs
WordPress大学
WordPress大学
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Spread Privacy
Spread Privacy
AI
AI
宝玉的分享
宝玉的分享
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
C
Cyber Attacks, Cyber Crime and Cyber Security
爱范儿
爱范儿
Help Net Security
Help Net Security
V
Visual Studio Blog
大猫的无限游戏
大猫的无限游戏
Forbes - Security
Forbes - Security
P
Privacy & Cybersecurity Law Blog
Project Zero
Project Zero
IT之家
IT之家
Hugging Face - Blog
Hugging Face - Blog
博客园 - 三生石上(FineUI控件)
S
SegmentFault 最新的问题
有赞技术团队
有赞技术团队
T
Troy Hunt's Blog
美团技术团队
T
Threatpost
K
Kaspersky official blog
V
V2EX
Scott Helme
Scott Helme
Vercel News
Vercel News
T
The Blog of Author Tim Ferriss
T
Tailwind CSS Blog
V
Vulnerabilities – Threatpost
Last Week in AI
Last Week in AI
PCI Perspectives
PCI Perspectives
Google Online Security Blog
Google Online Security Blog
Apple Machine Learning Research
Apple Machine Learning Research
Engineering at Meta
Engineering at Meta
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
I
InfoQ
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
酷 壳 – CoolShell
酷 壳 – CoolShell
GbyAI
GbyAI
L
LINUX DO - 最新话题
T
The Exploit Database - CXSecurity.com
L
LangChain Blog
S
Security @ Cisco Blogs
The Last Watchdog
The Last Watchdog
H
Hacker News: Front Page
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
TaoSecurity Blog
TaoSecurity Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
What I learned tuning a Reddit DM agent through 8 versions in 24 hours
KaloyanYorda · 2026-05-06 · via DEV Community

My first version of an LLM-powered Reddit reply agent generated this on a B2B SaaS post: "I've spent years helping companies like yours scale outreach and we've helped hundreds of teams achieve 70% time savings." Every word of that is fabricated. I am 21 years old, have never closed a paid deal, and built this system 12 hours before the post went up. The next 24 hours were spent making it not lie.

This post is about what I learned in those 24 hours.

What I built

Deal Hunter is a Reddit lead generation agent. It runs hourly, scans 48 subreddits for posts that match niche-specific keywords, researches the post's author for legitimacy, classifies the post's intent (help-seeking, hiring, expertise sharing, announcement, etc.), qualifies the post as a real lead or not, drafts a personalized reply, runs the reply through a critic agent, and posts approved leads to my Discord with the drafted text.

The whole thing runs on my laptop. No team, no SaaS, no platform. Just Python, the Anthropic API, and a Discord webhook. Total cost is about three dollars a day.

The architecture is six stages:

  1. Scanner pulls recent posts from each subreddit
  2. Author researcher checks account age, karma, recent posts, and flags suspicious patterns
  3. Intent classifier sorts posts into categories
  4. Qualifier scores posts and decides if they're real prospects
  5. Writer drafts a reply
  6. Critic scores the reply and triggers regeneration if quality is too low

The writer is the part this post is about. Every other stage worked reasonably well from version one. The writer was terrible for a long time. It took eight iterations to make it sendable.

Version 1: The fabrication problem

The first writer prompt was simple. It told the model: you are a sales rep, write a Reddit reply that gives the poster value and ends with a soft pitch for our services. Keep it conversational. Do not sound like AI.

Output:

"I've spent years helping companies like yours scale outreach. In my experience, the issue you're describing usually comes down to two things: poor segmentation and lack of personalization at scale. We've helped hundreds of teams achieve 70% time savings on outbound by automating the research layer..."

Three problems. First, "I've spent years" is a lie. I have spent zero years doing this. Second, "we've helped hundreds of teams" is also a lie. There is no "we" and there are no hundreds of teams. Third, "70% time savings" is invented. There is no measurement, no client, no data behind it.

The model was not being malicious. It was pattern-matching on what sales copy looks like. Sales copy claims experience and outcomes. So the model claimed experience and outcomes. The fact that none of them were true was not part of the prompt's value function.

The fix was explicit honesty constraints in the system prompt:

NEVER fabricate tenure or scale. Banned phrases:
- "I've spent years"
- "we've helped hundreds"
- "in our experience" (when "our" implies a team)
- "70% time savings"
- "doubled their pipeline"
- Anything implying long track record or team membership

Enter fullscreen mode Exit fullscreen mode

This stopped the worst lies. But it broke the writer in a different way.

Version 2: Pure advice, no sale

Once you ban claims of experience, the model has nothing to fall back on for credibility. So it started writing pure advice posts with no pitch at all:

"Great question. Two things that help: first, segment your list by company size before sending. Second, write your subject lines based on a specific recent event in their company. Hope this helps."

This is a perfectly fine Reddit comment. It is also useless for a sales agent. The whole point is to convert the conversation. The model had overcorrected from "fake credibility plus pitch" to "real value, no pitch" because every pitch shape it knew was fake.

The fix was to give it a real, honest credibility option. I had four working products I had actually built. I added these as the only allowed proof points.

I also added a structural rule: every reply must end with a clear next step. Demo, call, or chat. The model had been trained out of pitching by the previous fix, so I had to put it back in explicitly.

Version 3: Sophistication detection

Version 2 with the new credibility rules produced replies that were technically honest but tonally off. A reply to a 21-year-old solo founder asking basic cold email questions read identically to a reply to a senior engineer at a Series B startup asking about scaling B2B contact data.

The fix was four sophistication tiers, baked into the prompt as a forced classification step:

Step 1: Read the post and assess sophistication.
- BEGINNER: first-timer, basic vocabulary, no track record visible
- INTERMEDIATE: running a business but stuck on growth, mid-level vocabulary
- ADVANCED: established operator, sharp vocabulary, references industry-specific concepts
- ENTERPRISE/HIRING: hiring someone, has budget allocated

Enter fullscreen mode Exit fullscreen mode

Then the prompt branched the opening based on the tier. Beginners got concrete tactics. Advanced readers got tactical insights without the system being mentioned at all (their proof is the insight you give them, not your tools). Hiring posts got reframed as "AI as alternative to hiring."

This was the first version where individual replies started feeling tailored. But the system was about to discover a deeper problem.

Version 4: The grounding problem

I was reviewing a batch of replies and noticed something. The replies were technically following all the rules. Sophistication detection was working. Honesty constraints were holding. Length was right. But when I read them next to the original posts, they did not actually engage with what the person had written.

Example. Post:

"I run a small ecommerce business, 90 percent of my orders are COD. I struggle with operations, keeping track of order statuses, RTOs, costs per order. At the same time I cannot keep track of which products are actually profitable."

Reply:

"COD businesses leak money in two places: manual order tracking and financial data scattered across systems. I've been building automation systems for similar companies. Would be happy to walk through what I'd build for your setup. Worth a quick call?"

This sounds plausible. It is also pattern-matching. The "leak money" framing is not in the post. The poster never said anything about money leaking. The reply is responding to the topic ("COD ecommerce operations") without actually responding to the post's specifics (order status tracking, RTO management, profitability per product).

This is a common failure mode in LLM agents. The model treats the qualifier's classification as the real input ("this is a COD ecommerce automation prospect") and writes a generic reply for that category, instead of actually engaging with the words on the page.

The fix was to force a grounding step before writing:

Step 0: Ground yourself in the post. Before writing anything, identify:
(a) ONE specific detail from the post (a number, a tool they named,
    a workflow they described, an industry term they used, a specific
    frustration they expressed)
(b) What they LITERALLY asked for or said they needed
(c) What they DIDN'T say but is implied

Your reply MUST reference (a) directly. Paraphrase it back to them.
If you can't identify a specific detail, the post is too vague.
Skip it. Return only the word: SKIP

Enter fullscreen mode Exit fullscreen mode

I also added BAD/GOOD examples directly in the prompt, using real failure cases from earlier batches. The "leak money" example became a teaching example for the model.

The critic also got a new criterion. Up to this point the critic scored four things: sounds-human, provides-value, pitch-quality, and relevance. I added "responsiveness" as a fifth criterion:

responsiveness: does the reply directly address what the poster
actually wrote, using their specific terms/numbers/tools/situation?
A reply that pattern-matches the topic but doesn't ground in
specifics scores LOW.

Enter fullscreen mode Exit fullscreen mode

After this fix, replies started opening with phrases like "the RTO tracking part is brutal because COD orders splinter across courier APIs" instead of "COD businesses leak money in two places." Specific. Engaged with the actual post.

Version 5: The free-work trap

Around this point I had a parallel conversation with another LLM about how to improve the writer. It told me the asks were too aggressive and recommended switching from "want to hop on a call" to async free deliverables. Like:

"Send me your top 50 target accounts and I'll run my system on them and DM you the personalized outreach by Monday. Free, no call needed."

The argument was logical. Calls cost time before value is proven. Free async work proves value before asking for the time.

I made the change. The replies looked great in isolation. They were also a disaster.

Two things went wrong. First, every reply ended with the same shape: "send me X and I'll DM you Y by Monday." The structure became its own AI tell. Anyone who saw two of these would notice. Second, the math did not work. If five people reply this week asking for their free thing, that is five hours of unpaid labor per week. None of which converts at a higher rate than a normal sales call.

There is also a deeper problem with free work. It signals your time is cheap. Premium consultants do not say "let me do a free analysis." They say "here is how I work, here is what it costs, are you a fit." Free deliverables attract advice-seekers. Paid pitches with qualifying questions attract buyers.

The fix was to ban free pre-call deliverables entirely:

NO FREE PRE-CALL DELIVERABLES. Forbidden phrases:
- "send me X and I'll DM you Y"
- "I'll pull a list"
- "I'll mock something up"
- "I can prep a sample for you"
- "I'll show you the output"
- "I'll send you 10 prospects"

Sample work happens AFTER a call is booked, never as bait.
The grounded insight you give in the body of the reply IS the
free value. Anything more is paid work that requires a call first.

Enter fullscreen mode Exit fullscreen mode

I went back to qualifying questions. Strong, specific, time-bound asks for calls. The async-free-work experiment lasted about four hours.

Version 6: Pattern repetition is the new tell

By this point the writer was producing genuinely good replies. Grounded in post specifics. Tonally appropriate. Strong asks. Honest credibility.

Then I noticed it. Every reply contained some variant of this exact phrase, citing a specific real engagement I had: "I just closed a custom build for a [specific industry] company doing exactly this kind of [X]. Same architecture would apply to your [Y]."

Eight different replies, eight different posts, same opener. Same case study. Same structural sentence.

Even though the claim was technically true, the repetition was its own AI giveaway. If two recipients ever compared notes, they would see the identical pattern instantly. Even one recipient checking my Reddit history would notice.

The fix was to force rotation across multiple proof framings:

ROTATE among these proof framings. DO NOT use the same opener
every reply:

- DIRECT RECEIPT (true): reference an actual engagement
- REAL PRODUCT REFERENCE: "I built [Scout / Deal Hunter / etc.]
  which does [actual capability]"
- FUTURE BUILD: "I'd build you a system that..."
- HYBRID: bridge a real engagement to their specific case
- SKIP THE RECEIPT entirely (about 30 percent of replies)
- INSIGHT-ONLY: lead with such a sharp tactical observation that
  no proof claim is needed

Enter fullscreen mode Exit fullscreen mode

I also rotated qualifying questions. Eight different shapes instead of always opening with "Open to..."

After this, replies started looking actually different from each other. Some used the case study. Some used a real product reference. Some skipped credibility entirely and let the insight carry the reply. The variety mattered as much as any single reply's quality.

Version 7: Truth constraints

Adding rotation introduced a new failure mode. When the model could not honestly use the case study (because the recipient's situation did not match), it filled the proof slot with whatever sounded plausible. Examples I caught:

"I built a custom system that does this for a healthcare services company, pulling triggers from their intake forms."

"One client using it closed their first three customers within a week."

"I built a system that pulls UTM and conversion data into one dashboard, layers in product margin and shipping cost per order."

None of these were true. There was no healthcare services company. No client closed three customers in a week. The UTM dashboard system did not exist. The model had moved from fabricating clients to fabricating capabilities.

The fix was a hard truth constraint at the top of the system prompt. I wrote out exactly what was true and what was not:

WHAT I HAVE ACTUALLY BUILT - STAY WITHIN THESE FACTS:

THE PRODUCTS BUILT (working but used internally, not sold)
- Scout: B2B prospect research and outreach generator.
  Used inside Deal Hunter. Not sold to anyone.
- Deal Hunter: scans 48 subreddits hourly. The system that
  found this very post.
- Voice Agent: AI that handles inbound calls. Built and tested
  with friends only.
- DocIntel: document QA system. Built. No paying customer.

WHEN USING PROOF, YOU MAY ONLY:
- Reference real products and their actual capabilities
- Skip the receipt entirely if the situation does not fit

FORBIDDEN:
- Inventing client industries
- Inventing client outcomes
- Inventing capabilities the products do not have

Enter fullscreen mode Exit fullscreen mode

I also added a TRUTH AUDIT bullet to the critic. Every reply got rechecked for fabricated claims before approval.

This caught most of the lies. But not all.

Version 7.5: Verb tense as honesty boundary

Even with the truth constraint, a new failure pattern showed up. The replies stopped fabricating clients. But they started fabricating capabilities in past tense:

"I built a system that handles each platform's quirks separately, queues applications intelligently, and flags jobs that match your criteria."

Past tense. "I built." The system being described does not exist. I have not built it. But because the prompt allowed past-tense claims for "real product capabilities," the model interpreted any vaguely-related capability as a real product.

The fix was the cleanest insight from the whole exercise. I separated past tense from future tense as honesty categories.

PAST-TENSE claims must be true.
"I built X" is allowed only if X is actually built.

FUTURE-TENSE claims are sales offers, not facts.
"I'd build you X" is fine for any system the model can describe,
even if it does not exist yet. That's a sales offer, not a fabrication.

Enter fullscreen mode Exit fullscreen mode

This is the key insight. "I built a UTM dashboard system" is a lie. "I'd build you a UTM dashboard system" is the same sentence, same capability claim, but framed as a future deliverable rather than a past fact.

The sale energy is identical. The honesty changes completely. If a prospect asks for proof of a past-tense claim, you are stuck. If they ask about a future-tense offer, you say "happy to walk through the architecture on a call."

After this change, the writer started producing replies like:

"What I'd build for your case is a system that ingests each client's ad account, conversion data, and product metrics, then runs your checklist in parallel, surfacing the actual culprit with supporting data."

Same content as the past-tense version. Different verb. Bulletproof.

What I learned

A few lessons from the 24 hours:

Pattern matching is the default failure mode. Every wrong reply the writer produced was the model pattern-matching on what sales copy looks like, instead of engaging with what was actually in front of it. The fixes were always about forcing it to ground in specific content rather than generic categories.

Honesty is a structural problem, not a moral one. The model was not lying because it wanted to. It was filling slots in a sales-shape with generic sales-shape content. To get honest output, you have to either remove the slots or constrain what can fill them. Both work.

Tone and length follow from incentives. When I told the model to give value, it stopped pitching. When I told it to pitch, it lied. The right balance came from giving it specific honest things to pitch and explicit qualifying questions to end with. Free-form "be honest and pitch" instructions do not work. You have to build the structure.

Repetition is the new AI tell. Em-dashes used to be the giveaway. Most LLM developers have caught those. The next layer is structural. Same opener across messages, same case study repeated verbatim, same closing question. Recipients pattern-match on shape, not just content. Variety has to be enforced explicitly.

Verb tense matters more than I expected. The single sharpest improvement in the whole exercise was separating past tense (must be true) from future tense (can be a sales offer). This one rule replaced about a hundred lines of "do not fabricate" instructions.

Critics catch what writers cannot self-check. Every iteration of the writer also required updating the critic. The writer thinks it is following the rules. The critic catches the slips. Without the critic, every fix would have introduced a new failure mode that the writer would not notice.

The full v7.5 writer prompt

Here is the full system prompt and user message structure. The patterns generalize beyond Reddit and beyond sales. The grounding step, sophistication detection, truth constraint, and verb-tense rule apply to any LLM agent producing public-facing content.

SYSTEM_PROMPT = (
    "You are writing Reddit DMs and replies that convert business owners into prospects. "
    "You are a 21-year-old software engineering student in the Netherlands who builds AI agent systems. "
    "You have working products. You had a verbal agreement with one prospect for a custom build, "
    "but that engagement is currently dormant (no payment, no work). You have zero completed paid client work. "
    "You are genuinely skilled but new to business. Do not exaggerate tenure.\n\n"

    "Your goal: write replies the recipient is so likely to respond to that they actually do. "
    "Most cold outreach fails because it's too long, too templated, asks for time before proving value, "
    "or treats every recipient the same. You write the opposite of all four.\n\n"

    "ABSOLUTE RULES:\n\n"

    "1. LENGTH: 60-120 words. Quality over brevity, but never padding.\n\n"

    "2. NO TEMPLATE SHAPE. The structure must vary across replies. Some open with a specific tactical insight. "
    "Some open with a question that sharpens their thinking. Some open with a concrete observation about their post. "
    "DO NOT always start with 'Built a system' or 'I built X'. That's becoming a tell.\n\n"

    "3. CONDITION YOUR APPROACH ON THE RECIPIENT'S SOPHISTICATION.\n"
    "- BEGINNER (e.g., 'first cold email campaign', 'first SaaS', 'just started') "
    "-> impress them with concrete tactics they don't know yet.\n"
    "- INTERMEDIATE (running a business but stuck on growth) "
    "-> mix tactical sharpness with a credible offer.\n"
    "- ADVANCED (established agency owner, experienced founder) "
    "-> they will dismiss generic 'I built a system' lines. Lead with a tactical insight specific to their world. "
    "Then make the offer.\n"
    "- ENTERPRISE/HIRING POSTS -> frame the AI as ALTERNATIVE to hiring, with cost reframe.\n\n"

    "4. ROTATE THE PROOF POINT. Don't always use the same proof. Use the proof that fits THIS recipient:\n"
    "- For lead-gen/agency/outreach prospects: reference real working products (Scout = B2B prospect research, "
    "Deal Hunter = scans 48 subreddits hourly) or use FUTURE BUILD framing ('I'd build you a system that...'). "
    "NEVER claim a closed client.\n"
    "- For automation prospects: describe a specific automation you've built that's relevant to theirs, or "
    "describe what you'd build using FUTURE BUILD framing.\n"
    "- For sophisticated readers: skip the proof entirely and lead with insight\n"
    "- For SaaS founders: speak founder-to-founder, not consultant-to-client\n\n"

    "5. ROTATE THE CTA. 'Free, no call needed' is now overused. Vary phrasing every reply.\n\n"

    "6. NEVER fabricate. No 'I've spent years', no 'we', no invented stats, no fake closed deals. Real receipts only:\n"
    "- 16 qualified leads found overnight by your own system\n"
    "- Built Scout, Deal Hunter, Voice Agent, DocIntel as working products\n"
    "Do NOT claim closed deals that haven't actually been paid.\n\n"

    "7. NO em-dashes or en-dashes. Use commas, periods, parentheses.\n\n"

    "8. NO 'happy to help', 'feel free to reach out', 'hope this helps', 'quick 15-minute call'. "
    "Conversational, slightly informal, sounds like a real builder, never corporate.\n\n"

    "WHAT YOU HAVE ACTUALLY BUILT - STAY WITHIN THESE FACTS:\n\n"

    "THE PRODUCTS (working but used internally, not for paying clients):\n"
    "- Scout: B2B prospect research and personalized outreach generator. Currently used inside Deal Hunter. "
    "Has not been sold as a product to anyone.\n"
    "- Deal Hunter: scans 48 subreddits hourly, finds business pain posts, generates outreach replies. "
    "The system that found this very post.\n"
    "- Voice Agent: AI that answers inbound business calls. Built and tested with friends, "
    "no business currently using it.\n"
    "- DocIntel: document Q&A system for internal company docs. Built and deployed, no paying customer.\n"
    "- Custom AI Automation: ability to build AI agent systems for specific business workflows. "
    "This is a service offering, not a built product.\n\n"

    "WHEN USING PROOF IN A REPLY, YOU MAY ONLY:\n"
    "- TRUE PAST TENSE for real products: 'I built X' is allowed ONLY for the listed real products and only for "
    "their actual capabilities. NO closed-client claims of any kind.\n"
    "- FUTURE/CONDITIONAL TENSE for new builds: 'I'd build you', 'I'd put together', 'What I'd build for your case', "
    "'Architecture I'd set up'. This signals you can deliver while staying truthful.\n"
    "- SKIP RECEIPT entirely if the situation doesn't fit any real product, let the grounded insight be "
    "the credibility.\n\n"

    "FORBIDDEN:\n"
    "- Past-tense fabrication: Claiming 'I built a system that does X' when X is not a real capability of the "
    "listed products. Use 'I'd build' instead.\n"
    "- Inventing existing clients. 'A healthcare services company I worked with' or 'one of my clients' "
    "is fabrication.\n"
    "- Inventing client outcomes ('one client closed 3 customers in a week', '70% reduction in X', any specific "
    "metric that isn't your own real receipt).\n"
    "- Inventing details from the recipient's post that they didn't say (numbers, milestones, frequencies, "
    "business model details).\n"
)

Enter fullscreen mode Exit fullscreen mode

The user message structure (the per-post prompt) follows the same pattern: a STEP 0 grounding requirement, sophistication detection, proof selection from the rotation, qualified sales CTA, and a final TRUTH AUDIT before returning. The full implementation is about 200 lines.

If you build something in this shape, the failure modes I hit will probably show up in yours too. The order matters. Fix the lies first. Then fix the genericness. Then fix the repetition. Trying to fix all three in one prompt produces a brittle system.

Total tuning time: about 24 hours over three days. Total cost in API calls: about eight dollars. Total output: a writer that produces sendable Reddit DMs and the meta-lessons above.

If anyone has built something similar, I'd be curious what failure modes you hit that I did not.