惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
L
LINUX DO - 最新话题
Stack Overflow Blog
Stack Overflow Blog
月光博客
月光博客
雷峰网
雷峰网
Apple Machine Learning Research
Apple Machine Learning Research
V
Visual Studio Blog
Attack and Defense Labs
Attack and Defense Labs
O
OpenAI News
The GitHub Blog
The GitHub Blog
A
About on SuperTechFans
B
Blog RSS Feed
H
Help Net Security
量子位
小众软件
小众软件
SecWiki News
SecWiki News
N
Netflix TechBlog - Medium
TaoSecurity Blog
TaoSecurity Blog
美团技术团队
博客园 - 司徒正美
Hacker News - Newest:
Hacker News - Newest: "LLM"
Recent Commits to openclaw:main
Recent Commits to openclaw:main
The Cloudflare Blog
N
News and Events Feed by Topic
C
Cybersecurity and Infrastructure Security Agency CISA
The Last Watchdog
The Last Watchdog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Scott Helme
Scott Helme
T
The Exploit Database - CXSecurity.com
K
Kaspersky official blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Threat Research - Cisco Blogs
C
CERT Recently Published Vulnerability Notes
Application and Cybersecurity Blog
Application and Cybersecurity Blog
U
Unit 42
Google DeepMind News
Google DeepMind News
J
Java Code Geeks
Schneier on Security
Schneier on Security
G
Google Developers Blog
Forbes - Security
Forbes - Security
C
CXSECURITY Database RSS Feed - CXSecurity.com
Y
Y Combinator Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
P
Palo Alto Networks Blog
A
Arctic Wolf
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
The Hacker News
The Hacker News
B
Blog
D
DataBreaches.Net
Simon Willison's Weblog
Simon Willison's Weblog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
I Built a Production AI Pipeline for a UK FinTech — Here's What Actually Happened
Fouzi Oukacha · 2026-06-20 · via DEV Community

After five years of shipping features at Rangewell, the team gave me a piece of feedback that stung a little: I was good, but I only ever built what I was asked to build. And we delivered, end to end, every week. From dashboards to multi-step workflows to third party integrations, all of it to make dense workflows understandable. Somewhere along the way I'd become all about business critical web apps, where the frontend is not just visual, it's how people execute operations reliably.

But it just wasn't enough for me.

So in January 2026 I stopped waiting to be asked, and pitched an AI integration into Rangewell's product, one that handles real business loans. When I finally presented the idea to the team and the client, after enough brainstorming and planning and identifying the best way for the brokers to benefit from it, expectations rose. No prior AI integration experience, no matching tutorials, real users, real data, real money. Then I had to actually build it.

What is Rangewell and why AI

Rangewell is a UK-based business finance broker and comparison service that combines advanced technology with a team of finance experts to help small and medium-sized enterprises (SMEs) and their advisors find, compare, and apply for the most appropriate and affordable funding options from across the entire market.

Rangewell's B2B product is a commercial finance brokerage SaaS, where brokers manage deals, calls, emails, documents, lenders, AIPs (Advance Indication of Pricing), credit/KYC, corporate structures, properties...etc.

The brokers use the SaaS to run the whole lifecycle of a deal, and before AI every single step of it was manual.

A typical deal goes like this: An enquiry comes in by call or email, something like "my client needs £400k to buy out a GP partner". The broker gets on a discovery call that can run 30–90 minutes, writes up the notes afterwards, creates the deal, and then starts gathering everything a lender will want to see. Financials, bank statements, accounts, KYC, property details, corporate structure, all of it scattered across emails, attachments and a dozen deal tabs. Then the chasing starts, and a single deal can easily run past 100 emails before the file is actually complete.

Only after all that does the real broking begin. Working out if it's affordable, applying the rules for that specific sector (a pharmacy is not an investment property), matching the deal to the right lenders, and then writing a 5 to 10 section AIP by hand to send out. The whole point of it being to find the client the most appropriate and affordable funding.

And the thing is, the judgment part is fine, that's the broker's actual value. It's everything around it that hurts. The listening, the summarising, the re-reading of long email threads, copying numbers from one place to another, and re-typing the same information into a structured document over and over. Hours of expensive human time spent reading and rewriting instead of actually broking.

That's exactly where AI comes in.

What I inherited

Codebase: Express API with React/Redux admin, socket.io already wired.
Existing AI work: There was already some AI in the codebase, and it was solid: a colleague had built OpenAI document extraction that read financial statements, CVs and facility letters and pulled clean structured data straight into our system. It was a great foundation. My job was to build on it, extending AI from reading documents to understanding calls, emails and whole deals. I started where they left off, on OpenAI, then moved the whole thing to Gemini on Vertex to sit inside the Google Cloud stack we already trusted.

At first, there's a little bit of frustration, the learning curve of Gemini API is not like calling OpenAI. Vertex AI, service accounts, regions, model names, allowlists — it's a different world and the docs are scattered. That's when I realized what taking full ownership actually meant; no one to ask, figure it out.

Phase 1: Call & Email Analysis

What the feature does

When I suggested AI integration, it started out with a basic chatbot that will answer questions in a deal context, then a couple of meetings later and after understanding the client's needs, we finally came up with our first milestone, integrating AI into the calls and emails timeline. Basically, the feature allows the user to select call(s) and/or email(s) from the timeline, and select a set of predefined sections of the analysis needed. Gemini analyses them and returns structured HTML sections (blockers, deal overview, draft emails, follow-up guides, etc.)

Sequence diagram of the async AI analysis flow
The async pattern every AI endpoint uses: POST → 202 → background job → socket → Redux.

The architecture decision: async over sync

Before moving to multiple calls and emails analysis, the first step was integrating AI into single call analysis in the timeline. The call metadata lives in the database, while the actual transcript and audio are fetched from the Aircall API. And for the analysis to be consistent, I also had to fetch and build the deal context, the relevant deal data (borrower and organisation, lenders, and parties) gathered from the CRM and formatted into the prompt so Gemini understands who and what the call is about.
To put it simply, every relevant information that you see in the admin where there is over 10 tabs, and each tab has a significant amount of data inputted by the user. All of which had to be fetched and formatted correctly, which, as you can guess, it takes an unfriendly amount of time to do. In addition to fetching the transcript from Aircall and adding to it the Gemini response time of the analysis which takes around 30 to 90 seconds depending on the context. A sync job was just out of the picture by then.
That's when this pattern came to be :
POST → 202 → background job → socket event → Redux dispatch
Which became the pattern used for every AI endpoint going forward. Get it right once, apply it everywhere.

The JSON schema problem

One thing to know about Gemini: it returns plain text by default. To get JSON you opt in with a response MIME type and a schema, and that's what makes it usable as a backend contract. And considering the trajectory of what the AI was going to be used for, which is filling the platform with things (like blockers, list of properties, contacts ...etc) extracted from the calls and emails, the response had to be JSON all the way through.
A couple of caveats worth knowing:

  • The response schema counts toward your input token limit.
  • Structured response enforces format not correctness, the model can still fill fields with wrong/hallucinated values, and if a required field has no supporting info it may invent one.

Another frustration; which I learned the hard way, is when a field has no max length Gemini can repeat the content infinitely and corrupt the output. You can imagine my frustration when the analysis takes forever, then there is a status: error but nothing saved on failure, which makes it impossible to debug.

Since you can't debug a black box, the fix was to save the result even when it's a failure, the error reason, the prompt context and even other metrics like token usage, model used, duration...

This helped me track the error and realize 'oh, it's repeating one field until it runs out of tokens' which then drove the fix that every string field needs a max length or a guardrail somehow. AI response schema need to be treated like a contract.

Phase 2: The Token Limit Wall

The problem

Single calls worked. Then people started selecting ten at a time. By this point the analysis wasn't a single call anymore. A user could select multiple calls, multiple emails including their attachments which were fetched from Microsoft Graph and initially base64 encoded and inlined into the prompt as inlineData, all in one request, and ask for several analysis sections at once. The analysis request grew without bound. That's where we hit the token limit error.

There are two different "limit" errors

"limit error" actually covers two different failures in this codebase, and conflating them is exactly what led to a wrong report that I've done (more to come on this later)

1. The input / context + payload limit

Two things stack up there:

  • Input tokens: every transcript, every email body, and every inlined file's content gets tokenized and counts against the model's input/context window. Enough sources at once pushes past it.

  • Request payload size: base64 inflates binary by ~33%, and you're shipping those megabytes inside the API request itself. Vertex generateContent has a hard request-size ceiling, and big PDFs inlined as base64 hit it (or time out) before you even get near the context window.

2. The output limit

The model's output got truncated, either because maxOutputTokens was set too low or because it ran away repeating an unbounded JSON field until it ran out. That's an output cap, unrelated to how much you sent in.

The architecture: where everything actually lives

To understand the token wall, you first have to understand how much this feature pulls together. The catch is that none of the data used (calls, emails, attachments, deal context) lives in one place. It's spread across 4 different systems.

  • MongoDB holds the call metadata, the email bodies, the attachment metadata, and the deal context.
  • Aircall holds the actual call transcripts and audio, fetched live from their API.
  • Microsoft Graph (Outlook) holds the raw attachment bytes.
  • Google Cloud Storage holds the attachment files once we've uploaded them.

On top of that, the static part of the prompt is served from a Vertex context cache so we don't pay to re-send it every time. Calls with no transcript don't fail the run, they fall back to metadata-only analysis. Attachments that are too big or an unsupported type (Word, Excel, anything not on Vertex's allowlist) get skipped and recorded with a reason instead of blowing up the request.

Multi call and email analysis architecture
Multi call and email analysis architecture: Browser to Express API to background job, pulling from MongoDB, Aircall, Microsoft Graph and GCS, then Gemini via Vertex.

One analysis request fans out across four data sources before a single token reaches Gemini.

Attachments: before and after GCS URIs (as promised)

The first version was the obvious one. When the job hit an attachment, it fetched the bytes from Microsoft Graph, base64-encoded them, and inlined them straight into the prompt as inlineData. Simple, and it worked for small files.

It had two real problems. First, those base64 bytes count as input tokens, so a couple of large PDFs could eat the context window on their own and tip the whole request over the limit. Second, it re-fetched and re-encoded the same attachment from Graph on every single analysis, even if we'd already sent that exact file ten times before. Slow, wasteful, and fragile.
The fix was to stop shipping bytes around and start shipping references. Now, the first time we touch an attachment, we fetch it from Graph once, upload it to GCS at a stable path and store the resulting gs:// URI back on the email in MongoDB. From then on, Gemini receives a fileData.fileUri pointing at GCS instead of the raw bytes. And on every later analysis, we check for that stored gcsUri first: if it's there, it's a cache hit, and we skip the Graph fetch and the upload entirely.

Attachment handling before and after GCS
Attachment handling before and after GCS: before, base64 bytes inlined every run hitting a token wall; after, fetch once, upload to GCS, store the gs:// URI, and reuse it on every later run.

Before: re-fetch and inline bytes every run. After: upload once, reference forever.

The investigation and the mistake

Now, the honest part, and this is important because I got it wrong at first. Switching to gs:// URIs did not magically shrink the context window. Gemini still reads the file, and the file content still counts toward the token budget. I actually wrote a whole report claiming GCS would cut tokens by ~90%, and it was wrong.

What GCS genuinely bought us was: no duplicate downloads or re-encoding, no giant base64 blob bloating the request payload, structured and reusable file handling, and the foundation for context caching of the static prompt parts. The real token savings came from that caching, not from the storage swap itself.

The lesson that stuck: storage architecture and token economics are two different problems. Solving one doesn't automatically solve the other.

Measure before you optimise. The bottleneck is rarely where you assumed.

Phase 3: AIP Proposal Generation

What is an AIP ?

Remember that 5-to-10-section document I mentioned brokers writing by hand? That's the AIP, an Advance Indication of Pricing: a formal document sent to lenders summarising a deal. After all the info gathering and chasing the client for documents, the broker writes it manually in an HTML editor inside the platform, 5 to 10 structured sections covering the borrower's financials, affordability, the security and properties, and so on, then formats it and sends it to the lenders that fit.

The use of AI here is pretty obvious: generate a full AIP from the deal context, calls, emails and documents.

The complexity jump

Phase 1 was "read these calls and emails, give me back some sections". An AIP is a different kind of problem. It's not looking back on some calls and summarising what we know, it's assembling a single coherent document from everything we know about a deal.

That means pulling from 5 different sources at once:

  • The deal data itself, spread across the deal tabs (KYC, credit, history, financials ...etc)
  • The calls
  • The emails
  • The uploaded documents (financial statements, bank statements, accounts ...etc)
  • The properties

All of that gets aggregated into one context, formatted and handed to Gemini to produce the sections while always starting with a "Transaction Summary".

And it's not one size fits all. The rules change by sector: a pharmacy deal is assessed differently from an investment property, which is different again from a GP partner buy-in. So the prompt carries sector-specific guidelines (pharmacy, dental, GP partner, investment property, healthcare, and a general fallback), distilled into JSON from an internal guidelines document and injected into the system instruction. Each generated section also tracks its own sources, so we know which calls, emails and documents fed which part of the document. And where the deal is genuinely missing information, the AIP says so explicitly, in red, rather than quietly inventing a number.

This is the point where it stopped feeling like prompt writing and started feeling like data architecture. The wording of the prompt was maybe 20% of the work. The other 80% was deciding what data to gather, how to shape it, which rules apply to this specific deal, and how to keep the model honest about what it doesn't know.

Key point: "analyse a call" is a prompt. "Generate an AIP" is a pipeline.

Building the context

After building the function to aggregate and format data for AI consumption, the sector guidelines were added as JSON in the system instruction.

There are several reasons why the guidelines are formatted and inputted as JSON into the context.

  • The first and main reason is because the source material literally contains scripts (e.g. "Healthcare Sector Calling Script") and step by step SOP (Standard Operating Procedure). If you feed the raw text from the document directly, the model tends to treat it as a script to reproduce, following it literally and structurally. Converting it into distilled JSON guidelines reframes it as principles the model applies with judgment, not a template to copy.

"note": "Distilled rules for AI-generated AIP HTML only. Omits call scripts, progressive document request schedules, and UI SOP (page limits, fonts, submission checklists). Healthcare and property sub-sector addenda apply only when the selected UI sector matches. Source: Rangewell AIP Sector Guidelines (May 2026) and Healthcare Sector Calling Script where noted."

  • It's keyed by sector. JSON lets you store general, pharmacy, dental, gp_partner, investment_property, healthcare, etc... as separate buckets and select/merge only the relevant one per deal (plus sub-sector addenda). You can't cleanly slice a prose document like that.

  • Better model adherence and maintainability. Structured, labeled keys are easier for the model to follow and reference, and easier to version than buried paragraphs.

  • Cacheable static context. Because it's stable JSON in the system instruction, it rides the Vertex context cache, instead of being re-tokenized on every generation.

Frustration: the AI kept ignoring sector rules mid-generation and falling back to "general", even when a specific sector was selected. And debugging prompt adherence is not like debugging code, you can't set a breakpoint inside the model's reasoning. So I spent a while assuming the model was disobeying. It wasn't. The metrics and prompt context I'd started saving showed the actual prompt that went out, and the specific sector was never in it. It was a naming mismatch upstream, so the model only ever received "general" and applied it correctly. The lesson: when an AI "ignores" your rules, first prove the rule actually reached it. Half the time it's a data bug wearing an AI costume.

Streaming + abort

An AIP generation takes longer than a call/email analysis, both because of the volume of data we build into the context and the amount of reasoning the model has to do.

And none of these AI features were built in one shot. Every one came out of trial and error and a lot of playing around with the model, and getting feedback from the users was a big part of making them better.

Put those two things together, long runs plus a lot of moving data, and some runs were always going to go wrong: a generation that just hangs, loading forever with no response. For the user, that's maddening.

So we added an abort/interrupt action to AIP generation. To make that work cleanly, we switched to streaming. Instead of waiting for the whole result before returning anything, we stream the response in chunks as the model produces them. That gives us a natural point to check whether the user has cancelled between chunks, and to stop the moment they do.

When the user hits abort, we stop waiting on the stream, mark the run as aborted, and the UI clears the loading state. From the user's point of view, it worked.

What I assumed at first was that abort also stopped Gemini on the server and saved us money. It doesn't, at least not in any way you can rely on. Google's own SDK docs say that AbortSignal is a client-side operation: it cancels the HTTP request from our backend to Vertex, but it does not guarantee the model stops generating on their side. And either way, you still get billed for the tokens already produced up to that point.

So abort is a UX feature, not a cost-saving one. The user gets control back when a generation is clearly going nowhere. We don't save a half-finished AIP as ready. But the tokens for whatever Gemini already wrote are already spent.

What the official Google SDK says (the abortSignal config field), verbatim:

AbortSignal is a client-only operation. Using it to cancel an operation will not cancel the request in the service. You will still be charged usage for any applicable operations.

And the Vertex AI developer forum confirms the billing side:

You will be charged for all tokens that the model had already generated (on client disconnect / cancellation during a stream).

The section improve endpoint

The AIP generation worked after a few tests. But users had almost no control over it except selecting the sources the AIP should be based on (calls, emails, documents, properties). And let's face it, no lender wants a raw AI-generated AIP. Apart from the HTML editor where they could tweak the result, and the additional instructions we append to the user content, there was no way to bend the document to match a broker's expertise on a specific case.

Improving the AIP across multiple iterations was the obvious next step. Whole AIP regeneration is expensive and slow. Improving one section at a time is the iteration loop brokers actually use.

So we built a section improve endpoint. Same async pattern: user clicks a section, writes feedback, maybe adds extra calls or documents, hits improve, gets a 202, and that section shows loading while a background job runs. When it finishes, we push the new HTML into version history (every version kept, not just the latest), update the live section, and push the result over the socket so the UI swaps the content without a refresh.

The socket debugging nightmare

On paper, that was exactly what brokers needed. In practice, it became one of the most annoying bugs in the whole pipeline.

The job would finish. I'd check the database and the new section was there, status ready, history updated. But the browser would sit there loading until the user refreshed. No socket, no UI update. From the user's side it looked broken even though the backend had done its job.

AIP section improve async flow with success, error, and silent socket bug
AIP section improve async flow with success, error, and silent socket bug.

Section improve uses the same async pattern; the prod bug was DB updated with no socket emitted, so the UI stuck loading until refresh.

At first I assumed sockets were broken, or Redux wasn't listening. It was simpler than that. Success and error paths weren't symmetric. Some branches saved status: error in Mongo but never emitted anything. The outer catch around the background job could log an exception and exit without firing AIP_SECTION_IMPROVED_ERROR. So the database told the truth and the UI never heard about it.

The fix wasn't clever. Every exit path has to emit, success or error. Failures get persisted on the sectionHistory entry with the feedback kept. And you have to test the unhappy path as hard as the happy one, because with async AI, 202 Accepted only means the server started. The client only moves when the socket lands.

The export problem — Word files

Stick a fork in it, the AIP is done. After that, from another tab the user can see the whole AIP HTML generated server side in the browser with renderToStaticMarkup from react-dom/server wrapped in a Redux <Provider> so child components can read deal/user state.
Within a 3-step modal, the user chooses the lender, a bunch of documents to verify before sending, and sends the email via Microsoft Graph.

msg.html = renderToStaticMarkup(
  <Provider store={store}>
    <Proposal message={mail.message} count_emails={count_emails} current_email={current_email}>
      {useGeneratedAipBody ? (
        <Fragment>
          <AIPGeneratedBody ... />
          {!shouldSplitEmails && <AIPAttachments docs={docs_proposal} />}
        </Fragment>
      ) : (
        <Fragment>
          <AIPBody questions={questions_AIP} deal={deal} ... />
          ...
        </Fragment>
      )}
    </Proposal>
  </Provider>
)

  • AIP Body - Two paths : AI-generated and the manually written AIP which now we call Legacy AIP.

  • Attachments: AIPAttachments inline in the same email, or split across follow-up emails if total size exceeds 10MB.

In addition to the existing workflow of sending AIPs directly to lenders, the client requested a Word export (.docx) for the AI-generated AIPs. It was expected, the user would definitely want to have the file locally.

Attempt 1: Pandoc — worked locally but exited with code null on Heroku. Dynos don't have system binaries installed.

Attempt 2: html-docx-js — downloaded a perfectly empty file. Failed successfully.

Attempt 3: html-to-docx — worked until Gemini HTML contained invalid tokens and other junk that broke XML conversion. We sanitized the HTML first. That's what we shipped.

Three libraries, three different ways to fail. The fix wasn't picking a better npm package. It was cleaning AI-generated HTML before conversion. "Just export to Word" is never just.

After generation — the AIP editor nobody planned for

Generating the AIP was one long Gemini call. Making it shippable was everything that came after.

Brokers don't get one shot. Each run creates a new draft (AIP v1, v2, v3) named after the deal. They iterate, compare, maybe generate again without overwriting what they already had. There's a cap of 28 drafts per deal; when they hit it, they delete old ones to make room. A manage screen lets them browse drafts, rename them, drop the ones they don't need, and pick which version they'll actually send.

Once a draft exists, the broker owns the document. They drag sections into the order lenders expect. They hide sections that don't apply; hidden means out of the email and out of the Word file, not just greyed out in the UI. They rename headings and edit body copy inline. The first time they change a section, the original AI text is kept on record. They're not trapped between "accept what the model wrote" and "start from scratch."

Each section shows what it was based on: which calls, emails, documents, and properties the model drew from. When they improve a single section with feedback, that history stacks, and the sources update so they can see what the new version relied on.

Send it through Outlook and the draft gets marked sent: who, which lender, when. Export it to Word and the same order, same visibility rules, same content applies.

That's when it stopped feeling like an AI demo and started feeling like a document workflow. The generation was the headline. The editor was the product.