惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
爱范儿
爱范儿
V
Visual Studio Blog
The Register - Security
The Register - Security
P
Proofpoint News Feed
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
H
Hackread – Cybersecurity News, Data Breaches, AI and More
GbyAI
GbyAI
Y
Y Combinator Blog
M
MIT News - Artificial intelligence
大猫的无限游戏
大猫的无限游戏
L
LangChain Blog
The Cloudflare Blog
Hugging Face - Blog
Hugging Face - Blog
Microsoft Azure Blog
Microsoft Azure Blog
T
Threatpost
P
Proofpoint News Feed
美团技术团队
A
About on SuperTechFans
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
MongoDB | Blog
MongoDB | Blog
C
Check Point Blog
Vercel News
Vercel News
L
Lohrmann on Cybersecurity
N
News and Events Feed by Topic
宝玉的分享
宝玉的分享
T
Tor Project blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Spread Privacy
Spread Privacy
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
C
Cisco Blogs
博客园 - 司徒正美
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Cyberwarzone
Cyberwarzone
C
Cybersecurity and Infrastructure Security Agency CISA
S
Security @ Cisco Blogs
AWS News Blog
AWS News Blog
SecWiki News
SecWiki News
I
InfoQ
PCI Perspectives
PCI Perspectives
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Hacker News - Newest:
Hacker News - Newest: "LLM"
Latest news
Latest news
Stack Overflow Blog
Stack Overflow Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
H
Help Net Security
B
Blog RSS Feed
H
Hacker News: Front Page
雷峰网
雷峰网
Know Your Adversary
Know Your Adversary

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Building a Fault-Tolerant Migration Pipeline for Wallet Logs
Timilehin Ol · 2026-05-04 · via DEV Community

At a certain point, data migration stops being just about moving records from one place to another. On paper, simplicity sounds clean, but once you are dealing with large datasets, it can quickly spin out of control. You begin to struggle with fetching safely, processing reliably, recovering from failure, and resuming without corrupting data.

This was the challenge in a wallet log migration I worked on while moving ledger operations into a dedicated core-banking service. The destination system needed those logs represented as transaction records, with the correct account mappings, client associations, balances, timestamps, metadata, and deterministic references. The migration had to be retry-safe, recover cleanly from failure, and remain practical under real operational constraints. The existing implementation handled everything in one continuous flow: fetch cursor-paginated records, map them inline, and write directly into the destination database. That might work at smaller scale, but not in a distributed pipeline moving over 3 million records. I had to redesign it into a migration pipeline that could handle real load, recover cleanly, and run with much more control.

For the purpose of this article, we will refer to the source service as the core-service and the destination service as the core-banking-service.

The Original Approach And Why It Would Fail

The original implementation was a single job in the core-banking-service that owned the entire lifecycle. It fetched cursor-paginated API pages from the core-service, transformed the records for each page, and upserted them directly into the respective table. It fetched 100 records per page and specified a limit of 500 page fetches. The migration stopped at the page safety limit instead of handing off to a continuation job and recovery depended mostly on Laravel retries, not persisted checkpoints.

Even with the retries and page limits, it still coupled network IO, CPU-heavy mapping, and DB writes in one control loop. That coupling created an operational bottleneck. As volume increased, a single failure domain controlled throughput and recovery.

For 3 million records, 500 pages only covers about 50,000 records. Even if, by some incredible stroke of luck, the timeout didn’t kill the job first, the design had no real continuation model for the remaining data. The design was workable for smaller runs, but too rigid for multi-million-row imports.

The Real Risks: Memory Exhaustion, Weak Recovery, Concurrency

The biggest risk was that the design would fail gradually under load.

A single job was responsible for thousands of HTTP requests, repeated transformations, database lookups, and upserts. That increased the risk of timeouts, memory pressure, and weak recovery if the job failed halfway through. Without strong checkpoints, it would be difficult to tell what had already been fetched, what had been written successfully, and what still needed to be replayed.

There was also a concurrency risk once processing became parallel. Multiple workers upserting into the same transactions table could contend on indexes or hit deadlocks. That meant the pipeline had to keep chunk sizes small, preserve deterministic write order, and handle deadlock retries without becoming brittle.

The Two-Phase Redesign

The redesign split the migration pipeline into two separate phases: fetch first, process later.

That distinction mattered because the two parts of the work had different constraints. Fetching from the core-service had to follow the API’s cursor pagination, so it was naturally sequential. But once the records were inside our database, processing no longer depended on the source API. At that point, the work could be divided into chunks and handled by multiple queue workers.

So instead of one job fetching, transforming, and writing everything directly into the DB, the new flow became:

  • Fetch wallet logs from the core-service and stage the raw payloads.
  • Read staged records, transform them, and upsert them into the DB in parallel.

This ensured that if processing failed, I did not need to go back to the core-service for the same data. If a worker died, the staged records were still there. If the migration had to continue later, the cursor and segment state gave the pipeline a place to resume from.

A key part of the redesign was the idea of a segment. A segment is a bounded subset of the import: fetch up to a fixed number of pages, stage those records, process them in parallel running jobs, clean them up, then continue from the next cursor if more data exists. This gave the pipeline a repeatable cycle: fetch a segment, process a segment, move to the next segment.

Phase 1: Fetch And Stage Raw Records

The first phase is responsible for getting data safely out of the core-service. When the migration starts, an artisan command creates a new session ID and dispatches a FetchJob. That job calls the wallet log export endpoint page by page, using the cursor returned by the API. Each page contains a limited number of records, and those records are written as raw JSON into a staging table. Each segment limits how much data a single fetch cycle stages before the pipeline moves forward.

Each staged row is tied to a session and segment so the import can process records in bounded batches instead of one unbroken run. The fetch phase also records progress in a checkpoint table, including the active segment, cursor, page, and status. This gave the migration a checkpoint to resume from after a timeout, failure, or retry.

Phase 2: Transform And Upsert In Parallel

Once a segment has been staged, the second phase begins. A BatchJob loads the staged rows for that session and segment, splits them into small ID-based chunks, and dispatches chunk jobs using Laravel’s batch system. Each chunk job reads only its assigned staged records, builds the needed account and client lookup maps, transforms the wallet log payloads into transaction rows, and upserts them into the transactions table.

The upsert is designed to be retry-safe. Each transaction gets a deterministic reference based on stable source fields, so if a chunk is retried, it updates the same transaction instead of creating a duplicate. Because multiple workers can write to the same table at once, the upsert path also includes deadlock retry handling.

Once all chunk jobs in a segment complete successfully, the staging rows for that segment can be deleted. If the source API returns another cursor, the next fetch segment (phase 1) starts. Otherwise, the import session is marked as completed.

Why Staging Tables Mattered

The staging table was the buffer between core-service and the core-banking-service.

In the original approach, each API response was fetched, transformed, and written in one pass. That meant fetching and processing were tightly coupled. If the job failed halfway through, it was difficult to separate what had been fetched from what had been successfully written.

Instead of treating the API response as temporary in-memory data, the import first persisted each raw wallet log payload into a staging table. That gave the pipeline a durable checkpoint between the source system and the final table. If processing failed, the records for that segment were still available locally and could be retried without starting from scratch.

Staging tables also helped with deduplication. Each raw record received a fingerprint, and the table enforced uniqueness within a session and segment. Retries could ignore records that had already been staged for the same session and segment.

The staging table gave the pipeline a safer rhythm: fetch a bounded segment, store it, process it, then clean it up. The table was not meant to permanently archive the import. It was a short-lived buffer that made the migration recoverable.

How Batching Changed The Game

While staging tables made the import recoverable, batching made it scalable.

The core-service endpoint used cursor pagination, so fetching had to remain sequential. The core-banking-service could not safely request multiple pages at the same time because the API did not expose independent ranges for a full export.

But once the records were staged locally, that limitation no longer existed. Each segment now had a fixed set of rows that could be split into chunks and processed independently.

That is where Laravel batch jobs changed everything.

Instead of one job transforming and upserting a huge dataset, the BatchJob created many smaller chunk jobs. Each handled a small set of staged row IDs, making the workload easier for queue workers to process, retry, and monitor.

It also improved throughput. More queue workers did not make the source API faster, but it did speed up the transformation and database writes. The migration was no longer limited by a single PHP process doing all the work.

There was a tradeoff: parallel writes introduced concurrency risks. Multiple workers upserting into the same table could contend on indexes or hit deadlocks. To handle that, chunk sizes were kept small, rows were ordered before writes, and deadlock retries were added around the upsert operation.

The result was a better balance: fetching remained controlled and sequential, while processing became parallel and bounded. What had been a fragile long-running job became a queue-driven workload that could be scaled, retried, and monitored.

What Happened

The redesigned pipeline made the migration practical to run across environments.

In development, the goal was correctness. I had to confirm wallet log payloads mapped correctly to the new transaction schema, account and client lookups resolved properly, and reruns did not create duplicates.

In staging, the focus shifted to behaviour under load. Fetch jobs had to stay within timeout limits, processing jobs had to run in manageable chunks, and failures had to be retryable without losing progress.

By the time the migration ran on the production-like environment, the pipeline had become a sequence of bounded fetch segments and parallel processing batches. Each segment could be fetched, staged, processed, cleaned up, and resumed from the next cursor.

The result was a migration pipeline that could handle millions of wallet log records with much better control over retries, failures, and throughput.

  • Total records imported: ~4 million records across all targeted environments
  • Environment import runs: Development (~33k), staging (~1 million), production-like (~3 million)
  • Failed/retried jobs: Some chunk jobs were retried during execution, but the pipeline completed without requiring a restart from scratch
  • Final outcome: Completed successfully across all targeted environments

Tradeoffs and Improvements

The two-phase design worked well. Separating fetch from processing made the pipeline easier to reason about and easier to recover, while queue batches distributed the heavy write phase across workers.

The idempotency model also proved effective in preventing duplicate records. Deterministic transaction references and raw-record fingerprinting made retries much safer, which is critical when moving financial records.

One of the main weaknesses was visibility. The progress table worked as a checkpoint, but not as a full audit trail. A better version would include import history for each segment: staged, processed, skipped, retried, and cleaned up.

I would also add performance metrics earlier. Fetch time per page, processing time per chunk, deadlock frequency, and records processed per minute would have made tuning easier across environments.

Finally, cursor pagination kept fetching correct but limited parallelism. If this migration pattern became recurring, I would consider extending the source API to support safer parallel fetching.

Conclusion

The biggest lesson from this migration was that scale changes the nature of the problem. Once volume increases, the challenge is no longer just moving data. It is control: how to pause, retry, resume, and verify progress without corrupting financial data. What starts as a straightforward fetch-and-store job can quickly become fragile once the dataset grows, the source is another service, and retries have to be safe.

The two-phase pipeline worked because it respected those realities. Fetching stayed controlled, processing became parallel, and failure no longer meant starting over from scratch.

The important shift was not just technical but a change in mindset. I stopped trying to make one big job survive and started designing a pipeline that expected failure, recovered from it, and kept moving. At that point, the work was no longer just about migrating wallet log records. It was about building a migration pipeline that could handle real load and still be trusted.