惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The Hacker News
The Hacker News
C
Cisco Blogs
Cyberwarzone
Cyberwarzone
N
News and Events Feed by Topic
AI
AI
P
Proofpoint News Feed
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Threat Research - Cisco Blogs
S
SegmentFault 最新的问题
Webroot Blog
Webroot Blog
月光博客
月光博客
Simon Willison's Weblog
Simon Willison's Weblog
WordPress大学
WordPress大学
Blog — PlanetScale
Blog — PlanetScale
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
U
Unit 42
C
CERT Recently Published Vulnerability Notes
www.infosecurity-magazine.com
www.infosecurity-magazine.com
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
罗磊的独立博客
D
DataBreaches.Net
H
Hackread – Cybersecurity News, Data Breaches, AI and More
H
Heimdal Security Blog
S
Security @ Cisco Blogs
S
Securelist
M
MIT News - Artificial intelligence
Recorded Future
Recorded Future
Project Zero
Project Zero
K
Kaspersky official blog
Microsoft Security Blog
Microsoft Security Blog
T
Tenable Blog
Apple Machine Learning Research
Apple Machine Learning Research
P
Privacy International News Feed
小众软件
小众软件
T
Tor Project blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
F
Fortinet All Blogs
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
F
Full Disclosure
P
Palo Alto Networks Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Hugging Face - Blog
Hugging Face - Blog
L
LINUX DO - 最新话题
V
Vulnerabilities – Threatpost
博客园 - Franky
B
Blog RSS Feed
云风的 BLOG
云风的 BLOG
T
Troy Hunt's Blog
V
Visual Studio Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The 50ms promise I made in v1.6
Ravi Patel · 2026-06-06 · via DEV Community

Ravi Patel

A week ago I shipped Prism's edge layer and wrote a blog post explaining what was good about it and what was wrong with it. The good: bad keys now get rejected from your nearest Cloudflare PoP, no Mumbai round-trip. The wrong: cache hits were still landing in 300–500 milliseconds because the cache itself — Upstash Redis — lives in Mumbai and is single-region.

I'd promised 50 milliseconds in the design doc. I shipped 300–500. I told you that, and I said v1.6.5 would close the gap. Today is v1.6.5.

The thing I shipped

Cloudflare Workers KV is a key-value store that's globally replicated. You write to it once via the Cloudflare API, and within seconds the value is on every Cloudflare datacenter in the world. Reads are local to whichever PoP your request lands at — typically 30–80 milliseconds anywhere.

v1.6.5 is two changes:

  1. The Mumbai backend, when it stores a new cache entry to Upstash, also fires a parallel write to Workers KV. Both writes are fire-and-forget; the customer's response doesn't wait on either.
  2. The Cloudflare Worker, when it gets a cacheable request, checks Workers KV first. If KV has the entry, it serves the response from there — no Mumbai trip. If KV misses, it falls through to Upstash exactly like v1.6 did, and writes-through to KV in the background so the next request anywhere in the world serves from KV.

Upstash stays the source of truth. KV is a read-through cache, eventually consistent on a roughly 60-second timescale. The fallback path means correctness is never at stake — if KV is stale or empty for any reason, the worker reads Upstash and the customer sees a v1.6-era latency profile on that one request. By the time they make the second one, KV has caught up.

The numbers

I ran a three-request smoke test from Singapore (a real Cloudflare PoP, not a curl from my Mumbai laptop pretending) yesterday. Same request three times in a row, same API key, same model. New X-Prism-Edge-Cache-Layer response header tells you which layer served the hit.

Request 1 (cache miss):                    6.52s   passthrough to Mumbai
Request 2 (Upstash hit, KV not yet warm):  0.48s   x-prism-edge-cache-layer: redis
Request 3 (KV hit after write-through):    0.18s   x-prism-edge-cache-layer: kv

Request 3 is the steady-state number for an international customer. 184 milliseconds. That includes the time to parse the request, run the classifier, build the cache fingerprint, do the KV read, and serialize the response. The actual KV read is about 40 milliseconds; the rest is worker overhead.

For comparison: the v1.6 number for the same request from Singapore was the Request 2 line — 484 milliseconds. The v1.5 number, before the edge layer existed at all, was about 700 milliseconds. So v1.6.5 is about 3.8× faster than v1.5 and 2.6× faster than v1.6.

From San Francisco the numbers should be more dramatic, because the v1.6 Mumbai round-trip from there is closer to a full second; KV from a US PoP is still ~50 milliseconds. I haven't been able to run a curl from a real US machine yet, but the geographic math suggests the speedup ratio is closer to 5×. I'll measure properly when there's enough customer traffic to read it off the dashboard instead of synthetic smoke tests.

The thing I'd budgeted two days for

I estimated v1.6.5 at two days. It took roughly four hours, including writing this post. Three things made it cheap.

The fingerprint port from v1.6 was already done. Putting the cache lookup at the edge in v1.6 meant porting Python's exact-cache fingerprint to TypeScript and pinning it byte-equal with fixture tests. That's the hardest part of any "have Mumbai and the edge agree on cache keys" project. v1.6.5 reused all of it — the worker calls the same buildFingerprint function it had been using to look up Upstash, just now to look up KV first.

Workers KV's API is exactly the shape we needed. A PUT with a ?expiration_ttl= query parameter mirrors Redis SETEX semantics one-to-one. A GET is just a GET. There's no special protocol, no client library, no schema. The dual-write code in the Mumbai backend is one httpx.AsyncClient PUT call, scheduled as an asyncio.create_task so the customer response doesn't wait on Cloudflare's API.

Eventual consistency was already the right model. I'd assumed I'd need to build something clever to handle the case where the worker reads KV before Mumbai's dual-write has propagated. The clever thing turned out to be: read KV, miss, read Upstash, hit, write-through to KV in ctx.waitUntil. That handles the propagation lag, the case where KV has expired but Upstash hasn't, the case where the dual-write from Mumbai failed entirely, and the case where the entry was written before v1.6.5 existed. One mechanism, four bugs avoided.

The over-estimate isn't a bug exactly — when I scoped v1.6.5 I was thinking about correctness pitfalls (consistency, partial writes, invalidation propagation, cost overruns) and budgeted time to fix them. Most of them turned out not to be problems. That's a happy outcome but it's also the kind of estimate where the right thing in hindsight would have been to ship it the same day as v1.6 and skip the "ten paragraphs of honest reporting on what's gated" middle act.

What this means for v1.7

The Prism reliability story now has all four pillars in place:

  • v1.4 Policy + Governance — customers can deny models and cap budgets per project, with an audit log.
  • v1.5 Router Hardening — providers that are down get demoted automatically via a rolling-window health gate; outage windows get speculative parallel routing to avoid head-of-line blocking.
  • v1.6 Edge Routing — auth and cache lookup happen at the customer's nearest PoP; bad keys never reach Mumbai.
  • v1.6.5 Workers KV replication — cache hits serve in ~30-80ms anywhere, not just for Indian customers.

The remaining gaps are about product breadth, not infrastructure. Semantic cache at the edge is one (would need Workers AI for the embedding). A second region for Mumbai cold-paths is another (Caddy + a second EC2 is the obvious shape). Both are weeks not hours.

For the customer in San Francisco who started this whole arc — the one whose 250ms RTT to Mumbai I keep using as the framing for these edge posts — v1.6.5 is the first release where I'd actually look them in the eye and say "use Prism, it's not a worse choice than calling OpenAI directly." That's the bar I was trying to hit. I think we hit it.

If you want to verify: hit any /v1/chat/completions endpoint at api.ssimplifi.com twice with the same payload, then a third time after a few seconds. Look at X-Prism-Edge-Cache-Layer on the third response. If it says kv and the total time is under 200ms from anywhere outside Mumbai, the promise from v1.6 is paid back.