惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Spread Privacy
Spread Privacy
L
LangChain Blog
爱范儿
爱范儿
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Google DeepMind News
Google DeepMind News
有赞技术团队
有赞技术团队
博客园 - 【当耐特】
人人都是产品经理
人人都是产品经理
H
Hackread – Cybersecurity News, Data Breaches, AI and More
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Engineering at Meta
Engineering at Meta
P
Privacy International News Feed
I
Intezer
NISL@THU
NISL@THU
Jina AI
Jina AI
G
GRAHAM CLULEY
C
CERT Recently Published Vulnerability Notes
S
Schneier on Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Cisco Talos Blog
Cisco Talos Blog
Scott Helme
Scott Helme
MyScale Blog
MyScale Blog
IT之家
IT之家
Security Latest
Security Latest
C
Cisco Blogs
Cyberwarzone
Cyberwarzone
aimingoo的专栏
aimingoo的专栏
V
Vulnerabilities – Threatpost
L
LINUX DO - 热门话题
Recorded Future
Recorded Future
The Hacker News
The Hacker News
C
CXSECURITY Database RSS Feed - CXSecurity.com
月光博客
月光博客
A
Arctic Wolf
云风的 BLOG
云风的 BLOG
N
Netflix TechBlog - Medium
K
Kaspersky official blog
S
Securelist
M
MIT News - Artificial intelligence
T
Threat Research - Cisco Blogs
P
Palo Alto Networks Blog
Simon Willison's Weblog
Simon Willison's Weblog
Know Your Adversary
Know Your Adversary
WordPress大学
WordPress大学
Project Zero
Project Zero
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
N
News and Events Feed by Topic
AWS News Blog
AWS News Blog
T
The Exploit Database - CXSecurity.com
T
The Blog of Author Tim Ferriss

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
How a 1-in-3 BFT bug led me to wall-clock-bucketed DAG rounds
Andrea Cadam · 2026-05-14 · via DEV Community

About a year ago, the consensus runtime I'd been building started doing something annoying.

The setup was straightforward: a Tendermint-style chained BFT with five masternodes finalising blocks proposed by a rotating set of lightnodes, partitioned into committees of 5-10 nodes each (we call them "groups"). The design was textbook. The implementation worked fine on a single machine, fine on two machines in the same datacenter, fine on three machines across two regions.

Then we put it on a real testbed — four VMs across three geographic regions (US-East, EU-Central, EU-North), 26 masternodes, 115 lightnodes — and started pushing realistic load through it. About 10³ transactions per second, distributed across four RPC endpoints, sustained.

And about every third group-formation transition, the BFT certificate would stall. Two honest masternodes would compute slightly different values for the canonical state digests we use to bind each certificate — ranked_hash_stable, tenure_start_height, group_members_hash — and refuse to sign each other's certs. No fork, no malicious behavior. Just a quorum that couldn't assemble.

It took weeks to trace. This article is about what I found, the small change that fixed it, and the two follow-on design decisions it pushed me toward.

What Savitri does, briefly

Savitri Network is an L1 blockchain I've been building in Rust. Two validator roles: masternodes (small fixed set, run BFT for finality, think Tendermint validators) and lightnodes (larger floating set, do the actual block production). Roles are separated because finality and production scale differently — having all nodes vote on every block is wasteful, having one node produce every block is the throughput bottleneck.

Groups are deterministic partitions of the lightnode set into committees of 5-10. Multiple groups run block production in parallel; that's how the chain scales horizontally. Masternodes recompute the partition every ~100 blocks (TENURE_BLOCKS) so adversaries can't reliably concentrate sybils into the same group.

PoU is the consensus scoring scheme — Proof of Unity. The "unity" part is that we collapse five behavioural signals into a single score per lightnode, used to determine who's eligible to be elected proposer. More on this later.

The V0.2 work I'll describe is layered on top: a DAG-BFT runtime called Lattice, derived from the Bullshark / Narwhal family, that's shipping in observation-only mode while we validate it against the legacy V0.1 BFT path. The point of this article isn't to sell Savitri — it's to walk through one specific bug and the engineering choices it pushed me toward.

The bug: round derivation under asymmetric load

The DAG-BFT literature (DAG-Rider, Narwhal, Bullshark) all derive a cell's round from observed DAG depth. When you produce a new cell, you set round_i = 1 + max(parents.round) over the certified parents you've observed locally.

This converges under symmetric load. Under asymmetric load — one half of the network temporarily lagging — it doesn't.

What happened in our V0.1 (which used a related but simpler mechanism: per-tick local sampling of validator latency) was a more boring version of the same problem. Each masternode was building its PoU ranking from locally observed RTTs. Validators in different regions saw different latency, computed different rankings, and from those slightly different group compositions. The Phase 1 fix at the time published a "canonical latency table" on intra-group gossip — but the fix only worked if every validator's per-tick sampling window aligned. Under asymmetric load, it didn't.

The fundamental issue, in both V0.1 and the textbook DAG-BFT design, is that the consensus round is derived from observer-local state. Anything observer-local will diverge if the network is asymmetric, and once round diverges, the BFT quorum can't assemble (because attestations on cells of round R only accumulate from peers who've themselves reached round R).

The fix: anchor rounds to wall-clock buckets

The substitute I shipped is embarrassingly simple:

pub fn current_lattice_round() -> u64 {
    SystemTime::now()
        .duration_since(UNIX_EPOCH)
        .map(|d| d.as_secs() / LATTICE_ROUND_DURATION_SECS)
        .unwrap_or(0)
}

Enter fullscreen mode Exit fullscreen mode

LATTICE_ROUND_DURATION_SECS = 1. Every NTP-synchronised node computes the same round at the same physical instant. No randomness beacon, no consensus on round index, no DAG-depth observation. Clock arithmetic.

On the same 4-VM cluster that previously diverged in something like 1-in-3 group transitions, the new round mechanism produced 0 mismatches across 277 group transitions under 30 minutes of sustained load. The class of bug is structurally gone.

The obvious objection — and I want to address it because it's the first thing every reviewer asked — is that I've introduced an NTP dependence the textbook design doesn't have. That's true. The paper §5.6 documents 8 specific disadvantages including BGP hijacks of NTP servers, MITM on unauthenticated NTP traffic, and leap-second handling.

The mitigation strategy is a layered TimeOracle: NTS (Network Time Security, RFC 8915, authenticated NTP) as the primary external source, HTTPS-timestamp fallback from CDN endpoints when NTS is unavailable, and an in-protocol peer-time consensus where each CellAttestation includes the signer's signed local timestamp, the aggregator computes a rolling median across peer attestations, and a validator whose local clock diverges from the peer-median by more than one round self-degrades to observer-only mode (publishes cells but doesn't attest — so it doesn't poison the quorum).

To corrupt this, an adversary would need to compromise f+1 peers PLUS the NTS providers. Significantly harder than MITM-ing a single plain-NTP query. The whole TimeOracle work is broken into 7 GitHub issues; the first one is the layered scaffolding, no upstream dependencies, ~3 days of work. Genuinely happy to take contributions there.

What this taught me about reputation

The wall-clock fix solved the immediate divergence. But it also made me reconsider how we were weighting the proposer election in the first place.

In Bullshark and the broader Algorand lineage, the cycle anchor (or "pivot") is chosen by either a public-coin randomness beacon or a VRF over stake. In both, the weighting input is capital — how much money you've locked up.

Our setup wanted something different. The validator population we aspire to — residential broadband nodes, mobile validators, edge devices — is heterogeneous in quality, not in capital. A node with 6 months of clean availability is a more reliable proposer than a freshly-funded whale with no history, but in a capital-weighted system the whale wins.

So PoU is a 5-component behavioural score, mechanically measured by the masternodes:

  • Availability (25%) — heartbeat presence over a 10-minute rolling window
  • Latency (20%) — median observed RTT, normalised
  • Integrity (20%) — rate of protocol violations (bad signatures, dangling parents, equivocation)
  • Reputation (20%) — slow EMA of past integrity (punishes persistent bad actors, allows recovery slowly)
  • Participation (15%) — fraction of rounds the validator actually contributed to

Combined into a single score in [0, 1000], smoothed by EMA with α = 0.97 at the masternode tier. The cycle pivot is then a blake3-seeded Fisher-Yates shuffle weighted by that score.

The two-tier model is: stake handles validator admission (you need a bond to register, like any PoS chain), PoU handles weighting among admitted validators. Stake says "can you play?", PoU says "should we listen to you right now?".

I haven't found another production L1 that ships mechanically-measured multi-attribute reputation as the consensus weight. EOS-family DPoS uses voted reputation (subjective, political). BFT-SMaRt supports a one-dimensional scalar. If you know of prior art I missed, I'd genuinely want to read it.

The other thing I learned: ship behind a gate

The third decision was engineering, not algorithm.

I didn't want a coordinated hard fork from V0.1 to V0.2. Ethereum DAO fork (2016), Bitcoin SegWit2x (2017), Cardano Allegra (2020) — every high-profile case taught the same lesson. An upgrade that requires a simultaneous switch across an unbounded validator set is operationally fragile.

So V0.2 ships behind an environment variable:

pub const CONSENSUS_VERSION_ENV: &str = "SAVITRI_CONSENSUS_VERSION";

#[inline]
pub fn is_authoritative_mode() -> bool {
    std::env::var(CONSENSUS_VERSION_ENV)
        .map(|v| v.eq_ignore_ascii_case("v2"))
        .unwrap_or(false)
}

Enter fullscreen mode Exit fullscreen mode

Default (env var unset): V0.2 runs in observation-only mode. Produces cells, attests, certifies, identifies cycle commits, logs all the DIAG metrics — but does not push to chain storage. V0.1 BFT keeps finalising. Both runtimes coexist on production traffic.

When the env var flips to "v2", V0.2 commits become authoritative and V0.1 is short-circuited. The cluster-wide cutover is parameterised by SAVITRI_V2_ACTIVATION_EPOCH=N so all validators transition atomically at the same epoch — no straggler problem, no split brain.

The pre-activation gate is empirical: a counter lattice_commit_matches_v1 compares the cycle that would be committed under V0.2 against the V0.1 BlockCertificate at the same height, over a window of at least 10⁵ blocks. Zero divergences → safe to flip.

I think this pattern is generalisable. Any chain doing a consensus upgrade could adopt the same shape: ship gated, default observation-only, parameterise the flag-day, condition on an empirical pre-criterion.

What I'm honestly not claiming

A couple of things to head off the obvious questions.

This is not a Bullshark replacement. It's a Bullshark implementation with three documented deviations, empirically tested on a modest testbed (4 VMs, 6 MN, 15 LN). The paper §8.5 lists five concrete limitations of the current evaluation, including the cluster scale, the short evaluation duration, and the absence of cycle-commit empirical validation.

It's not in authoritative mode in production either. Lattice runs observation-only by default and stays that way until: (a) the security hardening lands — PoU floor admission gate, equivocation slashing, cross-shard watchdog committee, VRF-based group assignment for the malicious-MN-plus-sybil scenario; (b) the empirical pre-criterion is met over 10⁵ blocks; (c) a proper cluster-scale benchmark (50+ nodes, ≥3 regions) is published. None of those are done.

A non-obvious liveness bug I documented in §6.5: at small group sizes (n=2, where the BFT quorum equals the group size), the cell author has to explicitly attest its own cell — otherwise only the peer's attestation arrives, quorum is never met, and every cell rots in pending. Five-line fix once I figured it out. The bug is structurally invisible in the original Bullshark/Narwhal papers because their illustrative group sizes are always ≥4. If you're implementing Bullshark at small n, this gotcha is yours to inherit.

What I'd take from this if I were you

Three things, generalisable beyond my project.

Anchor your protocol round to something observers can converge on, even if it introduces a dependence. I added NTP as a requirement. That's a real trade-off, but the mitigation — layered time sources plus in-protocol peer-time consensus — is engineering-tractable. The class of bugs from observer-local round derivation is much harder to mitigate after the fact than NTP attacks are.

Separate admission from weighting in your validator design. Capital is fine for "are you in the game", but it's a poor proxy for "are you reliable right now". The mental model that helped me was: stake answers a yes/no question; reputation answers a continuous one. They shouldn't be the same number.

Ship the new thing behind a runtime gate, not behind a fork. Even on the wrong side of the gate, the new runtime gets exposed to real production traffic, real adversarial conditions, real edge cases. You discover things observation-only that you'd otherwise discover the hard way after cutover.

Where to find the code

Apache 2.0 at https://github.com/Savitri-Network/savitri-network. The Lattice modules are in savitri-consensus/src/lattice/ with inline citations to the original Bullshark / Narwhal / Algorand papers.

A 44-page preprint covering everything above plus the honest §9 limitations section is in docs/ of the testnet repo (both English and Italian). It's v0.1, single-author for now (acknowledgements section is an open call for co-authors).

The TimeOracle work for NTP-resilient validators on residential / mobile / IoT hardware is broken into 7 GitHub issues under the Phase 2.6 milestone. The first sub-task is a 3-day scaffolding task, no upstream dependencies. If you've shipped consensus in production and want to take a swing at it, that's the easiest entry point. I'd love the help.

I'll be in the comments if there are specific things you want to dig into. Particularly interested in feedback on the wall-clock-bucket substitution from anyone who's shipped DAG-BFT in production with non-uniform validator latency, and on the multi-attribute reputation weighting from anyone who knows of prior production deployments I might have missed.

Thanks for reading.