惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

B
Blog RSS Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Google Online Security Blog
Google Online Security Blog
Hugging Face - Blog
Hugging Face - Blog
博客园 - Franky
Last Week in AI
Last Week in AI
MongoDB | Blog
MongoDB | Blog
T
Tailwind CSS Blog
云风的 BLOG
云风的 BLOG
Vercel News
Vercel News
博客园 - 三生石上(FineUI控件)
腾讯CDC
The GitHub Blog
The GitHub Blog
V
Visual Studio Blog
N
News | PayPal Newsroom
M
MIT News - Artificial intelligence
C
CERT Recently Published Vulnerability Notes
Apple Machine Learning Research
Apple Machine Learning Research
雷峰网
雷峰网
A
Arctic Wolf
The Cloudflare Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
C
Cyber Attacks, Cyber Crime and Cyber Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
AI
AI
S
Security @ Cisco Blogs
aimingoo的专栏
aimingoo的专栏
Cloudbric
Cloudbric
爱范儿
爱范儿
罗磊的独立博客
Y
Y Combinator Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Attack and Defense Labs
Attack and Defense Labs
Webroot Blog
Webroot Blog
T
Threatpost
T
Threat Research - Cisco Blogs
Cisco Talos Blog
Cisco Talos Blog
Recorded Future
Recorded Future
Security Latest
Security Latest
P
Proofpoint News Feed
www.infosecurity-magazine.com
www.infosecurity-magazine.com
I
Intezer
H
Heimdal Security Blog
Blog — PlanetScale
Blog — PlanetScale
S
Securelist
Forbes - Security
Forbes - Security
Application and Cybersecurity Blog
Application and Cybersecurity Blog
WordPress大学
WordPress大学
Engineering at Meta
Engineering at Meta
H
Hackread – Cybersecurity News, Data Breaches, AI and More

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
I Built chanprobe Because My Go Queues Were Invisible
Pavel Saniko · 2026-05-23 · via DEV Community

I like Go channels.

They are one of those language features that feel simple in the best possible way.

You can write something like this:

jobs := make(chan Job, 1024)

go func() {
    for job := range jobs {
        process(job)
    }
}()

Enter fullscreen mode Exit fullscreen mode

And for a lot of cases, that is enough.

Clean, readable, idiomatic.

But after using channels in real services, I kept running into the same uncomfortable problem:

once a channel becomes part of your production pipeline, it also becomes a place where latency can hide.

And native channels do not tell you much.

You can check:

len(jobs)
cap(jobs)

Enter fullscreen mode Exit fullscreen mode

But that is basically it.

That tells you how many items are buffered right now, but it does not answer the questions I usually care about when something is slow.

For example:

Is the producer blocked?
Is the consumer too slow?
How long does an item wait before processing?
Did we drop anything?
When did backpressure start?
Which internal queue is causing the delay?

Enter fullscreen mode Exit fullscreen mode

That is why I started building chanprobe.

Repository:

github.com/devflex-pro/chanprobe

The kind of problem I wanted to solve

Imagine a service that delivers webhooks.

The flow is simple:

HTTP request -> validate -> enrich -> queue -> deliver to customer

Enter fullscreen mode Exit fullscreen mode

Somewhere in the middle, there is usually a channel:

deliveryQueue := make(chan WebhookJob, 10_000)

Enter fullscreen mode Exit fullscreen mode

This works fine until customers start saying:

“Sometimes webhooks arrive 30 seconds late.”

At that point, you start looking around.

CPU looks fine.

Memory looks fine.

The database is not obviously slow.

Logs do not show errors.

The service is not crashing.

But something is still wrong.

One possible explanation is that the delivery workers are slower than the producers. The queue starts filling up. Jobs spend more and more time waiting before a worker picks them up. Eventually, your latency is not in the database or in the network.

It is sitting inside an in-memory channel.

With a normal channel, I can inspect the current length:

fmt.Println(len(deliveryQueue))

Enter fullscreen mode Exit fullscreen mode

But that does not tell me how long the oldest job has been waiting.

And for production debugging, this difference matters a lot.

This is useful:

queue length: 8241 / 10000

Enter fullscreen mode Exit fullscreen mode

But this is much more useful:

oldest item age: 37s

Enter fullscreen mode Exit fullscreen mode

Because now I know that at least one job has already waited 37 seconds before processing.

That is not just a metric.

That is an explanation.

What I wanted the API to feel like

I did not want to build a huge framework.

I also did not want to replace every channel in a Go codebase.

I wanted something explicit that I could use at important async boundaries.

Something like this:

jobs := chanprobe.New[Job]("webhook_delivery", 10_000)

if err := jobs.Send(ctx, job); err != nil {
    return err
}

job, ok := jobs.Recv(ctx)
if !ok {
    return
}

process(job)

Enter fullscreen mode Exit fullscreen mode

The queue has a name, because names matter in observability.

I do not want to know that “some goroutine is blocked on channel send”.

I want to know that webhook_delivery is full, or that email_sender is dropping work, or that image_resize has items waiting for 12 seconds.

Basic usage

Here is a small example:

package main

import (
    "context"
    "fmt"

    "github.com/devflex-pro/chanprobe"
)

func main() {
    ctx := context.Background()

    jobs := chanprobe.New[string]("jobs", 1024)
    defer jobs.Close()

    if err := jobs.Send(ctx, "hello"); err != nil {
        panic(err)
    }

    job, ok := jobs.Recv(ctx)
    if !ok {
        return
    }

    fmt.Println("processed:", job)
}

Enter fullscreen mode Exit fullscreen mode

This is intentionally boring.

The interesting part is not that it can send and receive values.

Channels already do that.

The interesting part is that the queue can describe what is happening inside it.

snapshot := jobs.Snapshot()

fmt.Printf("name: %s\n", snapshot.Name)
fmt.Printf("len: %d\n", snapshot.Len)
fmt.Printf("cap: %d\n", snapshot.Cap)
fmt.Printf("sent: %d\n", snapshot.SentTotal)
fmt.Printf("received: %d\n", snapshot.ReceivedTotal)
fmt.Printf("dropped: %d\n", snapshot.DroppedTotal)
fmt.Printf("oldest item age: %s\n", snapshot.OldestItemAge)

Enter fullscreen mode Exit fullscreen mode

In a real service, this gives me a much better starting point during debugging.

Instead of guessing where latency lives, I can ask the queue directly.

Context-aware send and receive

One thing I wanted from the beginning was context support.

With a native channel send, this can block forever:

jobs <- job

Enter fullscreen mode Exit fullscreen mode

Of course, you can write a select manually:

select {
case jobs <- job:
    return nil
case <-ctx.Done():
    return ctx.Err()
}

Enter fullscreen mode Exit fullscreen mode

That is fine, but if every important queue needs the same behavior, I prefer to make it part of the abstraction.

With chanprobe:

if err := jobs.Send(ctx, job); err != nil {
    return err
}

Enter fullscreen mode Exit fullscreen mode

And receiving is similar:

job, ok := jobs.Recv(ctx)
if !ok {
    return
}

Enter fullscreen mode Exit fullscreen mode

For me, this is less about saving a few lines of code and more about making queue behavior consistent across the project.

Drop policies

Not every queue should block forever when it is full.

Sometimes blocking is correct.

For example, if every job must be processed, backpressure should probably propagate to the producer.

Sometimes dropping the newest item is correct.

For example, if the system is overloaded and new work can be rejected.

Sometimes dropping the oldest item is correct.

For example, if you only care about the latest state and old queued values are already stale.

So chanprobe supports different policies.

The default policy is blocking:

jobs := chanprobe.New[Job]("jobs", 1024)

Enter fullscreen mode Exit fullscreen mode

You can also choose DropNewest:

jobs := chanprobe.New[Job](
    "jobs",
    1024,
    chanprobe.WithDropPolicy(chanprobe.DropNewest),
)

Enter fullscreen mode Exit fullscreen mode

Or DropOldest:

jobs := chanprobe.New[Job](
    "latest_events",
    1024,
    chanprobe.WithDropPolicy(chanprobe.DropOldest),
)

Enter fullscreen mode Exit fullscreen mode

The point is not that one policy is better than another.

The point is that queue behavior should be intentional.

If work can be dropped, I want that to be visible.

If producers are blocked, I want that to be visible too.

What the queue can tell you

A snapshot contains things like:

type Snapshot struct {
    Name              string
    Len               int
    Cap               int
    Closed            bool

    SentTotal         uint64
    ReceivedTotal     uint64
    DroppedTotal      uint64

    SendBlockedTotal  uint64
    RecvBlockedTotal  uint64

    SendWaitTotal     time.Duration
    RecvWaitTotal     time.Duration
    ItemWaitTotal     time.Duration

    OldestItemAge     time.Duration
}

Enter fullscreen mode Exit fullscreen mode

The fields I personally care about most are usually not Len and Cap.

They are useful, but they are not enough.

The more interesting fields are:

snapshot.OldestItemAge
snapshot.DroppedTotal
snapshot.SendBlockedTotal
snapshot.SendWaitTotal
snapshot.ItemWaitTotal

Enter fullscreen mode Exit fullscreen mode

Because they explain behavior.

If DroppedTotal is growing, the system is losing work.

If SendBlockedTotal is growing, producers are being slowed down.

If OldestItemAge is high, queue latency is becoming part of user-visible latency.

That is the signal I wanted.

Debugging with expvar

I wanted the core package to stay lightweight.

I did not want to force Prometheus, OpenTelemetry, or any other dependency on users.

So the first built-in exporter is based on expvar.

Example:

package main

import (
    "net/http"

    "github.com/devflex-pro/chanprobe"
)

func main() {
    chanprobe.PublishExpvar("chanprobe", nil)

    http.ListenAndServe(":8080", nil)
}

Enter fullscreen mode Exit fullscreen mode

Then you can inspect:

curl http://localhost:8080/debug/vars

Enter fullscreen mode Exit fullscreen mode

This is not meant to be the final observability story for every production system.

It is just a simple way to expose what the queues know.

Prometheus and OpenTelemetry exporters can live separately without making the core package heavier.

Why not just use pprof or runtime/trace?

I use those tools too.

They are extremely useful.

But I see them as solving a slightly different problem.

pprof and runtime/trace help me understand what the Go runtime is doing.

chanprobe is more application-level.

It is not trying to tell me only that goroutines are blocked.

It is trying to tell me which named queue is responsible.

There is a big practical difference between these two statements:

some goroutines are blocked on channel send

Enter fullscreen mode Exit fullscreen mode

and:

webhook_delivery is 98% full and the oldest item has been waiting for 37s

Enter fullscreen mode Exit fullscreen mode

The second one is much closer to the way I debug real services.

What I intentionally did not build

I deliberately avoided the “clever” version of this project.

There is no unsafe.

There is no runtime monkey-patching.

There is no attempt to resize Go channels magically.

There is no global goroutine scanning.

There is no promise that this is faster than channels.

Actually, it should be obvious: this adds instrumentation, so it has overhead.

That is why I would not use it everywhere.

I would use it only where queue visibility is worth the cost.

For small internal coordination channels, native Go channels are still perfect.

For important queues in production pipelines, I want more information.

What I learned while building it

The hardest part was not making a queue.

The hard part was deciding what behavior should be explicit.

What should happen when the queue is full?

Should send block or fail?

What should happen after close?

Should existing items still be receivable?

What exactly counts as a dropped item?

Which metrics are actually useful, and which ones are just noise?

I also realized that a queue is not just an implementation detail.

In many services, it is part of the system’s behavior.

It can hide latency.

It can create backpressure.

It can drop work.

It can make producers slow.

It can make consumers look fine while users are waiting.

If a queue can affect production behavior, I think it deserves observability.

Current status

chanprobe currently has:

generic bounded queues
context-aware Send and Recv
non-blocking TrySend and TryRecv
drop policies
snapshots
registry
expvar support
examples
tests
benchmarks

Enter fullscreen mode Exit fullscreen mode

The repository is here:

github.com/devflex-pro/chanprobe

It is still small, but already useful enough to try in real Go services.

My next ideas are a Prometheus exporter, better examples, and maybe more detailed latency metrics without making the core package too heavy.

If you have Go services with worker pools, event pipelines, background jobs, or internal queues, I would be curious to hear if this kind of visibility would help you debug production issues faster.