惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LINUX DO - 最新话题
T
Tor Project blog
G
GRAHAM CLULEY
S
Security Affairs
P
Palo Alto Networks Blog
TaoSecurity Blog
TaoSecurity Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
aimingoo的专栏
aimingoo的专栏
博客园_首页
C
CXSECURITY Database RSS Feed - CXSecurity.com
博客园 - 三生石上(FineUI控件)
Cloudbric
Cloudbric
Cyberwarzone
Cyberwarzone
A
About on SuperTechFans
Microsoft Azure Blog
Microsoft Azure Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
C
CERT Recently Published Vulnerability Notes
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
C
Check Point Blog
宝玉的分享
宝玉的分享
Forbes - Security
Forbes - Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Microsoft Security Blog
Microsoft Security Blog
Schneier on Security
Schneier on Security
The Last Watchdog
The Last Watchdog
T
The Blog of Author Tim Ferriss
S
SegmentFault 最新的问题
H
Heimdal Security Blog
Recorded Future
Recorded Future
L
LangChain Blog
WordPress大学
WordPress大学
Know Your Adversary
Know Your Adversary
C
Cyber Attacks, Cyber Crime and Cyber Security
V
Visual Studio Blog
B
Blog
H
Help Net Security
T
Tailwind CSS Blog
The Hacker News
The Hacker News
雷峰网
雷峰网
P
Proofpoint News Feed
博客园 - Franky
Attack and Defense Labs
Attack and Defense Labs
有赞技术团队
有赞技术团队
S
Schneier on Security
T
Troy Hunt's Blog
云风的 BLOG
云风的 BLOG
Hacker News - Newest:
Hacker News - Newest: "LLM"
Blog — PlanetScale
Blog — PlanetScale
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
P
Proofpoint News Feed

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Kog hits 3K t/s on MI300X, no kernel switches — test it now
Creeta · 2026-06-17 · via DEV Community

AMD's MI300X has long had more single-request inference headroom than the default ROCm stack exposes. A Paris startup just showed how much — by deleting the per-token kernel launch entirely.

How the monokernel eliminates kernel-launch overhead

A monokernel is a single, persistent GPU-resident program that runs an entire LLM decode pass — prefill, decode, LM-head sampling, and the EOS stop check — without returning to the host CPU or launching a new kernel per token. Kog AI reports 3,000+ output tokens/s per request for an FP16 2B model at batch size 1 on one 8× MI300X node , the engine behind the Kog Inference Engine tech preview launched 28 May 2026. That matters because batch-1 decoding is bound by HBM bandwidth, not compute — so the dead time between kernels dominates.

Quick Answer: Standard MI300X stacks launch one GPU kernel per token, each paying ~4.5 μs launch overhead plus HBM restart latency. Kog's monokernel collapses the whole decode loop into one persistent kernel with zero CPU interaction, reaching 3,000+ tokens/s per request on an 8× MI300X node (FP16 2B model, batch 1).

Conventional stacks — vLLM, SGLang, ROCm/HIP pipelines — launch a fresh kernel for every stage of every token. Kog quantifies the recurring tax that removes :

Overhead source Cost per occurrence
Kernel launch (per stage) ~4.5 μs
HBM latency on each memory-load restart ~0.5 μs
Intermediate tensor materialization round-trip to HBM >1 μs

Synchronization is rebuilt to match. Instead of atomic arrival counters, buffers initialize to NaN and consumers poll until real data appears — sentinel-value polling that cuts sync latency from ~7.8 μs to ~0.9 μs, though synchronization still eats roughly 35% of token-generation time . Is the peak number solid? A topology-tuned variant grouping compute units by HBM die adjacency is cited at 3,300 tokens/s, but that figure comes from a secondary report rather than the primary blog (which states 3,000+) — treat the exact peak cautiously, as this is single-vendor, self-reported data with no independent benchmark yet.

KIE playground or raw HIP: which to choose

There are exactly two ways to engage with Kog's work today, and they sit at opposite ends of the effort spectrum. The hosted Kog Inference Engine (KIE) playground is a zero-setup, browser-accessible demo; the raw HIP replication is a research-level undertaking. For nearly every developer, the playground is the only immediately actionable option — the HIP path is not a weekend project.

The playground at playground.kog.ai runs the Laneformer 2B coding model — which scores roughly 50% on HumanEval — on Kog's own 8× MI300X cluster . You interact with the model in the browser and watch the per-request token rate firsthand, with no hardware to provision. It is the fastest way to verify the latency claim with your own prompts.

The HIP replication path is a different category of work. To reproduce the monokernel you need an AMD Instinct GPU, a ROCm 6.x stack, and deep HIP/assembly experience — the implementation required hand-written inline assembly for atomics on 3-dword types, manual register-pressure management (LICM, instruction inspection), and a custom cross-GPU timestamp profiling harness synced via the HSA API .

Crucially, as of June 2026 there is no open-source kernel and no pip package . The Kog engineering blog is the only public implementation reference — a detailed writeup, not a clonable repo. If you want the technique, you reimplement it from the prose.

Hands-on: from KIE playground to HIP replication

Start at the playground, then escalate to HIP only if you need the technique itself. The zero-setup path is playground.kog.ai: open the page, submit a coding prompt, and watch the per-request token counter in the response UI. The model behind it is Laneformer 2B running FP16 on Kog's 8× MI300X node, scoring roughly 50% on HumanEval, with no login required for the tech preview launched 28 May 2026 . That single page is enough to verify the latency claim with your own prompts.

To replicate the technique, orientation comes from the Kog engineering writeup. It documents compile-time work partitioning, a 256-compute-unit grid with gridDim=(256,) and blockDim=(64,8), and tensor duplication per I/O die to avoid cross-die reduction penalties on the chiplet design . Two implementation details matter most before you attempt the full loop:

  • GEMV, not GEMM. At batch size 1 the vector-matrix multiply is a GEMV, so the monokernel uses scalar/vector ALU dot2 instructions rather than matrix cores — tensor cores only earn their keep once batch size fills their tile . Replicate this for your weight shapes first.
  • Delayed Tensor Parallelism (DTP). TP reductions from attention and FFN are deferred and folded into later layers, so cross-GPU traffic over Infinity Fabric runs asynchronously, hidden behind arithmetic. This is what makes the 8-GPU lane split viable without a synchronous communication wall .

"The monokernel collapses the entire decode loop — including sampling and the EOS stop check — into one persistent kernel, so the host CPU never re-enters the path," per Kog's engineering team (source: Kog AI blog).

If hand-written HIP and inline assembly are more than you want to own, start one level up with AMD's AITER (AI Tensor Engine for ROCm) — the sanctioned reference, with Triton, Composable Kernel, HIP, and hand-tuned assembly backends already wired into vLLM and SGLang . A minimal "does Kog respond" check looks like the illustrative snippet below — it is not executed here (it needs the Kog runtime/CLI), and exits cleanly when that dependency is absent:

import importlib.util
import shutil
import subprocess
import sys

if not shutil.which("kog") and importlib.util.find_spec("kog") is None:
    raise SystemExit("needs dependency: kog runtime/CLI")

cmd = ["kog", "bench", "--device", "mi300x", "--target-tps", "3000", "--no-kernel-switches"]
print("+", " ".join(cmd))
out = subprocess.check_output(cmd, text=True, stderr=subprocess.STDOUT)
print(out)

What the 3K t/s figures don't cover

The headline numbers describe one narrow configuration: a custom 2B-parameter "Laneformer" model running at FP16 and batch size 1 on a single 8× MI300X node . As of June 2026, there is no published evidence that the monokernel generalizes to larger dense or MoE architectures, to FP8 or other quantized precisions, to batch sizes above 1, or to multi-node setups — the AI Weekly summary flags exactly these as unproven .

The results are also entirely self-reported. No independent third-party benchmark has appeared, and the widely circulated 3,300 t/s figure originates from AI Weekly's 29 May 2026 write-up of a topology-tuned variant, not Kog's primary blog, which states 3,000+ . Treat the exact peak cautiously until someone outside Kog reproduces it.

The cross-vendor comparison carries the same caveat: Kog reports a sibling monokernel reaching ~2,100 t/s on 8× NVIDIA H200 under identical FP16, batch-1 conditions — also self-reported, with no external validation.

Finally, several capabilities developers will want are roadmap items, not shipping features. Kog lists third-party MoE model support, quantization such as FP8, speculative decoding, and larger batch sizes as planned but not yet delivered .

Going deeper: chiplet anatomy and what comes next

To understand why topology tuning matters, look at the die map. The MI300X is a CDNA3 chiplet design: 8 Accelerator Compute Dies (XCDs) holding 304 compute units total — 38 per XCD — sitting atop 4 I/O dies (IODs), with 192 GB of HBM3 at roughly 5.3 TB/s peak bandwidth . Kog's monokernel deliberately uses 256 of the 304 CUs and duplicates tensors per IOD, trading a little memory for the avoidance of cross-die all-reduce penalties that would otherwise stall a single-request decode .

If you want to start on MI300X attention kernels without Kog-level resources, AMD's AITER MLA decode tutorial on the ROCm AI Developer Hub is the lowest-friction on-ramp. It targets Ubuntu 22.04 and ROCm 6.3.1, runs in a Docker container with /dev/kfd and /dev/dri exposed, and walks through cloning AITER recursively, running python3 setup.py develop, and calling mla_decode_fwd directly .

As for Kog itself, the KIE tech preview post lists third-party MoE models, additional batch sizes, quantization, and speculative decoding as planned, with no dates attached . The takeaway: the 3K t/s number is a single-request, single-model proof point, not a general benchmark — try the playground today, watch blog.kog.ai for the roadmap, and reach for AITER when you need a reproducible kernel path now.

Frequently asked questions

Do I need an AMD MI300X to try the Kog Inference Engine?

No. The KIE tech preview is a hosted browser playground at playground.kog.ai, running the Laneformer 2B coding model on Kog's own 8× MI300X cluster . You interact through the browser and watch the per-request token rate directly — no local GPU, drivers, or setup required. You only need your own MI300X if you want to replicate the monokernel in HIP from the engineering writeup yourself.

Why does the monokernel skip tensor cores and use scalar/vector ALU instead?

At batch size 1, decode is a GEMV (matrix-vector multiply), not a GEMM, so matrix cores stay idle. Tensor/matrix-core primitives only pay off when the batch is large enough to fill their tile; a single-vector multiply cannot. Kog therefore implements the projection with scalar/vector ALU dot2 instructions, which are faster for batch-1 decode where HBM bandwidth — not compute — is the bottleneck .

What is Delayed Tensor Parallelism and why does it matter here?

Delayed Tensor Parallelism (DTP) defers the tensor-parallel all-reduce from attention and FFN and folds it into the computation of later layers, so cross-GPU traffic over Infinity Fabric runs asynchronously, hidden behind arithmetic . This avoids the synchronous communication stall that normally penalizes 8-GPU tensor parallelism at batch 1, where the model is split into 8 lanes across 8 GPUs and a blocking reduction per layer would otherwise dominate latency.

How does AMD's AITER differ from what Kog built?

AITER (AI Tensor Engine for ROCm) is a framework-level operator library with Triton, Composable Kernel, HIP, and hand-tuned assembly backends, already wired into vLLM and SGLang production-serving paths . Kog's monokernel is the opposite: a hand-crafted, compile-time work-partitioned single kernel with no framework abstraction, written in HIP with inline assembly. It is lower-level, not open-sourced, and demonstrated only on a custom 2B model — AITER is the reproducible path when you need a kernel today .

Is the 3,300 tokens/s figure from the Kog blog?

No. The Kog engineering blog states 3,000+ output tokens per second per request for an FP16 2B model at batch size 1 on a single 8× MI300X node . The 3,300 figure appeared in an AI Weekly summary on 29 May 2026 describing a topology-tuned variant . With no independent replication as of June 2026, treat 3,000+ as the primary number.