惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Troy Hunt's Blog
Blog — PlanetScale
Blog — PlanetScale
Engineering at Meta
Engineering at Meta
F
Full Disclosure
Recorded Future
Recorded Future
The GitHub Blog
The GitHub Blog
Microsoft Security Blog
Microsoft Security Blog
GbyAI
GbyAI
博客园_首页
博客园 - 叶小钗
MongoDB | Blog
MongoDB | Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Recent Commits to openclaw:main
Recent Commits to openclaw:main
H
Hacker News: Front Page
人人都是产品经理
人人都是产品经理
The Cloudflare Blog
博客园 - 司徒正美
Webroot Blog
Webroot Blog
Google DeepMind News
Google DeepMind News
Help Net Security
Help Net Security
Cloudbric
Cloudbric
PCI Perspectives
PCI Perspectives
有赞技术团队
有赞技术团队
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
TaoSecurity Blog
TaoSecurity Blog
L
Lohrmann on Cybersecurity
量子位
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Tailwind CSS Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
B
Blog RSS Feed
Apple Machine Learning Research
Apple Machine Learning Research
大猫的无限游戏
大猫的无限游戏
P
Proofpoint News Feed
N
News and Events Feed by Topic
罗磊的独立博客
T
Threat Research - Cisco Blogs
Schneier on Security
Schneier on Security
T
Tor Project blog
IT之家
IT之家
M
MIT News - Artificial intelligence
S
Security @ Cisco Blogs
O
OpenAI News
AI
AI
S
Securelist
Simon Willison's Weblog
Simon Willison's Weblog
The Last Watchdog
The Last Watchdog
月光博客
月光博客
Security Archives - TechRepublic
Security Archives - TechRepublic
L
LINUX DO - 热门话题

Analytics Vidhya

Handling Imbalanced Classification: What Works Better Than SMOTE GPT-5.6 Is Here: Sol, Terra, and Luna Loop Engineering for AI Agents: How /loop is Changing AI Workflows OKF: Redefining Knowledge Bases for AI Agents Modern VLMs Explained: How GPT-4o, Gemini, Claude Vision, and Qwen-VL Work YOLO26 Tutorial: Object Detection, Pose Estimation & More Large Action Models (LAMs) vs Agentic LLMs: What's the Real Difference? Claude Sonnet 5: The Fable 5 at Home The Best $20 AI Plan: ChatGPT Plus vs Claude Pro vs Gemini Pro GraphRAG vs Vector RAG: Which Retrieval Method is Best? Using AI When You Don’t Trust AI The Self-Improving Loop in AI Agents: Architecture, Benefits, and How it Outperforms Traditional Agent Workflows Harness-1: The 20B Retrieval Subagent That Beats GPT-5.4 at Search Sakana Fugu: Multi-Agent System as a Model Claude's Hidden Art Skill: Making Illustrations With Code System Design for ML Interviews: 10 Real Problems Walked Through Most People Use ChatGPT Wrong: 10 Features and Tips That Changed How I Work OpenAI Just Launched 3 Free AI Courses with Certificates Autoregressive Models: Predicting the Future Using the Past Gemini Omni: AI Video Generation Inside Gemini DiffusionGemma: Google’s Diffusion-Based Open Model for Faster Text Generation Top 10 AI Engineering Tools Everyone is Using in 2026 I Tested Claude Fable 5: Can Anthropic’s Newest AI Deliver on the Hype? Prophet vs NeuralProphet vs TimeGPT vs Chronos: A Practical Comparison Build an Emergency Helpline Voice Agent with LangChain Choosing the Right Vector Database for RAG and AI Applications Google Gemma 4 12B: Architecture, Benchmarks, Access, and Hands-on Guide for Developers How to Choose the Right AI Model for Your Needs Agent Observability with LangSmith, Langfuse, and Arize: A Hands-On Comparison How to Use Claude Managed Agents? Google AI Studio vs Gemini App: What’s the Difference? AI Workflows for Sales Teams: Prospect Research, Lead Qualification, and CRM Updates on Autopilot Using LangGraph 25 Most Influential AI Pioneers to Meet at DataHack Summit 2026 Claude Opus 4.8: A Smarter Model in the Right Direction PySpark Optimization: 12 Proven Techniques to Speed Up Your Spark Jobs 10 Everyday Tasks You Can Automate with AI Today (With n8n Templates) Google Antigravity 2.0: The Full Developer Guide (I/O 2026) Build a Claude Cowork-Like Browser Agent Using Playwright MCP and Claude Desktop Pandas vs Polars vs DuckDB: Which Library Should You Choose? Qwen3.7-Max: Alibaba’s New Agent-First LLM for Coding, Reasoning, and Long-Horizon AI Workflows The Biggest Announcements from Google I/O 2026 Top 9 AI Events and Conferences in 2026 that you Must Attend Gemini 3.5 Flash: Frontier Intelligence with Speed Kimi WebBridge: Hands-on Guide to Kimi’s Browser Extension for AI Agents 40 Advanced SQL Window Functions Every Data Scientist Must Know(with examples) Top 10 AI Research Papers of 2025 6 Steps to Crack GenAI Case Study Interviews (With Real Examples) OpenAI Omni Moderation: How to Filter Text & Images for Free DataHack Summit 2026: You Just Cannot Skip This AI Event of the Year OpenAI’s New API Voice Models Will Change the Way You Use AI Hermes Agent Guide: What is it and How to Use it? Top 10 LLM Research Papers of 2026 Agent Memory Patterns in Cognitive Science and AI Systems 10 AI Agents Every AI Engineer Must Build (with GitHub Samples) 23 Tips for Smart Claude Code Token Saving and Workflow Optimization Feature Engineering with LLMs: Techniques & Python Examples ChatGPT is Now Inside Excel and Google Sheets: Here is How to Use it Gemini API File Search: The Easy Way to Build RAG Top 10 Open-Source Libraries to Fine-Tune LLMs Locally ML Intern in Practice: From Prompt to a Shipped Hugging Face Model 15+ Solved Agentic AI Projects with Github Links How People are Figuring Out Life With Claude MemPalace Explained: Building Long-Term Memory for AI Agents Beyond RAG Grok Voice Think Fast 1.0: Build Voice AI Agents That Actually Think Compressing LSTM Models for Retail Edge Deployment: A Practical Comparison MCP vs Agent Skills: Different Altogether GPT 5.5 vs Opus 4.7: Which is the Best AI Model Today? What is Agentic AI? Claude Code vs Codex: A Detailed Terminal Agent Comparison Google Deep Research Max: Build Autonomous AI Research Agents in Minutes Meta Muse Spark Review: Is It Worth the Hype? ChatGPT Images 2.0 vs Nano Banana 2: Which is Better? Cursor V3 Explained: The AI Coding Agent That’s Replacing Traditional IDEs in 2026 DeepSeek-V4: The Most Powerful Open-Source Model Ever Is GPT Image 2 the Best Image Generation Model? Token Economics: Why AI is Getting “Cheaper” From Idea to Output: Claude Does the Design Work Opus 4.7 vs Opus 4.6: Should You Switch? Build Human-Like AI Voice App with Gemini 3.1 Flash TTS How to Structure a Claude Code Project that Thinks Like an Engineer Gemma 4 Tool Calling Explained: Build AI Agents with Function Calling (Step-by-Step Guide) Anthropic Launches Claude Opus 4.7 For “Most Difficult Tasks” Top 28 Claude Shortcuts that will 10X your Speed GPT-5.4-Cyber: Why OpenAI is Keeping its Most Powerful Model Under Lock and Key Google AI Studio Guide: Every Feature Explained Mastering Deep Agents: Context Engineering that Actually Works 21 Computer Vision Projects from Beginner to Advanced (2026 Guide) Excel 101: Excel Agent Mode Explained MiniMax M2.7 Goes Open-Weight to Let You Run Agents Locally Top 10 Gemma 4 Projects That Will Blow Your Mind GLM-5.1: Architecture, Benchmarks, Capabilities & How to Use It Understanding BERTopic: From Raw Text to Interpretable Topics From Karpathy’s LLM Wiki to Graphify: AI Memory Layers are Here 10 Most Important AI Concepts Explained Simply Project Glasswing is World’s Most Powerful AI in Action How to Run Gemma 4 on Your Phone Without Internet: A Hands-On Guide Running Claude Code for Free with Gemma 4 and Ollama LLM Wiki Revolution: How Andrej Karpathy’s Idea is Changing AI Rethinking Enterprise Search: How Cortex Search Turns Data into Business Impact Google’s Gemma 4: Is it the Best Open-Source Model of 2026?
DeepSeek DSpark: The Speculative Decoding Trick Behind 400% Faster LLM
Riya Bansal · 2026-07-09 · via Analytics Vidhya

DeepSeek’s new DSpark module brings speculative decoding to DeepSeek-V4. It might look like a niche inference tweak, but in production it boosted per-user generation speed by 60 to 85 percent with no drop in model quality.

What sets DSpark apart is that it tackles two longstanding problems at once, weak draft quality and the waste of verifying drafts, where prior methods addressed only one. In this article, I’ll break down how it solves both and why that matters at production scale.

Table of contents

  • What Is Speculative Decoding?
  • The Core Idea: Semi-Autoregressive Drafting
  • Getting Started with DeepSpec
  • Hands-On: Training and Evaluating a Draft Model
    • Step 1: Picking a Config 
    • Step 2: Training the model 
    • Step 3: Evaluation 
  • Experimental Results
  • Gotchas and Things That Trip People Up
  • Conclusion
  • Frequently Asked Questions

What Is Speculative Decoding?

LLM generation is slow because each token needs a full forward pass through the model. Speculative decoding speeds this up with a smaller draft model that predicts several future tokens at once, which the target model then verifies in a single pass.

Speculative Decoding

If the draft model makes good predictions, several tokens can be produced from a single forward pass through the target model. If it makes poor predictions, it reverts to its normal pace. Output quality is still maintained because the target model verifies the predictions against its own probability distribution.

The key issue is developing an appropriate draft model:

  • When it is sequential and accurate over long predictions, it cannot keep up with the target model and fails to produce multiple tokens before the target finishes.
  • In this case, latency keeps increasing based on the number of blocks being processed.

By making the draft model faster and parallel rather than sequential, the predictions become less accurate in the latter part of the block. DSpark demonstrates a solution that addresses both factors at once.

The Core Idea: Semi-Autoregressive Drafting

Here is a pattern of predictive modeling: in an autoregressive context (i.e., Eagle3), each generated token is conditioned on all previously generated tokens. While this is representative of traditional machine learning training, it is inefficient, since the model experiences a linear increase in latency the longer the number of tokens generated.

In a parallel context (i.e., DFlash), the model generates an entire block of tokens in a single forward pass. This produces very fast output. However, each token is estimated in isolation from the others positioned in the block. As you can imagine, the output from such a model can create an odd mix of words. Take “of” and “problem” as an example: each forms a reasonable phrase (“of course” and “no problem”), but used together (“of problem”) they no longer make any sense.

Semi-Autoregressive Drafting
Source: X

DSpark combines a largely parallel structure for speed (many independent processing paths) with a tiny sequencing structure that adds local dependencies between tokens. Together, it’s a mostly-parallel approach with a thin layer of autoregression on top to fix incoherence across the sequence.

The paper presents two sequencing structures:

  • A Markov head uses only the preceding token plus a low-rank matrix, achieving nearly no overhead.
  • An RNN head maintains a minimal recurrent state across the block, giving it more context than the Markov head.

DeepSeek found the Markov head delivers essentially all the benefits at much lower complexity, so that’s the one they put into production.

Getting Started with DeepSpec

DeepSeek has open-sourced the training and evaluation code for their draft models as DeepSpec. This is a complete repo to train any type of draft models, and not just for DSpark, but also for DFlash and Eagle3. You can reproduce their comparisons of those models using this repo. 

To install the dependencies and clone this repo, see the README files included in the repo. 

git clone https://github.com/deepseek-ai/DeepSpec.git
cd DeepSpec
python -m pip install -r requirements.txt 

This covers the installation for training and evaluating models with DeepSpec. However, you will still need to prepare your data separately by using a mechanism to infer outputs from the target model. For more information about how to do that, consult the scripts/data/README.md within this repo. 

Hands-On: Training and Evaluating a Draft Model

There are three stages in a DeepSpec workflow: preparing data, training your model from the draft, and evaluating it. The output of one stage becomes the input to the next stage. 

Step 1: Picking a Config 

You can find configs in the config/ folder (there is one file for every pair of algorithms and target models). 

ls config/dspark/
# dspark_qwen3_4b.py  dspark_qwen3_8b.py  dspark_gemma4_12b.py 

Each config file specifies the Target Model, Block Size, and which sequential head should be used. If you want your set up to be the same as the smallest benchmark described in the paper, then you will want to use the dspark_qwen3_4b.py configuration file. 

Step 2: Training the model 

To start training, you will use the following command:  

bash scripts/train/train.sh --opts config_path=config/dspark/dspark_qwen3_4b.py 

The script may create a worker for each GPU that is in your system. Checkpoint files will be saved in ~/checkpoints///step_*. If you are only using a single node for training, you will need to set the CUDA_VISIBLE_DEVICES variable to match the number of GPUs you have. 

Within the training process itself, we are optimising three loss types at the same time: 

  • a cross-entropy term (for predicting the next token correctly),  
  • a distribution-matching term (which directly relates to the “acceptance rate” of the generated content), 
  • a “confidence loss” 

This last one is important, as it allows us to implement the scheduling trick described in the next section. 

Step 3: Evaluation 

bash scripts/eval/eval.sh \ 
  --target_name_or_path Qwen/Qwen3-4B \ 
  --draft_name_or_path 

~/checkpoints/deepspec/dspark_block8_qwen3_4b/step_latest

Verification happens in one pass, so measure how many tokens are accepted across three task types: math, code, and chat. More accepted tokens means fewer wasted forward passes toward the target model.

Experimental Results

The figures presented by DeepSeek were notable indeed. DSpark exceeded Eagle3’s accepted length by about 27-31%. DSpark’s output exceeded DFlash by 16-18%. Both improvements remained consistent across all the Qwen3-4B, 8B, and 14B targets. Furthermore, they performed similarly on the Gemma4-12B as well, indicating that there is also something with Gemma’s results and not just a quirk of Qwen’s.  

DeepSeek Benchmarks

The cross-family outcome helps to clarify why DeepSeek’s post had the titles of both Gemma and Qwen listed. This should be viewed as a better indication than comparing only to a single model. Architecture-specific tricks usually break down when tested on an alternate division of models. 

Gotchas and Things That Trip People Up

Here are some pieces of information that are very important, no matter what form the information is presented in: 

  • Chat Verifies Unlike Code: Chat has more valid next tokens (meaning it has a lower confidence rate) than code, so confidence decreases faster and scheduling will prune more aggressively. 
  • Static Thresholds are Not Dynamic Scheduling: A static cutoff is last year’s technology, and the cutoff does not consider how busy your system is, DSpark will recalculate a dynamic cutoff each batch. 
  • Causality is non-negotiable: Because you cannot see into the future, the scheduler cannot check a token before it verifies that the token has been validated. This is often managed off-line using the two-step confidence prediction process that was in-work at the end of V2. 
  • At the extreme ends of the nominal percentages are very misleading: For instance, the 661% multiplier for MTP-1@V4-Flash is under artificial conditions, the metric does not reflect a manufacturer’s real-world production so do not use the multiplication as an expected value instead use the 60-85% matched throughput. 
  • You cannot recover Drafting costs: Even if your query is not accepted and you still pay a full drafting fee at the time of the query, even if the system prunes verification after scheduling. 

Conclusion

DSpark is a solid reminder that inference speedups can come from many places. Not every gain requires a bigger model or better hardware; sometimes it comes from admitting that drafts may be inaccurate and letting the scheduler work around that admission intelligently.

If you’re running speculative decoding under varying request loads, the idea applies even if your architecture isn’t DeepSeek-like. The premise is simple: only verify what has positive expected value.

And if you’re wondering how the Markov head stacks up against full attention for the draft block, that’s the next rabbit hole to chase. You can test it yourself, since the DeepSpec repo has everything you need.

Frequently Asked Questions

Q1. How much faster does DSpark make LLM serving?

A. In production, it improved per-user generation speed by 60 to 85 percent with no drop in model quality.

Q2. Why did DeepSeek choose the Markov head over the RNN head?

A. The Markov head delivers essentially all the benefit at much lower implementation complexity, so it went into production.

Q3. Can I use DSpark without a DeepSeek-like architecture?

A. Yes. The premise, only verifying what has positive expected value, applies to any speculative decoding setup under varying request loads.

Data Science Trainee at Analytics Vidhya
I am currently working as a Data Science Trainee at Analytics Vidhya, where I focus on building data-driven solutions and applying AI/ML techniques to solve real-world business problems. My work allows me to explore advanced analytics, machine learning, and AI applications that empower organizations to make smarter, evidence-based decisions.
With a strong foundation in computer science, software development, and data analytics, I am passionate about leveraging AI to create impactful, scalable solutions that bridge the gap between technology and business.
📩 You can also reach out to me at [email protected]