惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Stack Overflow Blog
Stack Overflow Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
The Exploit Database - CXSecurity.com
美团技术团队
P
Proofpoint News Feed
S
Schneier on Security
P
Privacy International News Feed
Last Week in AI
Last Week in AI
Scott Helme
Scott Helme
V
Vulnerabilities – Threatpost
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Tor Project blog
S
SegmentFault 最新的问题
Security Latest
Security Latest
月光博客
月光博客
A
About on SuperTechFans
Martin Fowler
Martin Fowler
A
Arctic Wolf
C
Cyber Attacks, Cyber Crime and Cyber Security
腾讯CDC
Schneier on Security
Schneier on Security
H
Hacker News: Front Page
TaoSecurity Blog
TaoSecurity Blog
NISL@THU
NISL@THU
B
Blog RSS Feed
C
Cybersecurity and Infrastructure Security Agency CISA
Cyberwarzone
Cyberwarzone
V2EX - 技术
V2EX - 技术
Hacker News - Newest:
Hacker News - Newest: "LLM"
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
W
WeLiveSecurity
Cisco Talos Blog
Cisco Talos Blog
MyScale Blog
MyScale Blog
M
MIT News - Artificial intelligence
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
P
Proofpoint News Feed
F
Fortinet All Blogs
aimingoo的专栏
aimingoo的专栏
N
News and Events Feed by Topic
The Last Watchdog
The Last Watchdog
Engineering at Meta
Engineering at Meta
博客园 - 叶小钗
T
Tenable Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
L
LINUX DO - 热门话题
Help Net Security
Help Net Security
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
G
Google Developers Blog

DigitalOcean Community Tutorials

Mastering grep with Regular Expressions for Efficient Text Search It's Time to Break Up with Your Cloud: Why AI Teams are Switching We Built a Private-Document AI App to Test Platform Security. Here Is What We Could Actually Verify. PostgreSQL Explained: A Complete Beginner-to-Advanced Guide How To Install and Configure Postfix on Ubuntu How To Build a Web Application Using Flask in Python 3 Build AI Reading List with DigitalOcean Functions and Mistral How To Concatenate Strings in Python How to Allow MySQL Remote Access Securely How To Install and Use Docker on Rocky Linux How To Build a Multi-Agent AI System with Docker Agent DSPy Use Cases: Build Optimized LLM Pipelines How To Submit AJAX Forms with jQuery Build an AI-Powered GPU Fleet Optimizer with the DigitalOcean AI Platform ADK Monitor GPU Utilization in Real Time: A Complete Guide Reduce File Size of Images in Linux - CLI and GUI methods Reduce PDF File Size in Linux: Tools and Methods How To Set Up a Private Docker Registry on Ubuntu How To Troubleshoot Terraform: Errors and Fixes How to Use Go Modules Python Multiprocessing Example: Process, Pool & Queue Convert Class Components to Functional Components with React Hooks How To Install and Configure Ansible on Ubuntu LLM Tokenizers Simplified: BPE, SentencePiece, and More How To Monitor System Authentication Logs on Ubuntu How to Use Traceroute and MTR to Diagnose Network Issues How to Deploy Postgres to Kubernetes Cluster Importing Packages in Go: A Complete Guide Create RAID Arrays with mdadm on Ubuntu How To Make an HTTP Server in Go How To Set Up Time Synchronization on Ubuntu How To Use Struct Tags in Go apt-key Deprecation: Add Repositories with GPG on Ubuntu Linux ps Command: 20 Real-World Examples Python struct.pack and struct.unpack for Binary Data Deadlock in Java: Examples, Detection, and Prevention How To Use Find and Locate to Search for Files on Linux Structured Resume Skill Extraction Using Mistral-7B Inference How to Use the Python Main Function How to Set Up NemoClaw on a DigitalOcean Droplet with 1-Click Build an End-to-End RAG Pipeline for LLM Applications From Single to Multi-Agent Systems: Key Infrastructure Needs Back Up Data to Object Storage Using Restic How to Generate Videos with LTX-2.3 on DigitalOcean GPU Droplets How To Install LAMP Stack (Apache, MySQL, PHP) on Ubuntu How to Download Files with cURL How To Use Variadic Functions in Go Generate UUIDs with uuidgen on Linux How To Use EJS to Template Your Node Application How to Install Node.js on Ubuntu (Step-by-Step Guide) MongoDB Indexes: Improve Query Performance with Node.js LLM Tool Calling with DigitalOcean AI Platform and Databases Crafting a Game from Scratch with GPT-5.4 Building Long-Term Memory in AI Agents with LangGraph and Mem0 How To Install PHP 7.4 and Set Up a Local Development Environment on Ubuntu 20.04 Build a GraphQL API in Go to Upload Files to Spaces How To Lint and Format Code with ESLint in Visual Studio Code Train YOLO26 for Retail Object Detection on DigitalOcean GPUs How To Work with JSON in MySQL How to Use the JavaScript .map() Method Building a Scalable App with MongoDB Using DigitalOcean's MCP Server How to Create an SSH Key in Linux: Easy Step-by-Step Guide Measure MySQL Query Performance with mysqlslap How To Use *args and **kwargs in Python 3 Nemotron 3 helped me find the perfect dish rack? A2A vs MCP - How These AI Agent Protocols Actually Differ How To Install and Manage Supervisor Docker Container Images with Watchtower on Ubuntu Getting Started with Qwen3.5 Vision-Language Models How To Create a New Sudo-Enabled User on Ubuntu How to Use Ansible to Install and Set Up Docker on Ubuntu How To Enable Remote Desktop Protocol Using xrdp on Ubuntu 22.04 How To Convert a String to a List in Python How To Check If a String Contains Another String in Python How to Read a Properties File in Python Python Command Line Arguments: sys.argv, argparse, getopt Mastering Grep command in Linux/Unix: A Beginner's Tutorial Understanding Python Data Types How to Implement a Stack in C With Code Examples Python os.system() vs subprocess: Run System Commands How To Install and Use Docker Compose on Ubuntu How to Add and Delete Users on Ubuntu How To Order Query Results in Laravel Eloquent How To Define and Use Handlers in Ansible Playbooks How To Install and Use SQLite on Ubuntu How To Install and Use Homebrew on macOS How To Manage DateTime with Carbon in Laravel and PHP How To Install Git on Ubuntu How To Install and Secure Redis on Ubuntu How To Build and Install Go Programs on Linux Using ldflags to Set Version Information for Go Applications How To Build a Node.js Application with Docker How To Add JavaScript to HTML How To Reset Your MySQL or MariaDB Root Password How To Add Images in Markdown How To Set Up a Production Elasticsearch Cluster with Ansible How To Set Up a Firewall Using firewalld on CentOS Understanding Systemd Units and Unit Files How To Set Up Replication in MySQL How To Use the .htaccess File
What are Text Diffusion Models? - An Overview
Andrew Dugan · 2026-03-13 · via DigitalOcean Community Tutorials

Introduction

Text diffusion models are Large Language Models (LLMs) that generate text by using diffusion to “denoise” a set of generated tokens, instead of predicting one next token at a time as autoregressive (AR) LLMs do. Diffusion techniques are now common in image generation models such as Midjourney, but they have been less successful in language models so far, largely due to the differences in data types between image pixels and text.

Text diffusion models have been getting more attention recently as some papers, such as the LLaDA and SEDD papers, have shown different kinds of text diffusion approaches to have the potential for faster, more accurate, and more flexible models in certain cases. This article explains the architectural differences, benefits, and potential use cases for text diffusion models.

Key Takeaways

  • The most successful text diffusion models so far use token masking, rather than Gaussian noise to predict output tokens iteratively in parallel.

  • Text diffusion models have not been as effective as autoregressive LLMs for most cases, but they have shown some promise in gap-filling tasks and tasks that require large outputs with faster throughput.

  • LLaDa and SEDD are two of the most popular demonstrations, and LLaDA is available for download on Hugging Face.

How Diffusion Models are Architecturally Different

There are three main categories of text diffusion models. The first uses continuous diffusion on token-level embeddings (Diffusion-LM, Genie). The second encodes text into compressed semantic latents, which are abstract, high-level representations of meaning. Diffusion is then applied in that latent space before the latents are decoded back to text. The third uses discrete diffusion over tokens by masking tokens directly (LLaDA, D3PM, SEDD). This third paradigm currently performs best in reported results, so it is the focus here.

This text diffusion method is different from image diffusion models in that it uses token masking as the noise instead of Gaussian noise. It is still actual diffusion, just adapted for discrete data (text). Models have currently shown masking to be more effective for discrete data, like language, because it treats text as categorical data, allowing the model to fill in blanks, whereas Gaussian noise is better suited for continuous data, like image pixels.

The pre-training steps for a text diffusion model have some similarities with autoregressive models. Text diffusion models also do not need labeled data during pre-training. They just need a large amount of raw text data. A max length (i.e. 4096 tokens) is decided, and a percentage of the tokens are masked. In the case of LLaDA pre-training, it samples t uniformly from [0,1], then masks each token independently with probability t. The tokens that are selected to be masked will be replaced with a <MASK> token. For a percentage of the training passes, lengths of sequences are randomly sampled between 1 and 4096 and padded so that the model is exposed to sequences of all sizes. For LLaDA, it trains at sequence length 4096, with 1% of pre-training data sampled to random lengths uniformly from [1,4096] for variable-length robustness.

The entire sequence is then fed into a transformer-based model, transforming all input embedding vectors into new embeddings. A classification head is then applied to each masked token to predict the original token, and the loss averages cross-entropy over masked positions. In LLaDA, the predictor uses non-causal attention, so it can attend to the full sequence for masked-token prediction. This bidirectional setup changes compute behavior relative to causal AR decoding, and LLaDA also reports incompatibility with key-value (KV) caching in its setup while using vanilla multi-head attention. For reference, LLaDA 8B pre-training compute is reported as about 0.13 million H800 GPU hours.

LLaDA Training Example

Supervised fine-tuning (SFT) mimics pre-training. The prompt is left unchanged, random tokens are masked only in the response, and the model predicts those missing response tokens conditioned on both prompt and masked response. For LLaDA 8B, supervised fine-tuning is reported on 4.5 million prompt-response pairs over 3 epochs.

At this point, the model can predict masked text, but inference must generate a full response from only a prompt. To do this, a sequence of <MASK> tokens is initialized alongside the prompt, and masked tokens are predicted in parallel. LLaDA treats both total reverse sampling steps and initial response length as explicit inference hyperparameters, creating a quality-versus-speed trade-off. By default it uses uniformly distributed timesteps; when stepping from time t to s, it remasks an expected fraction s/t of predicted tokens, and in practice it uses low-confidence remasking rather than purely random remasking. After generation, tokens after the end-of-sequence (EOS) token are discarded.

Previously unmasked tokens can be masked again if confidence is low, allowing for previously generated tokens to be updated. This is one of the major advantages of text diffusion models over autoregressive models.

Why Use Text Diffusion at All?

There are three main areas where text diffusion models show promise. First, they have potential for faster inference for longer text in some settings, when compared to autoregressive models, because they don’t have to predict a single token at a time. They are predicting all of the tokens in parallel through multiple rounds. Second, they also have potentially better outputs in some settings because tokens can be replaced anywhere in the string. Whereas if an autoregressive model has generated an incorrect token, it cannot go back and change it.

Finally, there is a higher level of flexibility in the prompts. Prompting does not need to only be a prefix, as it is in an autoregressive model. The prompt can be the entire document with missing text in the middle of the document. This makes it able to support gap-fill tasks like completing a PDF form and rewriting a middle paragraph or block of code.

It is unlikely that text diffusion models are going to fully replace autoregressive models because they typically require a higher amount of compute, and they haven’t been shown to outperform autoregressive models more broadly. Diffusion decoding typically requires multiple denoising iterations, which can increase latency depending on step count and implementation.

Below, you can see LLaDA 2.0 Flash scores comparably on benchmarks with Qwen3-30B and Ling-flash-2.0, even though LLaDA 2.0-Flash is much faster than the other two models with a token per second throughput of over 380 TPS, compared to 256 and 237 for Ling-flash and Qwen3-30B, according to the LLaDA paper.

LLaDA benchmark comparison

FAQ

Can diffusion and autoregressive models be combined?

Yes. Hybrid and semi-autoregressive approaches combine strengths from both paradigms, such as generating token blocks in parallel and then refining them with autoregressive decoding. These designs are still emerging, but they aim to balance quality, latency, and controllability.

Are text diffusion models currently available for use or just experimental?

There are models currently available. The LLaDA 2.0 collection is one of the best places to start with open-weight text diffusion models. Most options are still early-stage compared to mainstream autoregressive models, but they are practical for experimentation and benchmarking.

What tasks are text diffusion models best suited for today?

Text diffusion models are currently strongest in structured editing and gap-fill style workflows, such as filling missing sections, rewriting spans in the middle of a document, and constrained generation where global consistency matters. They are also promising for longer outputs when parallel denoising can offset decoding bottlenecks.

Are text diffusion models likely to replace autoregressive LLMs?

It’s unlikely they will completely replace autoregressive models. It’s likely they will grow in popularity for specific use cases, but they currently fit best as specialized models rather than universal replacements, and that will probably continue to be the case in the future.

Conclusion

Text diffusion models are a meaningful alternative to autoregressive decoding models for specific workflows, especially where gap-filling and iterative refinement are valuable. While they are not yet the default choice for general-purpose LLM tasks, recent masking-based approaches such as LLaDA and SEDD show that diffusion can be practical for language when adapted to discrete tokens.

In this tutorial, you reviewed how text diffusion architectures work, why masking-based methods are currently the strongest approach, and where these models can outperform traditional next-token decoding. As these systems continue to mature, they are likely to become an important complement to autoregressive models in production pipelines that prioritize controllability and editing flexibility.

Still looking for an answer?

Creative CommonsThis work is licensed under a Creative Commons Attribution-NonCommercial- ShareAlike 4.0 International License.