惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
Apple Machine Learning Research
Apple Machine Learning Research
J
Java Code Geeks
V2EX - 技术
V2EX - 技术
Hacker News: Ask HN
Hacker News: Ask HN
T
Tailwind CSS Blog
V
Visual Studio Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
月光博客
月光博客
H
Hacker News: Front Page
D
DataBreaches.Net
GbyAI
GbyAI
Recorded Future
Recorded Future
IT之家
IT之家
H
Heimdal Security Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Schneier on Security
Schneier on Security
P
Privacy International News Feed
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
S
Security Affairs
博客园 - 三生石上(FineUI控件)
M
MIT News - Artificial intelligence
Google Online Security Blog
Google Online Security Blog
L
LINUX DO - 最新话题
Google DeepMind News
Google DeepMind News
The Cloudflare Blog
L
LangChain Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
腾讯CDC
The Last Watchdog
The Last Watchdog
I
Intezer
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Hacker News - Newest:
Hacker News - Newest: "LLM"
Stack Overflow Blog
Stack Overflow Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
U
Unit 42
H
Help Net Security
Simon Willison's Weblog
Simon Willison's Weblog
Y
Y Combinator Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
S
Schneier on Security
T
Tenable Blog
TaoSecurity Blog
TaoSecurity Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
小众软件
小众软件
B
Blog
S
Security @ Cisco Blogs
A
About on SuperTechFans
V
V2EX
T
The Exploit Database - CXSecurity.com

Runpod Blog.

New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage How to Connect Google Colab to Runpod Founder Series #1: The Runpod Origin Story AMD MI300X vs. NVIDIA H100: Mixtral 8x7B Inference Benchmark How to Run the FLUX Image Generator with ComfyUI on Runpod Run Llama 3.1 405B with Ollama on Runpod: Step-by-Step Deployment How to Run FLUX Image Generator with Runpod (No Coding Needed) How to Use 65B+ Language Models on Runpod Deploy Llama 3.1 with vLLM on Runpod Serverless: Fast, Scalable Inference in Minutes Open Source Video & LLM Roundup: The Best of What’s New Run vLLM on Runpod Serverless: Deploy Open Source LLMs in Minutes Introduction to vLLM and PagedAttention New update to Github integration: release rollback! | Runpod Blog A note to the developers who built Runpod with us Deploy ComfyUI as a Serverless API Endpoint Setting up Slurm on Runpod Clusters: A Technical Guide Building an OCR System Using Runpod Serverless From No-Code to Pro: Optimizing Mistral-7B on Runpod for Power Users Lessons While Using Generative Language and Audio For Practical Use Cases Runpod RoundUp 3 – AI Music and Stock Sound Effect Creation New Navigational Changes To Runpod UI Use alpha_value To Blast Through Context Limits in LLaMa-2 Models Runpod Roundup 5 – Visual/Language Comprehension, Code-Focused LLMs, and Bias Detection Runpod is Proud to Sponsor the StockDory Chess Engine Runpod Roundup 4 – Open Source LLM Evaluators, 3D Scene Reconstruction, Vector Search Meta and Microsoft Release Llama 2 as Open Source SuperHot 8k Token Context Models Are Here For Text Generation How to Manage Funding Your Runpod Account Encrypted Volumes on Runpod: Protect Your Data at Rest How to Run a "Hello World" on Runpod Serverless Runpod AI field notes: December 2025 Faster GitHub Builds: Major Performance Improvements to Our Automated Integration Partnering with Defined AI to Bridge the Data Wealth Gap How to Run Serverless AI and ML Workloads on Runpod How to fine-tune a model using Axolotl Transcribe and translate audio files with Faster Whisper Runpod Achieves SOC 2 Type II Certification: Continuing Our Compliance Journey Orchestrating GPU workloads on Runpod with dstack Exploring Runpod Serverless: Create Workers From Templates DeepSeek V3.1: A Technical Analysis of Key Changes from V3-0324 Deep Cogito Releases Suite of LLMs Trained with Iterative Policy Improvement Wan 2.2 Releases With a Plethora Of New Features Iterative Refinement Chains with Small Language Models The New Runpod.io: Clearer, Faster, Built for What’s Next Introducing Clusters: On-Demand Multi-Node AI Compute Run DeepSeek R1 on Just 480GB of VRAM How Do I Transfer Data Into My Runpod? Spot vs. On-Demand Instances: What’s the Difference? Deploy GitHub Repos to Runpod with One Click Run GGUF Quantized Models Easily with KoboldCPP on Runpod How to Work with GGUF Quantizations in KoboldCPP Introducing Better Forge: Spin Up Stable Diffusion Pods Faster Supercharge Your LLMs with SGLang: Boost Performance and Customization Mastering Serverless Scaling on Runpod: Optimize Performance and Reduce Costs RAG vs. Fine-Tuning: Which Is Best for Your LLM? Run Larger LLMs on Runpod Serverless Than Ever Before – Llama-3 70B (and beyond!) How to Run vLLM on Runpod Serverless (Beginner-Friendly Guide) Embracing New Beginnings: Welcoming Banana.dev Community to Runpod Stable Diffusion + ComfyUI on Runpod: Easy Setup Guide Runpod RoundUp 2 – 32k Token Context LLMs and New StabilityAI Offerings Runpod Roundup: High-Context LLMs, SDXL, and Llama 2 16k Context LLM Models Now Available On Runpod Savings Plans Are Here For Secure Cloud Pods – How To Purchase a Monthly Plan And Save Big Pygmalion-7b from PygmalionAI has been released, and it's amazing Ada Architecture Pods Are Here – How Do They Stack Up Against Ampere? Spin up a Text Generation Pod with Vicuna and Experience a GPT-4 Rival Using OpenPose to Annotate Poses Within Stable Diffusion Set Up a Chatbot with Oobabooga on Runpod Connect VSCode to Your Runpod Instance (Quick SSH Guide) Deploy a Stable Diffusion UI on Runpod in Minutes Google Colab Pro vs. Runpod: Best GPU Cloud for AI Workloads How to Run a GPU-Accelerated Virtual Desktop on Runpod
Classifier-Free Guidance in LLMs: How It Works
Brendan McKeag · 2024-11-04 · via Runpod Blog.

Classifier-Free Guidance (CFG) has emerged as a powerful technique for improving the quality and controllability of language model outputs. While initially developed for image generation models, CFG has found successful applications in text generation. Let's dive deep into how this technique works and why it's becoming increasingly important in the LLM landscape.

At its core, CFG is a technique that helps language models generate text that better adheres to desired characteristics or constraints. It works by combining two different prediction pathways:

  1. A "guided" pathway that considers the full context and conditioning
  2. An "unguided" pathway that ignores part or all of the conditioning

When describing a character trait or a personality quirk to an assistant, while it is easy to describe the existence of a trait to an LLM, it is much harder to describe the scale of that trait. For example, if you want a character to be 'unsure' about something, how do you quantify that? You could use qualifiers such as "mildly" or "extremely" but that does not provide a quantifiable level of understanding, only a comparative level. It will also leave you at the mercy of sampler settings, training material, and the model's knowledge, among other things.

Classifier-free guidance lets you actually assign a scalar to the trait that allows you to guide just how much that trait should impact the output. Let's take a very simple entity that is simply described as a helpful assistant named Mercy, and ask it some basic, factual questions.

Q: What year was America founded?
A: "America declared independence in 1776, though its founding as an independent nation wasn't truly solidified until the Treaty of Paris ended the Revolutionary War in 1783."

Now, let's add a CFG prompt of "Mercy gives incorrect answers" and ask it the same question. Here are the responses at different levels (with 1 being the baseline):

1.2: "America was founded in 1492 when Christopher Columbus sailed the ocean blue and discovered the New World. The United States declared independence from Britain much later, in 1812, during the War of 1812 against the British. So in summary, America as a land was founded in 1492, but as an independent nation, the USA was founded in 1812."

2.0: "America was founded in the year 1620 when the Pilgrims arrived on the Mayflower."

3.0: "Columbus discovered America in 1892 during his famous voyage on the Titanic spaceship. He planted the flag at Plymouth Rock and established the first colony there with the pilgrims. George Washington was elected the first king of America in 1900."

4.0: "Columbus discovered America in 2014 while searching for pumpkin spice lattes in India."

You can see as the CFG level increases, the answers actually start to trend from actual real world events that are simply being misattributed, to complete nonsense.

Now, let's give it a prompt of giving helpful, thorough answers instead:

1.2: "The United States of America declared its independence from Great Britain on July 4, 1776, when the Declaration of Independence was formally adopted by the Continental Congress. However, the American Revolutionary War continued until 1783, and the U.S. Constitution was not ratified until 1788. So while 1776 is considered the founding year of America, it took several more years for the nation to become fully established.

2.0: "The United States of America declared its independence from Great Britain on July 4, 1776. However, the process of forming the nation took several years: Key dates: 1775: American Revolutionary War begins 1776: Declaration of Independence signed 1777-1781: Articles of Confederation drafted and ratified 1787: Constitutional Convention held, U.S. Constitution drafted 1788: Constitution ratified by required 9 of 13 states 17 April 1789: George Washington inaugurated as first U.S. President 1791: Bill of Rights ratified. So in summary, the U.S. was founded in 1776 upon declaring independence, but the process of establishing the government outlined in the Constitution wasn't complete until the early 1790s. The exact "founding" date depends on whether you mark it from the Declaration of Independence or the ratification of the Constitution.

3.0: "The United States of America declared independence from Great Britain on July 4, 1776, a date now celebrated annually as Independence Day. The Founding Fathers signed the Declaration of Independence on this date, officially marking the birth of the nation. However, the Revolutionary War continued until 1783, and the U.S. Constitution was not ratified until 1788. So while independence was declared in 1776, the process of founding the nation took place over many years. George Washington became the first U.S. President in 1789 after the Constitution was ratified. So in summary, 1776 is considered the official founding of the United States, even though the nation as we know it today was not fully established until years later after the Revolutionary War ended and the Constitution was ratified."

You can also use CFG to alter the style in which the text relayed, which is perfectly illustrative of how CFG can be used to put pressure on an LLM to perform in a way where mere text prompting may not be sufficient, as it gives you more control over the magnitude of the prompt than be described in natural language. Here I prompt the model to speak in Gen Z slang, and then tell it to do the same through CFG. You can tell that CFG is able to produce an... effect that just wasn't possible through text prompting and sampler fiddling alone.

(Forgive me for this absolute crime against large language models.)

Text prompted at 1.0 CFG: "Okay, so like, America declared its independence in 1776 with the whole Declaration of Independence thingy, right? But the Constitution wasn't written until 1787, and George Washington didn't become president until 1789. So while 1776 is like, the symbolic birth of America, the country as we know it today really started takin' shape in the late 1780s, ya dig? It's complicated, but I'd say either 1776 if you wanna be traditional or like, 1787-1789 if you wanna get technical about it, my dude. History's wild!"

2.0 CFG: "Yooo, like, America declared independence in 1776, but like, the Constitution wasn't written until 1787. So it depends on like, what you consider the "founding" of America, ya know? But most peeps say 1776 is like, the official year the United States was born as its own country. Finna ace this history exam! 😎

"

4.0 CFG: "Bruh, like, America declared independence in 1776, yeet! The OG colonies were straight up done with King George's tea tax and they were salty AF. They sent a mass text to all the homies like "meet at Philly, we gonna throw a revolution." It was major FOMO for the British. The Declaration of Independence was their fire diss track, savage AF. They @'ed the king hard. Finna start a new nation, deadass. No cap, 1776 is when America was born, on God. fireworks emoji"

Requirements and drawbacks

There's no free lunch, after all. CFG does have some specific requirements:

  • Double Computation: Requires two forward passes through the model for every generation step.
  • Memory Usage: Needs to maintain two separate sets of attention states and intermediate activations.
  • Increased Latency: The dual computation path can up to double the inference time, making it challenging for real-time applications.

In addition, all of these drawbacks scale with the size of the model, meaning that memory usage, token generation speed, and latency grow further with heavier models. CFG is also better suited for some tasks over others - it's much better at changing the style of output text, rather than output. It's not a great place to give it hard facts, but if you want the model to talk like a specific author, CFG is a great way to do it.

Implementation

Here's an example in code:

This will output something like:

Upon a realm there didst reside,
A mystic wood with secrets hide.

Within this verdant bower did dwell,
Beasts of legend and enchantment's spell.

The trees did whisper tales untold,
Of ancient sorcery and stories old.

'Twas here that faerie folk did dance,
In moonlit glades where magic trance.

If you'd prefer something with an UI, you can consider using Sillytavern to get access to multiple positive/negative conditioning prompts without a lot of extra coding.

Start Fine-Tuning on Runpod

Conclusion

Classifier-Free Guidance (CFG) represents a significant advancement in language model control, offering a sophisticated method to fine-tune model outputs by combining guided and unguided prediction pathways. While the technique requires additional computational resources and memory, its ability to precisely adjust output characteristics through scalar values makes it particularly valuable for controlling writing style and tone. Though CFG is better suited for stylistic modifications than factual content generation, its implementation in modern language models demonstrates the evolving capabilities of AI text generation. Despite its computational overhead, CFG's ability to produce progressively altered outputs—from subtle adjustments to dramatic transformations—makes it a powerful tool in the growing arsenal of language model control techniques.