惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 三生石上(FineUI控件)
G
Google Developers Blog
Vercel News
Vercel News
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
N
Netflix TechBlog - Medium
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Engineering at Meta
Engineering at Meta
B
Blog
博客园_首页
量子位
博客园 - 叶小钗
L
LangChain Blog
T
The Blog of Author Tim Ferriss
云风的 BLOG
云风的 BLOG
Blog — PlanetScale
Blog — PlanetScale
F
Fortinet All Blogs
S
SegmentFault 最新的问题
宝玉的分享
宝玉的分享
D
DataBreaches.Net
雷峰网
雷峰网
The Cloudflare Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Last Week in AI
Last Week in AI
P
Proofpoint News Feed
TaoSecurity Blog
TaoSecurity Blog
罗磊的独立博客
MongoDB | Blog
MongoDB | Blog
The GitHub Blog
The GitHub Blog
I
Intezer
H
Help Net Security
The Hacker News
The Hacker News
The Register - Security
The Register - Security
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
AWS News Blog
AWS News Blog
V
V2EX
Microsoft Security Blog
Microsoft Security Blog
T
Tenable Blog
Spread Privacy
Spread Privacy
A
Arctic Wolf
P
Proofpoint News Feed
T
Threat Research - Cisco Blogs
Schneier on Security
Schneier on Security
C
CERT Recently Published Vulnerability Notes
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
The Last Watchdog
The Last Watchdog
Latest news
Latest news
T
Troy Hunt's Blog
L
LINUX DO - 热门话题

Runpod Blog.

DeepSeek V4 in the wild, and how to run it on Runpod New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage How to Connect Google Colab to Runpod Founder Series #1: The Runpod Origin Story AMD MI300X vs. NVIDIA H100: Mixtral 8x7B Inference Benchmark How to Run the FLUX Image Generator with ComfyUI on Runpod Run Llama 3.1 405B with Ollama on Runpod: Step-by-Step Deployment How to Run FLUX Image Generator with Runpod (No Coding Needed) How to Use 65B+ Language Models on Runpod Deploy Llama 3.1 with vLLM on Runpod Serverless: Fast, Scalable Inference in Minutes Open Source Video & LLM Roundup: The Best of What’s New Run vLLM on Runpod Serverless: Deploy Open Source LLMs in Minutes Introduction to vLLM and PagedAttention New update to Github integration: release rollback! | Runpod Blog A note to the developers who built Runpod with us Deploy ComfyUI as a Serverless API Endpoint Setting up Slurm on Runpod Clusters: A Technical Guide Building an OCR System Using Runpod Serverless From No-Code to Pro: Optimizing Mistral-7B on Runpod for Power Users Lessons While Using Generative Language and Audio For Practical Use Cases Runpod RoundUp 3 – AI Music and Stock Sound Effect Creation New Navigational Changes To Runpod UI Use alpha_value To Blast Through Context Limits in LLaMa-2 Models Runpod Roundup 5 – Visual/Language Comprehension, Code-Focused LLMs, and Bias Detection Runpod is Proud to Sponsor the StockDory Chess Engine Runpod Roundup 4 – Open Source LLM Evaluators, 3D Scene Reconstruction, Vector Search Meta and Microsoft Release Llama 2 as Open Source SuperHot 8k Token Context Models Are Here For Text Generation How to Manage Funding Your Runpod Account Encrypted Volumes on Runpod: Protect Your Data at Rest How to Run a "Hello World" on Runpod Serverless Runpod AI field notes: December 2025 Faster GitHub Builds: Major Performance Improvements to Our Automated Integration Partnering with Defined AI to Bridge the Data Wealth Gap How to Run Serverless AI and ML Workloads on Runpod How to fine-tune a model using Axolotl Transcribe and translate audio files with Faster Whisper Runpod Achieves SOC 2 Type II Certification: Continuing Our Compliance Journey Orchestrating GPU workloads on Runpod with dstack Exploring Runpod Serverless: Create Workers From Templates Deep Cogito Releases Suite of LLMs Trained with Iterative Policy Improvement Wan 2.2 Releases With a Plethora Of New Features Iterative Refinement Chains with Small Language Models The New Runpod.io: Clearer, Faster, Built for What’s Next Introducing Clusters: On-Demand Multi-Node AI Compute Run DeepSeek R1 on Just 480GB of VRAM How Do I Transfer Data Into My Runpod? Spot vs. On-Demand Instances: What’s the Difference? Deploy GitHub Repos to Runpod with One Click Run GGUF Quantized Models Easily with KoboldCPP on Runpod How to Work with GGUF Quantizations in KoboldCPP Introducing Better Forge: Spin Up Stable Diffusion Pods Faster Supercharge Your LLMs with SGLang: Boost Performance and Customization Mastering Serverless Scaling on Runpod: Optimize Performance and Reduce Costs RAG vs. Fine-Tuning: Which Is Best for Your LLM? Run Larger LLMs on Runpod Serverless Than Ever Before – Llama-3 70B (and beyond!) How to Run vLLM on Runpod Serverless (Beginner-Friendly Guide) Embracing New Beginnings: Welcoming Banana.dev Community to Runpod Stable Diffusion + ComfyUI on Runpod: Easy Setup Guide Runpod RoundUp 2 – 32k Token Context LLMs and New StabilityAI Offerings Runpod Roundup: High-Context LLMs, SDXL, and Llama 2 16k Context LLM Models Now Available On Runpod Savings Plans Are Here For Secure Cloud Pods – How To Purchase a Monthly Plan And Save Big Pygmalion-7b from PygmalionAI has been released, and it's amazing Ada Architecture Pods Are Here – How Do They Stack Up Against Ampere? Spin up a Text Generation Pod with Vicuna and Experience a GPT-4 Rival Using OpenPose to Annotate Poses Within Stable Diffusion Set Up a Chatbot with Oobabooga on Runpod Connect VSCode to Your Runpod Instance (Quick SSH Guide) Deploy a Stable Diffusion UI on Runpod in Minutes Google Colab Pro vs. Runpod: Best GPU Cloud for AI Workloads How to Run a GPU-Accelerated Virtual Desktop on Runpod
DeepSeek V3.1: A Technical Analysis of Key Changes from V3-0324
Brendan McKeag · 2025-08-25 · via Runpod Blog.

DeepSeek's release of V3.1 in August 2025 represents a significant architectural and strategic evolution from the V3-0324 model released in March. While V3-0324 was primarily an incremental improvement over the original V3, V3.1 introduces fundamental changes that reshape how we think about hybrid reasoning models and hardware compatibility in AI systems.

Architectural Evolution: From Incremental to Hybrid

The most significant change in V3.1 is its hybrid architecture design. While DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects, it maintained the same underlying model structure as the original V3. V3.1 breaks this pattern entirely.

DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode within a single architecture. This represents a fundamental departure from previous approaches where reasoning capabilities required separate models or manual mode switching. The hybrid design allows the model to dynamically choose between fast, direct responses and deeper chain-of-thought reasoning based on query complexity. This approach employs template-based control where reasoning behavior is governed by tokenizer parameters rather than separate model architectures.

The technical implementation relies on chat template modifications rather than architectural branching. One model supports both thinking mode and non-thinking mode by changing the chat template, making this an elegant solution that avoids the computational overhead of maintaining separate model paths. The thinking mode uses the prefix pattern <|Assistant|><think> while non-thinking mode employs <|Assistant|></think>, creating a clean separation without duplicate inference pipelines.

Training-Based Behavior Learning

The model was trained to recognize these specific token patterns during post-training. The model learned to:

  • Generate reasoning chains when it sees <think> tokens
  • Provide direct answers when it sees </think> tokens
  • Switch behavior based on these token cues

V3.1's hybrid approach enables flexible deployment scenarios - developers can use non-thinking mode for fast interactions and thinking mode when reasoning transparency is required. The model switches dynamically without architectural changes, making it ideal for applications requiring both speed and occasional deep reasoning.

In comparison, R1's dedicated approach provides consistent reasoning patterns but lacks flexibility. The model cannot provide quick responses for simple queries, always investing computational resources in detailed reasoning chains. There may be situations where you want this, but constantly loading and unloading models this large isn’t particularly feasible, and flexibility is often warranted to prevent unnecessary transition time between models or cold starts in serverless.

The comparison reveals two successful but distinct philosophies for AI reasoning implementation. DeepSeek R1 represents specialized excellence - a dedicated reasoning system that consistently provides transparent, deep analysis at the cost of speed and flexibility. DeepSeek V3.1 embodies adaptive intelligence - a hybrid system that dynamically balances reasoning depth with practical efficiency based on user needs and query complexity.

Parameter and Context Window Changes

The transition from V3-0324 to V3.1 involved significant changes in model sizing and context handling. V3-0324 maintained the 685B total parameter count of the original V3, while V3.1 adopts a more efficient structure with 671B total params, 37B activated params, 128K context length. For those on a budget and trying to run the model in, say, an 8xH100 NVL pod in 8 bit where you would have about 685GB occupied by the model, this represents a not insignificant amount of additional wiggle room for additional context length. Inference patterns between the two models do differ significantly. V3.1's hybrid design allows efficient resource allocation based on query complexity, while R1 consistently uses full reasoning overhead.

Performance Benchmarks and Real-World Impact

The performance improvements from V3-0324 to V3.1 are substantial across multiple domains. In mathematical reasoning, V3.1's thinking mode achieves AIME 2024 (Pass@1): 66.3 for non-thinking mode, 93.1 for thinking mode compared to V3-0324's 59.4. Code performance shows similar dramatic improvements, with LiveCodeBench (2408-2505) (Pass@1): 56.4 for non-thinking, 74.8 for thinking versus V3-0324's 43.0.

The agent capabilities demonstrate the practical impact of architectural changes. SWE Verified (Agent mode): 66.0 for V3.1 compared to 45.4 for V3-0324 represents a 45% improvement in software engineering tasks. This suggests that the hybrid architecture and enhanced tool calling provide tangible benefits for complex, multi-step reasoning tasks.

Here’s how the model shapes up compared to previous iterations in thinking and non-thinking modes:

__wf_reserved_inherit

Getting Started on Runpod

As always, the general rule for loading an LLM at full weights is 2 * the number of parameters, plus 10% for context and cache. So you could be looking at approximately 1500 GB of RAM required for this model. This will require an Clusters setup; we’ve previously gone into setting up Kimi K2 on our Youtube channel, and you can use the same process to run the new Deepseek in a cluster.

For more budget oriented setups you can always use a GGUF quantization such as those provided by Unsloth, and we’ve also got a video on how to get those up and running using KoboldCPP.

(FWIW, one amusing but fairly on the nose rule of thumb for quantizations is ‘the number of bits is how many hours of sleep the model got last night.’)

Looking Forward: The Evolution of Hybrid Models

The architectural patterns established in V3.1—hybrid reasoning modes, custom precision formats, and enhanced agent capabilities—likely preview the direction of next-generation language models. Rather than scaling parameters indefinitely, the focus appears to be shifting toward architectural efficiency, specialized capabilities, and hardware optimization. We’ve already seen this with the industry-wide shift away from dense models; most of the large closed-source models are known or theorized to be MoE models, and doubly true with the open-source community with very large dense models seeming to fall out of favor (such as the previous Llama-3 405b.)

The architectural patterns established in V3.1—hybrid reasoning modes, custom precision formats, and enhanced agent capabilities—likely preview the direction of next-generation language models. Rather than scaling parameters indefinitely, the focus appears to be shifting toward architectural efficiency, specialized capabilities, and hardware optimization. We’ve written before about the importance of speed when it comes to LLM inference — faster models incur less GPU time, and thus less billing, especially important in a serverless environment where you pay per second.

We’re super excited to see what you create with this new model! Feel free to hop onto our Discord if you have any questions.

Author profile: Brendan McKeag