惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
T
The Exploit Database - CXSecurity.com
C
CERT Recently Published Vulnerability Notes
Simon Willison's Weblog
Simon Willison's Weblog
T
Tor Project blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
D
DataBreaches.Net
The Hacker News
The Hacker News
有赞技术团队
有赞技术团队
Latest news
Latest news
T
Tailwind CSS Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
H
Help Net Security
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Threat Research - Cisco Blogs
G
GRAHAM CLULEY
G
Google Developers Blog
W
WeLiveSecurity
Project Zero
Project Zero
WordPress大学
WordPress大学
人人都是产品经理
人人都是产品经理
博客园 - 司徒正美
博客园 - 三生石上(FineUI控件)
MyScale Blog
MyScale Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
F
Full Disclosure
The Last Watchdog
The Last Watchdog
Security Archives - TechRepublic
Security Archives - TechRepublic
Attack and Defense Labs
Attack and Defense Labs
N
News and Events Feed by Topic
博客园 - 【当耐特】
Google DeepMind News
Google DeepMind News
V
Visual Studio Blog
Blog — PlanetScale
Blog — PlanetScale
F
Fortinet All Blogs
PCI Perspectives
PCI Perspectives
小众软件
小众软件
N
News | PayPal Newsroom
罗磊的独立博客
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
AI
AI
T
Tenable Blog
S
Schneier on Security
O
OpenAI News
The Register - Security
The Register - Security
Google DeepMind News
Google DeepMind News
Engineering at Meta
Engineering at Meta
T
Threatpost
Hacker News: Ask HN
Hacker News: Ask HN

Runpod Blog.

DeepSeek V4 in the wild, and how to run it on Runpod New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide 10 billion Serverless requests and counting Building for resilience: Runpod’s response to the AWS us-east-1 outage How to Connect Google Colab to Runpod Founder Series #1: The Runpod Origin Story AMD MI300X vs. NVIDIA H100: Mixtral 8x7B Inference Benchmark How to Run the FLUX Image Generator with ComfyUI on Runpod Run Llama 3.1 405B with Ollama on Runpod: Step-by-Step Deployment How to Run FLUX Image Generator with Runpod (No Coding Needed) How to Use 65B+ Language Models on Runpod Deploy Llama 3.1 with vLLM on Runpod Serverless: Fast, Scalable Inference in Minutes Open Source Video & LLM Roundup: The Best of What’s New Run vLLM on Runpod Serverless: Deploy Open Source LLMs in Minutes Introduction to vLLM and PagedAttention New update to Github integration: release rollback! | Runpod Blog A note to the developers who built Runpod with us Deploy ComfyUI as a Serverless API Endpoint Setting up Slurm on Runpod Clusters: A Technical Guide Building an OCR System Using Runpod Serverless From No-Code to Pro: Optimizing Mistral-7B on Runpod for Power Users Lessons While Using Generative Language and Audio For Practical Use Cases Runpod RoundUp 3 – AI Music and Stock Sound Effect Creation New Navigational Changes To Runpod UI Use alpha_value To Blast Through Context Limits in LLaMa-2 Models Runpod Roundup 5 – Visual/Language Comprehension, Code-Focused LLMs, and Bias Detection Runpod is Proud to Sponsor the StockDory Chess Engine Runpod Roundup 4 – Open Source LLM Evaluators, 3D Scene Reconstruction, Vector Search Meta and Microsoft Release Llama 2 as Open Source SuperHot 8k Token Context Models Are Here For Text Generation How to Manage Funding Your Runpod Account Encrypted Volumes on Runpod: Protect Your Data at Rest How to Run a "Hello World" on Runpod Serverless Runpod AI field notes: December 2025 Faster GitHub Builds: Major Performance Improvements to Our Automated Integration Partnering with Defined AI to Bridge the Data Wealth Gap How to Run Serverless AI and ML Workloads on Runpod How to fine-tune a model using Axolotl Transcribe and translate audio files with Faster Whisper Runpod Achieves SOC 2 Type II Certification: Continuing Our Compliance Journey Orchestrating GPU workloads on Runpod with dstack Exploring Runpod Serverless: Create Workers From Templates DeepSeek V3.1: A Technical Analysis of Key Changes from V3-0324 Deep Cogito Releases Suite of LLMs Trained with Iterative Policy Improvement Wan 2.2 Releases With a Plethora Of New Features Iterative Refinement Chains with Small Language Models The New Runpod.io: Clearer, Faster, Built for What’s Next Introducing Clusters: On-Demand Multi-Node AI Compute How Do I Transfer Data Into My Runpod? Spot vs. On-Demand Instances: What’s the Difference? Deploy GitHub Repos to Runpod with One Click Run GGUF Quantized Models Easily with KoboldCPP on Runpod How to Work with GGUF Quantizations in KoboldCPP Introducing Better Forge: Spin Up Stable Diffusion Pods Faster Supercharge Your LLMs with SGLang: Boost Performance and Customization Mastering Serverless Scaling on Runpod: Optimize Performance and Reduce Costs RAG vs. Fine-Tuning: Which Is Best for Your LLM? Run Larger LLMs on Runpod Serverless Than Ever Before – Llama-3 70B (and beyond!) How to Run vLLM on Runpod Serverless (Beginner-Friendly Guide) Embracing New Beginnings: Welcoming Banana.dev Community to Runpod Stable Diffusion + ComfyUI on Runpod: Easy Setup Guide Runpod RoundUp 2 – 32k Token Context LLMs and New StabilityAI Offerings Runpod Roundup: High-Context LLMs, SDXL, and Llama 2 16k Context LLM Models Now Available On Runpod Savings Plans Are Here For Secure Cloud Pods – How To Purchase a Monthly Plan And Save Big Pygmalion-7b from PygmalionAI has been released, and it's amazing Ada Architecture Pods Are Here – How Do They Stack Up Against Ampere? Spin up a Text Generation Pod with Vicuna and Experience a GPT-4 Rival Using OpenPose to Annotate Poses Within Stable Diffusion Set Up a Chatbot with Oobabooga on Runpod Connect VSCode to Your Runpod Instance (Quick SSH Guide) Deploy a Stable Diffusion UI on Runpod in Minutes Google Colab Pro vs. Runpod: Best GPU Cloud for AI Workloads How to Run a GPU-Accelerated Virtual Desktop on Runpod
Run DeepSeek R1 on Just 480GB of VRAM
Brendan McKeag · 2025-02-27 · via Runpod Blog.

Even with the new closed model successes of Grok and Sonnet 3.7, DeepSeek R1 is still considered a heavyweight in the LLM arena as a whole, and remains the uncontested open-source LLM champion (at least until DeepSeek R2 launches, anyway.) We've written before about the concerns of using a closed-source LLM if you are working with sensitive data, and it's unlikely those concerns over data usage and transmission will ever truly disappear in the AI world as we know it today. Data crossing borders may fall under different legal frameworks, and you can't fully audit what happens to your data on closed platform; it can become a challenge to provide evidence of data handling practices to regulators.

DeepSeek's large model size means it has been something of a challenge to host, however, compared to simply signing up for a paid closed model service. Thanks to the magic of quantization, though, you can now run a Q4 4-bit quantization on Runpod through a one-click template deploy which will spin up a pod running KoboldCPP, and between download and loading should get you up and running in approximately 20 minutes. This will only cost $10/hr in Secure Cloud on 6xA100s or $16/hr on 4xH200s.

To get started, you can go to our Template hub or deploy a pod, search for KoboldCpp - DeepSeek-R1 - 6x80GB of VRAM, and deploy on a pod configuration with at least 480GB of VRAM. This will set up an OpenAI-compatible endpoint on port 5001 in the pod that you can then send API requests to. (Note that community templates can have bugs or unintended operation - use with caution!)

If you'd like an alternate template for this package, you can try the official KoboldCPP package and setting up the download in the environment variables on your own. Here is the model quantization by Bartowski, and here is a guide to set up the template parameters.

Want to run it on even less hardware? Can do - you can offload some of the work to the system RAM in the pod. Change the --gpulayers argument in the environment variable to something below 61, which will start pushing the layers off of the GPU. Note that this will result in a sizable performance hit, but may be preferable for some use cases since you could run the model on even less of 480GB of VRAM and drop some cards from your system requirements.

Environment variables panel with KCPP_ARGS GPU layer flags and KCPP_DONT_REMOVE_MODELS set to true

As a reminder, here are the official suggested sampler and prompting settings:

DeepSeek Drops Repos For Open Source Week

While waiting for the next iteration of the model (practically an inevitability, given the success of R1) DeepSeek has open-sourced five repos on GitHub for Open Source Week. Here's a quick roundup:

Monday: FlashMLA

DeepSeek's FlashMLA is a significant open-source contribution that addresses a critical performance bottleneck in LLM deployment. It's specifically designed to handle variable-length sequences efficiently during model inference. FlashMLA builds upon innovations from FlashAttention 2 & 3 and CUTLASS projects, but extends them specifically for MLA decoding which is becoming increasingly important for efficient LLM inference (especially given the sheer size of the largest open source models; Llama-3 having a 405b variant and R1 having 671b parameters would have been absolutely unthinkable just a year or two ago. BLOOM having 176b was considered an extreme anomaly back then!)

Tuesday: DeepEP

Mixture of Experts models (like R1) are becoming essential for scaling AI capabilities beyond traditional dense models, but they've been bottlenecked by communication overhead. DeepEP directly addresses this limitation.

DeepEP is a specialized communication library designed for MoE architectures and expert parallelism. It provides high-performance GPU kernels for the critical "all-to-all" communication pattern that MoE models require when routing tokens to their appropriate experts. The library achieves near-theoretical maximum performance on both NVLink (~158 GB/s) and RDMA (~47 GB/s) connections, meaning it's extracting nearly all available performance from the hardware.

Wednesday: DeepGEMM

DeepSeek's DeepGEMM library represents a significant contribution to the AI infrastructure ecosystem by addressing the critical performance bottleneck of matrix multiplications, particularly for FP8 precision operations. The benchmarks are impressive, showing up to 2.7x speedups over existing carefully-optimized CUTLASS implementations, with some configurations reaching 1358 TFLOPS on NVIDIA Hopper GPUs.

Thursday: Optimized Parallelism Strategies: DualPipe, EPLB

DualPipe introduces a novel bidirectional pipeline parallelism algorithm that significantly improves distributed training efficiency by fully overlapping computation and communication phases between forward and backward passes. While lightweight in code, it demonstrates a significant architectural innovation that could influence how large language models are trained in distributed environments.

EPLB addresses a critical challenge in scaling MoE models, which are becoming increasingly important for achieving state-of-the-art performance while managing computational costs. While lightweight in code, it demonstrates a significant architectural innovation that could influence how large language models are trained in distributed environments.

Friday: Fire-Flyer File System (3FS)

The Fire-Flyer File System (3FS) from represents a significant contribution to AI infrastructure, addressing one of the most critical but often overlooked components: storage. It bridges traditional file system interfaces with modern hardware capabilities, leveraging advanced SSD storage and high-speed RDMA networking to create a cohesive, efficient storage layer. One of the biggest challenges in a serverless setup is how do you load a huge LLM like R1 responsively enough to avoid long cold starts? Moving all that data takes time, and by the time you even catch wind of a request, the user is already waiting. In large scale testing this system achieved 6.6 TiB - enough to completely move R1 in well under a second.

Spin Up a DeepSeek R1 Pod

Conclusion

The team at DeepSeek is clearly cooking some new innovations beyond releasing one of the most notable LLMs in history - and many of them help address the challenges involved moving the data required for such a large model. We'll keep you posted on any new updates on their new model!

Author profile: Brendan McKeag