惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Attack and Defense Labs
Attack and Defense Labs
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Recent Announcements
Recent Announcements
博客园 - 【当耐特】
博客园 - 三生石上(FineUI控件)
量子位
aimingoo的专栏
aimingoo的专栏
V
V2EX
Vercel News
Vercel News
B
Blog
M
MIT News - Artificial intelligence
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
The Cloudflare Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Hacker News: Ask HN
Hacker News: Ask HN
TaoSecurity Blog
TaoSecurity Blog
N
News and Events Feed by Topic
D
DataBreaches.Net
Blog — PlanetScale
Blog — PlanetScale
S
Secure Thoughts
U
Unit 42
博客园 - 叶小钗
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Hacker News - Newest:
Hacker News - Newest: "LLM"
N
News | PayPal Newsroom
Help Net Security
Help Net Security
S
Security Affairs
Microsoft Security Blog
Microsoft Security Blog
W
WeLiveSecurity
博客园 - Franky
Forbes - Security
Forbes - Security
Microsoft Azure Blog
Microsoft Azure Blog
博客园_首页
Schneier on Security
Schneier on Security
I
InfoQ
B
Blog RSS Feed
大猫的无限游戏
大猫的无限游戏
A
About on SuperTechFans
Webroot Blog
Webroot Blog
AWS News Blog
AWS News Blog
Last Week in AI
Last Week in AI
Security Archives - TechRepublic
Security Archives - TechRepublic
C
CERT Recently Published Vulnerability Notes
N
News and Events Feed by Topic
阮一峰的网络日志
阮一峰的网络日志
L
Lohrmann on Cybersecurity
SecWiki News
SecWiki News
Recent Commits to openclaw:main
Recent Commits to openclaw:main
J
Java Code Geeks

hsfzxjy 的博客

解决 VSCode + CMake + MSVC 编译器信息乱码的问题 如何在 VS Code DevContainer 中配置 HTTP 代理 如何在跳板机背后的服务器上使用 VS Code Remote - Containers Cohesive Digests for Ints and Floats Rust 中的隐匿概念 —— Place(位置) 美术馆 一尺之槌,日取其半,1075日而竭 老生常谈:使用 Cloudflare 自选 IP 加速站点访问 辩义 State、Nation 与 Country 将 Base64 编码的数据快速转换为 Uint8Array 折腾 NPU·第1章 —— 搭建 Level Zero 开发环境 折腾 NPU·第0章 —— Intel NPU 概述与 Level-Zero 新增域名 monad.run CSS 中为特定字符设置不同字体 Arbitary Lifetime Transmutation via Rust Unsoundness Dijkstra 算法的延伸 Manacher 回文计数算法 硬卧 Go Fact: Zero-sized Field at the Rear of a Struct Has Non-zero Size Display *big.Rat Losslessly and Smartly in Golang 代码的仪式 Building Electron From Scratch 中式亲属称谓研究之一:构建半群 Some Notes on Kotlin Coroutines Git sparse-checkout and partial clones for Mega-Repos 辩义“封建” Diving from the CUDA Error 804 into a bug of libnvidia-container Modern Cryptography, GPG and Integration with Git(hub) Move the Root Partition of Ubuntu A New Programmer Kicks a Roadblock Git-based Dependencies in Dart and Go Reversy Naming 人类一败涂地 Invalid Golang Pointers Can Bite You Even If You Don't Dereference Side Project(副业) A Flaw of Promoting Complex Trait Bounds in Rust Initialize Process Pool Worker with Individual Value Rust - Python FFI From Scratch [Extending Hexo For My Site] Part 1 [Extending Hexo For My Site] Part 0 Debug a 'torch.tensor(1).cuda()' hanging 不自由的互联网 Retrieve Contents over HTTP without curl or wget [Unravelling mocona] Part 1 - Verbosity or Anti-Pattern [Unravelling mocona] Part 0 - Preface Understanding pickle in Python Rough Notes on Deploying Vaultwarden & NextCloud Bookmarks 语言狂热者与实用主义者 Demystify the randomness in CUDA kernels Performant Bulk Mutations in IndexedDB Auto Rebuild .pyx Files with pyximport Cython and Threads Obtain a Random Available TCP Port with Bash Information Theory: KL Divergence Information Theory: Entropy and Mutual Information 铁板烧 西郊线 Proof of the Gumbel Max Trick Option::as_ref Rc, RefCell and Interior Mutability Visualizing Correlation 三月十日杂感 三月一日杂感 二月十一日杂感 一月二十六日杂感 SS Configuration 一月七日杂感 四月·病 Haskell 笔记:State Monad Haskell 笔记:Monad 引论 Haskell 笔记:Applicative Haskell 笔记:Category Theory and Functor Haskell 笔记:data, type, newtype Haskell 笔记:folds 使用 Aria2 在 Ubuntu 中下载百度云资源 从伪并行的 Python 多线程说起 一个 Reentrant Error 引发的对 Python 信号机制的探索和思考 Linux 文件权限 HSFZMUN 4.0 部署小记 午后雨·科大 最后的雨夜·广州 揭秘·变态的平方根倒数算法 神坑·Python 装饰类无限递归 Python“黑魔法”之 Encoding & Decoding Ubuntu 重新映射键盘布局 为什么我要翻墙 Python“黑魔法”之 Generator Coroutines 数学美 之 判断线段相交的最简方法 除夕杂感 17 行代码实现的简易 Javascript 字符串模板 Python“黑魔法”之 Meta Classes 诗集 生活,需要被“发现” 家书·十八岁成人礼 炫技?还是需求? 【译】响应式图片的现状 【译】“为什么有这么多的编程语言?” Wisecity 商赛总结——也谈前端自动化测试 记一次 DoS 诈骗网站的经历 那一年,我们望向星空
使用 3090 部署 1.58bit 动态量化版 DeepSeek R1 671b
2025-02-22 · via hsfzxjy 的博客

1.58 bit 量化技术通过将每个权值量化为仅有三个状态,最大程度地节约显存并加速推理。

Unsloth AI 在 Run DeepSeek R1 Dynamic 1.58-bit 一文介绍了 1.58 bit 动态量化版的 DeepSeek R1 671b(以下简称 unsloth 版)。我们知道 DeepSeek R1 70b 及以下的版本都是通过知识蒸馏得到的,唯有 671b 是满血的版本。然而原版 671b 的推理需要大量显存及算力,仅模型文件便有 404GB。

unsloth 版有选择地对部分权值作 1.58 bit 量化,将模型文件压缩至 131GB,同时也保持了不错的生成质量。结合 llama.cpp 实现的 CPU+GPU 混合推理,我们可在低成本的硬件上部署 671b 模型。

使用 llama-bench 测试生成速度

先说结论。笔者在实验室集群单个 3090 节点上尝试部署 unsloth 版 671b 模型,最大 token 生成速度可达 10 tokens/s,足够应付单人日常使用。节点的硬件配置如下:

  • CPU – AMD EPYC 7402 24-Core Processor
  • RAM – 32GB x 16, DIMM DDR4 Synchronous Registered (Buffered) 3200 MHz
  • GPU – 8x NVIDIA GeForce RTX 3090 (24GB VRAM)

以下是 token 生成速度与 GPU Layers 数量的关系曲线:

GPU Layers 指置于 GPU 上的网络层数,取值范围为 0~62,可由 CLI 参数 --n-gpu-layers 配置。其中取 0 时代表完全使用 CPU 运行,取 62 时代表完全使用 GPU 运行。图中还标出了不同 GPU Layers 数量所需的 GPU 卡数(顶部横轴),以方便读者根据自己的显存大小调整 GPU Layers。

值得注意的是,虽然 7 张 3090 可以放下整个模型,在实际推理时考虑到 Context 的额外开销,所需显存可能不止于此,比如我后文部署 --ctx-size=8192 的 llama-server 时就需要占满 8 张 3090。

以上测试数据由如下脚本得出:

export CUDA_VISIBLE_DEVICES=0
next_gpu=1

for i in {0..62}; do
mkdir -p benchout
while true; do
./llama.cpp/build/bin/llama-bench \
--model /models/DeepSeek-R1-GGUF/DeepSeek-R1-UD-IQ1_S/DeepSeek-R1-UD-IQ1_S-00001-of-00003.gguf \
--cache-type-k q4_0 \
--threads 64 --prio 2 \
--n-gpu-layers $i -n 128 2>&1 | tee benchout/bench-$i.log
if [ ${PIPESTATUS[0]} -eq 0 ]; then
break
else
echo "Retrying..."
export CUDA_VISIBLE_DEVICES=${CUDA_VISIBLE_DEVICES},${next_gpu}
next_gpu=$((next_gpu + 1))
fi
done
done

下面介绍如何借助 llama.cpp 运行一个基础可用的 Web 界面,以使用 unsloth 版 671b 模型。

llama-server 服务部署

模型下载

本次使用的模型位于 Huggingface 的 unsloth/DeepSeek-R1-GGUF 仓库。仓库提供了多种量化的版本,我们只需下载其中的 DeepSeek-R1-UD-IQ1_S 目录,置于 /mnt/models/ 路径,形成以下结构:

/mnt/models/DeepSeek-R1-GGUF/DeepSeek-R1-UD-IQ1_S
├── DeepSeek-R1-UD-IQ1_S-00001-of-00003.gguf
├── DeepSeek-R1-UD-IQ1_S-00002-of-00003.gguf
└── DeepSeek-R1-UD-IQ1_S-00003-of-00003.gguf

下载模型可使用 Python 的 huggingface_hub 库:

import os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from huggingface_hub import snapshot_download
snapshot_download(
repo_id = "unsloth/DeepSeek-R1-GGUF",
local_dir = "/mnt/models",
allow_patterns = ["*UD-IQ1_S*"],
)

当然,如果环境中不方便使用 Python,笔者推荐使用 HuggingFaceModelDownloader 下载。这是一个用 Go 编写的下载器,无需像 Python 那样配置环境即可使用:

curl -sSL https://g.bodaay.io/hfd -o hfd && chmod +x hfd
./hfd -m unsloth/DeepSeek-R1-GGUF:UD-IQ1_S -c 8 -s /mnt/models

编译 llama.cpp 及启动 llama-server 服务

运行 1.58 bit 的 DeepSeek R1 需要使用 llama.cpp。笔者选择使用 Docker 部署,以避免繁琐的环境配置。

为了方便读者使用,笔者将镜像的准备和启动过程整合到了 Docker Compose 文件中。读者只需将以下两个文件 compose.yamlllama_Dockerfile 放在同一个目录下,执行 docker compose up 即可启动 llama-server。llama-server 内置了一个简单的 Web 界面,可以通过 http://localhost:10000 访问。

services:
deepseek:
build:
args:
- PROXY=
- CUDA_ARCH=86
dockerfile: llama_Dockerfile
command: |
llama-server
--model /models/DeepSeek-R1-GGUF/DeepSeek-R1-UD-IQ1_S/DeepSeek-R1-UD-IQ1_S-00001-of-00003.gguf
--cache-type-k q4_0
--threads 12 --prio 2
--temp 0.6
--ctx-size 8192
--port 10000
--host 0.0.0.0
--n-gpu-layers 62
volumes:
- /mnt/models:/models
ports:
- 10000:10000
deploy: { resources: { reservations: { devices: { driver: nvidia, capabilities: [ gpu ],
count: 8
} } } }
FROM hsfzxjy/devkit:cuda11.7.1
ARG PROXY
ENV http_proxy=${PROXY} https_proxy=${PROXY}
RUN mkdir /app/ && cd /app && git clone https://github.com/ggerganov/llama.cpp
WORKDIR /app/llama.cpp
RUN apt-get install -y libcurl4-openssl-dev
ARG CUDA_ARCH
RUN cmake . -B build \
-DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON -DLLAMA_CURL=ON -DCMAKE_CUDA_ARCHITECTURES=${CUDA_ARCH}
RUN cmake --build build --config Release -j16 --clean-first --target llama-cli llama-server llama-bench
ENV PATH=/app/llama.cpp/build/bin:${PATH}
ENV http_proxy= https_proxy=

作者:hsfzxjy
链接:
许可:CC BY-NC-ND 4.0.
著作权归作者所有。本文不允许被用作商业用途,非商业转载请注明出处。