惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Engineering at Meta
Engineering at Meta
月光博客
月光博客
WordPress大学
WordPress大学
C
Cisco Blogs
Recent Commits to openclaw:main
Recent Commits to openclaw:main
博客园 - 【当耐特】
大猫的无限游戏
大猫的无限游戏
The GitHub Blog
The GitHub Blog
Google DeepMind News
Google DeepMind News
The Cloudflare Blog
有赞技术团队
有赞技术团队
Microsoft Azure Blog
Microsoft Azure Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
小众软件
小众软件
H
Heimdal Security Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
W
WeLiveSecurity
量子位
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Fortinet All Blogs
T
Threat Research - Cisco Blogs
Attack and Defense Labs
Attack and Defense Labs
P
Privacy & Cybersecurity Law Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
NISL@THU
NISL@THU
Forbes - Security
Forbes - Security
L
Lohrmann on Cybersecurity
C
CERT Recently Published Vulnerability Notes
L
LINUX DO - 热门话题
Google Online Security Blog
Google Online Security Blog
S
Security Affairs
V2EX - 技术
V2EX - 技术
TaoSecurity Blog
TaoSecurity Blog
N
News and Events Feed by Topic
N
News | PayPal Newsroom
S
Security @ Cisco Blogs
宝玉的分享
宝玉的分享
Project Zero
Project Zero
The Hacker News
The Hacker News
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
PCI Perspectives
PCI Perspectives
G
GRAHAM CLULEY
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Y
Y Combinator Blog
N
Netflix TechBlog - Medium
S
Schneier on Security
Application and Cybersecurity Blog
Application and Cybersecurity Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
博客园 - 聂微东

StudyingLover's Blog

Diffusion Policy笔记 rwkv笔记 act笔记 nanovllm-block_manager opencode多智能体 nanobot-pre-train nanobot-rl nanobot-sft nanobot-checkpoint_manager nanobot-gpt nanobot-mid-train BPE演示 最后一遍学习Transformer YOLOv5 目标检测笔记 下载根服务器解析记录 Dynaseal A Backend-Controlled LLM API Key Distribution Scheme with Constrained Invocation Parameters 判断链表有环 王道25数据结构勘误 关于perplexity的open-sourcing-r1-1776 AI为什么不像人类一样进行多轮对话 新博客改造日记和功能测试 linuxqq只显示登陆背景图 数字设计和计算机体系结构(机械工业出版社)勘误(自制) Dynaseal:面向未来端侧llm agent的llm api key分发机制 A Definitive Guide to Markdown Style This post is using MDX, Where you can embed JSX and Astro components RT-Patch学习 pydantic实现的LLM ReAct fastapi 和 uvicorn 设置监听 ipv6 pydantic+openai+json 控制大模型输出的最佳范式 解决 Matplotlib Scatter 不支持 Marker 列表的问题:mscatter 实现 roofline model zhipuAI接口兼容openai 在docker部署fastapi宝塔里使用nginx反代套上cloudflare获取请求的真实ip clion搭建libbpf-bootstrap开发环境 coze+coze-discord-proxy+ChatNextWebUI实现AI自由 安卓内核时间使用的是UTC时间 colab运行google最新开源模型Gemma Sora技术报告 视频生成模型作为世界模拟器 笔记 archlinux flutter开发踩坑 fastapi集成google auth登录 linux下NTFS磁盘报错输入输出错误 Venn-Abers 预测器 基于Venn-Abers预测器的系统日志异常检测方法_顾兆军 手机平板远程访问kvm虚拟机的windows phi-2弱智吧测评 poe的gemini pro或是百度开发 google gemini api使用 google gemini api申请 构建用于复杂数据处理的高效UDP服务器和客户端 matplotlib中文字体渲染 TruFor笔记和代码复现 深入分析:GitHub Trending 项目 "multipleWindow3dScene" pua大模型 ggml教程|mnist手写体识别量化推理 xgboost2.0最佳实践 xgboost使用GPU最佳实践 马踏棋盘 cloudlflare推理llama2 docker搭建elasticsearch并使用python连接 FreeU-文字生成图片的免费午餐笔记 使用xgboost的c接口推理模型 Archlinux使用CMake调用xgboost的c接口 m2cgen生成机器学习c语言推理代码 xgboost模型序列化存储并推理 speculative-sampling笔记 prompt2model笔记 RoboTAP笔记 自建obsidian同步服务 MediaPipe即将推出图像生成服务 Dual-Stream Diffusion Net for Text-to-Video Generation笔记 ViT在DDPM取代UNet(DiT) arch4edu搞崩了我的flutter LISA(推理分割)笔记 在终端绘制GPU显存使用曲线 GPTBot介绍 arch蓝牙无法连接 GPU部署llama-cpp-python(llama.cpp通用) 花式求GCD 使用llama构建一个蜜罐(前端) 使用llama构建一个蜜罐(后端) llama-cpp-python快速上手 快速上手llama2.c(更新版) Paper Gestalt笔记 DINO-v2笔记 快速上手llama2.c AnyDoor笔记 Archlinux安装scrcpy加载共享库出错 error while loading shared libraries:libusb-1.0.so.0:wrong ELF class:ELFCLASS32 npc_gzip笔记 python调用c++函数 Filesystem type ntfs3,ntfs not configured in kernel open_clip编码图像和文本 PicGo配置CloudflareR2图片储存 ArchlinuxGnome快捷键打开终端 clip-interrogator代码解析 GroundingDINO安装报错解决 2023华为鲲鹏畅想日暨西安高新国际会议中心零食午饭测评 RoboMaster开源仓库汇总(长期更新) 没有手都可以在腾讯云创建镜像 I3D笔记
Vision Mamba (Vim)笔记
About the Author StudyingLover · 2026-01-08 · via StudyingLover's Blog

Vision Mamba(ViM)和Vision Transformer (ViT) 大体上都是相同的,只有部分细节不同

双向机制的实现

vit是天生全局可见的,而 vim 是天生单向的

Mamba 本质上是像 RNN 一样“从左读到右”的,如果不做处理,它只能看到“之前”的像素,看不到“之后”的。

为了解决这个问题,Vim 在代码中强制实现了双向扫描,最后直接相加。

Vim 的 Mamba 层是成对设计的。假设模型有 24 层,它是把它们分成了 12 对(每对包含一个前向层和一个后向层)。

假设我们有一张极小的 2×22 \times 2 图片,切成了 4 个 Patch:

Patch0Patch1Patch2Patch3\begin{matrix} \text{Patch}_0 & \text{Patch}_1 \\ \text{Patch}_2 & \text{Patch}_3 \end{matrix}

  • 展平 (Flatten):ViT/Vim 第一步是把图片“拉直”。通常是按扫描(Row-Major):
    • 正向序列 (Forward): [Patch0,Patch1,Patch2,Patch3][\text{Patch}_0, \text{Patch}_1, \text{Patch}_2, \text{Patch}_3],左上角开始,一行行读到右下角。
  • 翻转 (Flip) 倒序序列, 这就是 hidden_states.flip([1]) 做的事情:
    • 倒序序列 (Backward): [Patch3,Patch2,Patch1,Patch0][\text{Patch}_3, \text{Patch}_2, \text{Patch}_1, \text{Patch}_0],从右下角开始,倒着一行行读回左上角。
for i in range(len(self.layers) // 2):
   if self.if_rope:
       hidden_states = self.rope(hidden_states)
       if residual is not None and self.if_rope_residual:
           residual = self.rope(residual)

   hidden_states_f, residual_f = self.layers[i * 2](
       hidden_states, residual, inference_params=inference_params
   )
   hidden_states_b, residual_b = self.layers[i * 2 + 1](
       hidden_states.flip([1]), None if residual == None else residual.flip([1]), inference_params=inference_params
   )
   hidden_states = hidden_states_f + hidden_states_b.flip([1])
   residual = residual_f + residual_b.flip([1])

CLS处理

CLS的处理vim也有所不同,有两种策略

  • 中间 CLS Token
    • 输入时:把 [CLS] Token 插在序列的正中间 (Index N/2N/2)。
    • 输出时:取中间那个向量。
    • 放在中间是物理上接收前向信息流和后向信息流距离最短、最均衡的位置。如果放在头部,它离“从后向前”的信息流终点太远了。
  • 双 CLS Token
    • 输入时:在头部尾部各放一个 Token。
    • 输出时:取这两个 Token 的平均值
    • 暴力解决递归模型的“遗忘”问题。前向扫描时,尾部 Token 拥有完整信息。后向扫描时,头部 Token 拥有完整信息。两者结合,确保万无一失。