惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

W
WeLiveSecurity
Jina AI
Jina AI
博客园 - 司徒正美
雷峰网
雷峰网
宝玉的分享
宝玉的分享
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园_首页
WordPress大学
WordPress大学
Google DeepMind News
Google DeepMind News
GbyAI
GbyAI
MyScale Blog
MyScale Blog
Apple Machine Learning Research
Apple Machine Learning Research
美团技术团队
I
InfoQ
博客园 - Franky
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
博客园 - 叶小钗
阮一峰的网络日志
阮一峰的网络日志
Cyberwarzone
Cyberwarzone
C
CXSECURITY Database RSS Feed - CXSecurity.com
S
Schneier on Security
P
Privacy & Cybersecurity Law Blog
T
Threatpost
Cloudbric
Cloudbric
D
Docker
M
MIT News - Artificial intelligence
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Vercel News
Vercel News
Martin Fowler
Martin Fowler
J
Java Code Geeks
AWS News Blog
AWS News Blog
The Cloudflare Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
L
Lohrmann on Cybersecurity
Hacker News: Ask HN
Hacker News: Ask HN
Last Week in AI
Last Week in AI
S
Security @ Cisco Blogs
Help Net Security
Help Net Security
C
Cisco Blogs
V
V2EX
博客园 - 【当耐特】
I
Intezer
爱范儿
爱范儿
F
Fortinet All Blogs
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
P
Privacy International News Feed
IT之家
IT之家
L
LINUX DO - 最新话题
B
Blog RSS Feed
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO

Cisco Blogs

Cisco Live 2026: Bringing the Future of Customer Experience to Las Vegas Edge opportunity for service providers: Turn infrastructure into new services MRC and SRv6: How Foundational Networking Innovations Are Enabling the Next Generation of AI Supercomputers The SMB Marketing Reset: Winning Customer Trust in a Digital-First Economy Inside the SOC: AI-powered DNS defense against ransomware Our Path Forward Securing the Federal Digital Experience with Cisco ThousandEyes for Government Cisco at ONUG Dallas 2026: Securing the AI Data Center in the Agentic Era Cisco and Red Hat are powering intelligent core to edge: Red Hat Summit insights Building the Capabilities That Win: How Cisco Partners Can Lead in the SMB & Mid-Market Era How Two Hours Felt Bigger Than My To-Do List Announcing Foundry Security Spec Ace the CCIE Collaboration Lab: Success Tips from a TAC Engineer Turned CCIE Protecting Agents with Cisco AI Defense and Google Agent Development Kit Powering an Inclusive Future: Your guide to the Purpose Pavilion at Cisco Live Las Vegas The Infrastructure Behind the Mission: SOF Week 2026 Cisco Networking App Marketplace Partners at Cisco Live 2026 Beyond the Pilot: Building the Clinical Data Fabric for the Agentic Era Benchmarking scale-out AI fabrics with Cisco N9000 + AMD Pensando™ Pollara 400 NICs Month of Developer Productivity: Build and Forget The race to autonomous transport networks: A new study Lean IT, future-ready: How to save time and simplify wireless management with AI Biochar’s triple win: Healthier soils, improved crops, and decarbonization Designing a Proactive Customer Journey Modernize your data center operations with Cisco Nexus Dashboard Why your automation stack needs Cisco Agentic Workflows Try Cisco AI Defense Explorer Edition in this hands-on lab From Bandwidth to Intelligence: How Cisco is Powering AI-Ready Networks Spotlight on digital transformation | FY25 Purpose Report Galaxy Mode is live: A limited-time look at what your Cisco AI Assistant and AgenticOps can already do Securing the Agentic Workforce: Cisco Announces Intent to Acquire Astrix Security Understanding CISA BOD 26-02: Mitigating Risk from End-of-Support Edge Devices Digging Deeper: The Future of Mining with Automation and Ultra-Reliable Wireless Voices from the field: Helping farmers build resilient local economies across rural America Built like a startup, scaled like Cisco: Transforming data center cooling for the AI era Defining Model Provenance: A Constitution for AI Supply Chain Safety and Security Introducing Model Provenance Kit: Know Where Your AI Models Come From Security Insights: A Threat-First View for the Platform That Enforces Access How I Turned My Curiosity into a Patent From Strategy to Architecture: How Cisco is Building a Quantum-Safe Future Maximizing Managed Security Services: A Strategic Guide to Optimizing Your Portfolio (Part 1 of 2) Simplify access control in five easy steps Trust: Why security is your next growth engine Cisco IQ is generally available. Here’s what that actually means. From Vision to Reality: Intelligence in Action with Cisco IQ How connectivity is shaping the future of surgical care The power of your network: Solving a physical security incident on Vision portal 5 signs your data center is holding your AI strategy back Stop Overthinking OT Security: The Total Cost of Ownership and Being Smart with Refreshes AI-Ready, Simpler, and More Secure WAN: Cisco SD-WAN Innovations Scaling the digital future: Why AI and skills investments matter for business and society Expanding our Product Organization Recap Scaling the Future: Reddit AMA on Network Automation at Scale Bringing Professional-Level Skills to Cisco Networking Academy Announcing Cisco Availability in Google Cloud Marketplace: A New Path to Scalable, Partner-Led Growth The Innovation Paradox: How We Reduced Incidents by 25% While Deploying Faster Funding the AI-ready data center: Why flexibility wins The switch that quantum networking has been waiting for From a Message I Couldn’t Believe to a Stage I’ll Never Forget The Hidden Bottleneck Slowing Down Manufacturing Transformation 30 Years as a CCIE: Why Certifications Matter in the AI Era Securing Enterprise AI: Cisco AI Defense Expands to Google Cloud How ThousandEyes Closed the Cloud Visibility Gap by Solving It Themselves First Energy Will Define the Scale of AI Introducing the AI Agent Security Scanner for IDEs: Verify Your Agents Stop Overthinking OT Security: People, Process and Technology Powering the Future of Research: Join Cisco at NLIT 2026 Building the Digital Foundation for a Smarter West Lincoln Memorial Hospital How Cisco built an AI-RRM that maximizes your wireless solution From Automation to Autonomy: Cisco and Rockwell Power a New Era for Manufacturing Unlocking the Future of Fan Engagement: The Power of VisionEDGE Find Yourself in the Future: AI Is the New Baseline—Here’s How to Build Your Skills One Day with Our Customers: Driving better outcomes through customer centricity What It Really Takes to Build an AI-First Workforce From Connectivity to Security: How E80 Future-proofed its AGV Operations with Cisco The Infrastructure of a Floating City: AIDA Cruises’ CX-Led Digital Transformation Scaling your network for AI without a forklift upgrade Why modern networks are moving DDoS defense to the edge Evolve IP Media to AI-Driven Media Fabrics: Future-Proof Broadcast with Cisco and NVIDIA Cisco and Generation are scaling AI-powered pathways to employment Reading Between the Pixels: Assessing Prompt Injection Attack Success in Images Lean IT, future-ready: Why Wi-Fi is your AI growth strategy Cisco Modeling Labs: Bringing the Network Digital Twin to Life AI on the Factory Floor: Why Manufacturing Requires a New Architecture with Cisco Unified Edge Designing for What’s Next: Securing AI-Scale Infrastructure Without Compromise Scaling the Future: Join Our Reddit AMA on Network Automation at Scale 5 wireless trends retail IT teams can’t ignore in 2026 Can your infrastructure management tools do that? Sustainability 101: Let’s talk about energy efficiency From Chai Breaks to Checkpoints: A Day at Cisco Bengaluru Preparing for Post-Quantum Cryptography: The Secure Firewall Roadmap Non-Obvious Patterns in Building Enterprise AI Assistants Making AI Trustworthy and Observable in Real-Time: Cisco Announces Intent to Acquire Galileo A simpler path to unified, AI-ready network operations Cisco Celebrates The Smart Industry Industrial Transformation Award Winners Mobile World Congress 2026: AI-powered Network Security Powering MWC Barcelona – Building a Unified SOC and NOC with Splunk in Record Time How New Data Streams Transformed Cisco Store’s Decision-Making AI-powered Network Security at the Mobile World Congress 2026 SNOC Inside the Mobile World Congress 2026 SOC: Detecting Shadow Traffic with Firepower 6100
Reading Between the Pixels: Failure Modes in Vision Language Models
Ravikumar Ba · 2026-05-06 · via Cisco Blogs

This post is Part 2 of a two-part series on multimodal typographic attacks.

In Part 1 of “Reading Between the Pixels,” we demonstrated that text–image embedding distance correlates with typographic prompt injection success: conditions that push a typographic image further from its source text in embedding space (small fonts, heavy blur, rotation) reduce attack success, and conditions that bring them closer increase it.

If embedding distance is a reliable predictor of attack success, we explored whether we could directly reduce the distance to make a failing attack succeed. We apply small, controlled changes to degraded typographic images so that a model interprets them as closer to their original text. The results reveal two distinct failure modes in vision language model (VLM) safety—readability recovery and refusal reduction—two issues that co-occur but vary depending on model and types of visual degradation.

The optimization we tested on images resulted in the effects of a successful typographic attack that evaded simple image filters, indicating a need for more robust defenses in the representation space.

From Correlation to Causation

Part 1 established a strong correlation between embedding distance and attack success rate (ASR) across four VLMs (r = −0.71 to −0.93, all p < 0.01). We investigated further and found that relationship between embedding distance and ASR is mediated by two factors:

  • Perceptual readability: Can the VLM parse the text in the image at all? A 6px font or heavy blur may render text unreadable to the model, causing it to fail before safety alignment even enters the picture.
  • Safety alignment: Even when the VLM can read the text, does it refuse to comply with the harmful instruction? Models like GPT-4o and Claude Sonnet 4.5 have strong safety filters that catch many harmful requests even when the text is perfectly legible.

We attempted to optimize these factors by reshaping the image’s representation in embedding space. Our research demonstrates that an attacker with the above goal could be recovering readability alone or also undermine the safety alignment depending on the robustness of the underlying model’s post-training alignment. 

The Approach

Our method is conceptually straightforward: take a degraded typographic image that currently fails as an attack (because the model cannot read it, refuses to comply, or both), and apply a small, bounded pixel-level perturbation that makes it “look more like the text” to an ensemble of multimodal embedding models (in our experiments, we used Qwen3-VL-Embedding, JinaCLIP v2, OpenAI CLIP ViT-L/14-336, SigLIP SO400M). Importantly, this optimization does not require access to the target VLM, its safety classifier, or any class labels—the text embedding of the attack prompt serves as a fixed target.

We adapt SSA-CWA (a technique that integrates a common weakness attack approach with spectrum simulation attack) to solve this optimization: meaning we allocate 100 steps to make perturbations to the input by a maximum of 12.5 percent (see Figure 1) to find the best way to alter an image’s pixels to elicit a successful attack.

Figure 1: Overview of embedding-guided adversarial optimization. A degraded typographic image is optimized via SSA-CWA across four surrogate embedding models. The resulting image is visually similar but semantically realigned, producing two co-occurring effects: readability recovery and refusal reduction.

What We Tested

We selected the same four VLMs from Part 1 as our target VLMs: GPT-4o, Claude Sonnet 4.5, Mistral-Large-3, and Qwen3-VL-4B using the same GPT-4o based refusal judge.

We evaluated across five degradation settings with low baseline ASR:

  • 6px font;
  • 8px font;
  • 90° rotation;
  • triple degradation (blur + noise + low contrast); and
  • heavy blur (σ=5).

And we selected 50 prompts where the text attack succeeds on both GPT-4o and Claude but the degraded image fails on both. Our selection focused on attacks that were successful in text form but blocked in image form, as well as heavily degraded images where OCR-based detection pipeline would struggle to flag a harmful intent.  A 6px font or heavily blurred image is hard for conventional text extraction to parse, so if a bounded perturbation makes the image readable to the VLM without restoring human- or OCR-legible text, the attack evades both the OCR filter and the model’s own safety alignment.

The results of our experiment are shown in Table 1 below. The optimization consistently increases ASR where baseline is lowest. Claude goes from 0% to 28% on heavy blur; GPT-4o from 0% to 16% on rotation. Interestingly, Mistral (already high baseline) sometimes drops — the perturbation can trigger safety mechanisms that were not previously engaged.

Table 1: Comparison of attack success rates before and after optimization under different degradation conditions (B: indicates before optimization, A: after optimization).

Two Failure Modes from One Objective

We observed two patterns during our research. The key finding is not just that the optimization works, but how. By classifying each failure as an explicit refusal, a readability failure, or other (misreading/tangential), we identify two distinct patterns:

Readability Recovery Dominates at 6px and Heavy Blur

At 6px font, 35 of GPT-4o’s 50 pre-optimization failures are readability-related, with only 10 refusals. After optimization, readability failures drop from 35 to 10, but refusals rise from 10 to 28, yielding only 1 net success. In other words: the perturbation makes the text readable, but GPT-4o’s safety filter catches the now-legible requests. GPT-4o’s safety alignment holds firm once readability is restored.

Claude on heavy blur tells a different story and shows the strongest overall gain (+28%). Before optimization, Claude returned empty responses on 39 out of 50 samples because it could not process the image at all. After optimization, empty responses drop from 39 to 14 as the perturbation makes the blurred content readable. Unlike GPT-4o, Claude’s safety filter does not fully catch the newly readable content, and a significant portion of these recovered readings result in compliance.

For Mistral, readability gains translate most directly to ASR: misreadings drop from 20 to 5 on heavy blur (+20%) and from 23 to 11 on 6px (+14%). This could be attributed to the safety alignment mechanism not being the bottleneck in the case of Mistral, wherein making the text readable is sufficient to succeed.

Mixed Readability and Refusal Reduction at Rotation and 8px

At 90° rotation, our observations are more nuanced. GPT-4o has 28 refusals but also 12 readability failures, indicating that 20px rotated text is not fully legible even at a reasonable font size. After optimization, readability failures drop (12 to 7) and 8 successes emerge (+16%). Claude gains +22%, with refusals dropping from 27 to 19 and empty responses from 13 to 8. At 8px, a similar pattern holds: GPT-4o and Claude gain +10% and +8% from a 0% baseline.

This is the most concerning pattern from a safety perspective: in these settings, the VLM can partially read the text but refuses to comply. The perturbations to the image do not improve visual legibility to a human observer, but the tweaked image confuses the model’s safety reasoning, causing the model to shift from refusal to compliance. The model’s safety decision breaks down when faced with small changes in the input representation.

What This Means for Practitioners

The optimization results reveal two distinct artifacts an attacker can recover through bounded perturbations, each exploiting a different gap in VLM safety:

  • Artifact 1: Degraded images that evade OCR detectors while remaining model-readable. When a model fails to read the original image (small font, heavy blur, rotation), a bounded perturbation can recover semantic content in the model’s internal representation without restoring visual legibility to a human. This means an attacker can craft images that look like noise or illegible distortion to any OCR-based content filter yet carry fully readable instructions to the target VLM.
  • Artifact 2: Refusal suppression transferred from successful prompt injections. When the VLM already reads the text but refuses, bounded perturbations can shift the safety decision boundary. The key insight is that the perturbation patterns learned from successful attacks in one configuration (e.g., a particular font size or degradation type) generalize — an attacker can exploit what works in one modality or setting and apply it to suppress refusals in others. This means safety alignment that holds for clean inputs can be systematically eroded by perturbations informed by prior successes, without requiring model internals.

These two artifacts compound: an attacker does not need access to the target model to elicit this response. By generating the perturbations on a diverse group of substitute models (Artifact 1), the attacks generated can transfer to proprietary target models without ever interacting with them (Artifact 2). Together, they form a pipeline from evading detection to achieving compliance.

Looking Ahead

Multimodal embedding distance emerges as a powerful lens for understanding typographic prompt injection. Part 1 showed it correlates with attack success; Part 2 shows it can be weaponized to expose two co-occurring fragilities. The finding that bounded perturbations can reduce refusal rates without improving visual legibility points to a need for safety mechanisms robust in representation space, not just the pixel domain.