惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
MyScale Blog
MyScale Blog
Recent Announcements
Recent Announcements
N
Netflix TechBlog - Medium
GbyAI
GbyAI
Vercel News
Vercel News
The GitHub Blog
The GitHub Blog
阮一峰的网络日志
阮一峰的网络日志
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
V
Visual Studio Blog
Martin Fowler
Martin Fowler
腾讯CDC
大猫的无限游戏
大猫的无限游戏
aimingoo的专栏
aimingoo的专栏
云风的 BLOG
云风的 BLOG
J
Java Code Geeks
WordPress大学
WordPress大学
P
Proofpoint News Feed
雷峰网
雷峰网
酷 壳 – CoolShell
酷 壳 – CoolShell
有赞技术团队
有赞技术团队
人人都是产品经理
人人都是产品经理
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Y
Y Combinator Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
I Ran Five Small Multimodal Models on a Jetson. The Faste...
Ryan Hsu · 2026-06-18 · via DEV Community

Ryan Hsu

I have been building WearEdge Pro, a wearable industrial edge AI runtime. Think of a frontline operator wearing a smart-glasses device, capturing a first-person image of a machine, and getting back a structured action card from a local Jetson box.

The key phrase is "structured action card." This is not a chat demo. In a factory setting, an answer needs an audit trail, a mode boundary, a human-confirmation gate, and a way to hand off to maintenance, quality, EHS, or work-instruction workflows.

I recently tested five compact multimodal models on the same Jetson path:

  • Gemma 4 E2B
  • Qwen2.5-VL-3B
  • SmolVLM2-2.2B
  • InternVL3-2B
  • Qwen2.5-Omni-3B

The goal was not to crown a universal benchmark champion. I wanted to know which model was the best current baseline for an industrial edge Agent runtime.

The Harness

Every model was exposed through a local OpenAI-compatible llama.cpp endpoint on the Jetson. Each model got the same five prompts and images:

  • maintenance
  • quality inspection
  • changeover
  • work instruction
  • hazard review

The main run used 560 image tokens, which matches the current WearEdge gateway budget. Qwen2.5-VL also got a 1024-image-token pass because grounding can improve with more visual tokens.

The Results

Model Completion Avg latency Takeaway
Gemma 4 E2B 5/5 37.51s raw Best product baseline
Qwen2.5-VL-3B 5/5 39.72s Best OCR challenger
SmolVLM2-2.2B 5/5 12.84s Fastest, but weak grounding
InternVL3-2B 5/5 only after ctx4096 80.35s Too slow/risky for baseline
Qwen2.5-Omni-3B 5/5 50.09s Interesting future audio/video branch

SmolVLM2 was the speed star. But the answers were often too generic for real operator guidance. In changeover and work-instruction tasks, it returned fields that looked more like placeholders than grounded industrial guidance.

Qwen2.5-VL was the most impressive challenger. It nailed a changeover OCR task with LABELER-FL1 and SKU-C500, where Gemma had a machine-label typo. It also produced useful IQC defect scores. If I were building a pure OCR or visual inspection assistant, I would take Qwen very seriously.

InternVL3 reminded me that token speed is not the whole story. At 2048 context it failed three of five tasks with context errors. At 4096 context it finished, but the latency was high and one raw IQC answer had unsafe release-style wording.

Qwen2.5-Omni ran cleanly, but its strongest value is probably a future audio/video workflow rather than this current image+text industrial baseline.

Why Gemma Still Won

Gemma 4 E2B did not win every micro-test. It stayed the baseline because it fit the product runtime:

  • local Jetson deployment
  • structured multimodal prompts
  • long-context workflow design
  • function-calling-oriented architecture
  • deterministic guards
  • human confirmation
  • action cards
  • audit logs

In an industrial setting, "fast and fluent" is not enough. The model has to behave inside a system that can say: this came from this image, this route, this required field, this action boundary, and this audit record.

That is why Gemma remained the WearEdge baseline, while Qwen2.5-VL became the serious A/B challenger for OCR-heavy branches.

Lesson Learned

Edge AI model selection is not just a leaderboard exercise. The right question is:

Can this model run locally, understand the evidence, obey the workflow boundary, and produce an action that the system can audit?

For WearEdge Pro today, the answer is Gemma 4 E2B as the baseline, Qwen2.5-VL as the next challenger, and a clear path to keep testing without pretending every benchmark cell means the same thing.

Public artifact link: Benchmark results and public discussion: https://www.hackster.io/ryanon2008/wearedge-pro-jetson-edge-ai-agent-50ec35