惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

I
InfoQ
H
Heimdal Security Blog
罗磊的独立博客
B
Blog RSS Feed
WordPress大学
WordPress大学
The Register - Security
The Register - Security
N
Netflix TechBlog - Medium
美团技术团队
量子位
GbyAI
GbyAI
Recent Announcements
Recent Announcements
博客园 - 叶小钗
D
DataBreaches.Net
S
SegmentFault 最新的问题
Hacker News - Newest:
Hacker News - Newest: "LLM"
T
Troy Hunt's Blog
The Last Watchdog
The Last Watchdog
O
OpenAI News
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
酷 壳 – CoolShell
酷 壳 – CoolShell
Webroot Blog
Webroot Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Last Week in AI
Last Week in AI
V
V2EX
N
News and Events Feed by Topic
Jina AI
Jina AI
Y
Y Combinator Blog
T
The Blog of Author Tim Ferriss
IT之家
IT之家
C
Check Point Blog
H
Hacker News: Front Page
爱范儿
爱范儿
Schneier on Security
Schneier on Security
Apple Machine Learning Research
Apple Machine Learning Research
P
Privacy & Cybersecurity Law Blog
L
LINUX DO - 最新话题
Forbes - Security
Forbes - Security
人人都是产品经理
人人都是产品经理
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Microsoft Azure Blog
Microsoft Azure Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
S
Secure Thoughts
The Cloudflare Blog
Simon Willison's Weblog
Simon Willison's Weblog
Stack Overflow Blog
Stack Overflow Blog
腾讯CDC
MongoDB | Blog
MongoDB | Blog
V2EX - 技术
V2EX - 技术
AI
AI

Crazyrouter Blog

Gemini CLI Complete Guide 2026: Repo Automation, CI Agents, and Multi-Model Routing Ideogram AI Guide 2026: Brand Design Automation, API Workflows, and Alternatives GLM 4.6 API Guide 2026: Agents, RAG, Tool Calling, and Bilingual Apps WAN 2.2 Animate Tutorial 2026: Character Consistency, Shot Control, and API Workflows Google Veo3 API Guide 2026: Production Video Pipelines, Prompts, Pricing, and Fallbacks AI API Pricing Comparison 2026: Text, Image, Video, Caching, and Router Costs Codex CLI Installation Guide 2026: Windows, macOS, Linux, Proxies, and CI Setup How to Get a Claude API Key in 2026: Secure Setup for Teams, CI, and Alternatives Gemini Advanced Review 2026: Is It Worth It for Coding, Research, and API Teams? Seedance 2.0 Pricing: Convert 46 CNY per Million Tokens to Cost per Second Seedance 2.0 计费详解:46元/百万Token换算成每秒多少钱 Seedance 2.0料金解説:100万Tokenあたり46元を1秒あたりコストに換算 Gemini CLI 使用教程 2026:安装、代码示例、代理环境与 API 接入 Gemini 是什么?2026 完整介绍、API 使用教程与价格对比 Qwen2.5-Omni Guide 2026: Real-Time Voice, Vision, and Multimodal Agents Kimi K2 Thinking Guide 2026: Reasoning Workflows, Evals, and Cost Control Google Veo3 API Guide 2026: Batch Video Pipelines, Pricing, and Fallbacks Codex CLI Installation Guide 2026: macOS, Linux, WSL, Proxies, and Dev Containers How to Get a Claude API Key in 2026: Safe Production Setup and Alternatives AI API Pricing Comparison 2026: GPT, Claude, Gemini, Video, and Agent Workloads Gemini Advanced Review 2026: Is It Worth It for Developer Teams? Claude Code Pricing Guide 2026: API Fallbacks, Team Seats, and Budget Control Seedream 4.0 API Tutorial 2026: Batch Image Generation, Product Creative, and Pricing Qwen2.5-Omni Guide 2026: Real-Time Voice, Vision, Text Agents, and API Integration Kimi K2 Thinking Guide 2026: Reasoning Agents, Evaluation Workflows, and API Cost Control WAN 2.2 Animate Tutorial 2026: Character Motion, Shot Control, API Pipelines, and Pricing Google Veo3 API Guide 2026: Production Video Workflows, Prompts, Pricing, and Fallbacks AI API Pricing Comparison 2026: OpenAI, Claude, Gemini, DeepSeek, and Router Costs How to Get a Claude API Key in 2026: Setup, Security, Rotation, and Alternatives Codex CLI Installation Guide 2026: macOS, Linux, WSL, Proxies, and Devcontainers Gemini Advanced Review 2026: Is It Worth It for Developers and API Builders? Claude Code Pricing Guide 2026: CI Agents, Team Seats, and API Budget Planning 一個 API Key 呼叫 GPT、Claude、Gemini:5 分鐘設定教學 AI API Gateway for Singapore and Malaysia Developers: One Endpoint for GPT, Claude and Gemini AI API Gateway for Thai Developers: Use GPT, Claude and Gemini with One Key Cómo usar GPT, Claude y Gemini con una sola API key One API Key for GPT, Claude and Gemini: A Practical Setup for Central Asia Developers Gemini 3.5 Flash vs Claude レスポンスティアモデル:開発者はどちらを選ぶべきか Gemini 3.5 Flash vs Claude Response-Tier Models: Какую модель выбрать разработчику? Gemini 3.5 Flash vs Claude Response-Tier Models: Which One Should Developers Use? Gemini 3.5 Flash vs Gemini 3 Flash vs Gemini 2.5 Flash:実運用APIベンチマーク Gemini 3.5 Flash vs Gemini 3 Flash vs Gemini 2.5 Flash: Real API Benchmark text-embedding-3-large 值不值得用?和 text-embedding-3-small 的成本、效果与选型对比 用 text-embedding-3-large 搭建 RAG 知识库:从切块、向量化到检索排序 text-embedding-3-large 是干什么的?Embedding 模型入门与 RAG 场景详解 AI 扩图 API 指南 2026:Uncrop、Outpaint、gpt-image-2 和 Nano Banana 路线怎么选 How to Test Multiple AI Image Models with One API Key "How to Test Multiple AI Image Models with One API Key" Codex CLI Installation Guide: Setup on macOS, Linux, Windows WSL and CI/CD Gemini CLI 使用教程:开发者终端 AI 助手完全指南 Grok 4 免费使用教程:合法体验路径、API 接入与替代方案 Seedream 4.0 API Tutorial: ByteDance Image Generation for Production Pipelines Kimi K2 Thinking Model: Complete Developer Guide for Reasoning Workflows Luma Ray 2 Review: AI Video Generation Quality, Speed, and API Guide Pika 2.2 New Features Review: Scene Director, Sound Design, and API Updates Google Veo 3 API Guide: Video Generation with Audio for Developers AI Lip Sync Tools Comparison 2026: Best APIs for Talking Avatars and Video Dubbing Gemini Advanced Review May 2026: Is It Worth $20/Month for AI Power Users? Claude Code Pricing in May 2026: Max Plan, Opus 4, and Real Cost Breakdown Hermes Agent + Crazyrouter: One-Click Setup for 627+ AI Models Text-Embedding-3-Small: Complete Guide to OpenAI's Most Popular Embedding Model (2026) Cursor 配置 Crazyrouter 教程:国内用上 GPT-5.4 / Claude 写代码 2026 年国内如何调用 Claude API?Claude Opus / Sonnet 接入完全指南 2026 年国内如何调用 GPT-5.4 API?完整接入指南(含代码示例) AI API 常见报错排查大全:401、429、500、timeout 一篇搞定 2026 年 AI API 中转站哪家好?六大平台横向对比评测 2026 年 DeepSeek R1 API 接入指南:国内最强推理模型怎么调用 Trình Tạo Meme & Sách Tô Màu Bằng AI Với GPT-image-2 — Những Dự Án Vui Mà Vẫn Kiếm Ra Tiền Dự Đoán Em Bé Tương Lai Bằng AI Với GPT-image-2 — Xem Con Bạn Có Thể Trông Như Thế Nào Chuyển Đổi Ảnh Sang Phong Cách Ghibli Với GPT-image-2 — Biến Mọi Bức Ảnh Thành Tranh Anime Tạo Mô Hình Nhân Vật Hành Động Bằng AI Với GPT-image-2 — Biến Bất Kỳ Ai Thành Đồ Chơi Trong Hộp GPT-image-2: Nhận Diện Khuôn Mặt Và Phân Tích Màu Sắc Bằng AI Xem chỉ tay với GPT-image-2 — Tạo bản phân tích chỉ tay chuyên nghiệp chỉ từ một bức ảnh GPT-image-2로 AI 밈 생성기 & 컬러링북 만들기 — 재미있고 수익도 되는 프로젝트 GPT-image-2로 AI 미래 아기 예측 — 우리 아이는 어떤 모습일까? GPT-image-2로 지브리 스타일 변환 — 사진을 애니메이션 아트로 바꾸기 GPT-image-2로 AI 액션 피규어 생성하기 — 누구나 박스형 피규어로 바꾸는 법 GPT-image-2로 AI 관상 분석 & 퍼스널 컬러 진단 — 두 가지 바이럴 활용법 완벽 가이드 GPT-image-2 실전 가이드:AI 손금 분석 — 손바닥 사진 한 장으로 전문 손금 인포그래픽 생성하기 GPT-image-2 で AI ミーム生成 & ぬりえブック制作 — 楽しくて本当に稼げるプロジェクト GPT-image-2 で AI 未来の赤ちゃん予測 — 将来の子どもの顔を見てみよう GPT-image-2 でジブリ風写真変換 — どんな写真もアニメアートに GPT-image-2 で AI アクションフィギュア生成 — 誰でもボックス入りおもちゃに変身 GPT-image-2 で AI 顔相診断 & パーソナルカラー分析 — 2つのバズ活用法を1本で解説 GPT-image-2 で AI 手相占い — 1枚の写真からプロ仕様の手相分析を生成 GPT-image-2 на практике: AI-генератор мемов и раскрасок — весёлые проекты, которые приносят деньги GPT-image-2 на практике: AI-предсказание будущего ребёнка — как будет выглядеть ваш малыш GPT-image-2 на практике: стиль Гибли — превратите любое фото в аниме-арт GPT-image-2 на практике: AI-генератор фигурок — превратите себя в коллекционную игрушку GPT-image-2 на практике: AI-физиогномика и анализ цветотипа — два вирусных кейса в одном гайде GPT-image-2 на практике: AI-хиромантия — генерация профессионального анализа ладони по фото GPT-image-2 实战:AI Meme 生成器 & 涂色书制作 — 好玩还能赚钱的两个项目 GPT-image-2 实战:AI 预测未来宝宝 — 看看你们的孩子长什么样 GPT-image-2 实战:吉卜力风格转换 — 把任何照片变成宫崎骏动画 GPT-image-2 实战:AI 手办生成器 — 把任何人变成盒装公仔 GPT-image-2 实战:AI 面相分析 & 个人色彩诊断 — 两大爆款玩法一文搞定 GPT-image-2 实战:AI 看手相 — 一张手掌照片生成专业手相分析图 AI Meme Generator & Coloring Book Creator with GPT-image-2 — Fun Projects That Actually Make Money AI Future Baby Prediction with GPT-image-2 — See What Your Child Might Look Like Ghibli Style Photo Transformation with GPT-image-2 — Turn Any Photo Into Anime Art
youtu-vita OCR Benchmark 2026: Live Test Results on Documents, Receipts, UI Screens, and Small Text
Crazyrouter Team · 2026-06-24 · via Crazyrouter Blog

youtu-vita OCR Benchmark 2026: Live Test Results on Documents, Receipts, UI Screens, and Small Text#

If you are evaluating OCR-capable multimodal models for production work, the question is not just whether a model can read text in an image. The real question is whether it can do it consistently, with structured outputs, across the kinds of inputs teams actually send in production: screenshots, receipts, tables, rotated documents, scene text, and low-resolution UI captures.

We ran a live benchmark for youtu-vita through our OpenAI-compatible API path and scored it on a controlled OCR test set. This article shares the actual test data, what the model passed, where it struggled, and what kind of OCR workloads it looks good at right now.

Test setup#

This benchmark was run on June 24, 2026.

We used a local generated OCR benchmark set with these eight cases:

  1. document_basic
  2. receipt_total
  3. ui_settings
  4. table_statement
  5. scene_text_signboard
  6. rotated_document
  7. low_res_small_text
  8. chart_with_legend

The benchmark images were generated locally so the test would be stable and reproducible, rather than depending on external image hosts. Each request used the same OpenAI-compatible image input shape and the same JSON-only output instruction. The model was asked to:

  • transcribe visible text
  • answer structured OCR questions
  • return machine-readable JSON

The benchmark files used for this run:

  • .tmp/ocr_compare_manifest.generated.json
  • .tmp/ocr_benchmark_assets/
  • .tmp/ocr_model_compare_20260624_153400.json

Scoring method#

Each case was scored on four dimensions:

  1. Format stability Did the model return valid structured JSON?
  2. OCR text match Did it correctly capture the required visible text?
  3. Regex-sensitive fields Did it preserve exact formats for fields like IDs or totals?
  4. Structured answers Did it answer the requested key fields correctly?

The final score per case is a weighted total. A score of 1.000 means the model fully passed that case under this benchmark.

youtu-vita live benchmark results#

The full 8-case run completed successfully.

Headline metrics#

  • Success rate: 100%
  • Average total score: 0.875
  • p50 latency: 3794 ms
  • p90 latency: 4088 ms
  • Slowest case: 15284 ms

Case-by-case results#

Case IDCategoryStatusScoreLatency
document_basicDocument OCR2001.0003852 ms
receipt_totalReceipt OCR2001.0003616 ms
ui_settingsUI screenshot OCR2001.0003794 ms
table_statementTable OCR2001.0003434 ms
scene_text_signboardScene text2001.0001661 ms
rotated_documentRotated document2001.0002908 ms
low_res_small_textSmall text / low resolution2001.0004088 ms
chart_with_legendChart reasoning2000.00015284 ms

What youtu-vita did well#

For this run, youtu-vita was strong on the OCR tasks most teams care about first:

1. Clean document OCR#

It correctly extracted:

  • title
  • date
  • document ID
  • paragraph text

On document_basic, it returned a full structured transcription and correctly answered:

  • title = Quarterly Operations Summary
  • date = 2026-06-24
  • document_id = AB-123456

2. Receipt reading#

It handled receipt-style layout correctly on receipt_total, including:

  • merchant name
  • total amount
  • receipt number

That matters because receipt OCR often breaks on spacing, alignment, or repeated numeric fields. In this run, youtu-vita passed the receipt case with a full score.

3. UI screenshot OCR#

On ui_settings, it correctly captured:

  • page title
  • button labels
  • error code
  • supporting text

It also returned the structured answers we asked for:

  • primary_cta = Continue
  • error_code = E102

That makes it promising for support automation, QA workflows, screen parsing, and screenshot-based extraction tasks.

4. Table OCR#

On table_statement, the model passed the table case with a full 1.000 score in this run.

That is important because table OCR is often where vision models look good at plain text but fail at row-column alignment. In this benchmark, youtu-vita handled the table extraction cleanly enough to pass both the visible text requirements and the structured answer checks.

5. Rotated documents#

On rotated_document, the model also scored 1.000.

That suggests it is not limited to perfectly upright scanned pages. If your OCR workflow includes phone photos, skewed uploads, or documents captured in the wild, this is a meaningful result.

6. Low-resolution small text#

One of the most practically useful passes in this run was low_res_small_text, which also scored 1.000.

That case is closer to real dashboard and UI OCR than a clean printed PDF. If you need to read release notes, settings screens, logs, or admin panels from screenshots, this is a positive signal.

Where youtu-vita was weak#

The weak spot in this run was not standard OCR. It was chart reasoning.

On chart_with_legend, youtu-vita returned HTTP 200 but scored 0.000. It also took much longer than the rest of the test set at 15284 ms.

That tells us two things:

  1. The model can complete the request, but this benchmark did not show reliable performance on chart interpretation.
  2. OCR and chart understanding should be treated as separate capabilities.

This matters because many teams group all “image understanding” into one bucket. That is too coarse. A model can be strong at:

  • OCR
  • receipt parsing
  • screenshot reading
  • text extraction

and still be weak at:

  • chart reasoning
  • visual analytics
  • higher-order graph interpretation

Practical interpretation#

Based on this live run, youtu-vita looks strongest for these workloads:

Good fit#

  • document OCR
  • receipt OCR
  • UI screenshot extraction
  • rotated page reading
  • small-text screenshot parsing
  • sign and scene text extraction

Not yet a strong conclusion#

  • chart understanding
  • graph interpretation
  • analytics-style visual reasoning

If your workload is mostly “read the text, extract the fields, give me clean JSON,” this benchmark suggests youtu-vita is already useful.

If your workload is “understand a chart, infer trends, compare the latest month, and reason visually,” this benchmark does not support calling it strong there yet.

Why this matters for production teams#

Many OCR evaluations are too soft. They say a model is “good at image understanding” after a single logo or document test. That does not help if you need to decide whether to route:

  • support screenshots
  • invoices
  • receipts
  • internal ops dashboards
  • phone photos of documents

into a production OCR pipeline.

This benchmark is more useful because it separates:

  • OCR stability
  • structured extraction
  • small-text robustness
  • rotated-input handling
  • chart reasoning

For this run, youtu-vita was clearly stable across the OCR-heavy categories.

Final verdict#

For this live benchmark, youtu-vita was the most stable OCR model we tested in the current environment.

The actual results were:

  • 100% request success across the full 8-case OCR set
  • 0.875 average total score
  • 1.000 on 7 of 8 OCR-oriented cases
  • failure only on the chart reasoning case

That makes youtu-vita a strong candidate if your main need is text extraction from images, especially for:

  • business documents
  • receipts
  • UI screenshots
  • low-resolution text
  • rotated pages

It does not mean it is the best choice for every kind of vision workload. But if your problem is OCR, not visual analytics, this is one of the clearest positive live results we have seen in this environment.

Reproduce the test#

If you want to run the same OCR benchmark shape yourself, we used:

The local result summary used for this article:

If you want a follow-up article comparing youtu-vita directly against Gemini or GPT-family vision models on the same OCR set, that should be written as a stability + latency + category-by-category comparison, not just a single overall score.