惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Engineering at Meta
Engineering at Meta
罗磊的独立博客
Apple Machine Learning Research
Apple Machine Learning Research
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
有赞技术团队
有赞技术团队
小众软件
小众软件
GbyAI
GbyAI
J
Java Code Geeks
B
Blog
Blog — PlanetScale
Blog — PlanetScale
宝玉的分享
宝玉的分享
M
MIT News - Artificial intelligence
T
The Exploit Database - CXSecurity.com
Security Latest
Security Latest
阮一峰的网络日志
阮一峰的网络日志
NISL@THU
NISL@THU
T
Tenable Blog
S
Schneier on Security
T
Tor Project blog
V
V2EX
S
Secure Thoughts
P
Privacy International News Feed
Spread Privacy
Spread Privacy
博客园 - 三生石上(FineUI控件)
博客园_首页
T
Threatpost
月光博客
月光博客
Know Your Adversary
Know Your Adversary
Cyberwarzone
Cyberwarzone
T
The Blog of Author Tim Ferriss
H
Help Net Security
The Hacker News
The Hacker News
A
Arctic Wolf
B
Blog RSS Feed
雷峰网
雷峰网
The Last Watchdog
The Last Watchdog
S
SegmentFault 最新的问题
博客园 - 【当耐特】
Cisco Talos Blog
Cisco Talos Blog
Vercel News
Vercel News
Microsoft Security Blog
Microsoft Security Blog
A
About on SuperTechFans
P
Palo Alto Networks Blog
Attack and Defense Labs
Attack and Defense Labs
N
News | PayPal Newsroom
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
S
Securelist
L
LangChain Blog
I
InfoQ

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Mobile AI Features That Never Send User Data to a Server: What Is Possible on iOS and Android in 2026
Mohammed Ali · 2026-04-26 · via DEV Community

This piece was written for enterprise technology leaders and originally published on the Wednesday Solutions mobile development blog. Wednesday is a mobile development staffing agency that helps US mid-market enterprises ship reliable iOS, Android, and cross-platform apps — with AI-augmented workflows built in.

Your CISO wants AI that stays on the device. Here is the complete list of what on-device AI can actually do in 2026, what it cannot do yet, and the accuracy data behind each claim.


Your CISO wants AI that never sends user data to a server. The product team wants AI that actually works. In 2026, these requirements are no longer in conflict for most enterprise mobile use cases. Here is exactly what on-device AI can do, on what hardware, at what accuracy level — based on production systems, not benchmarks.

Key findings
Current on-device models achieve 88-94% accuracy on enterprise text classification tasks. For field documentation, clinical note structuring, and internal productivity, that accuracy is production-grade.
On-device Whisper achieves 95%+ word accuracy on English speech. Wednesday shipped on-device voice transcription in Off Grid with no server dependency across 50,000+ users.
Wednesday's Off Grid ships text, voice, image generation, vision-language, and document Q&A entirely on-device — iOS, Android, and macOS from a single React Native app. These are not prototype features. They are production features with real users.
What on-device cannot do: real-time knowledge, very long document analysis, and the highest-accuracy multi-step reasoning tasks. For most enterprise mobile use cases, none of these limitations apply.

What on-device means in 2026

"On-device" means the model weights are stored locally on the device. All inference — the process of running user input through the model to produce output — happens on the device processor. No API call. No network request. No data leaving the device during AI use.

This was a theoretical statement for most enterprise use cases three years ago. In 2026, it is a practical one. The combination of capable open-source models (Llama 3, Phi-4, Gemma 2, Mistral 7B, Whisper), efficient inference frameworks (llama.cpp, Core ML, QNN, MNN), and hardware that has caught up with the model requirements (A15+ NPU on iOS, Snapdragon 8 Gen 1+ on Android) means that production-quality AI runs locally on current devices.

Wednesday built Off Grid to prove this with production users, not benchmarks. 50,000+ users run text AI, voice transcription, image generation, and vision features in Off Grid with no cloud inference calls. The capabilities described in this article are what those users experience daily.

Text AI on-device

Text AI covers the largest category of enterprise mobile AI use cases. On-device 7B parameter models handle:

Documentation assistance. A field technician describes a repair verbally and in short notes. On-device AI structures those notes into a formatted work order. A clinician speaks rough observations and the AI organises them into a structured clinical note format. Accuracy on enterprise text structuring tasks: 88-94%.

Summarisation. A sales rep reviews 20 customer interactions before a quarterly review call. On-device AI summarises each interaction into two sentences. A manager needs a summary of the week's field reports. On-device AI processes the reports and produces a summary. Context window limits apply — 4,000-8,000 tokens per call covers most individual document summarisation tasks.

Classification. Support tickets classified by urgency and category. Customer feedback classified by sentiment and topic. Work orders classified by job type and priority. Classification tasks are among the strongest on-device AI use cases — accuracy of 90-95% is achievable with appropriate model selection.

Named entity extraction. Extracting key information from unstructured text: customer names from calls, part numbers from field notes, medication names from clinical text, dates and amounts from financial documents. On-device models handle this well for common entity types.

Short-form drafting. Generating first drafts of standard emails, reports, or notifications from structured inputs. Response length and quality are appropriate for enterprise internal communication; on-device models are not suitable for polished long-form content generation.

Conversational Q&A over provided context. A technician asks "what is the maintenance interval for this equipment?" and the app retrieves the relevant manual section and passes it to the on-device model as context. The model answers the question using only the provided context, not general knowledge. This is the on-device RAG pattern and it works well for knowledge bases under 500MB.

Voice transcription on-device

On-device Whisper is the standard for enterprise voice transcription features that must not send audio to a server.

Whisper achieves 95%+ word accuracy on clear English speech at 1.5-3x real-time processing speed on current flagship hardware. A 5-minute dictation is transcribed in 100-200 seconds.

For enterprise use cases:

  • Field service documentation: strong accuracy in moderate noise environments. Background machinery noise reduces accuracy; extreme industrial noise environments require testing with representative audio.
  • Clinical documentation: strong accuracy on standard medical terminology with the medium or large model variant. Very specialised terminology (rare surgical procedures, uncommon drug names) may benefit from a domain-adapted Whisper fine-tune.
  • Sales and customer interaction: strong accuracy on phone-quality audio, clear English, and standard business vocabulary.
  • Legal dictation: strong accuracy on standard legal vocabulary. Unusual case citations or rare jurisdictional terminology may require domain adaptation.

Wednesday shipped on-device Whisper in Off Grid. The implementation handles variable noise environments, devices without NPU acceleration (using CPU inference at slower speed), and language variation. It is a production implementation across 50,000+ users, not a demonstration.

Image and vision AI on-device

Image AI runs on-device using hardware-accelerated backends: Core ML on iOS (Metal GPU), QNN on Snapdragon 8 Gen 1+ (NPU), and MNN on ARM64 Android (CPU, with SIMD optimisation).

Image generation. Diffusion models (LCM, SDXL Turbo) generate images in 10-30 seconds on current flagship hardware. Quality is suitable for product visualisation, simple illustrations, and content creation. Not suitable for photorealistic generation or complex compositions requiring very high resolution. Wednesday shipped three production image generation backends in Off Grid — the only known production implementation across all three hardware backends from a single React Native app.

Image classification. Classifying images into predefined categories — equipment condition ratings, product quality tiers, damage assessments, plant species identification — runs on MobileNet, EfficientNet, and similar lightweight architectures with 94-98% accuracy on well-defined classification tasks. This is among the most mature on-device AI capability.

Object detection. Identifying and locating specific objects in images — parts, products, text, barcodes — runs on YOLOv8 and similar architectures at real-time speed on current devices. Enterprise use cases: quality inspection, inventory management, equipment identification in the field.

Vision-language models. Models that can answer questions about an image — "what is wrong with this equipment?" or "what does this document say?" — now run on-device with 7B-class vision-language models. Accuracy is lower than GPT-4o vision for complex scenes but suitable for structured enterprise visual inspection tasks. Wednesday shipped vision-language features in Off Grid.

Face detection (local only). Detecting whether a face is present in an image, without identification or matching against any database, runs on-device without any privacy concern. This is the only face-related AI that should be on-device in enterprise apps — face recognition against external databases requires careful compliance review regardless of where inference runs.

Document analysis on-device

Document analysis covers features that extract, understand, or answer questions about documents loaded into the app.

Document Q&A. Load a PDF, specification, policy document, or manual. Ask questions about it. On-device embedding converts the document to a local vector index; on-device inference answers questions using retrieved passages as context. Works for documents up to approximately 500MB. Wednesday shipped document Q&A in Off Grid using this pattern.

Form and table extraction. Extracting structured data from scanned forms, tables, and documents. Combination of on-device OCR and on-device text processing. Accuracy: 88-93% on well-structured forms; lower on handwritten or degraded originals.

Document classification. Identifying the type, category, or routing of a document without reading its full content. Suitable for enterprise document management features where documents need to be sorted or routed automatically.

Translation. On-device translation for common language pairs (English-Spanish, English-French, English-Portuguese, English-Mandarin, and others in the top 20 language pairs) using NLLB and similar models. Quality is production-grade for standard business text. Less common language pairs have lower accuracy.

What on-device cannot do in 2026

Honesty about the limitations matters. These are the enterprise use cases where on-device AI is not the right answer today.

Real-time knowledge. On-device models have a training cutoff. They do not know about events, regulatory changes, or product updates that occurred after the training data was collected. A customer service AI that needs to answer questions about current product pricing or a compliance assistant that must reflect this week's regulation cannot be purely on-device.

Very long document analysis. Processing a 200-page contract or a full year of financial statements in a single inference call requires a 100,000+ token context window. On-device 7B models typically handle 4,000-8,000 tokens. For very large document tasks, cloud APIs with 128,000+ token context windows are required.

Complex multi-step reasoning. Tasks that require many steps of logical inference — "given these 10 constraints, determine the optimal allocation" — where current on-device 7B models lag GPT-4o class cloud models significantly. Most enterprise mobile tasks do not require this level of reasoning, but complex analytical tasks do.

High-accuracy multi-speaker real-time transcription. Identifying who said what in a live multi-person conversation, in real time, with high accuracy, is not reliably achievable on-device in 2026. Cloud APIs with speaker diarisation and streaming transcription remain ahead for this specific capability.

Full capability table

Capability On-device feasible Accuracy Best for Not suitable for
Text structuring and documentation Yes 88-94% Field notes, clinical documentation Very long documents
Text summarisation Yes (under 6,000 words) Strong Individual documents, reports Book-length analysis
Text classification Yes 90-95% Ticket routing, sentiment Ambiguous, multi-class edge cases
Named entity extraction Yes 88-93% Common entities (names, dates, numbers) Highly specialised terminology
Voice transcription Yes 95%+ English, clear speech Extreme noise, rare vocabulary
Image classification Yes 94-98% Defined category sets Open-world classification
Object detection Yes Strong Known object categories Unconstrained open-world
Image generation Yes (flagship, 10-30s) Suitable for illustration Product viz, simple content Photorealistic, high-res
Vision-language Q&A Yes Moderate Structured inspection Complex scene description
Document Q&A Yes (under 500MB) Strong Manuals, policies, specifications Very large knowledge bases
Translation Yes (top 20 pairs) Production-grade Standard business text Rare language pairs
Real-time knowledge No News, live data, current events
100K+ token context No Very long document analysis
Multi-speaker real-time transcription Not reliably Live meeting captioning

Read more case studies at mobile.wednesday.is/work

How Wednesday has shipped all of this in production

Off Grid is not a proof of concept. It is a production product with 50,000+ users, 1,700+ GitHub stars, and publicly auditable architecture.

Text AI: llama.cpp on CPU with Core ML and QNN acceleration where available. Multi-turn conversation, document Q&A, and text structuring all running locally.

Voice AI: on-device Whisper across the full device compatibility matrix. Handles variable noise environments and non-NPU devices via CPU inference fallback.

Image AI: three backend implementations — MNN for ARM64 Android, QNN with NPU for Snapdragon 8 Gen 1+, Core ML for iOS. All three running production image generation for 50,000+ users.

Vision AI: vision-language models for image understanding and Q&A, running on-device with the same backend infrastructure as image generation.

This is the reference implementation for enterprise teams whose CISO needs to understand what on-device AI is capable of, what it requires to ship, and how it handles the device matrix. It is not a vendor's capability claim. It is a public product with verifiable user numbers.


Want to go deeper? The full version — with related tools, case studies, and decision frameworks — lives at mobile.wednesday.is/writing/mobile-ai-features-no-server-data-ios-android-2026.