인셔셔RSS 관심 있는 블로그, 뉴스, 기술 정보를 효율적으로 추적하고 읽으세요
원문 읽기 InertiaRSS에서 열기

추천 피드

Microsoft Security Blog
Microsoft Security Blog
WordPress大学
WordPress大学
S
SegmentFault 最新的问题
爱范儿
爱范儿
B
Blog RSS Feed
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Blog — PlanetScale
Blog — PlanetScale
Vercel News
Vercel News
Jina AI
Jina AI
aimingoo的专栏
aimingoo的专栏
I
Intezer
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Attack and Defense Labs
Attack and Defense Labs
The GitHub Blog
The GitHub Blog
小众软件
小众软件
AI
AI
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
N
News and Events Feed by Topic
腾讯CDC
D
Docker
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
罗磊的独立博客
人人都是产品经理
人人都是产品经理
W
WeLiveSecurity
N
News and Events Feed by Topic
Security Archives - TechRepublic
Security Archives - TechRepublic
C
Check Point Blog
Webroot Blog
Webroot Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
H
Help Net Security
Recorded Future
Recorded Future
H
Hacker News: Front Page
T
Troy Hunt's Blog
V
V2EX
Forbes - Security
Forbes - Security
Stack Overflow Blog
Stack Overflow Blog
The Register - Security
The Register - Security
P
Palo Alto Networks Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
博客园 - 叶小钗
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
S
Security Affairs
The Hacker News
The Hacker News
Simon Willison's Weblog
Simon Willison's Weblog
博客园 - 三生石上(FineUI控件)
B
Blog
Apple Machine Learning Research
Apple Machine Learning Research
C
Cyber Attacks, Cyber Crime and Cyber Security
D
DataBreaches.Net

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The Brilliant Person in Your Pocket
sadiq mohamm · 2026-05-23 · via DEV Community

What Gemma 4 taught me about the real future of on-device AI, and why I owe Seyi an apology

This is a submission for the Gemma 4 Challenge: Write about Gemma 4

I poured cold water on an intern's excitement about a powerful AI model that could run entirely on a phone. That was a few weeks ago, and now, I am the intern.

Let me explain - Seyi is a mechatronics student at Bells University, Nigeria, doing an internship at my workplace. He came in buzzing one morning about Gemma 4, Google's latest open-weight language model, capable of running completely offline on consumer hardware. His eyes were wide, but mine were narrower. My internal monologue went something like: Why would I want to downgrade to a local model when I have cloud access to far more powerful ones? Privacy concerns didn't register for me. Usage limits weren't a real problem on my Google One plan, and an offline model, by definition, is frozen in time, limited to whatever it knew at training.

Where's the upside, I asked?

I said all of this out loud. Seyi nodded, deflated, and went back to his desk. The universe, as it often tends to do, had a rebuttal scheduled for later that week.

The Argument I Didn't Expect to Lose

A few days later, I was browsing Martin Sauter's blog (he's a telecom author whose book remains one of my favourites), and I check in occasionally even when the content goes well over my head. He'd started a thread on running local AI: terms like OpenWebUI and RAG were being thrown around, discussions about making offline models more dynamic through external connectivity.

Something clicked in my mind. I've been thinking for a while about using AI for specialised tasks at work, but corporate environments don't always play nicely with public cloud AI. Uploading proprietary documents to a public model is a legitimate concern. A capable model running on a personal device, answering to no server and leaving no data trail, was beginning to sound less like a downgrade and more like a completely different category of tool.

Then I saw a LinkedIn post about the AI Edge Gallery, an Android app from Google that lets you run Gemma 4 directly on your phone. No cloud, no subscription, no data leaving the device.

That was it for me. I was in.

More Than a Chatbot in Your Pocket

Before we go further, let me quickly explain what Gemma 4 on AI Edge Gallery actually is, because it is not simply "a smaller ChatGPT that works offline." That framing undersells what's actually available here. While Gemma 4 comes in four variants, two of them are light enough to run on a phone: the E2B, optimised for speed on mid-range devices, and the E4B, a smarter model targeted at modern phones with 8GB of RAM or more. I run a Google Pixel 8 Pro with 12GB of RAM, so the E4B was my lane.

The Edge Gallery app itself organises Gemma 4's E2B and E4B models' capabilities into distinct workspaces. There's AI Chat with Thinking Mode, which shows you the model's step-by-step reasoning in real time (a genuinely useful window into how it arrives at answers). There's Ask Image, which lets the model read and analyse photos from your camera or gallery, entirely offline. There's Audio Scribe for transcription and translation. And then there's Agent Skills, the feature that consumed most of my attention.

AI Edge Gallry App- Gemma 4 Use cases

Agent Skills are best explained with an analogy. Gemma 4 out of the box is a brilliant person who can hold a nuanced conversation. Skills are what you hand that person before the conversation starts, like a calculator, a specialist reference manual, or a set of behavioural instructions. The model reads a menu of available skills at the start of each session, determines which one is relevant to your request, and activates it. The core file that defines each skill is called a SKILL.md — a plain text document with a metadata block at the top (a name and a trigger description) and freeform instructions below (text-only skills). No coding required. Anyone who can write clearly can, in theory, build a custom AI skill.

That word "in theory" is pushing it a bit. More on that shortly.

My First Experiment: Teaching Gemma to Write Like Me

As a telecom professional who writes about the industry, I run a LinkedIn newsletter called Signal Over Noise, and my first instinct was to test Gemma's ability to replicate my writing style. Not to replace my writing, but to see if a local model could serve as a personal writing assistant, one that knows my voice well enough to be actually useful when I hand it raw material to reshape.

Teaching Gemma to write

The process of building the skill itself was fascinating. I worked with Claude to do a detailed stylistic analysis of my previous non-AI-assisted articles, creating twelve touchpoints covering voice, rhythm, structure, openings, analogies, vocabulary, signature moves, and more. The analysis surfaced consistent fingerprints: the conversational-analytical blend, the personal entry point as a default opening, the pop-culture analogies with Nigerian context, invented compound words, and the self-aware aside that steps just outside the narrative. That profile became a SKILL.md file - my personal, reusable, installable instruction set that could theoretically rewrite any source material in a documented style, calibrated by adherence level.

Then I tried to run it. And Gemma 4, to its credit, was very honest about its limitations.

Three Lessons From a Phone That Crashed, Looped, and Gave Up Mid-Sentence

Lesson One: Instruction Footprint Matters More Than You Think

The first time I tried to activate my style skill on the E2B and E4B variants, the app crashed immediately. Not "it gave a bad output." It crashed repeatedly.

The root cause, as I later understood it, is that local models operate within strict, hardware-enforced RAM allocations. The original instruction file was verbose: exhaustive descriptions for every adherence level of all twelve style touchpoints. Before the model could process a single user message, the system had to load all of those instructions into its working memory. The sheer volume saturated the phone's RAM allocation before the engine could even initialise, and the application stopped working.

RAM Overload

The fix was aggressive compression, stripping the descriptive matrices, keeping only the core tenets, reducing the file by roughly 60%, and bringing the system-level instructions under the memory threshold. It worked, sorta. The app stopped crashing.

But this was also the first clear signal: cloud-model habits do not translate to mobile hardware.

Lesson Two: Structured Instructions Apparently Confuse Small Models?

With a leaner file, the skill loaded. But when I prompted it to rewrite a long technical article on 5G network architecture, the model didn't write anything. It produced this instead:

<|tool_call>call:run_intent{intent:"dabs-style-rewriter", parameters:{"Source Material": "..."

The model had looked at an instruction file formatted with numbered steps, labelled inputs, and a structured schema and concluded it was not a writing persona but an API gateway. It was trying to route data to another application that didn't exist. Instead of channelling a writer's voice, it was behaving like a function router.

Confused by highly structured instructions...

The solution was to flip the entire framing from "Steps and Inputs" to "Be this person immediately." Move away from programmatic structure, toward a pure system instruction that establishes persona from the first line. The model responds to being told who to be far better than it responds to being told what to do in sequence.

Lesson Three: Small Models Have a Working Memory Ceiling…and It Shows

With the condensed, persona-first version of the skill, I finally got output. But it didn't quite sound like me, and it stopped mid-sentence abruptly. What was happening, as best as I can understand it, is context exhaustion, i.e., the model's working memory (its KV cache, if you want the technical term) was completely filled by the combination of the instruction file and the dense 700-word 5G article I'd handed it as source material. With no computational overhead left to manage, it began to stutter, outputting typographical duplicates like "all the* the*", before hitting its timeout and cutting out entirely.

Memory Cap

This is the fundamental gap between a local mobile model and a large cloud LLM. A cloud model can hold hundreds of complex rules, a dense technical source, and an elaborate creative brief simultaneously and produce something nuanced that balances all of them. A 4-billion-parameter model on a phone cannot. Not even close. The reasoning gap is not a flaw to be patched in a software update; it just a consequence of physics and silicon.

So What Are Small Local Models Actually Good For?

Here is the reframe that changed how I think about all of this and, I'd argue, the more interesting question...

Forcing a small local model to perform complex creative persona replication is the wrong brief entirely. It's like hiring a specialist fabricator to write you a novel. The skills aren't absent; they're just misallocated.

What local mobile models genuinely excel at is structured data, deterministic rules, and zero-hallucination utility work. They are fast, private, and always available. The question isn't "can they replicate a cloud model's creative output?" The question is "what tasks benefit from a local, lightweight, always-on intelligent layer?"

ILocal LLM on mobile as a key utility

One practical answer I landed on came from my own life. I use a solar-powered inverter setup at home. Managing backup power means keeping an eye on cloud cover, grid stability, and battery load, and I sometimes have to call my wife from work to ask her to check conditions and act accordingly. How about trying out a JavaScript-enabled Agent Skill on my phone, which could probably handle it elegantly.

Here's how: JavaScript Skills (as opposed to text-only skills, earlier described) give the AI the ability to actually execute code within a hidden browser environment on the phone. The model's job is reduced to one thing: capturing the user's natural language input, like a neighbourhood name.

A script then takes that variable, calls a live weather API, retrieves real-time cloud cover data, applies the relevant electrical formulas within JS, and returns a plain-language recommendation. "Turn on the deep freezer. Solar output is adequate." No creative reasoning required and no hallucination risk. The AI is the natural language interface; the code is the engine.

That division of labour, model as conversational intake and code as execution layer, in my opinion, is where mobile AI actually earns its keep.

What This Means for the Bigger Picture

Step back for a moment and consider what is actually happening here. A model capable enough to understand nuanced instructions, hold a meaningful conversation, and trigger appropriate specialised behaviours, all without sending a single byte of your data to any server, now fits inside a device you carry in your pocket.

That is not a minor development.

The immediate practical implications are clearest in environments where cloud AI is restricted: corporate networks, regulated industries, and low-connectivity regions. For anyone in those categories, Gemma 4 on a capable Android phone is not a compromise. It's the first viable option. And for the African context specifically, where data costs remain non-trivial, and cloud latency is a genuine UX concern, the case for capable on-device inference is stronger than it might appear from a high-bandwidth Western baseline.

AI in your hands

But the more interesting implication is what it signals about the direction of travel. The bottlenecks I encountered (memory saturation, intent looping, context exhaustion) are real, but they are also well-understood engineering problems.

Model quantisation is getting better. Mobile hardware is also improving. The techniques for writing efficient, structured skills that play to a local model's strengths rather than fighting its limitations are learnable, and they're becoming more documented. The floor on what a phone-based AI can do in 2026 is already higher than most people realise. In a few years, the question won't be whether it's capable enough, but whether the existing and popularly used cloud model is offering you enough additional value to justify the dependency.

I am not suggesting local models will replace cloud AI. That would be a lazy take, and I've already used up my quota of lazy takes in this article (see: Seyi, opening paragraphs). What I am saying is that the role of on-device AI is transitioning from novelty to infrastructure, from "cool demo on a tech blog" to a genuine utility layer for people who need private, fast, and accessible intelligence built into their workflows.

Closing Thoughts: Seyi Was Right

The arc of this whole experiment (from skeptic to cautious convert) tracks almost perfectly with how genuinely disruptive technologies tend to arrive. Not with an obvious, immediate use case that justifies the hype, but with a quiet accumulation of practical realisations that eventually tip into "oh, this is actually something."

My style-replication skill didn't work the way I hoped. The model crashed, looped, and truncated. But in failing those tests, it showed me exactly where its real value lies, and gave me a concrete mental model for building agent skills that actually perform. The solar power assistant is on my list. A RAN architecture reference tool for work is also on the list, one that accepts base station design inputs and returns equipment specifications based on local rules, privately, offline. These are not glamorous AI use cases. They are, however, genuinely useful ones.

On-device AI is not the future of AI. It's the future of how AI fits into ordinary life - quiet, local, and practical, like a good tool should be.

Seyi, if you're reading this: I owe you one.