ๆƒฏๆ€ง่šๅˆ ้ซ˜ๆ•ˆ่ฟฝ่ธชๅ’Œ้˜…่ฏปไฝ ๆ„Ÿๅ…ด่ถฃ็š„ๅšๅฎขใ€ๆ–ฐ้—ปใ€็ง‘ๆŠ€่ต„่ฎฏ
้˜…่ฏปๅŽŸๆ–‡ ๅœจๆƒฏๆ€ง่šๅˆไธญๆ‰“ๅผ€

ๆŽจ่่ฎข้˜…ๆบ

Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Vercel News
Vercel News
Microsoft Azure Blog
Microsoft Azure Blog
Stack Overflow Blog
Stack Overflow Blog
Martin Fowler
Martin Fowler
Hacker News - Newest:
Hacker News - Newest: "LLM"
Cyberwarzone
Cyberwarzone
Recorded Future
Recorded Future
H
Hackread โ€“ Cybersecurity News, Data Breaches, AI and More
T
Threat Research - Cisco Blogs
Know Your Adversary
Know Your Adversary
Recent Announcements
Recent Announcements
L
LINUX DO - ็ƒญ้—จ่ฏ้ข˜
D
DataBreaches.Net
K
Kaspersky official blog
T
Threatpost
F
Full Disclosure
T
The Exploit Database - CXSecurity.com
C
CERT Recently Published Vulnerability Notes
S
Securelist
I
Intezer
ๆœ‰่ตžๆŠ€ๆœฏๅ›ข้˜Ÿ
ๆœ‰่ตžๆŠ€ๆœฏๅ›ข้˜Ÿ
็ฝ—
็ฝ—็ฃŠ็š„็‹ฌ็ซ‹ๅšๅฎข
็ˆฑ่Œƒๅ„ฟ
็ˆฑ่Œƒๅ„ฟ
S
Schneier on Security
P
Privacy & Cybersecurity Law Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Cisco Talos Blog
Cisco Talos Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
L
LangChain Blog
็พŽ
็พŽๅ›ขๆŠ€ๆœฏๅ›ข้˜Ÿ
G
Google Developers Blog
T
Tor Project blog
Project Zero
Project Zero
ๅฅ‡ๅฎขSolidotโ€“ไผ ้€’ๆœ€ๆ–ฐ็ง‘ๆŠ€ๆƒ…ๆŠฅ
ๅฅ‡ๅฎขSolidotโ€“ไผ ้€’ๆœ€ๆ–ฐ็ง‘ๆŠ€ๆƒ…ๆŠฅ
The Hacker News
The Hacker News
W
WeLiveSecurity
Engineering at Meta
Engineering at Meta
Apple Machine Learning Research
Apple Machine Learning Research
aimingoo็š„ไธ“ๆ 
aimingoo็š„ไธ“ๆ 
PCI Perspectives
PCI Perspectives
L
LINUX DO - ๆœ€ๆ–ฐ่ฏ้ข˜
MyScale Blog
MyScale Blog
้˜ฎไธ€ๅณฐ็š„็ฝ‘็ปœๆ—ฅๅฟ—
้˜ฎไธ€ๅณฐ็š„็ฝ‘็ปœๆ—ฅๅฟ—
้…ท ๅฃณ โ€“ CoolShell
้…ท ๅฃณ โ€“ CoolShell
V
V2EX
Hugging Face - Blog
Hugging Face - Blog
P
Proofpoint News Feed
Webroot Blog
Webroot Blog
T
Troy Hunt's Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Donโ€™t Fail โ€” They Drift Spilling beans for how i learn for exam๐Ÿ˜"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" โ€” What Actually Happened Comfy Cloudโ€™s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions โ€” here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components โ€” Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cรณmo construรญ un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 ๐Ÿš€ I Built an Ethical Hacking Scanner Tool โ€“ Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points โ€” Here's What I Found About How Markets Really Move EcoTrack AI โ€” Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead โ€” I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve โ€” no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like Youโ€™re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace โ€” how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025โ€“62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D โ€” A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent โ€” It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly โ€” 2026/04/10โ€“04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI ้€ฑๅ ฑ โ€” 2026/04/10โ€“2026/04/17 ๆจกๅž‹ๅฐ้Ž–ๆฝฎไพ†ไบ†๏ผŒไฝ†ๅทฅๅ…ท้ˆๆ‰ๆ˜ฏ็œŸๆˆฐๅ ด Maybe this is how Open-Source apps are born... ๐Ÿš€ Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge โ€” $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase โ€” Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train โ€” Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extraรงรฃo de Vรญdeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life โ€” Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 โ€” Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows โ€” Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTrackingๅฎ‰่ฃ…ๅ’ŒiPhone้ขๆ•้…็ฝฎๆ•™็จ‹๏ผŒๆœ‰bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
๐Ÿ›ก๏ธ PromptGuard: I Built a Local AI Privacy Firewall That Sanitizes Your Prompts Before They Leave Your Machine
Nizzad ยท 2026-05-25 ยท via DEV Community

This is a submission for the Gemma 4 Challenge: Build with Gemma 4

Every time you paste sensitive data, legal documents, or personal details into ChatGPT or Claude, that data leaves your device. PromptGuard intercepts it first โ€” and Gemma 4 running locally does the redaction before the prompt ever touches a cloud server.


Table of Contents


The Problem Nobody Talks About

Every week, professionals across healthcare, law, finance, and government are doing something they probably shouldn't: pasting sensitive documents directly into public AI interfaces.

A lawyer drafting a brief pastes a client's NIC number, phone, and case details into ChatGPT to get a summary. A doctor asks Claude to help structure a patient report โ€” with the patient's full name and health data in the prompt. A developer pastes a production database dump to debug a query. A researcher uploads a compliance document containing employee records.

Each of those prompts is transmitted to a cloud server, processed, potentially logged for safety review, and retained under terms of service that most users haven't read carefully.

This isn't a hypothetical risk. Sri Lanka's Personal Data Protection Act No. 9 of 2022 (PDPA) imposes legal obligations on controllers who process personal data. Section 10 requires appropriate technical and organizational measures. Sections 13โ€“18 guarantee data subject rights that can be violated by unauthorized disclosure. Section 38 sets penalties up to Rs. 10 million per non-compliance.

GDPR, UAE PDPL, and equivalent frameworks carry similar โ€” or higher โ€” obligations.

The problem: there is no guard between the user's clipboard and the AI's cloud API.

PromptGuard is that guard.


What I Built

PromptGuard is a two-component, local-first privacy firewall:

1. promptguard/ โ€” Local Python backend
A FastAPI server running on localhost:8000 that receives raw prompts, runs a two-stage redaction pipeline using Gemma 4 via Ollama, and returns a sanitized version. Zero network calls. Everything on-device.

2. promptguard-extension/ โ€” Chrome Extension (Manifest V3)
Injects a "Sanitize Prompt" button into ChatGPT and Claude.ai. When clicked, it intercepts the current prompt, sends it to the local backend for sanitization, replaces the prompt in the input box with the cleaned version, and only then allows the user to submit.

User types prompt with PII
        โ†“
[Chrome Extension intercepts]
        โ†“
POST to localhost:8000/scan
        โ†“
[Regex pre-redaction: NIC, email, phone]
        โ†“
[Gemma 4:e4b on-device LLM redaction]
        โ†“
Safe prompt returned
        โ†“
Input box updated with sanitized version
        โ†“
User submits to ChatGPT / Claude โ€” clean

Enter fullscreen mode Exit fullscreen mode

The entire redaction process happens on your machine. The cloud AI never sees the original.


Why Gemma 4 Is the Right Model for This

This is the question the judges will ask โ€” so I want to answer it directly and honestly.

Why not GPT-4o or Claude for redaction?

Sending sensitive data to a cloud API to redact sensitive data before sending to a cloud API is circular and defeats the purpose entirely. The solution has to be local.

Why not a simple regex approach?

Regex handles known patterns โ€” NIC numbers (\d{9}[VvXx]), emails, phone numbers. But PII is contextual:

  • "Call me at the usual number" โ€” no regex catches this
  • "The patient presented at 14:30, John Smith, age 34" โ€” name + age in natural language
  • "My CNIC is written on the form I mentioned earlier" โ€” reference without the number itself
  • "Send it to the Gmail I use for work" โ€” implied email without the address

You need a model that understands intent and context, not just patterns. Gemma 4 provides that understanding at a scale that runs locally.

Why Gemma 4 specifically โ€” and why the e4b variant?

Gemma 4 model family:
  2B / 4B  โ†’ ultra-mobile, browser, edge (Pixel, Raspberry Pi)
  27B      โ†’ server-grade, high accuracy
  e4b (MoE) โ†’ efficient inference, advanced reasoning, local deployment

Enter fullscreen mode Exit fullscreen mode

gemma4:e4b is the Mixture-of-Experts variant โ€” it activates only the expert subnetworks relevant to the current task. For a redaction task that requires:

  • Named entity recognition in natural language
  • Context-aware sensitivity detection
  • Understanding of legal and medical terminology
  • Preservation of semantic meaning after redaction

The MoE architecture gives you reasoning quality close to the 27B model at a fraction of the inference cost. It runs comfortably on a machine with 16GB RAM via Ollama. The 2B/4B models were too aggressive โ€” they redacted useful context along with PII. The 27B model was too slow for real-time prompt interception. e4b was the right balance.

This wasn't a default choice. I tested all three and e4b was the only one that preserved readability while catching contextual PII that regex missed.


System Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                     USER'S MACHINE                       โ”‚
โ”‚                                                         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”‚
โ”‚  โ”‚  Chrome Browser  โ”‚     โ”‚   PromptGuard Backend  โ”‚     โ”‚
โ”‚  โ”‚                  โ”‚     โ”‚   (FastAPI :8000)       โ”‚     โ”‚
โ”‚  โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚     โ”‚                        โ”‚     โ”‚
โ”‚  โ”‚  โ”‚ ChatGPT /  โ”‚  โ”‚     โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” โ”‚     โ”‚
โ”‚  โ”‚  โ”‚ Claude.ai  โ”‚  โ”‚     โ”‚  โ”‚  Stage 1: Regex   โ”‚ โ”‚     โ”‚
โ”‚  โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚     โ”‚  โ”‚  NIC / Email /    โ”‚ โ”‚     โ”‚
โ”‚  โ”‚        โ”‚         โ”‚     โ”‚  โ”‚  Phone redaction  โ”‚ โ”‚     โ”‚
โ”‚  โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚     โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚     โ”‚
โ”‚  โ”‚  โ”‚PromptGuard โ”‚  โ”‚POST โ”‚           โ”‚            โ”‚     โ”‚
โ”‚  โ”‚  โ”‚ Extension  โ”œโ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚     โ”‚
โ”‚  โ”‚  โ”‚(content.js)โ”‚  โ”‚     โ”‚  โ”‚  Stage 2: Gemma  โ”‚  โ”‚     โ”‚
โ”‚  โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚     โ”‚  โ”‚  4:e4b via Ollamaโ”‚  โ”‚     โ”‚
โ”‚  โ”‚        โ”‚         โ”‚โ—„โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”ค  Contextual PII  โ”‚  โ”‚     โ”‚
โ”‚  โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚JSON โ”‚  โ”‚  redaction       โ”‚  โ”‚     โ”‚
โ”‚  โ”‚  โ”‚  Input box โ”‚  โ”‚     โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚     โ”‚
โ”‚  โ”‚  โ”‚  (cleaned) โ”‚  โ”‚     โ”‚                        โ”‚     โ”‚
โ”‚  โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                                   โ”‚
โ”‚                                                         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
โ”‚  โ”‚  Ollama Runtime  โ”‚  Gemma 4:e4b model weights    โ”‚   โ”‚
โ”‚  โ”‚  (local process) โ”‚  (on-device, no network)      โ”‚   โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
โ”‚                                                         โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                         โ”‚
                    ONLY sanitized
                    prompt leaves
                         โ”‚
                         โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚  ChatGPT / Claude    โ”‚
              โ”‚  Cloud API           โ”‚
              โ”‚  (never sees raw PII)โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Enter fullscreen mode Exit fullscreen mode


How It Works: The Full Pipeline

Stage 1: Regex Pre-Redaction (run.py)

Fast, deterministic, zero-latency redaction of known PII patterns:

def regex_redact(text):
    # Sri Lanka NIC: 9 digits + V/X suffix
    text = re.sub(r'\b\d{9}[VvXx]\b', '[REDACTED_NIC]', text)

    # Email addresses
    text = re.sub(r'[\w\.-]+@[\w\.-]+', '[REDACTED_EMAIL]', text)

    # Phone numbers (10 digits)
    text = re.sub(r'\b\d{10}\b', '[REDACTED_PHONE]', text)

    return text

Enter fullscreen mode Exit fullscreen mode

This catches the easy cases instantly before Gemma 4 even sees the text โ€” reducing both latency and the model's cognitive load.

Stage 2: Gemma 4 Contextual Redaction (run.py)

The partially-redacted text goes to Gemma 4:e4b with a precisely engineered system prompt:

def sanitize_prompt(prompt: str) -> str:
    partially_redacted = regex_redact(prompt)

    system_prompt = f"""
You are PromptGuard, a privacy-preserving AI firewall.

Redact sensitive information while preserving readability.

Text:
{partially_redacted}
"""
    response = ollama.chat(
        model="gemma4:e4b",
        messages=[{"role": "user", "content": system_prompt}]
    )
    return response['message']['content']

Enter fullscreen mode Exit fullscreen mode

Gemma 4 handles what regex can't:

  • Full names in natural language
  • Medical conditions and health data
  • Financial details described in prose
  • Implicit references to identifiable information
  • Sensitive context even without explicit identifiers

Stage 3: FastAPI Endpoint

The backend exposes a single clean endpoint:

# FastAPI backend (inferred from content.js calling /scan)
@app.post("/scan")
async def scan_prompt(payload: PromptRequest):
    safe = sanitize_prompt(payload.prompt)
    return {"safe_prompt": safe}

Enter fullscreen mode Exit fullscreen mode

The extension POSTs to http://127.0.0.1:8000/scan โ€” purely local, no TLS required, no external network call.


The Chrome Extension

The extension (promptguard-extension/) is a Manifest V3 Chrome extension with two files:

manifest.json โ€” declares permissions and injection targets:

{
  "manifest_version": 3,
  "name": "PromptGuard",
  "version": "1.0",
  "permissions": ["activeTab", "scripting"],
  "host_permissions": [
    "https://chatgpt.com/*",
    "https://claude.ai/*"
  ],
  "content_scripts": [
    {
      "matches": ["https://chatgpt.com/*", "https://claude.ai/*"],
      "js": ["content.js"]
    }
  ]
}

Enter fullscreen mode Exit fullscreen mode

content.js โ€” injects a persistent "Sanitize Prompt" button and handles the interception flow:

async function sanitizePrompt() {
    // Find the active prompt input (handles both contenteditable and textarea)
    const inputBoxes = document.querySelectorAll(
        '[contenteditable="true"], textarea'
    );
    let inputBox = null;
    for (let box of inputBoxes) {
        if ((box.innerText?.length > 0) || (box.value?.length > 0)) {
            inputBox = box;
            break;
        }
    }
    if (!inputBox) { alert("No prompt input found"); return; }

    const originalPrompt = inputBox.value || inputBox.innerText;

    // Send to local backend
    const response = await fetch("http://127.0.0.1:8000/scan", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ prompt: originalPrompt })
    });

    const data = await response.json();

    // Replace input with sanitized version
    if (inputBox.value !== undefined) inputBox.value = data.safe_prompt;
    else inputBox.innerText = data.safe_prompt;

    // Trigger React's onChange so the UI recognizes the update
    inputBox.dispatchEvent(new Event('input', { bubbles: true }));

    alert("Prompt sanitized โœ“");
}

// Inject the button and keep it alive through dynamic UI re-renders
function createButton() {
    if (document.getElementById("promptguard-btn")) return;
    const button = document.createElement("button");
    button.id = "promptguard-btn";
    button.innerText = "๐Ÿ›ก๏ธ Sanitize";
    // ... styling
    button.onclick = sanitizePrompt;
    document.body.appendChild(button);
}

setInterval(createButton, 2000); // Survives React re-renders

Enter fullscreen mode Exit fullscreen mode

The setInterval pattern is intentional โ€” ChatGPT and Claude.ai are React SPAs that frequently re-render the DOM, which can remove injected elements. The interval re-injects the button if it disappears.


The Local Backend

The promptguard/ folder contains the Python backend. To run it:

# 1. Install Ollama (https://ollama.ai)
ollama pull gemma4:e4b

# 2. Install Python dependencies
pip install fastapi uvicorn ollama

# 3. Start the backend
uvicorn main:app --host 127.0.0.1 --port 8000

Enter fullscreen mode Exit fullscreen mode

The backend stays running in the background. The extension talks to it automatically whenever you click "Sanitize."


Real-World Demo:

Here's the scenario this was built for. A legal professional in Sri Lanka is drafting a submission under the PDPA and wants AI assistance.

GitHub: The full code is available in two folders: promptguard/ (Python backend) and promptguard-extension/ (Chrome extension).

๐Ÿ›ก๏ธ PromptGuard

A local-first AI privacy firewall that sanitizes prompts before they reach the cloud.

PromptGuard intercepts prompts typed into ChatGPT or Claude.ai, runs PII redaction using Gemma 4:e4b entirely on your machine, and replaces the raw prompt with a sanitized version โ€” before anything leaves your device.

Built for the Gemma 4 Challenge on DEV.to


The Problem

Every day, professionals paste sensitive content into public AI interfaces:

  • Legal documents with client NIC numbers and case details
  • Medical records with patient health conditions
  • Financial data with account information and salary details
  • HR documents with employee personal data

This creates real legal exposure under Sri Lanka's PDPA No. 9 of 2022, GDPR, UAE PDPL, and equivalent frameworks. PromptGuard sits between your clipboard and the cloud โ€” nothing sensitive gets transmitted.


How It Works

You type a prompt with PII
        โ†“
[PromptGuard Extension intercepts on click]
        โ†“
POST โ†’

โ€ฆ

Raw prompt (what they typed):

My client John Doe, NIC 999995678V, reached out via 
john.doe@example.com about a data breach at ABCXYZ Pvt Ltd. 
Her phone is 0777654321. The breach exposed her health records 
including her HIV status from the XYZABC Hospital 
admission in March 2024. Draft a letter to the Data Protection 
Authority under Section 23 of the PDPA.

Enter fullscreen mode Exit fullscreen mode

After Stage 1 (regex):

My client John Doe, NIC [REDACTED_NIC], reached out via 
[REDACTED_EMAIL] about a data breach at XYZABC Hospital. 
Her phone is [REDACTED_PHONE]. The breach exposed her health records 
including her HIV status from the XYZABC Hospital 
admission in March 2024. Draft a letter to the Data Protection 
Authority under Section 23 of the PDPA.

Enter fullscreen mode Exit fullscreen mode

After Stage 2 (Gemma 4:e4b):

My client [REDACTED_NAME], NIC [REDACTED_NIC], reached out via 
[REDACTED_EMAIL] about a data breach at [REDACTED_ORGANIZATION]. 
Her phone is [REDACTED_PHONE]. The breach exposed her health records 
including [REDACTED_HEALTH_CONDITION] from a hospital admission in 
[REDACTED_TIMEFRAME]. Draft a letter to the Data Protection Authority 
under Section 23 of the PDPA.

Enter fullscreen mode Exit fullscreen mode

The cloud AI receives a complete, legally actionable task description. The client's identity, health condition, specific organization, and date are never transmitted. The AI can still draft the letter correctly.

What Gemma 4 caught that regex missed:

  • Full name (John Doe) โ€” natural language NER
  • Organization name (XYZABC Pvt Ltd) โ€” potential re-identification risk
  • Health condition (HIV status) โ€” special category data under PDPA Schedule II
  • Specific date (March 2024) โ€” temporal re-identification marker

What Gets Redacted

PII Type Detection Method Example
Sri Lanka NIC Regex 999995678V โ†’ [REDACTED_NIC]
Email addresses Regex user@mail.com โ†’ [REDACTED_EMAIL]
Phone numbers Regex 0777654321 โ†’ [REDACTED_PHONE]
Full names Gemma 4 (NER) John Silva โ†’ [REDACTED_NAME]
Health conditions Gemma 4 (context) HIV positive โ†’ [REDACTED_HEALTH]
Financial details Gemma 4 (context) Rs. 2.4M salary โ†’ [REDACTED_FINANCIAL]
Organization names Gemma 4 (risk assess) City Hospital โ†’ [REDACTED_ORG]
Dates + context Gemma 4 (re-id risk) March 2024 admission โ†’ [REDACTED_TIMEFRAME]
Implied references Gemma 4 (inference) my usual number โ†’ [REDACTED_REFERENCE]

Limitations and Honest Caveats

This is a v1 proof-of-concept. Here's what it doesn't yet handle well:

False positives. Gemma 4 occasionally over-redacts โ€” removing organizational names that are actually public information and don't need masking. The prompt engineering needs refinement for domain-specific contexts.

Latency. On a mid-range laptop, Gemma 4:e4b takes 2โ€“5 seconds per prompt. For short prompts this is acceptable. For multi-paragraph document pastes, it's noticeable. The regex pre-stage helps, but LLM inference time is the bottleneck.

No feedback loop. The current version replaces the prompt silently. A diff view โ€” showing the user exactly what was changed and why โ€” would significantly improve trust and usability.

Extension CSP constraints. Some AI interfaces (particularly enterprise versions) implement Content Security Policies that may block content script injection. The extension works on standard chatgpt.com and claude.ai but may not work on enterprise/team deployments.

It requires the backend to be running. If the Ollama server or FastAPI backend isn't started, the extension fails silently. Better error messaging and a backend health check are on the roadmap.


What's Next

  • Diff view โ€” show what changed before the user submits, not just "sanitized โœ“"
  • Domain profiles โ€” legal, medical, financial contexts each have different redaction thresholds
  • Firefox support โ€” MV3 is Chromium-specific; a MV2 variant for Firefox is straightforward
  • Offline indicator โ€” visual badge showing when the backend is active vs. unavailable
  • Fine-tuned Gemma 4 โ€” the PDPA document is already loaded into the RAG agent; fine-tuning Gemma 4 on Sri Lankan PII patterns (NIC format, address structures, Sinhala/Tamil name recognition) would significantly improve local-context accuracy
  • Auto-submit mode โ€” the app.py variant already implements auto-submit after sanitization; making this a configurable toggle is the next UX step

Key Takeaways

  • โœ… 100% local โ€” Gemma 4:e4b runs via Ollama on-device; the original prompt never leaves your machine
  • โœ… Two-stage pipeline โ€” regex catches known patterns instantly; Gemma 4 catches contextual PII that regex cannot
  • โœ… Model choice was deliberate โ€” e4b MoE architecture provides near-27B reasoning quality at local inference speeds; 2B/4B under-redacted, 27B was too slow
  • โœ… Works on ChatGPT and Claude.ai โ€” Chrome extension injects into both without modifying their code
  • โœ… PDPA-aligned โ€” the redaction taxonomy maps directly to Sri Lanka PDPA definitions: personal data, special categories, data subject identifiers
  • โš ๏ธ Latency is real โ€” 2โ€“5s per prompt on mid-range hardware; acceptable for sensitive workflows, not for casual use
  • โš ๏ธ False positives exist โ€” over-redaction is a known v1 limitation; domain profiles will address this
  • ๐Ÿ” The broader implication โ€” as AI becomes embedded in professional workflows, the question isn't "should we use AI?" It's "how do we use AI without creating PDPA/GDPR liability?" PromptGuard is one answer to that question.

Have you dealt with PII leakage in AI workflows? Particularly curious whether legal or healthcare professionals have built their own guardrails โ€” or just accepted the risk. Comments below.