惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
Y
Y Combinator Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Hugging Face - Blog
Hugging Face - Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
The Cloudflare Blog
L
LangChain Blog
美团技术团队
N
Netflix TechBlog - Medium
量子位
酷 壳 – CoolShell
酷 壳 – CoolShell
B
Blog
博客园 - 司徒正美
爱范儿
爱范儿
D
DataBreaches.Net
月光博客
月光博客
U
Unit 42
B
Blog RSS Feed
Engineering at Meta
Engineering at Meta
Apple Machine Learning Research
Apple Machine Learning Research
Jina AI
Jina AI
MongoDB | Blog
MongoDB | Blog
腾讯CDC

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
I Got Tired of AI Agents Having Root Access to Everything...
Aditya P Dixit · 2026-06-28 · via DEV Community

Aditya P Dixit

Everyone is building AI agents.

Very few people are building the thing that sits between an AI agent and a disastrous decision.

That's why I built XRisk.

XRisk is an open-source autonomous safety engine that acts as a decision layer between an AI agent and the real world.

Instead of blindly executing an action, an agent asks XRisk:

"Should I actually do this?"

XRisk responds with one of three deterministic decisions:

✅ Allow
⚠️ Confirm
❌ Block
Why I Started This Project

As I experimented with increasingly autonomous AI systems, I noticed the same pattern over and over again.

Most projects focused on making agents more capable.

Almost nobody was asking:

"What happens when the agent is wrong?"

Consider a few examples.

An agent accidentally leaks API keys.
A prompt injection convinces it to ignore previous instructions.
A model decides to execute a shell command.
An autonomous workflow loops forever and keeps calling expensive APIs.
A deployment bot pushes code without human approval.

Most agent frameworks assume the model behaves.

Reality doesn't.

I wanted something deterministic sitting between intention and execution.

Not another model.

Not another prompt.

An actual policy engine.

What XRisk Does

XRisk evaluates every proposed action before it's executed.

It combines multiple safety signals into a single explainable decision.

Some of the things it checks include:

Policy-as-code with layered precedence
Prompt injection detection
Sensitive data and secret detection
Capability token validation
Network egress restrictions
Circuit breakers for autonomous loops
Tamper-evident audit logs
Supply-chain verification
Policy conflict detection
Deterministic forensic replay

Instead of a mysterious "Safety Score: 67%," XRisk explains why it made a decision.

Example

Imagine an AI assistant wants to execute:

{
"tool": "deploy",
"actor": "release-bot",
"prompt": "Deploy production immediately."
}

Instead of sending that directly to your deployment system...

XRisk intercepts it.

It evaluates:

Does policy require approval?
Is the actor allowed to deploy?
Is the destination trusted?
Are capability tokens valid?
Does this resemble prompt injection?
Is this part of a dangerous execution loop?

Only then does it decide whether to:

Allow
Confirm
Block
One Design Decision I Feel Strongly About

I deliberately avoided using another LLM to make safety decisions.

LLMs are excellent at generating text.

Policy enforcement should be deterministic.

If an action is blocked, I want to know exactly why it was blocked.

Every decision should be reproducible.

Every audit should be explainable.

Every policy should be inspectable.

That's the philosophy behind XRisk.

What's Next

I'm currently working toward:

Threat intelligence correlation
Zero-trust workload identities
Autonomous containment
Adversarial simulation
Multi-party approval workflows

The long-term vision is to make XRisk a reusable security layer that can sit in front of any AI agent, regardless of framework.

I'd Love Feedback

This project is still evolving, and I'd genuinely appreciate feedback from people building AI systems.

Some questions I'm particularly interested in:

What attack vectors am I missing?
Which policies would you want in production?
What integrations would make this more useful?
How would you design a safety engine differently?

If you'd like to contribute, open an issue, suggest improvements, or submit a PR. Even small documentation fixes are welcome.

Thanks for reading—I hope XRisk becomes something that helps make AI systems not just more capable, but more trustworthy.

Link: https://github.com/Hootsworth/XRisk