惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Troy Hunt's Blog
Blog — PlanetScale
Blog — PlanetScale
Engineering at Meta
Engineering at Meta
F
Full Disclosure
Recorded Future
Recorded Future
The GitHub Blog
The GitHub Blog
Microsoft Security Blog
Microsoft Security Blog
GbyAI
GbyAI
博客园_首页
博客园 - 叶小钗
MongoDB | Blog
MongoDB | Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Recent Commits to openclaw:main
Recent Commits to openclaw:main
H
Hacker News: Front Page
人人都是产品经理
人人都是产品经理
The Cloudflare Blog
博客园 - 司徒正美
Webroot Blog
Webroot Blog
Google DeepMind News
Google DeepMind News
Help Net Security
Help Net Security
Cloudbric
Cloudbric
PCI Perspectives
PCI Perspectives
有赞技术团队
有赞技术团队
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
TaoSecurity Blog
TaoSecurity Blog
L
Lohrmann on Cybersecurity
量子位
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Tailwind CSS Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
B
Blog RSS Feed
Apple Machine Learning Research
Apple Machine Learning Research
大猫的无限游戏
大猫的无限游戏
P
Proofpoint News Feed
N
News and Events Feed by Topic
罗磊的独立博客
T
Threat Research - Cisco Blogs
Schneier on Security
Schneier on Security
T
Tor Project blog
IT之家
IT之家
M
MIT News - Artificial intelligence
S
Security @ Cisco Blogs
O
OpenAI News
AI
AI
S
Securelist
Simon Willison's Weblog
Simon Willison's Weblog
The Last Watchdog
The Last Watchdog
月光博客
月光博客
Security Archives - TechRepublic
Security Archives - TechRepublic
L
LINUX DO - 热门话题

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
We Wrapped an Open-Source Agent in GraphOS and Turned the Debugging Session Into a Story
Ahmed Fayyaz · 2026-04-26 · via DEV Community

GraphOS

A story-driven, hands-on report about taking an existing open-source LangGraph.js agent, wrapping it in GraphOS, and learning what observability actually feels like when an agent goes sideways.

There is a moment every agent project eventually reaches.

The demo works. The graph looks clean. The tools are wired up. The prompt feels smart.

And then one run goes sideways.

Not in a dramatic, movie-scene way. In the real way.

The assistant calls the same tool again. Then again. The state grows. The trace gets noisier. The budget keeps moving in one direction. And the hardest part is not even the cost. It is the feeling that you can no longer see the system clearly.

That is the moment GraphOS was built for.

This post is not just a feature announcement. It is a field report. We took a real open-source project, wrapped it with GraphOS, ran the integration end to end, and used that exercise to answer one question:

Can an existing LangGraph.js agent — one we did not write — be made easier to observe, safer to run, and easier to explain to other developers?

The short answer is yes.

The better answer is the story below.

Before and after, in one breath

Before wrapping the agent in GraphOS:

  • Every run was a black box. The first signal that something was wrong was the OpenAI bill or a stuck UI.
  • Loops only surfaced as "this got slow" or "this never finished."
  • Debugging meant reading log files after the fact and reconstructing a sequence the system had not preserved.

After wrapping the agent in GraphOS:

  • Every step of the graph was observable in real time.
  • A loop was caught at visit 7, not visit 700.
  • A budget ceiling halted a misbehaving run before the credit card did.
  • Every past session became time-travelable in a local SQLite-backed dashboard, no SaaS in the loop.

Same agent. Same model. Different blast radius.

The benchmark we chose

Instead of inventing a toy example, we used an existing open-source benchmark:

Inside this repository, that benchmark lives at:

  • benchmarks/agents-from-scratch-ts

It is a strong test case because it is not a toy. It already contains:

  • A working email assistant
  • A Human-in-the-Loop flow
  • A memory-enabled variant
  • Jest-based test suites
  • LangGraph wiring that looks like real application code, not a demo built for a tool launch

That makes it the right kind of pressure test for GraphOS. We did not want to ship a wrapper that only works on our own handcrafted demo. We wanted something that survives contact with someone else's architecture, state shape, tool conventions, and tests.

Why this matters

If you are reading this as a builder, imagine two options:

  1. Build a brand-new sample agent designed to make the tool look good.
  2. Take someone else's open-source agent and prove the tool against that.

We picked option 2.

That choice changes the tone of the work. Now the question is not "can GraphOS run a demo we wrote?" It becomes "can GraphOS survive contact with someone else's code?"

That is a much better story to tell — and a much better thing to ship.

What GraphOS adds to a graph

GraphOS is an observability and policy layer for LangGraph.js agents.

At a high level, it gives you three things:

  • A wrapper around any compiled graph
  • Composable policies (LoopGuard, BudgetGuard, more)
  • A local dashboard that shows what the agent did, step by step, and lets you scrub through past runs

GraphOS architecture: your code → @graphos-io/sdk → @graphos-io/dashboard, with SQLite persistence

The integration stays intentionally small:

import {
  GraphOS,
  LoopGuard,
  BudgetGuard,
  tokenCost,
  createWebSocketTransport,
} from "@graphos-io/sdk";

const managed = GraphOS.wrap(myCompiledGraph, {
  projectId: "my-agent",
  policies: [
    new LoopGuard({ mode: "node", maxRepeats: 10 }),
    new BudgetGuard({ usdLimit: 2.0, cost: tokenCost() }),
  ],
  onTrace: createWebSocketTransport(),
});

Enter fullscreen mode Exit fullscreen mode

That is the promise. But promises are cheap.

So we tested it against the benchmark.

How we brought GraphOS into the benchmark

There are two installation stories worth separating, because people often confuse "how we developed it inside the monorepo" with "how I should use it in my own codebase."

Story 1: how we integrated it inside this monorepo

Because the benchmark is checked into this repository, the local integration uses the built SDK directly:

import {
  GraphOS,
  LoopGuard,
  BudgetGuard,
  tokenCost,
  createWebSocketTransport,
  PolicyViolationError,
} from "../../packages/sdk/dist/index.js";

Enter fullscreen mode Exit fullscreen mode

That exact integration lives in:

  • benchmarks/agents-from-scratch-ts/graphos-wrap.ts

This is useful for development because it lets us iterate on GraphOS and immediately retest it against the benchmark without publishing a new package every time.

Story 2: how you would install it in any outside project

If you are doing this in your own LangGraph.js project, the install is the simple part:

npm install @graphos-io/sdk
# or
pnpm add @graphos-io/sdk

Enter fullscreen mode Exit fullscreen mode

Then replace the local import with the published package import:

import {
  GraphOS,
  LoopGuard,
  BudgetGuard,
  tokenCost,
  createWebSocketTransport,
  PolicyViolationError,
} from "@graphos-io/sdk";

Enter fullscreen mode Exit fullscreen mode

Same code, same wrapper. The only thing that changes is where the SDK comes from.

The wrapper we added

Here is the part of the benchmark integration that mattered:

const managed = GraphOS.wrap(graph, {
  projectId: "agents-from-scratch",
  policies: [
    new LoopGuard({ mode: "node", maxRepeats: 6 }),
    new BudgetGuard({
      usdLimit: 0.5,
      cost: tokenCost({ fallback: 0.05 }),
    }),
  ],
  onTrace: createWebSocketTransport(),
});

Enter fullscreen mode Exit fullscreen mode

Three details deserve a beat each.

1. We used LoopGuard in node mode

This benchmark is exactly why node mode exists.

In many real agents, the state changes every iteration because the messages array keeps growing. That means pure state-equality is not enough to detect a loop. The graph may be functionally stuck even though the raw state object is technically different each turn.

So instead of asking:

"Did we revisit the exact same state?"

we ask:

"Did we keep revisiting the same node too many times?"

That is the more practical safety rule for agents that keep appending messages as they reason.

2. We set a budget ceiling

BudgetGuard lets us cap cumulative spend per session.

new BudgetGuard({
  usdLimit: 0.5,
  cost: tokenCost({ fallback: 0.05 }),
})

Enter fullscreen mode Exit fullscreen mode

tokenCost() is a drop-in cost extractor that walks the state for LangChain messages, pulls usage from usage_metadata / response_metadata.usage / tokenUsage, and applies a built-in OpenAI + Anthropic price table. For unknown models you can pass a fallback (flat USD per step or a custom price entry).

This is not just observability anymore. The run has a real boundary.

3. We streamed telemetry to the local dashboard

This line is small:

onTrace: createWebSocketTransport()

Enter fullscreen mode Exit fullscreen mode

But it changes the experience completely. Instead of waiting for the final output and guessing what happened, you watch the run unfold — node by node — in the GraphOS dashboard.

The small but clever trick: a mock key path

One of the nicest touches in the integration is that graphos-wrap.ts checks whether OPENAI_API_KEY starts with sk-mock. If it does, it installs a fetch interceptor and simulates the OpenAI responses.

Why is that useful?

Because it gives us a reproducible benchmark run that is intentionally shaped to trigger the loop path. In the mock flow:

  • The triage step routes the email into the response subgraph
  • The agent keeps requesting schedule_meeting
  • The graph cycles through llm_call ↔ environment
  • LoopGuard halts at visit 7 with a clean policy reason

That is the kind of test harness you want when you are building safety infrastructure. You do not want to rely on "hopefully the model misbehaves today." You want a deterministic failure mode you can use on purpose.

Reproduce the setup

If you want to walk through this yourself, the full path is short.

1. Install the workspace

From the repository root:

pnpm install

Enter fullscreen mode Exit fullscreen mode

2. Build the SDK

pnpm --filter @graphos-io/sdk build

Enter fullscreen mode Exit fullscreen mode

3. Move into the benchmark

cd benchmarks/agents-from-scratch-ts

Enter fullscreen mode Exit fullscreen mode

4. Use the benchmark normally

The upstream benchmark documents its own workflow:

pnpm agent

Enter fullscreen mode Exit fullscreen mode

It expects a .env file with your API key if you want real model calls:

OPENAI_API_KEY=your_api_key_here

Enter fullscreen mode Exit fullscreen mode

5. Run GraphOS alongside it

In another terminal, start the dashboard:

npx @graphos-io/dashboard graphos dashboard
# open http://localhost:4000

Enter fullscreen mode Exit fullscreen mode

Then run the wrapped benchmark entrypoint:

OPENAI_API_KEY=sk-mock pnpm exec tsx graphos-wrap.ts

Enter fullscreen mode Exit fullscreen mode

To run against a real provider instead of the mock path, swap in a real key and keep the same wrapper.

What we actually verified

This is where the story becomes more than marketing. We did not just wrap the benchmark and eyeball the result.

GraphOS SDK verification

From the repo root:

pnpm --filter @graphos-io/sdk test

Enter fullscreen mode Exit fullscreen mode

All SDK tests pass. Coverage spans:

  • LoopGuard — both state and node modes
  • BudgetGuard — cumulative cost ceiling
  • tokenCost() — multiple LangChain message shapes, multiple price-table lookups
  • GraphOS.wrap() — session lifecycle, error handling, sessionId continuity, listener-throw resilience

Benchmark verification

From benchmarks/agents-from-scratch-ts, the benchmark's own Jest suites still pass:

pnpm test:base
pnpm test:hitl
pnpm test:memory

Enter fullscreen mode Exit fullscreen mode

That matters because it tells us something subtle but important:

GraphOS was developed alongside the benchmark without breaking the benchmark's behavior. We are not telling a story about observability by quietly degrading the agent underneath it.

What the benchmark actually exercises

This part is worth slowing down for. The benchmark is not one narrow happy path. It exercises:

  • Response quality
  • Expected tool calls
  • Human acceptance flow
  • Human edit flow
  • Human rejection with feedback
  • Memory persistence across later runs

So when we say we used agents-from-scratch-ts, we mean we used a compact but meaningful open-source application with real behavioral coverage — not "we ran one prompt once."

The human lesson

The benchmark is an email assistant, but the lesson is bigger than email.

Every agent team eventually needs answers to these questions:

  • What node did we get stuck in?
  • How many times did we visit it?
  • What tool calls were made before failure?
  • What did the state look like at that moment?
  • Was the run expensive because it was useful, or expensive because it was looping?

Without observability, those questions become archaeology.

With GraphOS, they become part of the normal debugging workflow.

A quick interactive moment

Imagine you are looking at a run and you see the same node lighting up over and over. Which of these do you want next?

  1. A bigger console log
  2. The final model output only
  3. A live graph, a session timeline, and a policy halt reason that says exactly which guard fired and why

That is the difference this project is trying to create. Not more noise. Better visibility.

Why this story is stronger than a basic product post

The first version of a launch blog usually says:

  • Here is what we built
  • Here is the API
  • Here is why it is useful

That is fine, but it is mostly a claim.

This story is better because it shows:

  • The open-source project we used
  • The exact link to it
  • Where it lives in our repo
  • How we installed GraphOS into the workflow
  • How we wrapped the graph
  • How we tested both the SDK and the benchmark
  • What safety behavior we specifically cared about

In other words, this is not just what GraphOS is. It is how GraphOS behaves when it meets a real agent.

The files behind this story

If you want to inspect the exact pieces mentioned above, start here:

Final takeaway

GraphOS becomes much easier to understand when you stop describing it as a package and start describing it as a moment in a developer's day.

An agent starts drifting.

A team needs answers.

A wrapper adds policies.

A dashboard turns hidden execution into something visible.

A benchmark proves the idea against real code.

That is the story. And that is why we used agents-from-scratch-ts.

If you want to try the same path yourself:

npm install @graphos-io/sdk
npx @graphos-io/dashboard graphos dashboard

Enter fullscreen mode Exit fullscreen mode

The next step is simple: wrap your graph, run your tests, and see what your agent was actually doing when nobody was watching.