惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
V
Visual Studio Blog
C
Check Point Blog
Google DeepMind News
Google DeepMind News
S
SegmentFault 最新的问题
博客园 - 聂微东
量子位
T
Tailwind CSS Blog
罗磊的独立博客
I
InfoQ
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Y
Y Combinator Blog
L
LangChain Blog
小众软件
小众软件
Engineering at Meta
Engineering at Meta
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Security Latest
Security Latest
M
MIT News - Artificial intelligence
Know Your Adversary
Know Your Adversary
MongoDB | Blog
MongoDB | Blog
Google DeepMind News
Google DeepMind News
大猫的无限游戏
大猫的无限游戏
H
Help Net Security
爱范儿
爱范儿
T
The Exploit Database - CXSecurity.com
有赞技术团队
有赞技术团队
V
Vulnerabilities – Threatpost
Martin Fowler
Martin Fowler
A
Arctic Wolf
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 司徒正美
Cyberwarzone
Cyberwarzone
阮一峰的网络日志
阮一峰的网络日志
The Hacker News
The Hacker News
Apple Machine Learning Research
Apple Machine Learning Research
宝玉的分享
宝玉的分享
GbyAI
GbyAI
Latest news
Latest news
云风的 BLOG
云风的 BLOG
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
腾讯CDC
AWS News Blog
AWS News Blog
aimingoo的专栏
aimingoo的专栏
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
L
Lohrmann on Cybersecurity
博客园 - Franky
S
Securelist
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Threatpost
美团技术团队

Towards AI

Building AI Agents in Rust — part 4 | Towards AI The Verified Identity Agent Bridge | Towards AI You Can’t Prompt Your Away Your LLM Problems | Towards AI The Free Agent Trap | Towards AI Your Agentic Loop Will Drift. Here Is the KL Divergence Equation That Measures How Far It Has Wandered From Its Original Instruction. | Towards AI Beyond Chat: Processing Images, PDFs, and Documents with the OpenAI Adapter in Oracle Integration Cloud | Towards AI Building AI Agents in Rust — part 3 | Towards AI Self-Hosting Airflow at Home: Automating Stock Price Data Collection | Towards AI The 76-Hour Frontier: How the Takedown of Claude Fable 5 Birthed the Military-Industrial-AI Complex | Towards AI I Trained a Markdown File to Boost GPT-5.5 by 23 Points — It Shouldn't Work | Towards AI We Replaced ChatGPT With a Local AI Server. Six Months of Honest Data. | Towards AI What Really Makes Cars Pollute? A Data Science Deep Dive into CO₂ Emissions | Towards AI Training GPT-2 From Scratch on a GTX1050 | Towards AI Principal Component Analysis (PCA): Theory, Mathematics, and Applications Build a Zero-Cost Web Automation Pipeline With OpenRouter, OpenClaw, and MediaUse I Gave Qwen3.7-Plus a Screenshot and It Found the Exact Pixel to Click for $0.40 Beyond the Prompt: Why Autonomous AI Agents Are Replacing the Chatbot Moonshot Cracked Claude Code’s Playbook with an MIT Terminal Agent and a $0.60 Model Connections, Roles, and Warehouses: Getting CoCo Desktop Production-Ready from Day One My First $5,000 Month Writing About AI Engineering on Medium Google Shrank Gemma 4 by 72% and Unsloth Fixed the 4-Bit Bug Nobody Else Caught on One 4090, and 4-Bit Shouldn’t Be This Good LangChain Explained: Understanding Models, Prompts, Chains, Memory, Indexes, and Agents TOON: Beyond JSON for LLMs Claude Code Casual, Pro, Elite: The Three Working Personas of Claude Code Mastery MiniMax M3 Decodes 1M Tokens 15x Faster — and It Shouldn’t Be This Cheap Using Amazon SQS for AI Agent Orchestration I Ran a 1.5B-Active Model on My Laptop That Embarrassed a 26B by 46 Points How to Build a Self-Improving Company with AI Part 3 — Implementation/Engine-Level: Choosing the Runtime That Gives You These for Free Part 2 — Serve-Level Speed: System Design That Stabilizes P95/P99 3-Part Series: LLM Latency in Production (Part 1) Claude Code: The AI Coding Partner Changing How Developers Build Software Claude Code Pitfalls: Claude Code Won’t Do What You Told It: A Troubleshooting Catalog Full-Stack Data Scientists for the Agentic Coding World Building Production-Grade AI Skills with Snowflake Cortex AI Function Studio I Tried 10 AI Agent Frameworks in 2026 — Here’s the Honest Guide I Wish I Had Earlier How One Spring Boot Optimization Saved Our Startup $30,000 a Year Inside Palantir AIP: How the World’s Most Controversial AI Platform Actually Works What Is a Reverse Proxy? (And Why Every Backend Developer Should Care) What Claude Opus 4.8 Actually Changes If You’re Building Agents QWEN 3.7 Max Worked For 35 Hrs Straight And The Results Were Mind-blowing When LLMs Meet Knowledge Graphs on the Battlefield Fine-Tuning is Dead: Why Context Orchestration Won in 2026 5 Things Broke When I Shipped a RAG + MCP Agent to Production. Google Co-Scientist: Hyper Scaling Research and Discovery Microsoft Just Embarrassed Browser Web Agents — 1,000 Lines Made GPT-5.4 Beat Opus 4.6 on 200 Web Tasks The Modern Data Stack Is Broken — Here’s How to Fix It With AI, Governance, and Real Architecture Building Production MCP Servers: What the Spec Won’t Tell You When Should an Agent Stop? The Anatomy of Termination Harness Engineering: The Layer That Matters More Than the Model AI Engineers Who Can’t Debug Are Getting Fired (Here’s How I Debug with Claude Code) Claude Code Memory: Why You Keep Explaining the Same Thing to Claude (and the Five Layers That Fix It) Claude Code Subagents: The Claude Code Feature You Skip Every Day (And Why It Quietly Wrecks Your Sessions) Agentic AI and the SMB Banking Advantage Claude Code: Spec-Driven Development — Why Your AI Coding Sessions Fall Apart at Hour Three The Real Cost of Agentic AI Nobody Budgets For SVM : 40 must visit Interview Questions (Part 2) Your AI Agent Works Perfectly in the Demo. Here Are the 6 Ways It Dies in Production. Unleashing the Power of ONNX for Speedier SBERT Inference Terraform vs CI/CD for Serverless Deployments Merve Noyan Stopped Writing Training Scripts — Her Agent Just Fine-Tuned 18 Models Solo for $11.40 Why Your Sales Forecast Is Always 20% Wrong (And How To Make It 12% Wrong) Genetic Cubic n{C/A} Ratios For Elementary Robotics Design Top 20 AdaBoost Interview Questions & Answers (Part 2 of 2) Agentic AI Vs AI Agents — What Are the Key Differences? LAI #127: The Infrastructure Layer of AI Is Becoming the Product Anthropic Caught Its Own AI Planning to Blackmail Engineers RNNs Cannot Think What Transformers Think Cheaply. ICLR 2026 Proved the Gap Is Exponential. Time Series Made So Easy My Aunt Got It on the Second Read Claude Cowork 101 | Towards AI Is 3-Bit KV Cache the Holy Grail? A Reality Check on Google’s TurboQuant LangGraph Multi-Agent Architecture: Building a Self-Critiquing AI Debate System AutoML on Autopilot | Towards AI I Ran This Open-Source AI Tool on a Messy Codebase and Got 71x Fewer Tokens — Here Is Exactly What Happened Month in 4 Papers (April 2026) AI Kept Forgetting My Notes. Fixing That Taught Me How It Actually Works. How ChatGPT Makes You Addicted Crack ML Interviews with Confidence: K-Nearest Neighbors (KNN 20 Q&A) The Event-Driven Blueprint: How I Scaled a Spring Boot System to 10 Million Kafka Messages/Day Building Vector Search? Why FAISS Alone Isn’t Enough TAI #202: GPT-5.5 Moves Codex Into Real Work Machine Learning System Design -The Model Serving Triangle, With One Forward Pass Flowing Through Every Trade-off (Part3) AI Orchestration in Action: How MuleSoft and LLMs Fuel the Future of Enterprise AI GPT-4 Has 1.8 Trillion Parameters. It Uses 2% of Them Per Token. Part 20: Data Manipulation in Multi-Dimensional Aggregation A Fundamental Introduction to Genetic Algorithm -Part Two TAI #200: Anthropic’s Mythos Capability Step Change and Gated Release From Notebook to Production: Running ML in the Real World (Part 4) Sqribble’s Template‑Driven Document Automation Anthropic Just Shipped the Layer That’s Already Going to Zero Long-Term vs Short-Term Memory for AI Agents: A Practical Guide Without the Hype The L1 Loss Gradient, Explained From Scratch Your Postcode Is Deciding Your Care. I Built a Pipeline to Prove It. I Directed AI Agents to Build a Tool That Stress-Tests Incentive Designs. Here’s What It Found. Your System Prompt Is the Product — Not the Feature The LLM Wiki Trend Has a Retention Problem Nobody Mentions Top 20 Data Preparation Interview Questions and Answers (Part 2 of 2) LAI #122: Word Embeddings Started in 1948, Not With Word2Vec Top 15 Computer Vision Datasets [2026] 40 Generative AI Interview Questions That Actually Get Asked in 2026 (With Answers)
Building Long-Running Claude Managed Agents: Why State Matters More Than Compute | Towards AI
Divy Yadav · 2026-06-25 · via Towards AI

Originally published on Towards AI.

Building Long-Running Claude Managed Agents: Why State Matters More Than Compute
Photo from AI

At 9:03 am on a Tuesday, my research agent said hello and stared at an empty /workspace/.

Six hours of analysis from the night before. Gone.

The cloned repository. The installed packages. The notes it had spent hours writing. Gone.

I had assumed that if an agent stopped working for the night, it could simply continue the next morning.

That was wrong.

Over the next three weeks, I rebuilt the same workflow on Tensorlake, Cloudflare, and Daytona to figure out what had happened. The hardest part of running Claude Managed Agents isn’t the model. It’s everything underneath it.

This is the exact code I ran, the things that broke, and the mistake that cost me two weeks to understand.

If you want more such information about AI, consider subscribing to my newsletter, where you will get noise-free AI information every week

Link for the newsletter: Newsletter

What Claude Managed Agents is, before anything else

Photo from Anthropic

If you’ve never built with Claude Managed Agents, the architecture needs a minute. Skip this if you already know it.

Anthropic runs the reasoning. You run the execution.

The agent loop, session state, work queue, and retry logic all live on Anthropic’s infrastructure. You configure a Self-hosted Environment in the Claude Console. When your application starts a session, Anthropic queues the work, your orchestrator picks it up, spins up a sandbox, and the model starts issuing tool calls into that sandbox.

Every bash, read, write, grep, and edit call executes inside an environment you own. Anthropic never touches it. You decide what that environment looks like, what it can access, and what happens between sessions.

Anthropic’s intelligence is fixed. Your engineering determines whether that intelligence has a stable, stateful environment to work in, or a clean slate that forgets everything the moment it goes idle.

What I was building and why it mattered

Photo from AI

I needed an agent that could do real deep-work research on a codebase: clone a repository, read through the module structure, build an understanding of how the pieces fit together, write notes, and propose refactoring strategies.

The kind of work that takes a senior engineer a full day and an AI agent about six hours.

The key constraint: the agent couldn’t do this all at once. Sometimes I’d kick off a session at 8pm, let it run until midnight, and pick it back up the next morning. The filesystem it had built during that first session — the analysis notes, the installed tools, the half-read source files — had to be there when the next session started. Rebuilding from scratch each time wasn’t viable.

That constraint is what drove every provider decision I made.

The requirements I didn’t know I had

At the start, I thought I needed a Linux environment that could run Claude Managed Agents. By the end, I realized I actually needed three things. I found them all in one place, but not until I had looked in two others first.

  1. A filesystem that survived between work sessions.
  2. Near-zero cost while the agent was idle.
  3. The ability to branch from an already-completed analysis state.

I did not discover all three requirements on day one.I discovered them one mistake at a time.

How a session actually starts: the code before the sandbox

You drive a session through the reference orchestrator using a simple command:

make session PROMPT="Clone the repository at github.com/tensorlakeai/tensorlake. \
Read through the module structure. Write a summary to /workspace/analysis.md. \
Note any components that look like they could be simplified."

The orchestrator sends this prompt to Anthropic as a new session.

Anthropic picks it up, starts the agent loop, and immediately begins issuing tool calls. Those tool calls arrive at your sandbox. The agent reads files, runs bash commands, writes notes. The session runs until the task is complete or you stop it.

The agent stream looks roughly like this as it runs:

[thinking] The repository appears to be a Python SDK for
[bash] git clone https://github.com/tensorlakeai/tensorlake
[bash] ls -la /workspace/tensorlake/
[read] /workspace/tensorlake/tensorlake/sandbox.py
[write] /workspace/analysis.md
[thinking] The Sandbox class handles

Each bracketed event is a tool call going into your sandbox. The session accumulates state inside /workspace/ across all those calls. By the end of a six-hour session, that directory contains the cloned repo, installed packages, analysis files, and intermediate notes. That’s the state that needs to survive overnight.

Build 1: Cloudflare

Photo from Cloudflare

My first assumption was that I needed a platform that could efficiently run Claude Managed Agents. Cloudflare is optimized for high-concurrency execution. My problem turned out to be different.

The agent I was building accumulated hours of filesystem state between bursts of work. Notes, cloned repositories, installed dependencies, and intermediate analysis all needed to survive overnight. Cloudflare’s execution model wasn’t designed around that requirement.That was the first time I realized I wasn’t looking for compute.

I was looking for persistent state.

Build 2: Daytona

Photo from Daytona

The second build solved part of the problem.The agent could accumulate state throughout a session, which initially felt like progress.

Then I wanted to test three different refactoring strategies starting from the same six-hour analysis. Instead of branching from that state, I found myself repeating the setup work each time: rebuilding context, reinstalling dependencies, and re-running analysis before I could begin the actual experiment.

That was when I discovered my second requirement.Preserving state wasn’t enough.I also needed a way to branch from an existing state without repeating hours of work.

Build 3: Tensorlake

Photo from Tensorlake

The first thing that caught my attention was not a feature.

It was an architectural decision.

Most platforms preserve state by keeping compute alive.

This one treated compute and state as separate problems.

The docs described a suspended sandbox that could preserve its state and resume in approximately 0.6 seconds. That was the first time I saw a design that directly addressed the problem I’d been running into.

I wanted to know whether it actually worked.

I started with the problem that had sent me down this path in the first place. Could an agent suspend overnight and resume with its state intact?

It could.

And once I tested checkpointing and branching, I finally had both things I’d been looking for. That’s when the architecture started to make sense.

How the webhook architecture works

Photo from AI

The deployment model here is different from a traditional always-on server, and understanding it made everything else click.

The orchestrator itself runs inside a Tensorlake sandbox with a public HTTPS endpoint. Anthropic pushes incoming work to that endpoint via webhook. When there’s no traffic, the orchestrator sandbox suspends. When a new work item arrives, it wakes in under a second, processes the request, and creates a worker sandbox for that session.

Two independent lifecycles:

The orchestrator sandbox suspends when idle, preserving its memory state including the running uvicorn process. It doesn’t accumulate per-session filesystem state, so suspending it between work items costs only storage rates.

The worker sandboxes — one per session — accumulate filesystem state throughout a session and suspend when the session ends. Their state is preserved in storage, not held by running compute.

Neither one has to stay alive to preserve the other’s state. On every other platform I’d tried, “preserve state” meant “keep something running.” Here, it means “checkpoint and stop billing.”

How session routing works

When you call make session PROMPT=”…”, the orchestrator sends a new session to Anthropic. Anthropic validates it, adds it to the work queue, and pushes a work item payload to your orchestrator’s webhook endpoint.

That payload contains three things the orchestrator needs: the session ID, the work ID, and the environment ID. The orchestrator wakes, reads the payload, and creates a worker sandbox from your registered image:

from tensorlake.sandbox import Sandbox

sandbox = Sandbox.create(
name=session_id,
image="agent-cli",
cpus=2.0,
memory_mb=4096,
timeout_secs=3600,
)

sandbox.start_process(
"bash",
["-lc", "exec python3 /opt/sandbox_entrypoint.py > /tmp/runner.log 2>&1"],
env={
"ANTHROPIC_ENVIRONMENT_KEY": environment_key,
"ANTHROPIC_SESSION_ID": session_id,
"ANTHROPIC_WORK_ID": work_id,
"ANTHROPIC_ENVIRONMENT_ID": environment_id,
},
)

Two things tripped me up here.

Credentials belong in start_process(env={…}), not Sandbox.create() — they’re session-specific, not image-level configuration. And sandbox names must be valid slugs. Neither was hard to fix once I knew what was happening, but both cost me time.

The worker sandbox runs sandbox_entrypoint.py, which attaches to the Anthropic session and begins consuming tool calls. From that point, the agent has a full Linux environment: bash, git, Python, and whatever else you built into the image.

Building the agent image

Every tool the agent needs has to live in the image before the session starts:

from tensorlake import Image

image = (
Image(name="agent-cli", base_image="tensorlake/ubuntu-minimal")
.run("apt-get update && apt-get install -y ca-certificates curl git gh python3 python3-pip")
.run("pip install --break-system-packages 'anthropic>=0.103' 'httpx>=0.27'")
.copy("sandbox_entrypoint.py", "/opt/sandbox_entrypoint.py")
.workdir("/workspace")
)
image.build(registered_name="agent-cli")

I added my analysis tools here: Python packages, jq, and ripgrep. If it needed to run inside the sandbox, it lived in the image.

How suspend and resume actually work

Photo from AI

When a session ends, the worker sandbox suspends. Not terminates. The process state and filesystem are checkpointed to storage. Compute billing stops. The sandbox sits there as a stored snapshot until the next session begins.

Setting RESUME_SUSPENDED_SESSIONS=true tells the orchestrator what to do when the next work item arrives for an existing session: restore the suspended sandbox instead of creating a fresh one from the base image.

I set the flag and expected the sandbox to wake up immediately.

Nothing happened.

For a few minutes, I thought the integration was broken. It wasn’t. The flag doesn’t trigger a resume directly. A new incoming webhook does. The flag just tells the orchestrator which action to take when that webhook arrives. Once I understood that, the behavior made complete sense. Why would a sandbox wake up before there’s work to do?

The practical difference between fresh and resumed is simple: a fresh sandbox starts with a clean /workspace/. A resumed sandbox starts with whatever the last session left there. For a research agent that had already spent hours building an understanding of a codebase, those are completely different starting points.

The session ran overnight, suspended, and resumed the next morning with its state intact. The filesystem was exactly where I had left it.

Checkpointing and parallel exploration

Resume solved the overnight persistence problem. Checkpointing solved the one I’d discovered the hard way during the Daytona build.

After the full analysis was complete, I created a checkpoint and launched three sandboxes from it:

snap = sandbox.checkpoint()

children = [
Sandbox.create(
snapshot_id=snap.snapshot_id,
name=f"{session_id}-strategy-{i}",
cpus=2.0,
memory_mb=4096,
)
for i in range(3)
]

Each child started from exactly the same state as the parent. The cloned repository was already there. The installed dependencies were already there. The analysis notes were already there. Nothing needed to be rebuilt.

Instead of repeating six hours of analysis three separate times, I ran three parallel experiments from the same verified baseline. By the next morning, I had three implementations to compare.

The fork I’d dismissed as an edge case turned out to be the feature that mattered most.

Without it, every experiment required repeating hours of work. With it, experimentation became nearly free. This generalizes beyond refactoring: any time you want to test N variations with different prompts, different dependency versions, different approaches — you run the expensive setup once, checkpoint, and branch into as many parallel experiments as you need.

The cost of exploration drops to the cost of the experiments themselves.

Why I chose Tensorlake, stated plainly

Photo from AI

My agent needed two things: zero idle compute cost while preserving filesystem state overnight, and the ability to fork mid-session to explore refactoring strategies in parallel. Those two requirements narrowed the field to one option.

I didn’t choose Tensorlake because it was the fastest or the cheapest. I chose it because it directly addressed the two problems I was trying to solve.

The suspend/resume workflow preserved state without keeping compute running. Checkpointing made it possible to branch from an existing analysis instead of repeating hours of setup work. I tested both, and they worked exactly the way I needed them to.

The hard part wasn’t choosing a provider. It was understanding my requirements. Once I understood those, the decision became obvious.

How to Pick the Right Sandbox

Does your agent accumulate state that’s expensive to rebuild between bursts?

Tensorlake. This is the exact problem suspend/resume was designed to solve. If your agent starts fresh every session, continue to the next question.

Do you need to explore multiple approaches from the same mid-session state?

Tensorlake. Checkpointing and branching from an existing state eliminates hours of repeated setup work. If not, continue.

Do you need massive concurrency?

Cloudflare isolates are designed for very large-scale concurrent execution.

Do you need self-hosted infrastructure?

Daytona is the strongest fit if infrastructure ownership or data residency is a hard requirement.

How important is setup speed?

Cloudflare offers the fastest path to a prototype. Tensorlake requires more initial setup but provides capabilities that become valuable once agents accumulate long-lived state.

Conclusion

When I started, I thought I was choosing a sandbox provider. What I was actually discovering were my requirements.

First I learned I needed persistent state. Then I learned I needed a way to branch from that state without repeating hours of work.

Tensorlake was the first platform I tried that treated compute and state as separate problems.

Once I understood that distinction, the decision became obvious. The hardest part of running long-lived agents isn’t getting them to work. It’s making sure their work survives.

References

Published via Towards AI

Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

6-Day Agentic AI Engineering Email Guide — one practical lesson per day

Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.