惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

O
OpenAI News
GbyAI
GbyAI
人人都是产品经理
人人都是产品经理
Last Week in AI
Last Week in AI
F
Fortinet All Blogs
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
罗磊的独立博客
爱范儿
爱范儿
B
Blog
C
Check Point Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园_首页
A
About on SuperTechFans
Engineering at Meta
Engineering at Meta
V
Visual Studio Blog
P
Proofpoint News Feed
小众软件
小众软件
Google DeepMind News
Google DeepMind News
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
G
Google Developers Blog
Y
Y Combinator Blog
Recorded Future
Recorded Future
博客园 - 聂微东
WordPress大学
WordPress大学
博客园 - 【当耐特】
腾讯CDC
T
Tailwind CSS Blog
The Register - Security
The Register - Security
V
V2EX
S
SegmentFault 最新的问题
IT之家
IT之家
D
Docker
I
InfoQ
大猫的无限游戏
大猫的无限游戏
云风的 BLOG
云风的 BLOG
Microsoft Azure Blog
Microsoft Azure Blog
月光博客
月光博客
Stack Overflow Blog
Stack Overflow Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Jina AI
Jina AI
The Cloudflare Blog
量子位
Microsoft Security Blog
Microsoft Security Blog
aimingoo的专栏
aimingoo的专栏
博客园 - 叶小钗
H
Help Net Security
MongoDB | Blog
MongoDB | Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
美团技术团队
B
Blog RSS Feed

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
GitHub - AhmadHammad21/OpenDevOps
ahmadhammad0 · 2026-06-14 · via Hacker News - Newest: "AI"

OpenDevOps Agent

Open-source multi-cloud DevOps agent (AWS + Azure). Bring any LLM via LiteLLM — OpenAI, Anthropic, OpenRouter, Groq, Gemini, Mistral, Ollama for air-gapped / regulated environments, or reuse your existing Claude Code subscription (auto-detected). Investigates incidents, finds root causes, and gives actionable mitigation plans — without the cloud-vendor DevOps-agent price tag.

📊 Benchmarked, not just demoed

On a reproducible 10-incident suite (real AWS + Azure resources, scored against ground truth), running on a commodity open model (gpt-oss-120b — no frontier model required):

Root causes found Median time Cost / investigation vs. AWS DevOps Agent vs. manual triage
9 / 10 (90%) ~52 s ~$0.03 ~10× cheaper¹
(~$0.03 vs ~$0.43)
~1,000× cheaper²

~$0.03 of compute replaces ~$50 of engineer toil — and costs a fraction of a managed cloud DevOps agent — while returning the answer in under a minute instead of half an hour. Reproduce it with make evalfull benchmark & methodology.

¹ vs. AWS DevOps Agent — its per-second rate applied to the same wall-clock time (~$0.43/investigation); verify against AWS's published pricing. ² Illustrative unit economics vs. ~20–40 min of on-call triage. Cost shown is the provider-dashboard actual; see caveats.

Cloud setup: AWS (IAM) · Azure (service principal / login)

Demo

Autonomous Lambda error-spike investigation

Autonomous incident detection — a crashing Lambda is caught automatically, the agent reads the traceback from CloudWatch Logs, finds the root cause, surfaces it on the Monitoring dashboard, and posts the mitigation to Slack. No human in the loop.

Amazon Q Developer and the AWS DevOps Agent are excellent if you live entirely inside the AWS Console with Bedrock-managed models. OpenDevOps is the open-source alternative for everyone else:

  • Any LLM, not just Bedrock. LiteLLM-compatible — OpenAI, Anthropic direct, OpenRouter, Groq, Gemini, Mistral, or run Ollama locally for air-gapped / regulated environments. Auto-detects your existing Claude Code subscription so you pay zero incremental LLM cost if you're already on a Max/Pro plan.
  • Multi-cloud out of the box. AWS + Azure investigations in the same chat (one organization can connect both clouds at once). AWS-only agents stop at the AWS perimeter.
  • Your data stays in your database. Investigations, prompts, and tool outputs persist in your Postgres or SQLite — your VPC, your retention, your encryption. Matters for HIPAA, PCI, FedRAMP, and EU AI Act audits.
  • Fully auditable. Every prompt, tool call (args + result), and token is open and streamed live to the UI; nothing is hidden. AWS Agent is a closed black box.
  • Customizable. Add tools as plain Python functions, add runbooks by dropping a SKILL.md file, modify the system prompt. Fork it if you need to.
  • Investigate from anywhere. Built-in MCP server makes it usable from Claude Desktop, Cursor, or any MCP client — not just the AWS Console.
OpenDevOps AWS DevOps Agent / Q Developer
LLM Any (LiteLLM, Claude Code, Ollama) Bedrock-managed only
Cloud coverage AWS + Azure (more coming) AWS only
Data location Your DB / VPC AWS-managed, not portable
Customization Open source — modify anything Closed product
Pricing LLM at retail (or $0 via Ollama / Claude Code) Per-investigation + Bedrock markup
Self-host Docker / Railway / on-prem / air-gapped No

When AWS is the better pick: if you're 100% AWS, never plan to leave, and want zero infrastructure to run, Amazon Q Developer's native Console integration and AWS-only signals (Trusted Advisor, AWS Config, Compute Optimizer) are hard to beat. OpenDevOps is for everyone else.

What's inside

  • LangChain DeepAgents as the agent framework — planning, tool orchestration, and session memory out of the box
  • 21 read-only AWS tools across CloudWatch (6), CloudTrail (2), ECS (4), Lambda (4), EC2 (2), RDS (2), IAM (1), plus bash escape hatch, cross-session history analytics, skills, and submit_investigation — plain Python functions, schemas inferred automatically
  • Azure support (CLI-first) — investigates Azure through the read-only az CLI + kubectl (for AKS) and a set of Azure runbook skills (AKS debugging, App Service errors, Azure Monitor/KQL, VM diagnostics) — no separate SDK tools needed. Read-only; connect via a service principal or az login — see apps/documentation/azure_setup.md
  • Sandboxed bash execution tool — agent can run whitelisted read-only AWS CLI (aws), Azure CLI (az), kubectl, and docker commands as a last resort when the structured tools fall short; every command validated against an allowlist before execution; never uses shell=True; hard 30-second timeout
    • Includes CloudWatch Logs Insights (query_logs_insights) — full query language support: fields, filter, stats, sort, limit; results include scanned MB
  • Streaming responses — FastAPI SSE endpoint streams agent tokens in real time as the LLM reasons; tool calls appear as they complete
  • Event-driven incident detection — EventBridge → SQS → long-poll consumer; 9 EventBridge rules cover CloudWatch alarms, ECS task failures, Lambda async errors, RDS events, EC2 state changes, CodePipeline failures, and AWS Health events; uses a DLQ plus database-backed incident claims to avoid duplicate investigations; runs alongside the metric poller — see apps/documentation/event_detection.md
  • Context enrichment — before the LLM runs, deterministic boto3 calls fetch facts about the affected resource (alarm details, recent logs, function config, etc.) to reduce tool call count and speed up investigations
  • Monitoring dashboard — live incident feed showing all event-driven investigations: confidence level (or FAILED badge), affected service, root cause summary; each alert links back to its original investigation session via View investigation so you can follow up without losing context; real-time SSE push keeps the page live without polling — see apps/documentation/monitoring.md
  • AWS Configuration settings tab — admin-only editable tab in Settings for SQS Queue URL and AWS Region; shared org-wide via database-backed app config; includes an inline IAM permission checker per service
  • Web UI — React + Vite SPA served by FastAPI:
    • Chat page — streaming responses, collapsible tool call inspector, cost/latency card, stop button; supports ?prompt= deeplink for pre-seeded investigations from the Monitoring dashboard
    • Session history sidebar — lists all past conversations; click any to resume with full tool call inspector and cost card restored; new chat and delete (soft) buttons
    • Monitoring page — live incident feed from event-driven detection; alert detail with investigate deeplink
    • Dashboard — session counts, tool call stats, cost/latency, context saved, activity chart, service breakdown, root cause distribution, recent sessions
    • History page — keyword search across all past sessions
    • Settings page — AWS Configuration (editable, admin-only), Environment (read-only env vars), Agent config, Integrations
    • Team page — admin-only user management: add, remove, and change roles
  • Auth & RBAC — optional password-based auth with admin and user roles; JWT tokens; first registered user auto-becomes admin; disabled by default (set JWT_SECRET to enable) — see apps/documentation/auth.md
  • Three storage backends — pick one via CHECKPOINT_BACKEND in .env; see apps/documentation/databases.md
    • memory — zero config, no persistence; great for CI and quick testing; autonomous polling/event monitoring is disabled in this mode
    • sqlite — local file, no external services; recommended for single-server and personal use
    • postgres — full production persistence via psycopg3 + AsyncPostgresSaver
    • Schema: users, sessions, messages, tool_calls, usage_events — see apps/documentation/schema.md
    • Soft delete — deleted sessions are hidden immediately but data is preserved for the 30-day cleanup job
  • Structured logging via Loguru — used consistently across all modules (tools, agent, API, CLI); every request shows agent reasoning, tool calls with args/results, and a done summary with latency + token counts
  • CLIdevops-agent investigate, ask, and report commands powered by the same agent
  • Any LLM via LiteLLM — OpenAI, Anthropic, OpenRouter, Groq, Gemini, Mistral, Ollama (local / air-gapped), or any OpenAI-compatible endpoint. Auto-detects local Claude Code subscription (~/.claude OAuth) so a Max/Pro plan can power the agent at zero incremental cost. Swap models via a single env var (LLM_MODEL) — no code changes

Quick Start

1. Install dependencies

cd apps/backend && uv sync

2. Configure environment

cp .env.example .env
# Edit .env — add your OPENROUTER_API_KEY and set AWS_PROFILE

3. Set up AWS profile

aws configure --profile devops-agent-readonly
# AWS Access Key ID:     your_key_id
# AWS Secret Access Key: your_secret_key
# Default region:        us-east-1
# Default output format: json

# Verify
aws sts get-caller-identity --profile devops-agent-readonly

4. Choose a storage backend

Three options — pick one and add it to .env. Full details in apps/documentation/databases.md.

Memory (default — zero config, nothing persists on restart)

CHECKPOINT_BACKEND=memory

SQLite (recommended for local dev — persists to a file, no external service needed)

CHECKPOINT_BACKEND=sqlite
SQLITE_PATH=./data/agent.db   # created automatically on first start

PostgreSQL (recommended for production)

# Start Postgres with Docker
docker run -d --name opendevops-pg \
  -e POSTGRES_DB=opendevops \
  -e POSTGRES_USER=dev \
  -e POSTGRES_PASSWORD=dev \
  -p 5433:5432 \
  postgres:16

# Add to .env
CHECKPOINT_BACKEND=postgres
DATABASE_URL=postgresql://dev:dev@localhost:5433/opendevops

# Create app tables (safe to re-run)
cd apps/backend && uv run migrate

5. Run

Option A — Docker Compose (recommended, AWS CLI included)

docker compose -f deployment/docker-compose/docker-compose.yml up --build
# Backend: http://localhost:8000
# Frontend: http://localhost:80
# Postgres (host): localhost:5433

The backend image installs AWS CLI v2 automatically — the bash execution tool works out of the box. Host AWS credentials (~/.aws) are mounted read-only into the container. For production on AWS, remove the volume mount and attach an IAM role to the instance/task instead.

Option B — Local dev (two terminals)

# Terminal 1 — FastAPI backend with hot reload
cd apps/backend && uv run dev
# Terminal 2 — React frontend (Vite dev server with HMR)
cd apps/frontend && npm run dev
# Open http://localhost:5173

Note: local dev requires aws CLI installed on your machine for the bash tool to work. Install it from https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html

CLI

cd apps/backend

# Investigate an incident
uv run devops-agent investigate "high error rate on my payment Lambda"

# With alarm and service hints
uv run devops-agent investigate "latency spike" --alarm HighLatencyAlarm --service api-service

# Freeform Q&A
uv run devops-agent ask "why would a Lambda function suddenly start throttling?"

# Daily ops health report
uv run devops-agent report

AWS IAM Setup

The agent needs read access across your AWS account, plus optional write access scoped to opendevops-* resources if you use the event-driven monitoring setup wizard. Two least-privilege policies (Operational + Setup) and full step-by-step instructions are in apps/documentation/iam_setup.md.

Project Structure

apps/
├── core/                  # Installable package `opendevops-core` — the shared agent brain
│   └── src/opendevops_core/
│       ├── agent/         # DeepAgents setup, prompts, LLM wiring, DB layer (backends + ABC)
│       ├── tools/         # bash, history, skills, final-answer + response cap/cache
│       ├── providers/     # AWS provider — tools, context, poller, event consumer
│       ├── models/        # Pydantic models: agent, chat, sessions, users
│       ├── skills/        # Markdown runbooks (lambda-throttling + add your own)
│       ├── integrations/  # slack_webhook.py, telegram.py
│       ├── migrations/    # Numbered baseline SQL migrations (001–013) — bundled with the wheel
│       └── config.py      # CoreSettings + get_settings()/configure() injection hook
├── backend/               # OSS web app + CLI — depends on opendevops-core via uv path source
│   ├── src/
│   │   ├── api/
│   │   │   ├── app.py     # FastAPI app factory — mounts routers, serves frontend
│   │   │   ├── auth.py    # JWT helpers + FastAPI auth dependencies
│   │   │   └── routers/   # chat, sessions, users, settings, history, dashboard, monitoring
│   │   ├── cli/           # Typer CLI commands
│   │   ├── config/
│   │   │   └── appsettings.py  # Settings(CoreSettings) — adds web/auth-only fields, calls configure()
│   │   └── mcp_server.py  # MCP server (stdio / HTTP+SSE)
│   ├── migrations/        # OSS-app-only migrations (currently none; all schema is core baseline)
│   ├── tests/
│   └── pyproject.toml
├── frontend/
│   └── src/
│       ├── pages/         # ChatPage, DashboardPage, HistoryPage, SettingsPage, UsersPage, LoginPage
│       └── components/    # Sidebar, Header, ProtectedRoute, AgentMessage, ...
└── documentation/         # Feature reference — auth, schema, skills, databases, UI, ...
deployment/
├── docker-compose/        # docker-compose.yml (PostgreSQL + backend + frontend)
└── railway/               # Dockerfile.railway + railway.toml (combined single-image deploy)

Configuration

Variable Default Description
LLM_MODEL openrouter/openai/gpt-4o LiteLLM model string — provider/model format; see apps/documentation/llm_providers.md
LLM_API_BASE none Custom base URL for OpenAI-compatible endpoints (e.g. Ollama, vLLM)
LLM_API_KEY none API key for custom endpoints; standard provider keys (e.g. ANTHROPIC_API_KEY) are read automatically
OPENROUTER_API_KEY none Required when using any openrouter/ model
CHECKPOINT_BACKEND memory Storage backend: memory · sqlite · postgres — see apps/documentation/databases.md
SQLITE_PATH ./data/agent.db SQLite file path — only used when CHECKPOINT_BACKEND=sqlite
DATABASE_URL none PostgreSQL connection string — only used when CHECKPOINT_BACKEND=postgres
AWS_REGION us-east-1 AWS region
AWS_PROFILE none AWS named profile (e.g. devops-agent-readonly)
MAX_TOOL_CALLS 20 Hard cap on tool calls per investigation
INVESTIGATION_TIMEOUT 120 Timeout in seconds
TOOL_RESPONSE_MAX_CHARS 40000 Truncate tool responses larger than this before feeding to the LLM; 0 disables
SLACK_WEBHOOK_URL none Slack incoming webhook URL; leave unset to disable notifications
TELEGRAM_BOT_TOKEN none Telegram bot token from @BotFather; leave unset to disable
TELEGRAM_CHAT_ID none Target chat/group/channel ID (negative number for groups)
POLL_INTERVAL_SECONDS 0 Proactive polling interval in seconds; 0 disables the poller
POLL_ERROR_THRESHOLD 5.0 Lambda error rate % that triggers an automatic investigation
POLL_REINVESTIGATE_HOURS 1 Cooldown period — skip re-investigating the same alarm within N hours
SUMMARIZATION_ENABLED true Auto-compact sessions when they exceed the threshold
SUMMARIZATION_THRESHOLD_CHARS 60000 Total session chars before compaction fires (~15K tokens)
SUMMARIZATION_KEEP_CHARS 20000 Recent chars to preserve intact during compaction (~5K tokens)
JWT_SECRET none Secret key for JWT signing; leave unset to disable auth entirely
JWT_EXPIRE_MINUTES 1440 JWT token lifetime in minutes (default 24 h)
SNS_TOPIC_ARN none SNS topic to publish investigation findings to after each event-driven run
SQS_QUEUE_URL none SQS queue URL for the event consumer to poll; also set via Settings → AWS Configuration
EVENT_CONSUMER_ENABLED false Explicitly enable the SQS event consumer (also auto-starts if SQS_QUEUE_URL is set)
DATA_DIR data Reserved data directory setting; init state is stored in the selected database backend

TODO / Roadmap

Near-term

  • Cache layer — in-process TTL cache (cachetools) on all 19 AWS tool functions; 2-minute TTL, 256 entry max, AWS profile+region included in cache key
  • Schema / models layer — centralized src/models/ package for all Pydantic models: agent domain, memory state, and API request/response schemas
  • Soft-deleted session cleanup job — product version only; OSS users manage their own DB
  • Investigation history skill — cross-session analysis: recurring errors, most-triggered alarms, patterns across all past sessions for a user
  • User rolesadmin / user roles with JWT auth, first-user bootstrap, admin-only user management UI; optional (disabled when JWT_SECRET unset) — see apps/documentation/auth.md

Medium-term

  • React frontend — rewrite the single-file HTML UI in React; component-based architecture, proper state management, hot reload
  • Dashboard — summarized view of troubleshooting activity, recurring incidents, query breakdown by service
  • Multi-provider LLM support — 100+ providers via LiteLLM; swap models with a single LLM_MODEL env var change; supports OpenRouter, Anthropic, OpenAI, Groq, Ollama, and any OpenAI-compatible endpoint; see apps/documentation/llm_providers.md
  • MCP integration — expose the agent as an MCP server (devops-agent mcp); investigate, ask, and list_sessions tools available in Claude Desktop, Cursor, or any MCP-compatible client; stdio and HTTP+SSE transports; see apps/documentation/mcp_server.md
  • Multi-backend storagememory (zero config), sqlite (local file, no external service), postgres (production); switch with one env var; see apps/documentation/databases.md
  • Skills system — on-demand investigation skills loaded from src/skills/*/SKILL.md; skill names injected into system prompt at startup, full content loaded only when agent calls use_skill(name); ships with lambda-throttling skill; add your own by dropping a SKILL.md into src/skills/<name>/
  • Custom tools via URL — register external tools by pointing at an OpenAPI/HTTP endpoint; agent discovers and calls them alongside built-in AWS tools
  • Bash CLI escape hatch (Phase 1)run_bash_command is implemented for read-only AWS CLI, kubectl, and docker commands with strict allowlist validation and timeout.
  • Bash sandbox Phase 2 — run each bash command in an isolated throwaway container (--network none, read-only FS, non-root, resource limits).
  • Tool response capping — truncates oversized AWS tool responses (CloudWatch logs, CloudTrail events) before they reach the LLM context window; configurable via TOOL_RESPONSE_MAX_CHARS (default 40 000 chars ≈ 10 K tokens)
  • Conversation summarization — automatically summarize old messages when the session approaches the model's context limit; preserves recent exchanges and injects a structured summary so long investigations never fail mid-session; summarization events tracked in usage_events.metadata and surfaced in the dashboard
  • Optimize tool loading — pass only relevant tools per investigation context instead of the full 19-tool set
  • Message middleware pipeline — compaction, summarization, intent detection, context trimmer
  • Guardrails — input/output validation, PII scrubbing, query scope enforcement
  • Multi-model escalation — route simple queries to cheaper/smaller models, escalate hard investigations to larger ones
  • Fun streaming labels — contextual loading copy ("Digging through CloudTrail…", "Lemonizing metrics…", "Cooking up a root cause…")
  • Slack & Telegram notifications — reactive: posts after every investigation to Slack (Block Kit) and/or Telegram (HTML bot message); proactive: background poller checks CloudWatch alarms and Lambda error rates, auto-investigates, and delivers to both channels; set SLACK_WEBHOOK_URL and/or TELEGRAM_BOT_TOKEN+TELEGRAM_CHAT_ID in .env; see apps/documentation/telegram.md
  • Event-driven incident detection — EventBridge → SQS → long-poll consumer; 9 EventBridge rules covering CloudWatch alarms, ECS, Lambda, RDS, EC2, CodePipeline, and AWS Health; runs in parallel with the metric poller; see apps/documentation/event_detection.md
  • Context enrichment — deterministic boto3 calls per event type before LLM runs; reduces tool call count by front-loading relevant resource facts
  • Monitoring dashboard — live incident feed with real-time SSE push, per-service health summary (DB-backed, survives restarts), alert detail page; "View investigation" opens the original agent session for follow-up; failed investigations flagged separately; see apps/documentation/monitoring.md
  • AWS Configuration settings tab — admin-only editable tab for SQS/region config; shared org-wide via database-backed app config; inline IAM permission checker

Later

  • Observability — OpenTelemetry traces for agent steps, tool call latency, LLM token usage
  • Follow-up question suggestions — after each investigation completes, generate 3 suggested follow-up questions in the background and surface them in the UI as clickable chips
  • Session / user feedback loop — thumbs up/down on investigations, feed signals back to the agent and to an internal quality dashboard
  • Knowledge base — attach internal runbooks, post-mortems, and architecture docs so the agent grounds answers in org-specific context
  • Multi-account AWS — support multiple AWS profiles per org via aws_profiles table (schema already in place)
  • Multi-cloud support — extend tooling to GCP (Cloud Monitoring, Cloud Logging, GKE) and Azure (Monitor, Log Analytics, AKS); unified incident investigation across providers
  • Bash sandbox Phase 2 — isolated Docker container — each run_bash_command call runs inside a throwaway container: --network none, --read-only filesystem, --memory 256m, --cpus 0.5, non-root user; container destroyed immediately after the command completes; IAM read-only role remains the last line of defense

Product (SaaS)

  • Redis cache — replace in-process cachetools with Redis; shared across workers, survives restarts, per-org cache namespacing to prevent data leakage between tenants
  • Soft-deleted session cleanup — scheduled job (Inngest or APScheduler) to purge is_deleted = TRUE sessions older than a configurable retention window (default 30 days); GDPR right-to-erasure compliance
  • Org-scoped AWS credential management — per-org credential vault; agents use org-scoped profiles instead of a single global AWS_PROFILE
  • Per-org AWS credential store — encrypted credential vault per organization; agents use org-scoped profiles instead of a single global AWS_PROFILE
  • Billing & usage metering — track token usage and tool calls per org/user; expose cost dashboards; integrate with Stripe for usage-based billing

Security & Sandboxing

The bash execution tool runs whitelisted read-only commands only. Every command is validated against an allowlist before execution — anything not on the list is rejected immediately and logged.

Current (Phase 1): allowlist validation + subprocess with hard timeout. No write commands permitted under any circumstances. shell=True is never used.

Phase 2 — Isolated sandbox (planned):

  • Every bash command runs inside a throwaway Docker container
  • --network none — no internet access from inside the sandbox
  • --read-only filesystem — container cannot write to disk
  • --memory 256m and --cpus 0.5 — resource caps
  • Non-root user inside the container
  • Container is destroyed immediately after the command completes
  • Even if the LLM misbehaves, the IAM read-only role is the last line of defense

The agent never executes fixes automatically. It investigates, suggests, and waits for human approval before anything changes.

Development

All backend commands run from apps/backend/, or use root make targets.

# Run tests
cd apps/backend && uv run pytest      # or: make test

# Lint + format
cd apps/backend && uv run ruff check src/ tests/
cd apps/backend && uv run ruff format src/ tests/   # or: make lint / make lint-fix