惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

量子位
博客园 - 三生石上(FineUI控件)
D
DataBreaches.Net
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
GbyAI
GbyAI
P
Proofpoint News Feed
Microsoft Security Blog
Microsoft Security Blog
月光博客
月光博客
I
InfoQ
V
Visual Studio Blog
罗磊的独立博客
Engineering at Meta
Engineering at Meta
Vercel News
Vercel News
Jina AI
Jina AI
L
LangChain Blog
Apple Machine Learning Research
Apple Machine Learning Research
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
B
Blog
The Cloudflare Blog
小众软件
小众软件
雷峰网
雷峰网
V
V2EX
人人都是产品经理
人人都是产品经理
Stack Overflow Blog
Stack Overflow Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
OpenKairos: Open Implementation of the Leaked KAIROS Arch...
prabhdeep · 2026-04-28 · via DEV Community

prabhdeep

The Context

The most interesting part of the leak wasn’t model weights or APIs—it was architecture.
Specifically, the idea of a persistent daemon: a system that observes, reacts, and schedules actions without explicit user prompts. Think less “chatbot,” more “background intelligence layer.”
That concept stuck with me.

The Build Timeline

I started building on April 19 with a simple constraint:
No massive infra. No hidden magic. Just reproducible components.
The goal wasn’t to copy anything—it was to see if the pattern could be rebuilt from scratch.
Nine days later, I had a working prototype.

The Stack (What actually matters)

This is where things get real for dev.to:
Python + Asyncio → event loop for continuous execution
Watchdog → filesystem + environment triggers
Ollama → local model inference (no external API dependency)
Task Scheduler Layer → priority + interrupt handling
3-Layer Memory System:
Short-term (context window)
Mid-term (session logs)
Long-term (vector store)
Everything runs as a daemon process—not a request/response server.

Core Design Idea

Instead of:

User → Prompt → Response

It works like:

System Loop → Observe → Decide → Act → Store → Repeat

That shift changes everything:

  • 1. latency expectations
  • 2. memory handling
  • 3. failure modes
  • 4. resource management

The Weird Part: “AutoDream”

The hardest problem wasn’t inference—it was memory.

I ended up building something I call AutoDream:

  • Runs periodically (or during idle windows)
  • Compresses recent interactions
  • Promotes useful patterns into long-term memory
  • Drops noise

The constraint:

Must complete within ~15 seconds or get killed by the scheduler.

This forced aggressive tradeoffs:

  • summarization vs fidelity
  • frequency vs cost
  • stability vs adaptability
  • Still not fully solved.

What Broke (and why it matters)

  • Long-running loops drift without strong constraints
  • Memory systems become garbage collectors if unmanaged
  • Background agents need interruptibility, not just intelligence

This isn’t just “LLM engineering”—it’s closer to OS design.

Call to Action
The full implementation is open source:

If you’re exploring persistent agents, daemonized LLMs, or memory systems—I’d be interested in what approaches you’re taking.