惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
爱范儿
爱范儿
D
Darknet – Hacking Tools, Hacker News & Cyber Security
C
CERT Recently Published Vulnerability Notes
P
Palo Alto Networks Blog
博客园 - 司徒正美
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
美团技术团队
罗磊的独立博客
阮一峰的网络日志
阮一峰的网络日志
The Register - Security
The Register - Security
D
DataBreaches.Net
A
Arctic Wolf
C
Cyber Attacks, Cyber Crime and Cyber Security
P
Privacy & Cybersecurity Law Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
B
Blog
V
Vulnerabilities – Threatpost
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
G
Google Developers Blog
aimingoo的专栏
aimingoo的专栏
T
Tor Project blog
GbyAI
GbyAI
Recent Announcements
Recent Announcements
T
The Blog of Author Tim Ferriss
Simon Willison's Weblog
Simon Willison's Weblog
Cyberwarzone
Cyberwarzone
C
Cisco Blogs
G
GRAHAM CLULEY
宝玉的分享
宝玉的分享
T
Threat Research - Cisco Blogs
C
Check Point Blog
W
WeLiveSecurity
F
Fortinet All Blogs
P
Proofpoint News Feed
Security Archives - TechRepublic
Security Archives - TechRepublic
月光博客
月光博客
Project Zero
Project Zero
Know Your Adversary
Know Your Adversary
V
Visual Studio Blog
H
Help Net Security
H
Hacker News: Front Page
Webroot Blog
Webroot Blog
S
Securelist
酷 壳 – CoolShell
酷 壳 – CoolShell
O
OpenAI News
The Cloudflare Blog
Attack and Defense Labs
Attack and Defense Labs

Show HN

GitHub - villagesql/villagesql-skills: Agent skills for VillageSQL - gemini-cli-extension; claude-code-plugin GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
GitHub - Tako-Research/TakoQA: A swarm of browser agents that breaks your web app before your users do.
sakuraiben · 2026-06-25 · via Show HN

A swarm of browser agents that breaks your web app before your users do.

Plain-language missions in, real bugs out.

takoqa drives a real Chromium browser against your running app, perceives each page the way a person does, and decides its next action with an LLM — clicking, typing, uploading, and exploring toward a goal you describe in plain language. Along the way it watches for broken behavior and reports what it finds, with screenshots, a video, and a step-by-step replay.

The engine knows nothing about any specific product. Everything app-specific lives in a single profile file, so pointing takoqa at a new app is just writing a new profiles/*.yaml.

How it works

Each step runs a four-beat loop:

  1. Observe — tag every visible interactive element with a ref number, plus a screenshot and the page text.
  2. Decide — the LLM is given that list (and the screenshot) and picks one human action, addressing elements by ref — never by CSS selector.
  3. Act — Playwright performs the action; the target is highlighted on-page first so the recording shows exactly what was clicked.
  4. Check — captured console errors, uncaught exceptions, and HTTP responses run through the oracles. A finding is raised when something looks broken.

At the end of each mission an LLM judge decides whether the user's goal was actually met and flags UX/quality issues even when the flow technically worked.

What it catches

  • Functional bugs — JS exceptions, 5xx responses, console errors, crash text.
  • Exploratory/edge cases — give it a goal and no script; it wanders.
  • UX/quality — the judge flags confusing or degraded flows.
  • Regressions — every run is saved (JSON, screenshots, video, trace) for run-to-run comparison.

Self-improvement

takoqa gets smarter the more it runs, without anyone editing the profile:

  • Known-bugs baseline (--baseline) classifies each finding new / known / muted so a repeat run reports only what changed.
  • Learned store — during --loop the harness distills durable app facts from what it saw (routes that turned out to be gated, controls that never did anything, what each page actually offers, missions already tried) into a per-profile JSON sidecar. The next run merges the confident subset into the app map it hands the acting agent, so it stops re-discovering the same things. Facts need ≥2 sightings to count and decay if not re-seen, so a one-off flake never ossifies. Learnings inform the agent only — never the judge.
  • --mute "<kind|title>" --as "<reason>" marks a finding a known non-bug. It is dropped from the report and the CI gate, and the reason is fed to the LLM judge as a "do not flag" exclusion next run — so a triaged non-bug stops coming back. (The reason is the only feedback signal allowed to reach the judge.)

The baseline (baseline/), recipes (recipes/), and learned store (learned/) are plain, human-inspectable JSON — delete an entry to forget it.

Quick start

npm install
npx playwright install chromium

# Copy the template and point it at your app:
cp profiles/example.yaml profiles/myapp.local.yaml   # *.local.yaml is gitignored

ANTHROPIC_API_KEY=sk-... npx tsx src/run.ts \
  --profile profiles/myapp.local.yaml --tag smoke

Outputs land in runs/<profile>-<timestamp>/:

  • index.html — self-contained replay: step timeline, screenshots, embedded video, and findings. Open it in any browser.
  • findings.txt / run.json — human- and machine-readable results.
  • missions/<id>/video.webm and trace.zip — per-mission recordings (npx playwright show-trace <path> for the time-travel viewer).

Useful flags

Flag Effect
--headed Watch the browser live
--tag <t> Run only missions with this tag
--base-url <url> Override the profile's baseUrl (local/staging/prod)
--no-record Skip video/trace for fast headless runs
--mock Run the loop with a scripted client (no API key)

Writing a profile

A profile declares intent and failure conditions, not clicks. See profiles/example.yaml for a documented template: baseUrl, an auth strategy, personas (who's driving), invariants (what counts as a bug), and missions (goals + success criteria the LLM judge uses).

Testing takoqa itself

takoqa is verified against a deliberately-buggy fixture app — no real app or API key needed:

npm test          # oracle unit tests + engine integration tests
npm run test:unit # fast, browserless oracle tests only
npm run selfeval  # absolute gate: does it catch the planted bugs? (see below)
npm run eval      # comparative gate: did it regress vs the previous state?
npm run metaeval  # meta gate: is every detector exercised AND protected?

These are three gates on three different questions. selfeval asks do we catch the planted bugs (absolute recall/precision). eval asks did we get worse than last time (comparative, per-case). metaeval asks would we even notice if a detector broke (coverage + mutation) — the question the other two can't answer.

Self-eval

npm run selfeval is the regression gate on takoqa's own coverage. It runs the real engine over the planted-bug fixture in two passes (functional + security), scores the findings against a co-located ground-truth manifest (test/fixture-manifest.ts), and asserts full recall over every must-catch case with zero false positives on the clean routes. A refactor that stops an oracle from firing — or starts crying wolf on a clean page — fails this gate and names the exact case. Adding a planted route to the fixture forces a matching manifest entry, so coverage can't silently rot.

Comparative eval

npm run eval goes one step further than the absolute self-eval gate: it scores the harness against the planted-bug fixture and diffs that score against the previous committed record (eval/eval_ledger.jsonl) — reporting the delta, not just the value. A per-case regression (a bug caught before, missed now) fails the gate even when aggregate recall is unchanged, which the absolute recall gate can't see. Each record stamps git provenance + a byte-hash of the fixture, so a stale baseline over a different fixture simply stops being comparable. npm run eval -- --record appends a new record, so every accepted improvement becomes the prior state the next change is measured against.

Meta-eval (test the tests)

npm run metaeval gates the gate itself. The self-eval proves takoqa catches the planted bugs, but it can't tell you whether every detector takoqa ships is actually exercised — a detector with no fixture case, or one always co-caught by another kind, could quietly stop firing and both gates above would stay green. The meta-eval answers two questions:

  • Coverage — is every deterministic detector kind exercised by a fixture case? KIND_CLASS (in src/metaeval.ts) classifies every FindingKind as a detector or an LLM/agent judgment; because it's an exhaustive map, adding a new kind is a compile error until it's classified, so a detector can't ship without a coverage decision.
  • Mutation / ablation — would the self-eval actually fail if a detector broke? For each detector it drops that kind's findings from a passing report and re-scores: if a previously-caught case now misses, the detector is protected; if the case stays caught (some other kind covers it), it's shadowed — covered on paper but the eval is blind to it breaking.

Like the comparative eval, it records to eval/eval_ledger.jsonl (as the harness_meta task) and diffs against the previous state, so a detector going protected → unprotected fails the gate. npm run metaeval -- --record appends a new baseline.

Pluggable route discovery

Route discovery is pluggable, so takoqa points at any app — not just Next.js. --explore/--matrix accept --app-dir <path> (read a Next.js app-router tree), --routes a,b,c (an explicit, app-agnostic list), or --sitemap <url> (extract same-origin paths from a sitemap.xml). A profile can pin the same via explore.source (or keep the explore.appDir shorthand).

Docker

docker build -t takoqa .
docker run --rm --network host -e ANTHROPIC_API_KEY=sk-... \
  -v "$PWD/runs:/app/runs" takoqa --profile profiles/example.yaml --tag smoke

See docker-compose.example.yml for wiring takoqa into an app's compose stack.

License

MIT — see LICENSE.