惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 三生石上(FineUI控件)
WordPress大学
WordPress大学
阮一峰的网络日志
阮一峰的网络日志
大猫的无限游戏
大猫的无限游戏
T
Tailwind CSS Blog
S
SegmentFault 最新的问题
The Hacker News
The Hacker News
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
小众软件
小众软件
Google DeepMind News
Google DeepMind News
腾讯CDC
博客园 - 司徒正美
Cisco Talos Blog
Cisco Talos Blog
Apple Machine Learning Research
Apple Machine Learning Research
The Cloudflare Blog
博客园 - 聂微东
博客园 - 【当耐特】
Project Zero
Project Zero
有赞技术团队
有赞技术团队
量子位
P
Privacy International News Feed
博客园_首页
酷 壳 – CoolShell
酷 壳 – CoolShell
J
Java Code Geeks
IT之家
IT之家
SecWiki News
SecWiki News
H
Hacker News: Front Page
PCI Perspectives
PCI Perspectives
L
Lohrmann on Cybersecurity
宝玉的分享
宝玉的分享
Cloudbric
Cloudbric
雷峰网
雷峰网
月光博客
月光博客
Cyberwarzone
Cyberwarzone
S
Securelist
Hugging Face - Blog
Hugging Face - Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
博客园 - Franky
T
Threat Research - Cisco Blogs
罗磊的独立博客
Forbes - Security
Forbes - Security
NISL@THU
NISL@THU
N
News and Events Feed by Topic
T
Troy Hunt's Blog
Jina AI
Jina AI
Hacker News - Newest:
Hacker News - Newest: "LLM"
C
Cyber Attacks, Cyber Crime and Cyber Security
The Last Watchdog
The Last Watchdog
V2EX - 技术
V2EX - 技术

Show HN

GitHub - villagesql/villagesql-skills: Agent skills for VillageSQL - gemini-cli-extension; claude-code-plugin GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
GitHub - Approxima-AI/Approxima-OSS: Open-source, agentic web testing platform.
Intragalacti · 2026-06-14 · via Show HN

Open-source, agentic web testing platform.

CI Status License TypeScript

Journeys dashboard

Overview

Approxima is an open-source, agent-based platform for end-to-end testing web applications. Write test journeys in English and an LLM-driven browser agent runs them against your live app. No selectors to maintain, no scripts to babysit.

Some cool features included in the platform:

  1. Goal Mode: Agent generates the steps needed to accomplish a goal (goal also specified in English.)
  2. Self-healing: With every pass the agent attempts to refine the steps it's working with so that the next pass takes fewer tokens/time.
  3. Streaming + captions: See the agent's thought process as it navigates your website. Helpful for discovering potential bad UX.
  4. Skills: Reusable journeys that can be put together like building blocks. For example if you have 3 steps needed to login and setup the workspace in your app, you can turn it into a skill that all your journeys use.
  5. Agent Fine-tuning: A/B test multiple versions of the web agent by tweaking the prompt it's given.

Why

Good testing will be the last part of software development to be fully automated. Having good verifiable end-to-end tests allow you to ship faster and worry less about the code you ship especially since everyone is shipping more than ever.

Today you get two options:

  1. Maintain a scripted E2E suite yourself. Tests point at the DOM through hardcoded selectors. The product ships daily, selectors drift, tests fail when nothing is broken. Coverage can't keep up with AI-assisted development speed.
  2. Hand your testing to a hosted AI QA platform. The flakiness problem gets better, but now your tests, their history, and your release confidence live on someone else's servers, completely at their mercy.

We think both halves are fixable at once. Tests written as intent ("add an item to the cart, verify the total updates") instead of selectors don't drift when the DOM changes. Plus, Approxima is open source and MIT, so you can self-host your tests and keep them running yourself forever.

Features

Journeys

A journey is an ordered list of plain-English steps ("Click Sign in", "Verify the dashboard shows 3 projects"). The agent executes them one at a time in a real browser, taking a screenshot after every action and visually verifying each step before moving on. Group journeys into suites and run them on a cron schedule, or one-off from the dashboard.

Editing a journey's steps

Live screencast

Every run streams a live screencast of the browser into the dashboard, with the agent narrating what it's doing as closed captions over the video: its reasoning ("the cart icon shows 0, looking for the add button"), the action it's taking, and each check as it passes or fails. When something breaks, you see exactly what the agent saw and why it made the call it made, without digging through logs.

Goal mode

Don't know (or don't care) what the exact steps are? Give a journey just a goal like "Sign up, create a project, and invite a teammate" and run it in explore mode. The agent explores your app until it accomplishes the goal, then writes the steps it discovered back into the journey and immediately triggers a validation run to prove they're reproducible. From then on it runs as a normal deterministic journey.

Skills

A skill is a reusable step sequence (e.g. "Login": enter email, enter password, click submit) that journeys reference as a single step. Skills are expanded inline at run time, so updating a skill updates every journey that uses it. Goal-mode runs are skill-aware too: the agent is handed your skill library, and when its discovered steps match an existing skill, they're collapsed back into a single skill reference instead of duplicated steps.

Creating a reusable skill

Self-healing

When your app changes and a step's wording no longer matches reality, the agent doesn't just fail — it works out what the step should be and reports refined steps alongside the run results. These show up in the dashboard as inline suggestions on the journey editor; accept them with one click (or dismiss them). A vague step can expand into several precise ones. Suggestions are LLM-reviewed before being surfaced so trivial rewordings and selector-ish noise get filtered out.

Variables & secrets

Steps can reference variables with $NAME syntax: "Log in as $TEST_EMAIL with $TEST_PASSWORD". Values are stored per app and resolved at dispatch time.

Mark a variable as secret and it gets encrypted at rest (AES-256-GCM), masked in the dashboard (••••••••), and scrubbed from stored run logs (the persisted run record shows •••, never the value). Step suggestions are generated from the unresolved $NAME labels rather than the substituted values, so self-healing doesn't bake secrets into your journeys. For the agent to log into your app it has to type the actual value, and a cloud LLM is what decides what to type, so the secret necessarily leaves your machine. It appears in the prompt the agent receives, in the screenshots it reviews to verify each step, and therefore at your configured LLM provider. Redaction only controls where the value comes to rest (database, dashboard, suggestions); it cannot un-send the value to the model. Concretely:

Secrets used in Approxima should be for test workspaces that don't matter to you and which you wouldn't mind exposing.

Shadow runs: A/B test the agent itself

The agents are versioned: each one (journey, explore) lives in a version folder under web-runner/src/agent/, and every run records exactly which versions it used. Working on a v2 prompt? Register it in shared/agent-versions.ts as the beta version and flip AUTO_SHADOW_ENABLED. Every real run then spawns a paired shadow run on the beta agent against the same journey, and the admin panel compares the two populations with proper paired statistics.

Architecture

Directory What it is
frontend/ Next.js dashboard: create apps, journeys, suites; watch live runs via screencast
api/ Hono API on Cloudflare Workers: apps, journeys, run queue, scheduling
web-runner/ Node service that runs the browser agent (Playwright + LLM loop)
shared/ Shared types/config (agent versions, run statuses)

The API queues runs in Postgres and dispatches them to the web-runner, which launches a local Chromium, drives it step-by-step with the configured LLM, and reports results back via callback.

Agents

Every run is driven by one of two versioned LLM agents, each living in a version folder under web-runner/src/agent/:

  • Journey agent (agent/journey/): the default. Executes a fixed list of plain-English steps in order, one at a time: take an action, screenshot the result, and visually verify the step passed before moving to the next. This is what runs a normal journey or suite, and what produces self-healing step suggestions when the wording drifts.
  • Explore agent (agent/explore/): powers goal mode. Given a goal instead of fixed steps, it explores the app until it accomplishes the goal, then writes back the concrete steps it discovered (skill-aware: matching sequences collapse into skill references). After discovery, the journey runs deterministically on the journey agent from then on.

Both agents share the same browser tools (click, type, screenshot, …) and the same LLM loop; they differ in their prompt and their terminal condition (every step verified vs. goal accomplished). Each run records the agent version it used, so prompt iterations stay comparable across runs. That's what the shadow-runs comparison is built on.

How the agent drives the browser

The agent works from screenshots, but it acts on the DOM. Two mechanisms are worth understanding.

Clicking: element-first, coordinates to disambiguate

Agents are bad at clicking by raw coordinates and often miss, so the Approxima web agent targets elements first and uses pixels to disambiguate. When the agent clicks, it calls the click tool with two things read off the screenshot: a target (visible text, aria label, or CSS selector) and the approximate x, y pixel coordinates of that element. Targets are resolved against the DOM in priority tiers (web-runner/src/browser/local.ts):

  1. Role: getByRole() for button, link, menuitem, tab, checkbox, radio matching the target's accessible name.
  2. Text: getByText() (case-insensitive substring) if no role matches.
  3. CSS selector: the target is treated as a raw selector as a last resort.

The actual click is Playwright's locator.click(), a real element click with auto-waiting and actionability checks, followed by a wait for domcontentloaded. Hover resolves targets the same way.

Screenshots: viewport by default, region crop on demand

  • Full viewport (takeScreenshot): captures the current viewport (fullPage: false), default 1280×720 (or 390×844 in mobile mode via set_viewport). The runner waits up to 3s for networkidle first, then reports the coordinate system back to the agent ((0,0) to (width, height)) so its click coordinates line up with what it sees.
  • Region crop (screenshot_regiontakeRegionScreenshot): a 400×400 crop centered on a given x, y, for inspecting small details without re-reading the whole page. clampRegion() (web-runner/src/browser/region.ts) keeps the crop inside the viewport: the center is pushed inward so the box never runs off an edge, and if the viewport is smaller than 400px the crop shrinks to fit.

Quick start

Prerequisites: Node 20+, a Postgres database, an S3-compatible bucket for screenshots, and at least one LLM API key (OpenAI, Anthropic, or Gemini).

1. Install

git clone <this-repo> && cd Approxima-OSS
npm install
npx playwright install chromium
npm run build

2. Create the database tables

Point Drizzle at your Postgres connection string and apply the migrations in api/drizzle/:

DATABASE_URL="postgresql://user:password@host/dbname" npm run db:migrate -w api

That's the whole schema: apps, journeys, runs, suites, variables, suggestions, daily costs. If you later change api/src/db/schema.ts, regenerate with npm run db:generate -w api and re-run migrate.

3. Configure the API

cp api/.dev.vars.example api/.dev.vars

Fill in:

Var What
DATABASE_URL same Postgres connection string as above
WEB_RUNNER_URL http://localhost:3002 for local dev
WEB_RUNNER_API_KEY any random string; must match the runner's API_KEY
ENCRYPTION_KEY optional, openssl rand -hex 32; needed for secret variables

4. Configure the web runner

cp web-runner/.env.example web-runner/.env

Fill in:

Var What
API_KEY same value as the API's WEB_RUNNER_API_KEY
CALLBACK_URL http://localhost:8787/api/internal/run-callback for local dev
LLM_PROVIDER + key openai/anthropic/gemini and the matching *_API_KEY; any extra keys you set become automatic fallbacks
S3_BUCKET + AWS creds where step screenshots are uploaded

5. Run it

npm run dev:api          # API on :8787
npm run dev:web-runner   # runner on :3002
npm run dev:frontend     # dashboard on :3000

Open http://localhost:3000 (no login needed), create an app pointing at the URL you want to test, and write your first journey (or just give it a goal and let the agent figure out the steps).

Logging into the app under test

Journeys authenticate the same way a user would: add steps like "Enter $TEST_EMAIL in the email field, enter $TEST_PASSWORD, click Sign in." Store credentials as app Variables (Settings → Variables & Secrets); secrets are encrypted at rest, masked in the UI, and redacted from run logs.

Development

npm run typecheck         # all workspaces
npm run test:api          # API unit tests (vitest)
npm run test:web-runner   # web-runner unit tests (vitest)
npm run test:frontend     # frontend unit tests (vitest)
npm run build             # build frontend into api/static for single-worker deploys

Iterating on an agent: copy web-runner/src/agent/<type>/v1/ to v2/, add it to the versions map in <type>/index.ts, and register it in shared/agent-versions.ts (AVAILABLE_VERSIONS, plus BETA_VERSIONS to shadow-test it). Runs record the versions they used, so results stay comparable across iterations.

Deployment

Note: there is no user authentication. The dashboard and API are open to anyone who can reach them, so run them on localhost, a private network, or behind your own auth proxy.

  • API + dashboard: deployed together as one Cloudflare Worker (npm run build && npm run deploy). Configure api/wrangler.toml with your own URLs; secrets are set with wrangler secret put.
  • Web runner: any Node host that can run Chromium. Point the API's WEB_RUNNER_URL/WEB_RUNNER_API_KEY at it.
  • Database migrations: npm run db:generate -w api / npm run db:migrate -w api (Drizzle).

What we run it on

For reference, the stack we built it on and run it on day to day:

Piece Service
Postgres Neon
API + dashboard Cloudflare Workers
Web runner Railway
Screenshot storage AWS S3
LLM Gemini (primary), with Anthropic as fallback

None of these are required — anything Postgres-compatible, any Worker-compatible host, and any Node box with Chromium will do.