惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

B
Blog
V
Vulnerabilities – Threatpost
P
Proofpoint News Feed
Google DeepMind News
Google DeepMind News
Y
Y Combinator Blog
V
Visual Studio Blog
阮一峰的网络日志
阮一峰的网络日志
腾讯CDC
月光博客
月光博客
T
Troy Hunt's Blog
博客园_首页
H
Hackread – Cybersecurity News, Data Breaches, AI and More
N
Netflix TechBlog - Medium
Microsoft Security Blog
Microsoft Security Blog
Recorded Future
Recorded Future
Blog — PlanetScale
Blog — PlanetScale
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Scott Helme
Scott Helme
T
Threat Research - Cisco Blogs
P
Palo Alto Networks Blog
T
The Exploit Database - CXSecurity.com
Simon Willison's Weblog
Simon Willison's Weblog
Know Your Adversary
Know Your Adversary
SecWiki News
SecWiki News
Security Archives - TechRepublic
Security Archives - TechRepublic
T
Threatpost
Forbes - Security
Forbes - Security
S
Schneier on Security
P
Proofpoint News Feed
T
Tor Project blog
Cyberwarzone
Cyberwarzone
The Hacker News
The Hacker News
Cloudbric
Cloudbric
S
Security @ Cisco Blogs
Webroot Blog
Webroot Blog
Attack and Defense Labs
Attack and Defense Labs
Hacker News: Ask HN
Hacker News: Ask HN
Google DeepMind News
Google DeepMind News
Hacker News - Newest:
Hacker News - Newest: "LLM"
C
CERT Recently Published Vulnerability Notes
The Last Watchdog
The Last Watchdog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
S
SegmentFault 最新的问题
V
V2EX
量子位
B
Blog RSS Feed
宝玉的分享
宝玉的分享
T
The Blog of Author Tim Ferriss
罗磊的独立博客
J
Java Code Geeks

Show HN

GitHub - steveking-gh/firmion: Firmion is DSL and engine for firmware image generation. GitHub - villagesql/villagesql-skills: Agent skills for VillageSQL - gemini-cli-extension; claude-code-plugin GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
Brontosaurus: A Voice-Driven Generative AI Canvas | Thomas Hughes
thomasdhughe · 2026-06-02 · via Show HN

Click for audio.

Brontosaurus is a web-based generative canvas, where you speak aloud what you want to see, and Bronto builds a widget of it in under a second. The underlying agents run on OpenAI’s gpt-oss-120b, served by Cerebras at a blistering 3,000 tokens/second, which makes the whole thing feel like magic :)

Brontosaurus is very much in development - if you’ve got ideas or things you’d want to see, send them my way, I’d love to hear them! (hello@thomasdhughes.com)

Contents:

  • The Inspiration
  • The Technical
  • The Future
  • Acknowledgments

The Inspiration

Two phenomenal blog posts inspired this project.

The first was Thinking Machines’ release of an Interaction Model.

On the technical side, I liked the architecture of a multi-modal (voice+vision+text) model hooked into a more powerful quasi-background agent which quietly executes requests without interrupting the flow of conversation. There are certainly other labs doing this today, but Thinking Machines takes a unique approach - Sean Goedecke does a great job of breaking down that novelty in his own blog post.

On the philosophical side, Thinking Machines argues that current discussions around AI agents mistakenly center agentic autonomy (the ability for an agent to receive a task, then work for hours uninterrupted) as opposed to human-AI collaboration (the ability for an agent to work on a task in tandem with a person). They assert that this is a mistake - that designing models this way leads to “humans increasingly get[ing] pushed out not because the work doesn’t need them, but because the interface has no room for them.” You put in your prompt and get out of the way. This new family of Interaction Models, on the flip side, operates more in the way you’d work with a teammate - you can talk, type, point at things, interrupt with new ideas. You can collaborate.

Yes to this!

Both the technical and philosophical pieces of this approach spoke to me, and they together made me want to build something of the same ethos. In Bronto, that manifests as prioritizing creation at the speed of thought above all else. It is an argument that more than capability, more than intelligence, more than long-running task autonomy, the ability to speak something into existence in under a second makes you feel like anything is possible.

The second blog post was from Ink & Switch, an independent research lab with a focus on malleable software, a concept I explored in Modifying Websites with LLM-Generated Javascript Bookmarks. The post is called “chitter chatter”, and it outlines a vision for a generative canvas. It reads like somebody’s diary, and makes the proposed software sound friendly and warm, which I always admire when people can do - code is too often cold and numbers. I will not apologize for that phrasing because I am prideful, but I do not stand by it let us not mention it going forward.

I read this piece (along with Thinking Machines’), loved it, and wanted to build something like it. Brontosaurus was born.

The Technical

Under the hood, there’s some multi-agent orchestration going on.

There are two agent types at play: Conductor and Builder. Both run on OpenAI’s gpt-oss-120b, served by Cerebras at 3,000 tokens/second. To put this in perspective, ChatGPT in the browser responds at ~50 tokens/second.

When you tap the space bar, the web app starts listening, and when you tap again, it does speech-to-text with Chrome’s built-in Web Speech API.

The text of what you said gets passed to the Conductor agent, along with a JSON array describing the widgets currently on the canvas, each of which have

  • an id (unique identifier for tool calls),
  • a title (what you see at the top of each widget),
  • a description (internal-facing explanation of what the widget contains), and
  • a rect (the widget’s size and position on the canvas).

The Conductor agent then makes tool calls. It can

  • arrange - move or resize a widget by id, without changing the contents (this is done by updating the rect value for that widget),
  • delete - remove a widget by id,
  • clear - remove all widgets at once,
  • create and edit.

The first three tool calls are handled deterministically. The final two - create and edit - send instructions to a Builder agent.

When createing, the Builder agent receives just the requested widget’s description. When editing, it receives the full HTML of the current widget along with the change instructions.

The Builder agent then returns a complete, self-contained HTML document, which is subsequently cleaned and rendered in an iframe.

Additional design choices which add to the magic:

  1. The Conductor agent can make multiple tool calls with a single instruction - this makes it possible to say “delete the piano, make a calculator, put it where the piano was” and it all happens at once.
  2. arrange calls don’t need to wait for the Builder agent to finish building - as soon as the create is run, an id exists, so widgets which are still populating can be moved around.
  3. Builder agents run in parallel so multiple widgets can be made at once.

The Future

There’s lots of room for improvement here.

For one, gpt-oss-120b is 9 months old and just 120B parameters. This means it’s dirt cheap - I used Brontosaurus nonstop for over an hour and spent less than a dollar - and that the ceiling for quality of output is way higher. If we used a model like GLM 4.7 from Z.ai (also served by Cerebras), it’d be 4x the cost and 1/3rd the speed, but 3x the parameters so could likely build far more complex widgets. The question is if the speed tradeoff would be worth it.

On this note, I initially added live search capabilities via Exa AI so that Bronto could pull things like weather and live stock price, but in my evals it added a delay of ~0.9s, which stings once you’ve gotten used to the sub-second generation speed.

Finally, the major one - a virtual file system! At the moment, all the widgets Bronto makes are ephemeral, single-use HTML files in iframes, but a VFS would allow for surfacing pre-existing documents and previously-built widgets to iterate on, as well as make it possible for Bronto to selectively pull the contents of widgets into its context window, so commands like “I checked off what I already have from the ingredients list, please remove those” would work.

The above are all technical changes. But I know there is a lot of interesting things that can be done just with the current architecture. The 8 row step sequencer (beat maker) at the end of the demo initially came from me saying “I want to make some music” and Bronto produced three widgets, that being one of them. It totally blew my mind. Which is why I put at the top: if you’ve got ideas, please send them my way! (hello@thomasdhughes.com) I’ll try them out and send back a video I promise.

Goodbye!

-Thomas


Thank you Thinking Machines and Ink & Switch for inspiring me with your work and writing.

Thank you Tristan and Steven for pressure testing early versions of Brontosaurus with requests I hope it never fulfills.