惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
Security Latest
Security Latest
博客园_首页
宝玉的分享
宝玉的分享
人人都是产品经理
人人都是产品经理
罗磊的独立博客
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Jina AI
Jina AI
爱范儿
爱范儿
小众软件
小众软件
IT之家
IT之家
Hugging Face - Blog
Hugging Face - Blog
博客园 - 三生石上(FineUI控件)
博客园 - 聂微东
博客园 - Franky
S
SegmentFault 最新的问题
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
大猫的无限游戏
大猫的无限游戏
Apple Machine Learning Research
Apple Machine Learning Research
量子位
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
月光博客
月光博客
NISL@THU
NISL@THU
博客园 - 司徒正美
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
AWS News Blog
AWS News Blog
有赞技术团队
有赞技术团队
V
Visual Studio Blog
雷峰网
雷峰网
C
Cybersecurity and Infrastructure Security Agency CISA
美团技术团队
The Cloudflare Blog
P
Privacy & Cybersecurity Law Blog
Latest news
Latest news
S
Securelist
C
CERT Recently Published Vulnerability Notes
C
CXSECURITY Database RSS Feed - CXSecurity.com
P
Palo Alto Networks Blog
Last Week in AI
Last Week in AI
V
V2EX
Know Your Adversary
Know Your Adversary
酷 壳 – CoolShell
酷 壳 – CoolShell
T
Threat Research - Cisco Blogs
T
Tailwind CSS Blog
J
Java Code Geeks
I
Intezer
Recent Commits to openclaw:main
Recent Commits to openclaw:main
博客园 - 【当耐特】
Schneier on Security
Schneier on Security

Show HN

GitHub - villagesql/villagesql-skills: Agent skills for VillageSQL - gemini-cli-extension; claude-code-plugin GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
Orion 2: Frontier Visual Agents with Code Execution
Dinesh Reddy · 2026-06-18 · via Show HN

We are excited to announce Orion 2: our most capable visual agent, now with code execution. Since our Orion 1 launch in November last year, our customers have run hundreds of thousands of requests spanning millions of tool-calls in production every month. Serving at this scale taught us that the bottleneck in production visual agents isn't just accurate perception – it's reliable orchestration.

Orion 2 generates and executes computer-vision code on the fly, expanding capabilities well beyond Orion 1 while being significantly faster, cheaper, and more reliable. Try it now at chat.vlm.run, every example in this post is a live chat thread you can inspect and re-run.

Visual Intelligence, now with Code-Mode

Complex computer vision tasks are inherently compositional: detect, crop, draw, measure, reason. Most vision agents are restricted to a predefined tool surface and a sequential tool-calling loop. With Orion 2, we combine deterministic tool-calling with dynamic code generation – the best of both worlds.

Code as the orchestration substrate (aka code-mode) has several properties that matter in production:

  • Reusable: generated programs are artifacts; scripts can be saved, versioned, and re-run on new inputs without re-invoking the LLM agent.
  • Inspectable: the executed code is something you can read, debug, diff, and verify; critical for our customers in regulated industries.
  • Composable: loops, conditionals, fan-out, and parallel tool calls live in a single program instead of N model round-trips.
  • Deterministic: computation (counting, measuring, normalizing, geometry) runs in code, not in the model's head; eliminating an entire class of numeric hallucinations.

Architecture: Code as the Agent Harness

Orion 2 is a visual agent harness: a planner and a code runtime wrapped around a vision-language model. At its core is a custom visual DSL and runtime we built to expose every tool we introduced with Orion 1 – detection, OCR, segmentation, cropping, image generation, and more – as native primitives the model can compose in code. Most notably, the entire Orion 1 tool surface is now fully programmable and deterministic.

Architecture diagram of Orion 2. Four input types on the left (text, image, video, and document) flow into the Orion 2 agent. Inside, the model issues tool calls to three subsystems: Visual Tools & Runtime, Code Gen & Execution, and Visual Skills. Outputs on the right span five capability classes: Describe, Tag, Detect, Generate, and Act.

Orion 2 accepts text, images, video, and documents, compiles each request into an executable program, and dispatches visual tools, code execution, and learned skills from a single harness.

Here is how a new request is handled in Orion 2:

  1. Prompt → Spec: An ambiguous request is compiled into an exact, executable program – written in our visual DSL, which reads like idiomatic Python.
  2. Execution: The program runs in a sandboxed runtime with async-native parallelism: independent operations dispatch concurrently via asyncio, with no per-step model round-trips. The runtime ships with OpenCV for classical computer vision and the vlmrun package, which exposes the full Orion tool surface – detection, OCR, segmentation, image generation – as importable primitives.
  3. Self-correction: Execution results return to the harness, which repairs and re-executes until the program runs to completion. Each turn also yields a (program, trace, outcome) record – fully verifiable supervision for our continual-learning flywheel.

That record is also the groundwork for what comes next: close the loop with a visual judge, a harness that can score its own outputs, prefer better programs over worse ones, and improve recursively with every workload it serves. More on that soon.

Orion 1 vs. Orion 2

The difference is easiest to see on a concrete task: the virtual try-on from the examples below, which composes detection, cropping, and image generation across two input images.

Orion 1's orchestration is lazy and interactive: one tool call, one LLM round-trip, repeat. Five visual operations means five round-trips, and every intermediate result (bounding boxes included) passes through the model's context before the next step:

Orion 1: Sequential Tool-Call Illustration

# Tools are called sequentially, with LLM reasoning at each step
boxes      = tool_call("detect", image, target="person")                  # call 1
person     = tool_call("crop", image, xywh=[0.22, 0.35, 0.04, 0.15])      # call 2
garment    = tool_call("detect", dress_img, target="garment")             # call 3
garment    = tool_call("crop", dress_img, xywh=[0.33, 0.41, 0.05, 0.13])  # call 4
result     = tool_call("generate", person, garment)                       # call 5
# ...the model parses every intermediate output before deciding the next tool.

Notice the coordinates in calls 2 and 4: the LLM reads them out of the detection results and calls them into the next tool call. Every hop like that costs a round-trip of latency and is a chance to transcribe something wrong.

Orion 2: Efficient Code-Mode Execution

Orion 2's orchestration is code. The model writes one program up front; intermediate results flow through variables instead of the model's context, independent operations run in parallel, and the whole workflow executes end-to-end before returning:

# Tools can be called efficiently via native async python ops
import asyncio

async def process(ctx, person_image, dress_img):
    vlmrun = ctx.import_lib("vlmrun")
    
    def crop(img, d):
        bx, by, bw, bh = d["xywh"]; W, H = img.width, img.height
        return img.crop(int(by * H), int((by + bh) * H), int(bx * W), int((bx + bw) * W))
    
    # Detect person and garment in parallel, take top crop of each
    p_det, g_det = await asyncio.gather(
        vlmrun.image.detect(person_image, "person"),
        vlmrun.image.detect(dress_img, "garment"),
    )
    person_crop = crop(person_image, p_det["detections"][0])
    garment_crop = crop(dress_img, g_det["detections"][0])

    # Composite the try-on and return the composite image
    (composite,) = await vlmrun.image.generate(
        "virtual try-on", images=[person_crop, garment_crop]
    )
    return {"composite": composite}

When orchestration is code, verifiable correctness and determinism is a property of the program, even if a statistical model generated it.

Under the Hood: Model-Agnostic, Purpose-Built Runtime

Orion 2 is a visual agent harness, not a model. Any multimodal model with strong code generation can drive the planner → execution → self-correction loop: open-weight VLMs like Gemma4-26B-A4B and Qwen3.6-35B-A3B, or frontier models like Gemini 3.5 Flash. Same harness, same runtime, different engine.

The benchmark below runs all three backbones through the identical harness. Each brings different strengths – Gemma4 leads on localization, Gemini 3.5 Flash on segmentation and video – and Orion 2 inherits every improvement in vision-grounded code generation, open or closed, for free.

The default, vlmrun-orion-2:auto, routes each request to the best backbone for the job, so you get the frontier of all three without choosing. Open-weight backbones run on our purpose-built inference runtime: favorable GPU economics at volume, deployable in isolated cloud environments. Pin a fixed backbone through the gateway when compliance or reproducibility demands it.

Examples

See Orion 2 in action in the following chat thread examples and inspect the artifacts.

Virtual try-on

Given an image of a dress and an image of a model, Orion 2 creates a realistic virtual try-on in a single turn. It detects the person and the garment in parallel, crops each subject in-process, generates the composite with one image-generation call, and asks a vision-LLM whether the fit looks natural – composing four distinct visual operations as a single program, with the intermediate crops, composite, and verdict all available as inspectable artifacts. See chat.

One pipeline, multiple visual operations, parallelized — the cleanest illustration of code-mode speed.

Robotics & Physical AI

Orion 2 extracts a representative frame from a robotics video, segments every object in the scene in parallel, and hands off the per-object masks to generate an interactive 3D reconstruction. The frame sampling, fan-out segmentation, and 3D-reconstruction handoff all live in a single program, so the intermediate frame, masks, and reconstruction are produced together in one turn and available as inspectable artifacts. See chat.

Video transcription, key-point tracking, and motion analytics in one end-to-end program.

Multi-document workflow

Given a multi-page healthcare PDF, Orion 2 splits a batch of healthcare documents into per-page images and classifies each one – claim form, instructions, insurance ID card, medical history, physician referral form etc. It then focuses on the referral form, helping the user localize the patient and physician names, and returns the full structured extraction grounded to the source page – every page split and bounding box available as inspectable artifacts. See chat.

Parse, align, and summarize across multiple documents in one reusable pipeline.

Manipulating Images

Orion 2 manipulates images on-the-fly with native OpenCV – Gaussian blur, Canny edge detection, color pop, and other cv2 compositions are written directly into the script and run in parallel where independent operations can be composed, instead of being constrained to a fixed, pre-defined tool surface. See chat.

Detect → tile → count → annotate, with the counting done deterministically in code.

Benchmarks

When we launched Orion 1, we showcased a benchmark dataset of 30+ multi-modal, multi-turn tasks involving complex reasoning and agentic actions on images, documents, and videos.

With Orion 2's code-execution capabilities, we expand the benchmark to 250+ multi-turn cases across image, document, audio, video, and multi-file inputs, covering perception, counting, OCR, grounding, and cross-turn reasoning. The hard tier introduces generative tool-use actions – cropping, masking, blurring, redaction, segmentation, keyframe extraction – that exercise Orion 2's code-mode through the VLM Run Chat Completions API.

Although state-of-the-art multimodal models perform well on many vision tasks, they do not cover the full spectrum of visual capabilities and are largely constrained to text-based outputs. Orion extends beyond traditional MLLMs by providing both pixel-level understanding and spatial reasoning capabilities, enabling richer interactions with visual content and more precise grounding in the visual world.

Get started

Try Orion 2 now at chat.vlm.run. Bring your own favorite images, documents, or videos and inspect the programs it writes. When you're ready to build, the same agent is available through the VLM Run Chat Completions API, and our team can help you evaluate it on your production workloads.