惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

SecWiki News
SecWiki News
罗磊的独立博客
U
Unit 42
I
InfoQ
B
Blog RSS Feed
Google DeepMind News
Google DeepMind News
J
Java Code Geeks
Blog — PlanetScale
Blog — PlanetScale
The GitHub Blog
The GitHub Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
B
Blog
S
SegmentFault 最新的问题
V
Visual Studio Blog
Engineering at Meta
Engineering at Meta
Microsoft Security Blog
Microsoft Security Blog
月光博客
月光博客
Vercel News
Vercel News
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
A
About on SuperTechFans
博客园 - 三生石上(FineUI控件)
博客园_首页
腾讯CDC
F
Fortinet All Blogs
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Hugging Face - Blog
Hugging Face - Blog
MongoDB | Blog
MongoDB | Blog
阮一峰的网络日志
阮一峰的网络日志
D
Docker
N
Netflix TechBlog - Medium
云风的 BLOG
云风的 BLOG
Apple Machine Learning Research
Apple Machine Learning Research
Microsoft Azure Blog
Microsoft Azure Blog
Martin Fowler
Martin Fowler
人人都是产品经理
人人都是产品经理
酷 壳 – CoolShell
酷 壳 – CoolShell
爱范儿
爱范儿
大猫的无限游戏
大猫的无限游戏
V
V2EX
Last Week in AI
Last Week in AI
博客园 - 司徒正美
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
IT之家
IT之家
L
LangChain Blog
WordPress大学
WordPress大学
Y
Y Combinator Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
M
MIT News - Artificial intelligence
The Cloudflare Blog
T
The Blog of Author Tim Ferriss
宝玉的分享
宝玉的分享

Show HN

GitHub - villagesql/villagesql-skills: Agent skills for VillageSQL - gemini-cli-extension; claude-code-plugin GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
GitHub - juancgarza/cody: A repository for cody
jcgr · 2026-06-21 · via Show HN

Neovim-first voice control for developers who want hands-free, low-latency command of their editor without leaving it.

Cody deliberately stays inside the editor: no screen overlay, no mouse pointer, no general desktop assistant. It lives inside Neovim and turns short voice or text commands into editor actions.

Examples:

:CodyDo go to line 48
:CodyDo go to file src/server.ts
:CodyDo edit this line to return early when request.user is missing

Cody should not rebuild editor primitives. It should route voice intent into the editor command surface developers already use, and install missing command providers only when that is explicitly supported by the user's setup.

Architecture

Neovim Lua plugin          Node Realtime bridge          OpenAI Realtime
----------------          --------------------          ---------------
:CodyDo / voice cmds  ->   JSONL over stdio        ->    WebSocket session
editor command adapter <-   function-call router    <-    gpt-realtime-2
buffer/cursor context  ->   prompt + tool schemas   ->    text/audio input

Rather than capturing the screen and pointing at UI elements, Cody sends editor state:

  • current file, filetype, cursor line/column
  • current line and nearby buffer lines
  • available editor commands from native Neovim, LSP, and installed plugins

Command Adapter

The important layer is not "go to line" itself. Neovim already has that. The useful layer is:

  1. detect what the editor can already do
  2. expose those capabilities to GPT Realtime as callable tools
  3. install a missing provider when the user's plugin manager supports it
  4. route the spoken command to the best existing command

Initial command providers:

  • Native Neovim: line jumps, file edits, buffers, windows, quickfix
  • LSP: rename, code actions, references, definitions
  • Pickers: Telescope, fzf-lua, Snacks picker, mini.pick
  • AI/code edit plugins: CodeCompanion, Avante, Copilot Chat, or a Cody-owned Realtime edit fallback

This means there is no separate Phase 1 for proving basic editor commands. We start at the adapter.

Setup

Requirements:

  • Neovim 0.10+
  • Node.js 20+
  • OPENAI_API_KEY for intelligent commands
  • sox for voice input: brew install sox

Install dependencies and build the local bridge:

npm install
npm run build

Install with your plugin manager. The Node bridge must be built, so use a build hook. With lazy.nvim:

{
  "juancgarza/cody",
  build = "npm install && npm run build", -- compiles the Node bridge (dist/)
  opts = {
    -- enable_shell = true,    -- on by default once setup() runs
    -- enable_commands = true, -- on by default once setup() runs
    -- tts_enabled = true, tts_voice_id = "<elevenlabs-voice-id>",
  },
  -- lazy.nvim calls require("cody").setup(opts) automatically.
}

Then export OPENAI_API_KEY (and ELEVENLABS_API_KEY for TTS) in the shell you launch Neovim from, and run :CodyStart.

Without a plugin manager, or during development:

set runtimepath^=/path/to/cody
runtime plugin/cody.lua
lua require("cody").setup()

If you pass the runtimepath before Neovim starts, the plugin/ file is sourced automatically:

nvim --cmd 'set runtimepath^=/path/to/cody'

For nvim -u NONE, plugin loading is disabled; use the explicit runtime plugin/cody.lua form above.

Optional quick-command routing:

require("cody").setup({
  quick_commands = "fallback", -- "fallback" | "always" | "off"

  -- Shell command tool (lets Cody run allowlisted terminal commands via
  -- vim.system). ON by default once setup() runs; set false to disable.
  enable_shell = true,
  shell_skip_confirm = true,        -- default true (no prompt); set false to confirm each command
  -- shell_allowlist = nil,         -- list of allowed executables; nil = built-in default set
  -- shell_timeout_ms = 15000,      -- per-command timeout (clamped 1000..120000)
  -- shell_output_max_bytes = 8000, -- cap stdout+stderr returned to the model

  -- Ex-command tool (lets Cody run :CodyTranscript, :split, and change settings
  -- via :CodySet). ON by default once setup() runs; set false to disable.
  enable_commands = true,
  -- commands_confirm = false,      -- ask before each command (default off; allowlist is the guard)
  -- commands_allowlist = nil,      -- list of allowed command names; nil = built-in default set

  show_assistant_messages = true,
  feedback_panel = true,
  feedback_auto_open = true,
  feedback_height = 16,
  feedback_width = 96,
  feedback_recent_lines = 4,
  feedback_conversation_items = 12,
  context_max_lines = 2000,
  context_max_bytes = 240000,

  -- Optional spoken feedback (ElevenLabs). Off unless tts_enabled = true.
  tts_enabled = false,
  tts_provider = "elevenlabs",
  tts_voice_id = nil,          -- falls back to $ELEVENLABS_VOICE_ID
  tts_model_id = nil,          -- falls back to $ELEVENLABS_MODEL_ID, then eleven_flash_v2_5
  tts_speak_phases = false,    -- opt-in "Listening." / "Thinking."
  tts_speak_actions = true,    -- "Editing range." / "Renaming."
  tts_speak_results = true,    -- "Done." / "Failed: <reason>."
  tts_speak_messages = true,   -- short final assistant replies
  tts_message_max_chars = 160, -- skip spoken messages longer than this
  tts_request_timeout_ms = nil, -- falls back to $CODY_TTS_REQUEST_TIMEOUT_MS, then 10000
})

fallback is the default: typed :CodyDo commands go through GPT Realtime when OPENAI_API_KEY is set, and simple local regex commands are used only when the key is absent. Assistant messages are shown by default and truncated to fit the command line; set show_assistant_messages = false to suppress prose during voice sessions. Cody sends the active buffer with line numbers and a cursor marker when it fits the context limits above; larger buffers are cursor-centered and marked as truncated with omitted-line counts. The feedback panel is enabled by default and auto-opens on Cody activity. It shows the current phase, intent, transcript, selected tool/action, result, assistant message text, and a short recent event stream.

Then start it:

Spoken Feedback (TTS)

Cody can optionally speak short, high-signal confirmations using ElevenLabs. It is off unless you opt in, and it is deliberately terse: it never reads back your command, streamed transcript, or streamed assistant text.

What gets spoken, by category (each toggleable):

  • tts_speak_phases: Listening., Thinking. (off by default)
  • tts_speak_actions: the selected tool, e.g. Editing range., Renaming. (read-only locator/context tools stay silent)
  • tts_speak_results: Done. on success, Failed: <reason>. on failure
  • tts_speak_messages: a short final assistant reply, only when it fits tts_message_max_chars

Speech is cancelled immediately on a new turn, :CodyVoiceStop, an interruption, or a failure, so stale audio never trails the current action.

Enable it and provide a voice:

require("cody").setup({
  tts_enabled = true,
  tts_voice_id = "<elevenlabs-voice-id>",
  -- tts_model_id = "eleven_flash_v2_5", -- optional; this is the default
  -- tts_request_timeout_ms = 10000,     -- optional; default is 10s
})
export ELEVENLABS_API_KEY="..."
# Optional, can be set here instead of in setup():
export ELEVENLABS_VOICE_ID="<elevenlabs-voice-id>"
export ELEVENLABS_MODEL_ID="eleven_flash_v2_5"
export CODY_TTS_REQUEST_TIMEOUT_MS="10000"

The API key is read from the shell environment in Node and is never passed from Lua. The voice and model fall back to ELEVENLABS_VOICE_ID / ELEVENLABS_MODEL_ID when set; otherwise Cody uses eleven_flash_v2_5, the ElevenLabs low-latency model for real-time use. Cody also defaults to the smaller mp3_22050_32 output format to reduce response payload size. Override that with ELEVENLABS_OUTPUT_FORMAT if you prefer higher-bitrate audio.

Playback uses macOS afplay on a temporary mp3 file. On other platforms, set CODY_TTS_PLAYER_COMMAND to an audio player that accepts a file path argument (for example mpg123 or ffplay). If ELEVENLABS_API_KEY or the voice id is missing while tts_enabled is true, Cody reports it once and continues without spoken feedback.

Useful live checks:

npm run tts:voices
npm run tts:smoke -- "Cody spoken feedback is working. Done." "<elevenlabs-voice-id>"

Inside Neovim, these check the same bridge process used by :CodyVoiceSession:

:CodyTtsStatus
:CodyTtsSmoke Cody spoken feedback is working. Done.

If the shell smoke test works but :CodyTtsStatus says TTS is disabled or the API key is missing, restart Neovim from the shell that exports the variables, or run :CodyStop then :CodyStart after changing require("cody").setup(...). An ElevenLabs 402 response means the request reached ElevenLabs but failed due to billing, quota, or plan/voice access.

Shell Commands

Cody can run terminal commands from inside Neovim via vim.system, exposed to GPT as the editor_run_command tool. It is on by default once setup() runs (set enable_shell = false to disable; a bare plugin load with no setup() stays off). It is gated several ways:

  • the bridge only advertises the tool when enable_shell is on (which sets CODY_ENABLE_SHELL=1 for the Node bridge);
  • the Lua handler refuses when enable_shell = false, even if the tool is somehow advertised (the bridge env is captured at start, so the two layers can briefly disagree until a restart);
  • every command is checked against an allowlist of executables. The per-command vim.fn.confirm prompt is off by default (shell_skip_confirm = true); set shell_skip_confirm = false to be asked before every command.

Commands run without a shell (argv only), so pipes, globs, redirection, and ; & | are rejected — pass an argv array like ["git", "status", "--short"] for anything with spaces in arguments. Output (stdout+stderr) is capped before being sent to the model, and execution is asynchronous, so a slow command never freezes the editor; it is killed at shell_timeout_ms.

The allowlist binds the executable name only. Some allowed tools are general interpreters or build drivers (node -e, python -c, make, npm run, cargo) that can run arbitrary code, so treat the allowlist as a convenience filter, not a sandbox — the per-command confirmation is the real authorization boundary. Set shell_skip_confirm = true only when you trust the session.

require("cody").setup({
  enable_shell = true,
  -- shell_skip_confirm = true,                            -- skip the per-command prompt (use with care)
  -- shell_allowlist = { "npm", "git", "make", "cargo" },  -- replaces the built-in default set
})

Then ask, for example, :CodyDo run the tests or say "git status". Changing enable_shell requires restarting the bridge (:CodyStop then :CodyStart).

Editor Commands

Cody can also run Ex commands (the kind you type after :) as the editor_command tool — so voice/text like "open the transcript" or "split the window" maps to :CodyTranscript / :split. Like the shell tool it is on by default once setup() runs (set enable_commands = false to disable), and only allowlisted command names run; :!, :lua, the ! variant, and | chaining are rejected.

require("cody").setup({
  enable_commands = true,
  -- commands_confirm = true,                      -- ask before each command (default off; allowlist is the guard)
  -- commands_allowlist = { "CodyTranscript", "split", "MyCmd" }, -- replaces the built-in set
})

The default allowlist covers safe Cody/display/navigation commands plus the built-in netrw file explorer (CodyTranscript, CodyFeedbackOpen, CodyCapabilities, split, vsplit, only, close, wincmd, nohlsearch, redraw, Explore, Lexplore, Sexplore, Vexplore, …). File-writing/buffer-loading commands (write, update, edit, tabnew, …) are intentionally excluded — with a path argument they write or load arbitrary files — so add them via commands_allowlist only if you want that (ideally with commands_confirm = true). Then say things like "open the transcript", "open the file tree", or :CodyDo show the feedback panel.

To change a setting by voice, Cody runs :CodySet <key> <value> (also usable directly):

:CodySet feedback_height 30
:CodySet show_assistant_messages false

:CodySet only changes live-applicable display keys (feedback_height, feedback_width, feedback_recent_lines, feedback_conversation_items, context_max_lines, context_max_bytes, show_assistant_messages) which take effect immediately. Env-derived flags (enable_shell, enable_commands, tts_*) and the confirm-guard toggles (commands_confirm, shell_skip_confirm) are not runtime-settable — set them in setup() (and restart the bridge for the env-derived ones).

Commands

:CodyStart
:CodyStop
:CodyDo go to line 48
:CodyDo go to file lua/cody/init.lua
:CodyDo edit this line to handle nil paths
:CodyCapabilities
:CodyCapabilities json
:CodyFeedback
:CodyFeedbackOpen
:CodyFeedbackClose
:CodyFeedbackClear
:CodyTranscript
:CodySet feedback_height 30
:CodyInstall
:CodyInstall telescope.nvim
:CodyInstall json
:CodyStartTsLsp
:CodyVoiceStart
:CodyVoiceSession
:CodyVoicePress
:CodyVoiceRelease
:CodyVoiceStop
:CodyTtsStatus
:CodyTtsSmoke

:CodyInstall explains missing installable providers and renders lazy.nvim specs when lazy.nvim is detected. :CodyInstall <provider> asks for explicit confirmation, then copies the suggested spec to a register; it does not edit plugin configuration or install anything silently.

For local TypeScript/JavaScript testing without your own LSP config, Cody includes an explicit helper that starts Neovim's built-in LSP client against the repo-local typescript-language-server:

:e src/realtime-session.ts
:CodyStartTsLsp
:CodyCapabilities

CodyStartTsLsp is opt-in and only attaches to the current JS/TS buffer. Use :CodyStartTsLsp! to force it for an unusual filetype.

Feedback panel controls:

:CodyFeedback       " toggle
:CodyFeedbackOpen
:CodyFeedbackClose
:CodyFeedbackClear
:CodyTranscript     " full conversation in a scrollable window (q to close)

The feedback panel is a compact, non-focusable HUD: it shows only the most recent lines that fit feedback_height and redraws on every event, so you cannot scroll it. To read or scroll a long (or streamed) assistant message, open :CodyTranscript — a focusable, wrapping window with the full conversation (q to close, normal motions / <C-d>/<C-u> to scroll). To make the inline panel itself taller, raise feedback_height (and optionally lower feedback_recent_lines to give the conversation more room):

require("cody").setup({
  feedback_height = 30,      -- default 16; capped to the editor height
  feedback_recent_lines = 2, -- default 4; fewer event lines = more message room
})

Suggested push-to-talk mapping:

vim.keymap.set("n", "<leader>vs", "<cmd>CodyVoiceStart<cr>")
vim.keymap.set("n", "<leader>vl", "<cmd>CodyVoiceSession<cr>")
vim.keymap.set("n", "<leader>vp", "<cmd>CodyVoicePress<cr>")
vim.keymap.set("n", "<leader>vr", "<cmd>CodyVoiceRelease<cr>")
vim.keymap.set("n", "<leader>ve", "<cmd>CodyVoiceStop<cr>") -- cancel/stop fallback
vim.keymap.set("n", "<leader>cd", ":CodyDo ")

Voice uses Realtime server-side VAD. Normal flow is :CodyVoiceStart, speak a short command, then stop speaking; Cody stops recording and submits the turn when the server reports speech has ended. For explicit turn boundaries, bind :CodyVoicePress to key down and :CodyVoiceRelease to key up where your keymap layer supports that shape. :CodyVoiceStop cancels the current recorder, model response, and queued tool results.

For a persistent microphone session:

:CodyVoiceSession
" speak commands one at a time
" say: stop listening

CodyVoiceSession keeps the recorder open across VAD turns. When you say "stop listening", the model should call Cody's cody_stop_voice_session tool and the bridge stops the recorder.

General search uses whichever picker Cody detects as available. For example, if Telescope is loaded:

:CodyDo find auth service
:CodyDo search for auth token in the workspace

The generated picker tool receives mode = "files" for file/path search and mode = "grep" for workspace text search.

Environment

export OPENAI_API_KEY="sk-..."
export OPENAI_REALTIME_MODEL="gpt-realtime-2"
export CODY_AUDIO_DEVICE="" # optional sox device override
export CODY_ENABLE_SHELL="1" # shell tool; normally set via enable_shell in setup() (default on)
export CODY_ENABLE_COMMANDS="1" # Ex-command tool; normally set via enable_commands in setup() (default on)

# Optional spoken feedback (see "Spoken Feedback (TTS)")
export ELEVENLABS_API_KEY="..."
export ELEVENLABS_VOICE_ID="<elevenlabs-voice-id>"
export ELEVENLABS_MODEL_ID="eleven_flash_v2_5" # optional
export ELEVENLABS_OUTPUT_FORMAT="mp3_22050_32" # optional
export CODY_TTS_REQUEST_TIMEOUT_MS="10000"     # optional
export CODY_TTS_PLAYER_COMMAND="afplay"        # optional, non-macOS players

gpt-realtime-2 is the default because the current OpenAI Realtime docs use it in the WebSocket and session examples.

Evals

Local deterministic checks:

npm run typecheck
npm test
lua test/adapter_spec.lua
nvim -l test/tts_env_spec.lua
nvim -l test/shell_handler_spec.lua
nvim -l test/command_handler_spec.lua

npm test covers the TypeScript bridge, including the TTS feedback-to-speech mapping and cancellation. test/tts_env_spec.lua runs under Neovim's LuaJIT (not the system lua) and checks the bridge environment built from the TTS config.

Live router evals use the actual Realtime model with fake editor context/capabilities and check the first selected tool:

export OPENAI_API_KEY="sk-..."
npm run eval:router

If OPENAI_API_KEY is not set, eval:router skips without failing. Search evals use fixture files under test/fixtures/search.

Product Direction

Cody is deliberately narrow:

  • It should feel like a modal editor command layer, not a chat sidebar.
  • Voice commands should be short and imperative.
  • Navigation should reuse native editor commands or the user's preferred picker.
  • The model should use tools, not narrate pretend actions.
  • Write tools are limited to the active editor buffers.

Phases

  1. Command adapter: detect native/LSP/plugin commands and expose them as Realtime tools.
  2. Provider installer: install missing command providers through lazy.nvim or another detected package manager.
  3. Realtime text loop: typed commands call the adapter through GPT Realtime.
  4. Push-to-talk voice: voice commands flow through the same adapter.
  5. Smarter edits: add Tree-sitter context and stricter write guardrails.

Useful later steps:

  • Add a native push-to-talk key listener for press/release instead of two Vim commands.
  • Add a panel focus/scrollback mode and a copy/export command.

Done:

  • Optional ElevenLabs spoken feedback driven by the feedback event stream (see "Spoken Feedback (TTS)").
  • Tree-sitter context for function/class-aware edits.
  • A small eval suite for command parsing and tool selection.