惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

W
WeLiveSecurity
T
Troy Hunt's Blog
S
Schneier on Security
C
Cybersecurity and Infrastructure Security Agency CISA
C
CXSECURITY Database RSS Feed - CXSecurity.com
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Security Latest
Security Latest
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
MyScale Blog
MyScale Blog
Recorded Future
Recorded Future
A
About on SuperTechFans
PCI Perspectives
PCI Perspectives
H
Help Net Security
量子位
Blog — PlanetScale
Blog — PlanetScale
云风的 BLOG
云风的 BLOG
S
Security @ Cisco Blogs
The Hacker News
The Hacker News
P
Privacy International News Feed
Hacker News: Ask HN
Hacker News: Ask HN
阮一峰的网络日志
阮一峰的网络日志
博客园_首页
N
Netflix TechBlog - Medium
N
News and Events Feed by Topic
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Threatpost
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Simon Willison's Weblog
Simon Willison's Weblog
有赞技术团队
有赞技术团队
博客园 - 司徒正美
J
Java Code Geeks
S
Securelist
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Security Archives - TechRepublic
Security Archives - TechRepublic
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 叶小钗
S
Secure Thoughts
Latest news
Latest news
S
Security Affairs
T
The Exploit Database - CXSecurity.com
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 聂微东
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Last Week in AI
Last Week in AI
I
Intezer
雷峰网
雷峰网
Hacker News - Newest:
Hacker News - Newest: "LLM"
Engineering at Meta
Engineering at Meta
Hugging Face - Blog
Hugging Face - Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org

Hacker News - Newest: "LLM"

GitHub - lechmazur/position_bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders. Flex routing (EU and EFTA) Dark Factories: Retooling for LLM Velocity Ask HN: What would be the impact of a LLM output injection attack? GitHub - AronDaron/dataset-generator: No-code desktop app for generating high-quality synthetic datasets to fine-tune LLMs — plan-then-execute pipeline, LLM-as-judge, HuggingFace upload. GitHub - Oaklight/llm-rosetta: Production-ready LLM API translation layer for Python — bidirectional conversion between OpenAI, Anthropic & Google formats via hub-and-spoke IR. Optional API gateway. Streaming & non-streaming. Zero core deps. Contributions welcome! GitHub - browser-use/browser-harness: Self-healing browser harness that enables LLMs to complete any task. GitHub - moeen-mahmud/remen: Remen turns thoughts into something you can return to Analyzing 156 LLM Launch Posts on Hacker News ChatGPT vs Gemini vs Claude: The Best LLM Subscription You Should Buy GitHub - salaamalykum/quran-semantic-search: High-density RAG Semantic Search Engine & Quran Corpus (GEO/SEO Architecture) GitHub - NVIDIA/TensorRT-LLM: TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. The State of LLM Bug Bounties in 2026 Operational Readiness Criteria for Tool-Using LLM Agents Meshcore: Architecture for a Decentralized P2P LLM Inference Network How an LLM becomes more coherent as we train it GitHub - seetrex-ai/laimark GitHub - Jossifresben/BibCrit: AI-assited biblical textual criticism GitHub - wastedcode/memex: File system based wiki, maintained by Claude 99helpers.com GitHub - cliver-project/AITrigram GitHub - unbody-io/adapt: A self-evolving memory layer for AI agents. GitHub - hb20007/awesome-gen-ai-fails: A list of incidents where reliance on generative AI and LLMs resulted in harm to companies, individuals, or society GitHub - nevenkordic/localmind: Run any local LLM with persistent memory and context. CLI agent over Ollama with SQLite-backed hybrid recall. No cloud. Ask HN: What are the machine requirements for a LLM like Llama-3.1-8B? Faster LLM Inference via Sequential Monte Carlo grpo explained: group relative policy optimization for llm finetuning - cgft Stop comparing price per million tokens: the hidden LLM API costs · TensorZero Andrej Karpathy's LLM Wiki Is a Bad Idea GitHub - GG-QandV/mnemostroma: Offline RAM-first cognitive leer/coprocessor for AI agents and robotics. Solves "Context Abandonment" with 20-80ms latency using a dual-thread biomimetic memory architecture (ONNX + SQLite WAL). mempalace/agent at agent · skorotkiewicz/mempalace GitHub - Nyquest-ai/nyquest-rust-fullstack-pub: Nyquest — Semantic Compression Proxy for LLMs. 350+ rules, local LLM stage, 15-75% token savings. Full Rust stack. GitHub - TheoV823/mneme: Enforce architectural decisions in AI-assisted development. GitHub - klemenvod/TokenBrawl: A 1v1 Bomberman-style game where two LLM agents play autonomously against each other. No human plays — you watch the AIs fight. Each agent receives a text description of the board state, reasons about it, and outputs a move as JSON. The game engine executes it. Introducing the Common AI Provider: LLM and AI Agent Support for Apache Airflow Power Circuit AI: Designing Power Electronic Circuits for Motor Drives with Generative Artificial Intelligence Ask HN: How to program with IDE and LLM on CPU locally? Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis Bonsai 1-bit WebGPU - a Hugging Face Space by webml-community The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows Ask HN: Simple tooling for local LLM code critique without IDE integration? Can a General LLM Diagnose a DICOM Slice? A 10-Case Public Benchmark Charts-of-Thought: Enhancing LLM Visualization Literacy (PDF, 2026) GitHub - Mesh-LLM/mesh-llm: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. GitHub - seamus-brady/springdrift: A persistent runtime for long-lived LLM agents Writing an LLM from scratch, part 32k -- Interventions: training a better model locally with gradient accumulation Ask HN: Which LLM model and agentic CLI are you using for local development? GitHub - wayneColt/modelcascade: Route local. Escalate smart. Never overspend. Open-source multi-model cascade routing for autonomous agents. LLM pricing is 100x harder than you think GitHub - asakin/llm-primer: Pre-warmed Claude Code sessions in tmux. No startup wait. GitHub - EggerMarc/chat-rs: A multi-provider LLM framework for Rust. GitHub - SynapseKit/SynapseKit: Minimal, async-first Python framework for production LLM apps- 2 hard deps, no magic, no SaaS. A Claude Skill that Makes LLM Paragraphs More Bearable Does Gas Town 'steal' usage from users' LLM credits & paid services to improve itself? What's Claude Code Actually Doing? Open the Black Box with the Arthur Engine Milla Jovovich's New Open Source LLM Memory App and the Dark Code Problem Your intuition of LLM token usage might be wrong Show HN: Bloomberg Terminal for LLM ops – free and open source GitHub - 0xchamin/mcptube: Transform YouTube videos into a compounding knowledge base with transcripts, vision analysis, and agentic search. Works as an MCP server for Claude, Copilot & more. Show HN: Open KB: Open LLM Knowledge Base Your LLM is a compiler, not a runtime GitHub - sapountzis/Unslop: A Web Feed That Deserves You crates.io: Rust Package Registry Beyond Karpathy's LLM-Wiki: The Necessity of Cognitive Governance GitHub - amitshekhariitbhu/llm-internals: Learn LLM internals step by step - from tokenization to attention to inference optimization. GitHub - parallem-ai/parallem: An expressive library for running agents with the Batch API. GitHub - stfurkan/pi-llm LLM-Wiki Show HN: Formal – Formal verification for AI-generated code using Lean 4 LRTS – Regression testing for LLM prompts (open source, local-first) LLM Wiki Skill: Build a Second Brain with Claude Code and Obsidian I built an LLM Wiki and RAG solution: here's a demo for a security KB The biggest advance in AI since the LLM Predict-Rlm: The LLM Runtime That Lets Models Write Their Own Control Flow the-synthetic-library/the-synthetic-mind at main · joshferrer1/the-synthetic-library GitHub - yisding/reviewwiggum GitHub - Donnyb369/mcp-spine: Context Minifier & State Guard — Local-first MCP middleware proxy GitHub - Beledarian/wgpu-llm: A from-scratch LLM inference engine that uses wgpu (the cross-platform WebGPU implementation) to dispatch WGSL compute shaders for every math operation a Transformer needs. No CUDA. No Python. No massive framework dependencies. Just Rust, raw shaders, and your GPU. GitHub - anitiue/Hindsight: An experience-driven self-improvement framework for LLM agents — 基于经验的 LLM Agent 自我改进框架 GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. GitHub - alainnothere/AmdPerformanceTesting: Amd Performance Testing Ask HN: Is a purely Markdown-based CRM a terrible idea? Optimized for LLM agents Context Engineering - LLM Memory and Retrieval for AI Agents | Weaviate little_helper_tui/letter.md at main · sleepyeldrazi/little_helper_tui GitHub - EvanZhouDev/umr: The Unified Model Registry for all your local AI apps. GitHub - JordanCT/VigIA-Orchestrator Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain A Taxonomy of RL Environments for LLM Agents Llama LLM Network Feture GitHub - genedeng-ca/ai-mac-migration: AI-powered Mac-to-Mac migration tool - replace Apple Migration Assistant with intelligent, selective transfer using local LLMs GitHub - lunargate-ai/gateway: High-performance self-hosted AI gateway (OpenAI-compatible) with routing, retries, and streaming GitHub - AuthBits/webmcp: A lightweight, prompt-driven MCP web research server for high-quality LLM powered information extraction. Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Springdrift: An Auditable Persistent Runtime for LLM Agents with Case-Based Memory, Normative Safety, and Ambient Self-Perception High-Stakes Personalization: Rethinking LLM Customization for Individual Investor Decision-Making From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
GitHub - AlphaBitCore/nexus-gateway: Enterprise AI traffic gateway — unified compliance, routing across 20+ LLM providers, semantic cache, quotas, and audit. SDK / network / OS-layer intercept.
jjrhodes · 2026-05-27 · via Hacker News - Newest: "LLM"

CI Go CI Coverage gate Status: Pre-GA License: Apache 2.0

Make AI safe to use across the enterprise.

Nexus Gateway intercepts enterprise LLM traffic at three layers and runs all of it through one compliance engine, one audit pipeline, and one control plane.

Mode Where it intercepts Code
🔑 AI Gateway SDK layer — virtual keys on /v1/chat/*, /v1/responses, /v1/embeddings, /v1/messages packages/ai-gateway/
🌐 Compliance Proxy Network layer — transparent TLS bump (CONNECT + MITM) packages/compliance-proxy/
💻 Desktop Agent OS layer — macOS / Linux / Windows builds all in development, awaiting QA packages/agent/platform/{darwin,linux,windows}/

The three pipes are independent: AI Gateway, Compliance Proxy, and Agent each run the full hooks pipeline on their own traffic (packages/shared/policy/hooks/, plus the per-service compliance pipeline — e.g. packages/agent/internal/compliance/pipeline.go). The Agent always egresses directly to the upstream provider — it does not care whether enterprise network policy then routes that traffic through the Compliance Proxy.

When it does — Agent stamps an Ed25519-signed X-Nexus-Attestation header on the outbound request (E60, packages/agent/internal/identity/attestation/). The Compliance Proxy peeks this header before the TLS bump (packages/shared/transport/tlsbump/forward_handler.go:119); if the signature verifies, the CONNECT becomes pure passthrough — no MITM, no hooks, no audit on that flow, since the Agent already ran them.


What Nexus does

🔁 Write once in OpenAI shape, route to 20 in-tree adapter codecs

Applications speak the OpenAI SDK. Nexus normalises every request to a canonical OpenAI shape, then translates wire format on the way to the actual provider. Shipped adapter codecs today (packages/ai-gateway/internal/providers/specs/):

  • First-class codecs (11): openai, anthropic, gemini, vertex, azure, bedrock, cohere, minimax, glm, replicate, voyage.
  • OpenAI-compatible passthrough (9): deepseek, moonshot, mistral, groq, fireworks, together, perplexity, xai, huggingface — all under packages/ai-gateway/internal/providers/specs/compat/.

Reasoning tokens, function calls, vision inputs, structured outputs are carried through the translation. Adding a new provider is a documented procedure under .claude/skills/add-provider-adapter/.

🧊 Multi-tier cache

  • Exact-match response cache — Valkey-backed, Redis-wire-compatible.
  • Provider-native cache accounting — surfaces Anthropic cached_tokens and Gemini cachedContentTokenCount in billing when the provider reports them.
  • Semantic vector cache via the valkey-search module — packages/ai-gateway/internal/cache/semantic/ (lookup, writer, client, circuit breaker, singleflight, poison guard, index lifecycle).
  • In-flight singleflight — concurrent identical prompts fold into one upstream call.

💰 Cost & quota control

  • Multi-axis quotas — per organization, per virtual key, per provider, per model. Each axis has its own budget and sliding-window enforcement.
  • Token-based or USD-based budgets.
  • Hard limits and soft limits — soft fires an alert; hard rejects with 429.
  • Real-time accounting — counters update on every traffic event, no batch lag.
  • Routing strategies in packages/ai-gateway/internal/routing/strategies/: single, fallback, loadbalance, conditional, absplit, policy, smart.

🛡 Compliance pipeline

PII detection · data classification · keyword filtering · content safety · rate limiting · IP allowlists · request-size validation · webhook forwarders · per-stage audit (request hooks and response hooks recorded independently) · body capture (256 KiB inline + spillstore for the rest, see packages/shared/storage/spillstore/) · SIEM forwarder (packages/compliance-proxy/internal/siem/) · three-tier kill switch · emergency passthrough (bypassHooks / bypassCache / bypassNormalize).

🎨 Modalities

Chat · Embeddings · Structured outputs · Function / tool calling · Vision input · Reasoning tokens. Multimodal (epic E62) in development.

🏢 Enterprise governance

  • IAM — RBAC + ABAC with an NRN resource model (packages/shared/identity/iam/).
  • Virtual keys with per-key model scope.
  • OIDC federation with JIT user provisioning (packages/control-plane/internal/identity/authserver/login/oidc.go, JIT flag in scim_store.go).
  • Organization / project hierarchy with per-org quota.
  • Credential vault — AES-256-GCM (packages/control-plane/internal/platform/crypto/aes_gcm.go, packages/ai-gateway/internal/credentials/decrypt/decrypt.go) with key rotation.
  • Agent fleet management — Hub CA, Thing-based config sync, drift detection.

Architecture in one minute

Five Go services + one React control console. The diagram below shows only the traffic plane — the three independent intercept pipes and where each one egresses. Control plane (Hub-centric) and storage are summarized in the component table immediately after.

flowchart TB
    SDK["SDK app<br/>(OpenAI SDK)"]
    HTTPS["HTTPS app<br/>(network-proxied)"]
    Endpoint["Developer endpoint<br/>(Cursor / Claude Code / …)"]

    AIGW["AI Gateway :3050<br/>routing · cache · quota<br/>+ hooks pipeline"]
    CPProxy["Compliance Proxy :3128<br/>MITM TLS<br/>+ hooks pipeline"]
    Agent["Desktop Agent · local<br/>OS-level intercept<br/>+ hooks pipeline"]

    Provider["LLM Provider<br/>(OpenAI / Anthropic / Gemini / …)"]

    SDK ==>|"/v1 + VK"| AIGW
    HTTPS ==>|HTTPS via proxy| CPProxy
    Endpoint ==>|OS-level capture| Agent

    AIGW ==> Provider
    CPProxy ==> Provider
    Agent ==> Provider

    Agent -. "X-Nexus-Attestation verified<br/>→ passthrough" .-> CPProxy
Loading

The lateral dotted arrow is the attestation handoff: the Agent always egresses directly, but when enterprise network policy happens to route Agent traffic through the Compliance Proxy, the Agent's Ed25519-signed X-Nexus-Attestation header (E60, packages/agent/internal/identity/attestation/) is verified at TLS-bump time (packages/shared/transport/tlsbump/forward_handler.go:119); on success the CONNECT becomes pure passthrough — no MITM, no hooks, no audit on that flow, since the Agent already ran them on its end.

Control plane (out-of-band). All four Go services register with Nexus Hub as Things via packages/shared/transport/thingclient/ (WebSocket primary, HTTP fallback) and pull configuration from the Hub's device shadow on boot and on change-signal — the Hub never pushes full state. The Control Plane admin API (:3001) and the React UI (:3000) sit alongside, talking to the Hub the same way.

Component Port Code
Nexus Hub 3060 packages/nexus-hub/ — Thing Registry, Device Shadow, config sync, jobs, agent CA, SIEM bridge
Control Plane 3001 packages/control-plane/ (Echo) — admin API / BFF, IAM, SSO, analytics
AI Gateway 3050 packages/ai-gateway//v1 AI traffic, provider adapters, routing, quota
Compliance Proxy 3128 packages/compliance-proxy/ — CONNECT, MITM, compliance pipeline
Agent local packages/agent/ — macOS uses pf packet filter (packages/agent/internal/platform/darwin/pfintercept/); Linux uses iptables; Windows uses WinDivert. The legacy NETransparentProxyProvider path (packages/agent/platform/darwin/NexusAgent/NexusAgentExtension/) is still in the repo behind interceptMode=ne, but new builds default to pf. All three platforms are development-complete, not yet QA-signed-off.
Control Plane UI 3000 packages/control-plane-ui/ — React + Vite + TypeScript

Storage stack

  • PostgreSQL 16 — durable storage. Prisma schema in tools/db-migrate/ is the source of truth for dev-time migrations; runtime code reads via hand-written SQL + pgx (no sqlc).
  • Valkey 8 — Redis-wire-compatible, pinned to valkey/valkey-bundle:8-trixie in docker-compose.yml for BSD-license parity; the valkey-search module ships in the bundle image and backs the semantic vector cache. Pure cache only — no pub/sub.
  • NATS JetStream — event streaming and Hub coordination via packages/shared/transport/mq/.

Deployment

Form factor How Status
AWS Marketplace AMI / single-instance appliance cd nexus-ami && ./build.sh — bakes binaries + UI + Prisma + nginx + Postgres + Valkey + NATS into one AL2023 image via Packer nexus-ami/README.md for build steps, docs/developers/architecture/cross-cutting/deployment/ami-appliance-architecture.md for design
Local development docker-compose + ./scripts/dev-start.sh (Postgres + Valkey + NATS) and per-service go run ./cmd/<svc>/ See Quick start below
VMware / KVM image / bare-metal appliance Reuses the same install.sh + harden.sh from nexus-ami/scripts/ under a different Packer builder Future
Container / Kubernetes Out of scope for the appliance form factor — separate product line Future

Quick start (local development)

Prerequisites

Tool Version Notes
Node.js 20+ npm workspaces require npm 10+
Go 1.25+ All Go modules share go.work at the repo root
Docker any recent Hosts PostgreSQL, Valkey, NATS via docker-compose.yml

One-shot bootstrap

The script:

  1. Verifies prerequisites (Node 20+, Go 1.25+, Docker, OpenSSL).
  2. Auto-creates repo-root .env from .env.example with safe dev defaults for CHANGE_ME_* secrets (INTERNAL_SERVICE_TOKEN, ADMIN_KEY_HMAC_SECRET, CREDENTIAL_ENCRYPTION_KEY = openssl rand -hex 32, …). All four Go services read this via packages/shared/core/bootenv/ at boot.
  3. Starts PostgreSQL + Valkey + NATS via docker-compose.yml.
  4. Runs npm install.
  5. Auto-creates tools/db-migrate/.env and propagates CREDENTIAL_ENCRYPTION_KEY into it so prisma db seed can re-encrypt the seed credentials.
  6. Applies the Prisma schema (db push) and seed under tools/db-migrate/.
  7. Auto-generates the Compliance Proxy dev CA at packages/compliance-proxy/dev-certs/{ca.crt,ca.key} so the TLS-bump cert issuer can boot.
  8. Prints the per-service go run … -config <svc>.dev.yaml commands.
  9. Finally starts the Control Plane UI dev server.

Flags:

  • --force-reset — DESTRUCTIVE: wipe local Postgres / Valkey / NATS volumes + the entire nexus_gateway database before re-applying the schema.
  • --no-dev — bootstrap only; print the per-service commands and exit instead of starting the UI dev server.

Start the services

Open one terminal per Go service after the bootstrap finishes:

cd packages/nexus-hub         && go run ./cmd/nexus-hub/         -config nexus-hub.dev.yaml          # port 3060
cd packages/control-plane     && go run ./cmd/control-plane/     -config control-plane.dev.yaml      # port 3001
cd packages/ai-gateway        && go run ./cmd/ai-gateway/        -config ai-gateway.dev.yaml         # port 3050
cd packages/compliance-proxy  && go run ./cmd/compliance-proxy/  -config compliance-proxy.dev.yaml   # port 3128
npm run dev:control-plane-ui                                                                          # port 3000

The -config <svc>.dev.yaml flag is required — each binary defaults to <svc>.config.yaml, which is the prod-shape template and is intentionally missing dev-only fields like hub.id. Without the flag the service fails fast at boot.

Each Go service tees logs to packages/<service>/logs/<service>.log in dev mode (configured in the service's *.dev.yaml). Override the path with LOG_FILE=/path/to/file.

Open the console

Browse to http://localhost:3000 and sign in as the seeded super-admin:

admin@nexus.ai / admin123

Additional seeded roles (alice@nexus.ai, carol@nexus.ai, bob@nexus.ai, diana@nexus.ai) are defined in tools/db-migrate/seed/seed.ts.

Prefer the terminal? nexus (packages/nexus-cli) is a single Go binary that operates and observes the gateway from the command line — a Bubble Tea TUI (health overview, live traffic radar, event drill-down with an LLM "explain this event", SLO, cost, a chat playground, the kill switch and emergency passthrough, alerts and nodes, a k9s-style command palette, and an "Ask Nexus" natural-language bar that turns a plain-English question into a view jump or a data answer), a scriptable CLI (nexus <noun> <verb> --output json), and an MCP server (nexus mcp serve) that exposes the read/analyze tools to agents with an opt-in mitigate tier. It is a pure client over the same admin API + /v1/*, governed by the same IAM. See docs/users/features/operator-toolkit.md.

Try it

After the stack is up, walk through examples/01-hello-world/ — a 3-minute curl-through-the-gateway demo that ends with you reading the resulting traffic_event Postgres row.

Admin-API debugging from the shell

The Control Plane uses OAuth + PKCE bearer tokens. Helpers wrap the flow:

cp tests/.env.local.example tests/.env.local      # gitignored; edit if you need to override defaults
source tests/lib/loadenv.sh local                  # picks up tests/.env.local + tests/.env.local.example defaults
source tests/lib/auth.sh

cp_login                                       # idempotent; caches token at /tmp/nexus_test_token_local
cp_curl /api/admin/analytics/cost?groupBy=device
cp_curl -X POST /api/admin/routing-rules -d @rule.json

For direct DB inspection in dev:

docker exec $(docker ps --filter "name=postgres" -q | head -1) \
  psql -U postgres -d nexus_gateway -c "SELECT ..."

🧪 …and one more thing: this repo is also an AI vibe-coding workbench

You came for an AI gateway. You also get the disciplined AI pair-programming setup that built it. CLAUDE.md, .cursor/rules/, .claude/skills/, and the scripts/check-* lint suite form a fork-adoptable methodology:

  • Binding rules in CLAUDE.md plus 35 .cursor/rules/ entries (ls .cursor/rules/).
  • 26 invocable skills under .claude/skills//prod-deploy, /smoke-gateway, /spec-writing, /add-provider-adapter, hardened runbooks for repeatable procedures.
  • 23 scripts/check-* lint scripts — every binding rule has a mechanical gate; pre-commit + CI dual layer.
  • 95% per-package coverage gate enforced by scripts/check-go-coverage.sh + scripts/.coverage-allowlist.
  • 2-round completion self-audit before claiming "done" (see CLAUDE.md → Mandatory rules → Workflow discipline → Self-audit).

Repository layout

packages/
  nexus-hub/         Go — Thing Registry, Shadow, config sync, jobs, SIEM bridge, agent CA
  control-plane/     Go + Echo — admin API / BFF, IAM, SSO, analytics
  ai-gateway/        Go — /v1 AI traffic, provider adapters, routing, quota
  compliance-proxy/  Go — transparent TLS proxy, CONNECT, compliance pipeline
  agent/             Go — desktop traffic interception (macOS / Linux / Windows;
                     all builds in development, awaiting QA)
  shared/            Go — cross-service business logic (hooks, traffic, configtypes,
                     mq, thingclient, cache, …)
  control-plane-ui/  React + Vite + TypeScript — admin dashboard
  ui-shared/         Shared design tokens, chart colors, i18n bundles

tools/db-migrate/    Prisma schema + migrations + seed (dev-time only)

scripts/             dev-start.sh + check-* lint scripts
tests/               Test harnesses, .env.local.example, auth.sh helper, smoke scripts
examples/            Self-contained demos (01-hello-world, …)

docker-compose.yml   Local PostgreSQL + Valkey + NATS
go.work              Go workspace (one module per package + tools)
Makefile             build / test targets per service

Tech stack

  • Go services — Go 1.25+ with go.work; Echo on Control Plane / Nexus Hub / AI Gateway (labstack/echo/v4 v4.15.2); structured logging via log/slog; metrics via Prometheus promauto; Redis-wire client redis/go-redis/v9 v9.19.0; WebSocket via coder/websocket v1.8.14.
  • Control Plane UI — React + Vite + TypeScript (strict mode); React Query via the useApi hook; layered design tokens in packages/ui-shared/src/styles/ (global.css raw → light.css / dark.css semantic, flipped by data-theme); i18n with react-i18next (en / zh / es under packages/control-plane-ui/public/locales/ and src/i18n/locales/); tests via Vitest.
  • Database — PostgreSQL 16. Prisma is the dev-time source of truth (tools/db-migrate/); runtime queries use hand-written SQL + pgx.
  • Cache — Valkey 8 (Redis-wire-compatible, BSD-licensed valkey/valkey-bundle:8-trixie image). Pure cache only — no pub/sub anywhere.
  • MQ — NATS JetStream behind the packages/shared/transport/mq/ interface.
  • Monorepo — npm workspaces (packages/control-plane-ui, packages/agent/ui/frontend, tools/db-migrate) + go.work for Go.

Go workspace — what every build context must carry

Every Go module under packages/ references its sibling workspace packages by require github.com/AlphaBitCore/nexus-gateway/packages/<sibling> v0.0.0-<timestamp>-<commit>. Those pseudo-version requires are only there to make each module syntactically valid on its own — real resolution comes from go.work at the repo root.

This has one consequence: if go.work is missing from the build context, Go falls back to the literal pseudo-version in require and tries to fetch the module from GitHub instead of using the local source tree. The build "succeeds" against an old remote snapshot, masking local changes.

Rules for every build environment:

  • Fresh clonegit clone already includes the committed go.work and go.work.sum. Run go build from inside the repo.
  • Docker — copy go.work + go.work.sum and every packages/<module> directory the service transitively depends on, not just the service's own folder. Minimum viable layout:
    WORKDIR /build
    COPY go.work go.work.sum ./
    COPY packages/shared       packages/shared
    COPY packages/<svc>        packages/<svc>
    WORKDIR /build/packages/<svc>
    RUN go build -o /out/<svc> ./cmd/<svc>/
  • CI — use full actions/checkout (default fetch-depth, no sparse-checkout).
  • Sanity probeGOWORK=off go build ./cmd/<svc>/ from inside a workspace package should refuse to build or pull a remote snapshot.

If a contributor reports "Go keeps downloading our own modules from GitHub", the answer is always: their build context is missing go.work (or they have GOWORK=off set).


Common commands

Command Purpose
./scripts/dev-start.sh One-shot bootstrap (Docker + DB + seed + UI)
npm run dev:control-plane-ui Start the UI dev server only
make build-all Build the Go services + UI. Go binaries land in dist/bin/<service>/<binary>.
make test-all Run go test -race -count=1 for every Go module + UI Vitest
make clean Remove dist/bin/ and packages/control-plane-ui/dist/. Platform agent packages under dist/{macos,linux,windows}/ are preserved — clean those via the per-platform targets (agent-clean-macos, agent-clean-windows).
npm run check:all Run every pre-commit lint (i18n parity, design tokens, terminology, migration timestamps, useApi keys, sidebar icons, …). CI runs the same set.
npm run db:migrate Create a new Prisma migration in tools/db-migrate/

To build, sign, notarize, or package the macOS Agent (.app / .pkg), always invoke the build-agent Claude Code skill — not the raw wails / codesign / notarytool commands. See CLAUDE.md → "macOS Agent builds MUST go through Skill('build-agent')" binding rule for why.


Authoritative documents

  1. CLAUDE.md — binding charter. Plan + Todo gate, English-only artifacts, IAM impact review, macOS NE fail-open, pre-edit reading, completion-time self-audit, real-implementation-only, development-phase greenfield policy.
  2. CONTRIBUTING.md — workflow summary, pre-commit checks, high-blast-radius surfaces, review pointers.

Acknowledgments

  • Steve — the original idea behind Nexus Gateway came from him, and he stayed hands-on throughout: code, tests, design reviews, architectural decisions.
  • The wider team — engineers, code reviewers, QA, design folks, and the people running prod. The architecture decisions, design reviews, code-review catches, and prod incidents that shaped this codebase all came from team collaboration.
  • Claude Code — Anthropic's CLI assistant did the lion's share of the implementation work, side-by-side with the human maintainers.

AI is already here. Keep learning, keep adapting.