惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
News and Events Feed by Topic
T
The Exploit Database - CXSecurity.com
P
Palo Alto Networks Blog
T
Threat Research - Cisco Blogs
Cloudbric
Cloudbric
Recent Commits to openclaw:main
Recent Commits to openclaw:main
I
Intezer
Attack and Defense Labs
Attack and Defense Labs
P
Privacy International News Feed
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
L
Lohrmann on Cybersecurity
C
Cybersecurity and Infrastructure Security Agency CISA
V2EX - 技术
V2EX - 技术
AWS News Blog
AWS News Blog
O
OpenAI News
L
LINUX DO - 最新话题
N
News | PayPal Newsroom
PCI Perspectives
PCI Perspectives
www.infosecurity-magazine.com
www.infosecurity-magazine.com
T
Troy Hunt's Blog
Latest news
Latest news
D
Darknet – Hacking Tools, Hacker News & Cyber Security
A
Arctic Wolf
Spread Privacy
Spread Privacy
G
GRAHAM CLULEY
T
Tor Project blog
博客园_首页
Know Your Adversary
Know Your Adversary
有赞技术团队
有赞技术团队
S
Secure Thoughts
美团技术团队
Apple Machine Learning Research
Apple Machine Learning Research
爱范儿
爱范儿
T
Tailwind CSS Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
V
Visual Studio Blog
J
Java Code Geeks
Cisco Talos Blog
Cisco Talos Blog
Schneier on Security
Schneier on Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
S
Security Affairs
Jina AI
Jina AI
人人都是产品经理
人人都是产品经理
雷峰网
雷峰网
宝玉的分享
宝玉的分享
量子位
Last Week in AI
Last Week in AI
月光博客
月光博客
罗磊的独立博客
S
SegmentFault 最新的问题

Hacker News

GitHub - ronak-create/FableCut: Zero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI GitHub - BuceaGeorgia/VIRENA: A minimal Vision-Language-Action model you can read: frozen CLIP + a tiny head on ManiSkill PickCube. Runs on a Mac, no GPU. Fehu - Apps on Google Play GitHub - fresswolf/Slopera: The browser for the slop era Noema — AI Company Analysis GitHub - hamidi-dev/opentab: 📊 Browse your AI coding spend in the terminal — OpenCode, Claude Code, Codex & more GitHub - sgInnora/wc2026-prediction-ledger: Receipt-verified AI prediction ledger for World Cup 2026: pre-kickoff sha256-locked forecasts scored vs results, with calibration + market baselines. Live: goalpulse.io/open-data GitHub - michaelwrites67-ctrl/yogen: 予言 Yogen — a 500-agent AI swarm that debates your idea and predicts the outcome. Self-host free with your own Anthropic key. GitHub - teddytennant/wizard: Self-extending autonomous agent in one Rust binary. One-line install, any provider (OpenAI-compatible, Anthropic, xAI) or fully local via llama.cpp, live /evolve self-modification, MCP, messaging gateway, built-in bench arcaide.foo Show HN: Android Developer Verification Package blacklisted in Aurora Store OpenDescent: Private messaging for normal people GitHub - tarunlnmiit/autopilot-jobhunt: AI job agent: scans 130+ careers pages nightly, scores every role against your resume with an LLM (0–100), alerts you on Telegram, and drafts tailored cover letters + resumes. Free & open source. 18 Words - Daily Word Challenge The State of US Local Government Accessibility 2026 flow - Real-time network throughput dashboard for the terminal. - Terminal Trove Battle LLM Robots GitHub - robesris/ffvii-realtime: Speed up Final Fantasy VII (Rebirth / Remake / Revelation) Tactical Mode slow-motion so combat plays at real-time speed Agent Sessions - Local History for AI Coding Agents Scrutora — code, cloud & consent compliance in one platform SoulOS Tutorials GitHub - gaemi/agentic-fc: Open-source football management simulation played by AI agents through MCP and watched through a TUI console. GitHub - talalalrwas/ocr-grab: Flameshot clone that adds OCR. GitHub - saifmukhtar/kinetic figment computer linear.gratis - Free Linear Client Feedback Forms GitHub - atelier-ws/atelier: Runtime for coding agents. Models are getting smarter, but a model is only as capable as the environment supporting it. Atelier is that environment. https://atelier.ws Atlas · Tomesphere | Tomesphere Codenames Generator GitHub - lyfeninja/lyfeninja_blkseal_python_sdk: Lightweight Python client for signing and verifying digital content using lyfe.ninja's BlkSeal product powered by BlkBolt™. Designed for zero dependencies, simple integration, and exact content verification. DateTimeMate OpenScreenShot — Full-page screenshot & annotation tool for Chrome 38-0 — Build Your Premier League Dream Team Cyrinx — data over sound, measured GitHub - BhaveshThapar/mcp-audit Probed — Talk to Your People HN Work GitHub - Northwood-Systems/foreman: Self-hosted LLM gateway. Cost effective, deterministic, and fast. Secure and private by default. A Visualization Language for the AI Era GitHub - weirdGuy/kastor: Declarative language and toolchain for AI agents: define agents, tools and prompts in HCL, then compile to frameworks or manage them on hosted platforms with plan/apply semantics. Abralo - Run multiple Claude Code agents in one window GitHub - mehranzand/repofleet: RepoFleet is an issue-centered CLI tool for managing Git workflows across multiple repositories. GitHub - exmergo/dex: Dex is the agent-native analytics engineering toolkit. Point it at your warehouse and your dbt project. It learns the landscape, authors your transformations, and tells you exactly what to fix when the schema drifts. Built for analytics engineers and data engineers who want more out of their coding agent. Pug — Open Source Product Analytics GitHub - instavm/tarit: A hypervisor and sandbox cloud for self-hosted AI agents and RL Chiptune Radio — Aleph Void, LLC Free Mermaid Live Editor & Diagram Maker GitHub - Salnika/dejavu: Stop showing coding agents the same command output twice. GitHub - hirasso/html-obfuscator: Obfuscate emails, phone numbers, and other sensitive data in PHP. Invisible to humans, hidden from bots until they interact. Davit — a native macOS UI for Apple containers Fenzo AI - The perfect course, every time. HTML Drive — Edit and Publish HTML from Google Drive GitHub - rowboatlabs/rowboat: Open-source AI coworker, with memory ZeroGate | Automated Cluster Scaling A tiny scale-free kernel language — Joa Ebert GitHub - arman-jalili/guardian-framework: Architecture Enforcement Framework for AI-Assisted Development Yamanote.fun PostgreSQL on AWS: Size & Benchmark EC2 Instances GitHub - Rodiun/frugon: Free, local, open-source LLM cost analyzer — see where your LLM bill leaks, on your machine. Artificiety — A Fantasy World for AI Agents Ex Situ WhimFiles - Find Any File in Seconds GitHub - josephsenior/Grinta-Coding-Agent: Local-first autonomous coding agent that plans, executes, validates, and finishes software tasks end-to-end. Nectar — The Web Without JavaScript Agent Draw: An agent draws while you talk, built on TLDraw Captchainbox - Make senders work to get into your inbox GitBiased — your whole engineering org, on one calm dashboard Neil the Seal GitHub - animesh-94/Onboard-CLI: An AST-powered, local-first CLI that visualizes complex system architectures and enforces architectural boundaries via instant Git hooks. ridealong — live London trains GitHub - rubix-studios-pty-ltd/rubix-redis-bridge: Secure production-hardened Rust HTTP bridge for Redis with Upstash-style API compatibility, command allowlisting, hard-denied dangerous Redis commands, Docker deployment, and SDK compatibility tests. Chauffeur – Deine Arbeitsumgebung mit einem Klick zurück Clusy | Agent-Native Notebook for ML and Data Science GitHub - therepanic/openleetcode: we have democratized the LeetCode tests Habit Pocket — your good days aren’t random GitHub - dogtorjonah/context-warp-drive See what your community is paying attention to. GitHub - puffinsoft/peek-cli: Let coding agents see your browser. Daily Vocabulary | Quizingo GitHub - khalid-src/corv-client: Corv Client is an SSH client for AI agents and humans. GitHub - vicmaster/framesmith: Open-source MCP server that gives AI assistants a visual design canvas, rendering HTML/CSS scene graphs to PNG via headless Chromium. GitHub - gojargo/jargo: A WebRTC-native, audio-first conversational-AI framework for Go. snowscroll · Instagram, without the spiral. The Tree of "Tree" GitHub - Arthur-Ficial/translate: On-device translator for macOS Tahoe — UNIX filter + drop-in HTTP server compatible with DeepL, LibreTranslate, and Google v2. 100% on-device, no cloud, no LLM, no API keys. GitHub - palmier-io/palmier-pro: macOS video editor built for AI The New Wave of Remote Work: An Async-First Job Board GitHub - agenthatch/agenthatch: Where agents hatch. pacwich — Monorepo tooling for Bun, npm, and pnpm workspaces | Documentation GitHub - openwong2kim/wmux: Windows tmux alternative for AI agents — split terminals for Claude Code, Codex, Gemini CLI with MCP browser automation. No WSL required. GitHub - loopgain-ai/loopgain: An open-source cost controller for AI agent loops — stops a loop when it's actually converged and rolls back before it degrades, instead of running to a fixed max_iterations cap. Real-time loop-gain (Aβ) bands + best-so-far rollback. Adapters for LangGraph, CrewAI, AutoGen, LangChain, OpenAI Agents, and Claude Agent SDK; raw API for custom stacks. GitHub - vishal-dehurdle/state-harness: Runtime safety net for LLM agents. Detects token spirals, kills doomed tasks early, tells you exactly why. Rust core, Python SDK. pip install state-harness GitHub - the0cp/pico: A small, compact, register-based scripting language and virtual machine implemented in C. Inspired by clox. GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. GitHub - pileax-ai/pileax: PileaX is an all-in-one AI knowledge base system. 🍀 GitHub - DO-SAY-GO/freelang: I love freelang GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland
FlexInference: Drop your AI costs today
FlexInference · 2026-07-07 · via Hacker News

A router that drops cost by 45%

We find cheaper inference within the SLA you provide.

−45%costs

input + output / 1M

−34%latency

time to first token

// median cost and latency across 22,138 requests

Some of the ways you can drop your costs by half

  • Gemini Image Classification

    +20.2% Latency-38.9% Cost

  • OpenAI Deep Research

    -30.1% Latency-44.8% Cost

  • OpenAI Browser Agent

    +9.7% Latency-51.5% Cost

Built so you aren't debugging at 2am

  • 3 ms routing across 300 cities.

    When you make a request it is fulfilled using edge compute. We use Cloudflare workers which is deployed in 300+ cities. Our own routing adds 1-5 milliseconds on a cold start.

  • Your prompts are never stored or read.

    Your prompt and its reply just pass through. We never store them and we never read them. Your provider key is envelope-encrypted with AES-256-GCM and locked to your exact org and provider slot, so it decrypts nowhere else, and a key that will not decrypt is treated as missing. So your production key is safe to drop in.

  • Your discounts, credits, and API tiers stay yours.

    You bring your own key, so the provider bills you directly. We also don't have a fee per request. By bringing your own key you also get the benefits of using your credits, discounts, negotiated rates, and API tiers.

  • We fail fast and loud.

    If you send a wrong parameter we don't quietly strip it to force a success. When the provider rejects a request, we send the status code and the error back to you. That way you can debug based on your intentions and not find a bug weeks after launching.

  • Your existing client works unchanged.

    Point the base URL at us and pass your key. We work with OpenAI, Anthropic, and Gemini clients. To use FlexInference you just need to add a single new field, start_within. If you decide to use our Python or TypeScript SDKs then you also get the benefits of strict types.

  • Errors your agent can fix itself.

    Every error comes back in the shape of the SDK you called, so your client parses it unchanged. It carries a machine-readable code, the exact fix, and a doc_url. Our MCP server goes further. Coding tools like Claude and Cursor can search the docs, look up any error code, and manage your keys over OAuth. So your agents fix their own mistakes, and you are not stuck babysitting them.

Only pay when we save you money.

Base price

5.5% per request$0per request

  • No fee per standard request
  • No sales call
  • No seats or tiers
  • No trial

Routing is free. When you give us an SLA to find a cheaper way to run your request, we hunt for cheaper inference. If we find you cheaper inference, we keep 20% of what we save you. If we cannot find anything cheaper, we escalate your request to a standard request and it is once again free.

If we take a $10 request down to $5, we take a $1 fee and you pay $6. That is 40% savings.

No new SDK semantics to learn.

Point your base URL at FlexInference, pass your FlexInference key, and add start_within to start saving.

curl https://api.flexinference.com/v1/responses \
  -H "Authorization: Bearer flex_live_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.2",
    "input": "Summarize this thread.",
    "start_within": "00h-00m-30s"
  }'

Why I built FlexInference?

Hey! My name is Adi.

I grew up wanting to help others. Actually, that's not always been true. I went through a tough time in college and it made me want to help people. Help people get educated. Help people get better healthcare. Help people with the basics, you know? At first I tried to do that directly. Made a few free websites/tools/videos. And it definitely helped. But, not the scale I wanted.

Then AI hit an inflection point. It started making education better because it could answer a billion questions happily. It made healthcare better by personalizing it to everyone. I realized something then. I hadn't built the models, but I could help make them easier to reach. The labs want the same thing. They're working to make inference cheaper, faster, and open to everyone. The problem is all the layers sitting on top. They slow it down, mark it up, and gate who gets in.

So I made FlexInference as a crack at that. It makes getting to these providers cheaper, faster, and more reliable. The first way it does that is by hunting for cheaper inference, and escalating to standard when it can't find any. By making it cheaper to work with these models, ideally we can get more ideas out there. More incredible products. And as such, help more people.

That's why I created FlexInference. Hopefully, it helps you all.

Sincerely,
Adi

P.S. It's free for all requests. And I only really charge when I save you money, so it's profitable to use FlexInference.

P.P.S I tried to make it super easy to use FlexInference. It's strictly typed. It also won't quietly fall back to parameters you never set just to force a request through. You also don't need to use the FlexInference SDK and can just change the URL and key instead.

Connect your agent and let it do the rest.

FlexInference runs an MCP server, so Claude, Cursor, and any other MCP client can check your usage, manage your API keys, and search the docs without leaving your editor. Point it at https://mcp.flexinference.com/mcp and you are set.

The docs tools need no login. Anything that touches your account runs over OAuth and asks first, so a client reads your usage or changes a key only after you approve it. It never runs inference and never takes a raw provider key.

See the MCP guide ->

Frequently Asked Questions