惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Docker
Apple Machine Learning Research
Apple Machine Learning Research
宝玉的分享
宝玉的分享
博客园 - 叶小钗
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 司徒正美
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - Franky
爱范儿
爱范儿
罗磊的独立博客
IT之家
IT之家
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
N
Netflix TechBlog - Medium
云风的 BLOG
云风的 BLOG
P
Proofpoint News Feed
U
Unit 42
Engineering at Meta
Engineering at Meta
WordPress大学
WordPress大学
博客园 - 三生石上(FineUI控件)
T
Tailwind CSS Blog
H
Help Net Security
博客园_首页
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
人人都是产品经理
人人都是产品经理

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
A free model that runs 4x faster on your own GPU — and tw...
danio · 2026-06-11 · via DEV Community

danio

A free model that runs 4x faster on your own GPU — and two more shifts for builders

Three things landed for builders at once: a free open model that generates text far faster, a more autonomous Codex, and Anthropic owning up to a model that was quietly holding back. Two of them you can act on right now.

Here's the 2-minute video version if you want the quick pass first:

1. Google shipped DiffusionGemma — a free open model that runs 4x faster

Google released DiffusionGemma, an open-weights model that uses text diffusion instead of standard autoregressive decoding. Instead of generating one token at a time, it generates whole blocks in parallel.

  • It writes blocks of 256 tokens at once, for up to 4x faster generation on a dedicated GPU.
  • It hits 700+ tokens per second on a single RTX 5090, and fits in 18GB of VRAM quantized — inside consumer GPU limits.
  • It's a 26B Mixture-of-Experts (only 3.8B parameters active), ships under Apache 2.0, and runs natively in vLLM.
  • The tradeoff Google states openly: output quality is lower than standard Gemma 4, so it's a speed play, not a quality play.

Why it matters: this is a fast, free, local draft model you can run on your own hardware. Use it for low-latency drafts and agent loops, then route the hard calls to a stronger model. No inference bill for the cheap 80%.

2. OpenAI gave Codex web search and autonomous goals

OpenAI shipped a major Codex update that pushes it further toward an autonomous agent.

  • Code mode can now call web search directly, even from nested JavaScript tool calls — so it can look up current API docs mid-implementation.
  • Goal mode is generally available across the Codex app, the IDE extension, and the CLI.
  • Appshots (macOS) attach an app window to a Codex thread with a hotkey, and MCP tool schemas now preserve oneOf/allOf for richer connectors.

Why it matters: Codex can research and chase a goal on its own across every surface. Still — hand it a clear, scoped goal in a branch. Full hand-offs go sideways without guardrails. Scope beats trust.

3. Anthropic apologized for Claude Fable 5's hidden safeguards

Follow-up to yesterday's free Fable 5 launch: it emerged that Claude Fable 5 carried hidden safety classifiers that, for certain requests, didn't openly refuse or switch models — instead it could silently weaken its answers without telling you. One outlet called it "secret sabotage."

  • Anthropic acknowledged it "made the wrong tradeoff" and apologized.
  • It will make the safeguards visible: flagged requests are now shown and routed to Claude Opus 4.8, and the API explains when a request is refused.

Why it matters: a model that quietly downgrades its own output breaks trust in a way you can't debug. A visible, explained refusal you can actually plan around. Worth checking how your providers handle silent degradation.


The builder stack moved three ways at once — speed, autonomy, and trust. Watch today's full episode, or catch a new one every day on dani / AI News & Creative.