惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
云风的 BLOG
云风的 BLOG
Microsoft Azure Blog
Microsoft Azure Blog
F
Fortinet All Blogs
A
About on SuperTechFans
月光博客
月光博客
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 司徒正美
P
Proofpoint News Feed
D
Docker
Jina AI
Jina AI
Apple Machine Learning Research
Apple Machine Learning Research
The Cloudflare Blog
I
InfoQ
Recorded Future
Recorded Future
爱范儿
爱范儿
Last Week in AI
Last Week in AI
J
Java Code Geeks
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
M
MIT News - Artificial intelligence
L
LINUX DO - 热门话题
腾讯CDC
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
IT之家
IT之家
博客园_首页
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
C
CXSECURITY Database RSS Feed - CXSecurity.com
L
Lohrmann on Cybersecurity
The Last Watchdog
The Last Watchdog
V
Visual Studio Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
P
Privacy International News Feed
博客园 - 三生石上(FineUI控件)
Schneier on Security
Schneier on Security
Simon Willison's Weblog
Simon Willison's Weblog
Martin Fowler
Martin Fowler
雷峰网
雷峰网
Latest news
Latest news
Scott Helme
Scott Helme
T
Tenable Blog
Vercel News
Vercel News
宝玉的分享
宝玉的分享
PCI Perspectives
PCI Perspectives
Help Net Security
Help Net Security
L
LINUX DO - 最新话题
Attack and Defense Labs
Attack and Defense Labs
Spread Privacy
Spread Privacy
量子位
H
Heimdal Security Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Mistral Medium 3.5 Review: A 128B Open-Weight Model With a Coding Agent That Opens PRs For You
Marcus Rowe · 2026-05-03 · via DEV Community

TL;DR: Mistral Medium 3.5 is a 128B open-weight model released April 29, 2026, with a 256K context window, configurable reasoning, and native multimodal input. It scores 77.6% on SWE-Bench Verified — close but not ahead of Claude Sonnet 4.6 — and ships alongside Vibe, a cloud coding agent that submits pull requests directly to GitHub without you babysitting it. API pricing is $1.50 per million input tokens. Open weights under a modified MIT license that routes high-revenue enterprises through Mistral's paid channel. It's a serious model with a real killer feature. But the pricing math is complicated.


77.6%. That's Mistral Medium 3.5's score on SWE-Bench Verified.

For context: Claude Sonnet 4.6 scores 79.6% on the same benchmark. So Medium 3.5 is close — very close — but it doesn't actually beat the current best. And it costs $1.50 per million input tokens on Mistral's API, while cheaper alternatives like Qwen 3.6 (27B, free to self-host under Apache 2.0) are nipping at its heels with 72.4% on the same benchmark at a fraction of the model size.

So this is either a strong open-weight option for developers who need local control, or an overpriced mid-tier API play, depending entirely on what you're trying to do. I'll tell you which one it is for you — but you need to actually read the answer, because "it depends" is doing real work here.


What Is Mistral Medium 3.5?

Mistral is a Paris-based AI lab founded in 2023, and they've been steadily building out their model lineup: Small, Medium, Large, with various numbered releases layered on top. Medium 3.5 is their April 2026 flagship — released April 29, 2026.

What's notable about the architecture is that Medium 3.5 consolidates three previously separate Mistral models into a single unified set of weights. You used to have to choose between Mistral Medium 3.1 (instruction-following), Magistral (reasoning), and Devstral 2 (coding). Now it's one model with a toggle.

128B parameters. All dense — meaning all 128 billion parameters activate on every token, not a mixture-of-experts setup where only a fraction fire. Dense models are generally more predictable in output quality, which matters if you're running production workloads.

256K context window. That's a big number. Larger than Claude Sonnet 4.6 (200K) and twice GPT-4o (128K). Whether you'll actually saturate a 256K window in practice is a separate question, but having the headroom is nice when you're feeding it a whole codebase.

Configurable reasoning. You can toggle between fast reply mode and deep reasoning mode per request. This is Mistral's version of extended thinking — the model applies more test-time compute to harder problems when you ask it to. It worked well for Magistral, and the integration here is smoother than flipping between two separate models.

Multimodal input. Text and images in, text out. The vision encoder was trained from scratch to handle variable image sizes and aspect ratios, which is a technical choice that often shows up as better handling of screenshots and diagrams compared to models that bolt on a fixed-size vision module.


The Vibe Remote Agent — This Is the Actually Interesting Part

OK, so the benchmark numbers are solid-but-not-exceptional. What actually got my attention when I read the launch announcement was Vibe.

Vibe is Mistral's cloud-based coding agent. And the way it works is fundamentally different from local coding assistants.

Standard coding agents — whether it's Cursor, GitHub Copilot, or a local Claude-powered setup — require you to be in the loop. The model suggests, you approve, the model writes, you review, repeat. It's faster than doing it yourself, but you're still the bottleneck. You have to babysit the session.

Vibe's remote agents change that model. You describe the task, hand it to Vibe, and it runs in an isolated cloud sandbox. You can walk away. It can handle long-running tasks — the kind that take hours — without needing you to keep a terminal open. When it's done, it doesn't just output code to a terminal. It opens a pull request on your GitHub repo, ready for review.

I want to be clear about what that means: you can queue up three coding tasks before bed, wake up the next morning, and have three open PRs waiting. You look at the diffs, approve what's good, request changes on what isn't. The edit-review loop compresses dramatically.

The integrations are solid too. Native GitHub support is table stakes. But Vibe also connects to Linear, Jira, Sentry, Slack, and Microsoft Teams. So in theory you could wire it to your issue tracker and have it automatically pick up tickets, implement them, and open PRs. That's not theoretical — it's the designed workflow.

One thing worth noting: Vibe requires explicit approval for any "sensitive actions," meaning it surfaces what it's about to do before doing it. That's the right design. Autonomous agents that go quiet and then surprise you are a support ticket waiting to happen.


The Benchmark Reality Check

Here's the honest picture of where Medium 3.5 stands.

SWE-Bench Verified (coding, real GitHub issues):

  • Mistral Medium 3.5: 77.6%
  • Claude Sonnet 4.6: 79.6%
  • Qwen 3.6 (27B): 72.4%

τ³-Telecom (agentic tool-use): 91.4% — a strong number, and this benchmark specifically tests the kind of multi-step tool-calling that Vibe relies on. So the model is well-suited to its headline use case.

What's notably absent from Mistral's launch materials: MMLU, GPQA, AIME, HumanEval, MATH. These are the standard general knowledge and reasoning benchmarks that every other frontier model publishes on release. Mistral hasn't published them. I don't know why.

That gap is actually meaningful. If you're using this model for coding and agentic tasks, the published SWE-Bench number is the one you care about. But if you're trying to evaluate Medium 3.5 for general knowledge queries, document analysis, or reasoning over complex text — you're flying without instruments. The predecessor, Medium 3, scored 92.1% on HumanEval and 57.1% on GPQA Diamond, but those numbers don't carry over automatically to a new architecture.

Developers on Hacker News noticed. The launch thread had a notably critical tone, with one developer summarizing it as: "their new flagship model is basically 'not the best' on any benchmark, yet costs multiple times more than most competitors." Harsh, but not unfair based on the published data.


Running It Locally: The 4-GPU Reality

The "runs on 4 GPUs" headline is accurate. But let's talk about which 4 GPUs.

The recommended production setup is 4× NVIDIA H100 80GB, running the model in FP8 precision. That gives you ~320GB of VRAM total, with FP8 cutting the model weight footprint to roughly 128GB, leaving headroom for the KV cache at the 256K context window.

H100 80GB GPUs cost about $30,000 each new, and rent for roughly $2-3/hour per GPU on cloud providers. So 4-GPU production inference on rented hardware runs $8-12/hour. At that compute cost, the API at $1.50 per million tokens starts looking attractive unless you're doing very high volume.

For teams that already own the hardware — inference labs, larger enterprise ML teams, research groups — self-hosting at this scale is very much a real option. Mistral officially supports vLLM and SGLang for production inference, both of which have well-developed deployment paths for models this size.

If you want to run it on something more accessible: Q4-quantized versions drop the VRAM requirement to roughly 70GB, which is approaching Mac Studio territory (128GB unified memory, around $3,500 at current pricing). You'll pay a quality penalty versus full precision, but for dev/test workloads it's workable.

For inference optimization, Mistral offers EAGLE speculative decoding via a separate draft head model that adds about 4GB of overhead but delivers 1.41× output throughput and 29% lower latency. Worth enabling in production if you're squeezing performance.


License: Open-Weight, With a Catch for Big Companies

The modified MIT license Mistral ships with isn't quite as clean as the Apache 2.0 on earlier models. The core terms: free to use for individuals, startups, mid-market companies, and most universities. No restrictions on commercial applications within that population.

The catch: companies above a revenue threshold need to negotiate a separate commercial arrangement with Mistral. The threshold isn't published explicitly, but the structure creates a two-tier situation — the model is effectively open for everyone except large enterprise.

For most developers reading this, that's a non-issue. If you're solo or at a startup, the modified MIT terms are plenty permissive. If you're at a Fortune 500, talk to a lawyer before deploying at scale.


How It Stacks Up Against the Competition

Let me just run through the practical comparisons.

vs. Claude Sonnet 4.6: Sonnet 4.6 wins on SWE-Bench (79.6% vs 77.6%) and has a comparable context window (200K vs 256K). Claude is proprietary API-only — you can't self-host it. If local deployment or open weights matter to you, Medium 3.5 is the obvious choice. If API-only is fine and you want the highest coding benchmark, Sonnet edges it out. Pricing: Medium 3.5 at $1.50/million input vs. Sonnet 4.6 at $3.00/million — half the cost.

vs. GPT-4o: Medium 3.5 has a larger context window (256K vs 128K) and is ~40% cheaper on input tokens. GPT-4o has better published general-knowledge benchmark coverage, and it's API-only. Medium 3.5 wins on the open-weight dimension.

vs. Qwen 3.6 (27B): This is the uncomfortable comparison. Qwen 3.6 is 27B parameters — roughly 1/5 the size — and still hits 72.4% on SWE-Bench Verified, under Apache 2.0, for free. Medium 3.5 beats it by 5.2 points on that benchmark. Whether 5.2 points of SWE-Bench is worth $1.50/million tokens vs. $0 depends on your workload. For a high-volume API use case, the math probably favors Qwen. For local enterprise deployment where you want a single consolidated model with reasoning mode built in — Medium 3.5 has a legitimate case.

vs. Kimi K2.6: Moonshot's K2.6 (our most recent open-weight review) is 1 trillion total parameters in a MoE architecture with 58.6% SWE-Bench Pro — a harder benchmark than Verified. Kimi K2.6 starts at $0.60 per million input tokens and can run 300-agent swarms. Medium 3.5 is considerably smaller (128B dense), more accessible to self-host, and has the Vibe async agent as a practical differentiator. Different use case: K2.6 for maximum agentic scale, Medium 3.5 for single-model unified deployment with coding agents.


Who Should Actually Use This

Self-hosting teams with H100s already in their stack. This is Medium 3.5's strongest use case. If you've got the hardware, you get a consolidated model (no more managing three separate Mistral deployments), a 256K window, configurable reasoning, and multimodal input. The open-weight terms let you fine-tune it on proprietary data. That's a legitimate enterprise play.

Developers who want Vibe's async coding agents. The "queue tasks, get PRs in the morning" workflow is genuinely compelling, and it's built on Medium 3.5. If Vibe's cloud agent pipeline fits how you work, you're getting a capable model as the backend.

Cost-sensitive API users coming from GPT-4o or Claude. At $1.50 input vs. $3.00 (Sonnet) or $2.50 (GPT-4o), there's real savings potential if you don't need the absolute top-of-benchmark model for your task. The performance delta is small for most practical applications.

Who should pass: If you're choosing a pure API model with no interest in self-hosting, and you need the best coding benchmark performance, Sonnet 4.6 still edges it at 79.6% SWE-Bench. If you need maximum open-weight model for the lowest possible cost, Qwen 3.6's Apache 2.0 license and 72.4% benchmark score make it very hard to justify the price gap.


Bottom Line

Mistral Medium 3.5 is a genuinely capable model with a smart consolidation story — three specialized models folded into one dense 128B architecture with a toggle for reasoning. The Vibe async coding agent is the most differentiated thing they shipped alongside it, and the "PR while you sleep" workflow is real and useful.

The friction points are real too. The benchmark coverage gaps matter if you're trying to assess general-purpose performance. The pricing faces legitimate pressure from efficient smaller models. And the enterprise license terms are worth reading carefully before you commit.

But for a European lab staying in the frontier race — open weights, self-hostable on accessible hardware, with a cloud agent product that integrates into actual developer workflows — Mistral Medium 3.5 is a meaningful release. It's not the best model on any individual benchmark. It might be the best practical open-weight model for teams that need the whole package.

Verification sources for this article: Mistral AI official announcement, HuggingFace model page, Mistral official documentation at docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04.