惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
F
Fortinet All Blogs
量子位
G
Google Developers Blog
J
Java Code Geeks
N
Netflix TechBlog - Medium
博客园 - 聂微东
宝玉的分享
宝玉的分享
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
月光博客
月光博客
The Cloudflare Blog
Apple Machine Learning Research
Apple Machine Learning Research
爱范儿
爱范儿
雷峰网
雷峰网
M
MIT News - Artificial intelligence
T
Tailwind CSS Blog
V
Visual Studio Blog
阮一峰的网络日志
阮一峰的网络日志
博客园 - 三生石上(FineUI控件)
Microsoft Azure Blog
Microsoft Azure Blog
aimingoo的专栏
aimingoo的专栏
Martin Fowler
Martin Fowler
有赞技术团队
有赞技术团队
T
The Blog of Author Tim Ferriss

NVIDIA Blog

Cute Critters Come to the Cloud: ‘Aniimo’ Launches on GeForce NOW NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers University of Manchester Uses NVIDIA Earth-2 to Forecast Air Pollution Across the UK ‘Now We Can Know Everything and Do Anything,' Jensen Huang Says at Dreamforce From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories Heart of the Matter: How a Major Children’s Hospital Uses Open Source NVIDIA AI for Cardiac Care Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026 ‘NBA 2K27’ With NVIDIA DLSS 5 Leads 26 New Games Coming to GeForce NOW NVIDIA to Acquire Hugging Face NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026 NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark How XPUs Meet a World-Class AI Factory With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents Bring the Fire: Play Games on GeForce NOW With New Firefox Browser Support Securing the Infrastructure of Intelligence Universitas Gadjah Mada, Indosat and NVIDIA Open Indonesia’s First University AI Center to Develop Local AI Talent Class Is in Session: GeForce NOW Levels Up Linux, Chromebooks and More NVIDIA CEO Tops Glassdoor’s 2026 List of Best CEOs NVIDIA AI Factory Compute Is Becoming an Investable Asset Class
NVIDIA and Local AI Community Fuel Open Source Models and...
NVIDIA Writers · 2026-08-11 · via NVIDIA Blog

The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. 

Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools emerging across the ecosystem. That includes NVIDIA’s latest open models, software and developer tools, plus the accelerated computing, libraries and educational resources that help users get started.

It’s shaping up to be a big month for agents. Follow along for the latest developments in this special-edition NVIDIA Local AI blog series, with new updates added over the coming weeks.

Follow NVIDIA RTX Spark on X, Instagram, TikTok and Facebook — and stay informed by subscribing to the RTX AI PC newsletter. Follow NVIDIA Workstation on LinkedIn and X


Tuesday, Aug. 11, 6:00 a.m. PT 🔗

NVIDIA Introduces Nemotron 3.5 Lightning for Fast, Specialized Agentic Tasks

Today, NVIDIA expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, a customizable open 30B mixture-of-experts (MoE) model for always-on agents. 

Nemotron 3.5 Lightning delivers up to 4x faster token generation and 30% faster time to completion compared to open models in its class. 

And because Nemotron 3.5 Lightning is open weights, AI enthusiasts and developers can fine-tune it with their own examples to better match specific tasks, interests and workflows. For example, they could train the model to: 

  • Write in a preferred style: Follow established tones, formats and terminology when writing emails, reports or other documents.
  • Learn a specialty: Better understand the language and common tasks associated with areas such as photography, gaming or 3D design.
  • Code a certain way: Follow preferred coding conventions, frameworks and testing approaches when writing, reviewing or refactoring code.

Paired with access to apps, files and other tools, these fine-tuned models can power more personalized local agentic AI experiences — from an assistant that helps manage email and calendars, to a smart-home agent that handles everyday routines, to a coding companion that works alongside developers on a local codebase.

NVIDIA collaborated with vLLM, Ollama, llama.cpp and LM Studio to provide the best local deployment experience for Nemotron 3.5 Lightning models — offering developers choice of NVFP4 and GGUF format of models. Unsloth also provides day-one support with optimized and quantized models for efficient local deployment via Unsloth Studio. 

Nemotron 3.5 Lightning runs locally on NVIDIA RTX PCs, NVIDIA DGX Spark and OEM GB10 systems, and NVIDIA Jetson, and scales up to RTX PRO workstations, NVIDIA DGX Station and GB300 deskside systems, data centers and cloud environments. With NVIDIA Blackwell systems available from Acer, ASUS, Dell Technologies, Exxact, GIGABYTE, HP, Lenovo, MSI and Supermicro, users can choose from a wide range of devices and form factors to fit their needs.

As generative AI adoption grows, enterprises are looking for ways to keep rising token costs in check without sacrificing access to frontier intelligence. NVIDIA NeMo Switchyard, an open source routing library, automatically directs each step of an agent workflow to the best-fit model based on accuracy, speed and cost. It also gives developers the flexibility to work across models and providers for different tasks. 

Internal benchmarks show that NeMo Switchyard, by routing each step across a system of models, helped maintain frontier-level task completion while reducing benchmark completion cost to roughly one-third of Opus 4.8 alone. NeMo Switchyard is available on GitHub.

Visit the Nemotron 3.5 Lightning, NeMo Switchyard and Jetson AI technical blogs to get started. And to build at the edge, start with Jetson AI Lab tutorials and discover real-world Jetson projects. Nemotron 3.5 Lightning is also available through OpenRouter, on build.nvidia.com as an NVIDIA NIM microservice, and through a broad ecosystem of NVIDIA Cloud Partners, post-training platforms, inference platforms and cloud service providers. NeMo Switchyard is available on GitHub.