惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 三生石上(FineUI控件)
U
Unit 42
人人都是产品经理
人人都是产品经理
罗磊的独立博客
Recent Announcements
Recent Announcements
云风的 BLOG
云风的 BLOG
aimingoo的专栏
aimingoo的专栏
T
Tailwind CSS Blog
GbyAI
GbyAI
Blog — PlanetScale
Blog — PlanetScale
I
InfoQ
Last Week in AI
Last Week in AI
宝玉的分享
宝玉的分享
B
Blog RSS Feed
WordPress大学
WordPress大学
腾讯CDC
H
Help Net Security
博客园 - Franky
博客园 - 【当耐特】
博客园 - 聂微东
Stack Overflow Blog
Stack Overflow Blog
B
Blog
Vercel News
Vercel News
博客园 - 司徒正美

Developer tools

Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe DevFest is back The latest AI news we announced in August 2026 Pairing Google Antigravity with Gemini 3.7 Flash solves notable multi-agent math and engineering problems. Gemini Omni 1.1 Flash lets you build with more control How developers build AI for good with Gemma 4 Inside the Gemmaverse: Celebrating one billion Gemma downloads Inside our 353,000-person vibe coding course Introducing Gemini Robotics ER 2 Gemini API Managed Agents: 3.6 Flash, hooks, and more We're rolling out AlphaEvolve widely to solve Google Cloud customers' hardest problems. Expanding Managed Agents in Gemini API: background tasks, remote MCP and more The latest AI news we announced in June 2026 Ask an AI expert: What exactly is the full stack? Interactions API: our primary interface for Gemini models and agents DiffusionGemma: 4x faster text generation See what 3 builders are making with Gemma 4 Bringing the latest Gemini models to Apple developers Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency Introducing Gemma 4 12B: a unified, encoder-free multimodal model How we used Gemini to build Google I/O 2026 Take our I/O 2026 quiz, vibe coded in Google AI Studio. Here's what developers can do with the latest Google Play updates. Building the agentic future: Developer highlights from I/O 2026 I/O 2026 Introducing Managed Agents in the Gemini API Bring any idea to life: Google AI Studio at I/O 2026 Gemini API File Search is now multimodal: build efficient, verifiable RAG Accelerating Gemma 4: faster inference with multi-token prediction drafters The latest AI news we announced in April 2026
Kaggle is making AI benchmark creation effortless
Nicholas Kang · 2026-06-05 · via Developer tools

Your browser does not support the audio element.

Listen to article

This content is generated by Google AI. Generative AI is experimental

[[duration]] minutes

As AI models evolve from simple chatbots into reasoning agents that write code, use tools and solve complex problems, traditional benchmarks are no longer enough. The community needs dynamic, rigorous evaluations — built by the people who use these models in the real-world.

That’s why we launched Kaggle Benchmarks. Since then, the global AI community has created more than 10,000 evaluation tasks, creating the trustworthy, transparent public leaderboards that help labs measure and accelerate AI progress.

Today, we are taking the next step by launching local development for Kaggle Benchmarks.

Use Kaggle Benchmarks from your local development environment

Until now, creating evaluation tasks meant working exclusively in Kaggle's web-based notebook editor, instead of developers’ preferred stack to build with.

Our new update enables developers to create, validate, push, run and download tasks directly from their local development environments like Antigravity, VSCode, Cursor and coding agents. This update is designed to meet developers where they work, making the journey from idea to evaluation faster and more intuitive.

Build evaluation tasks in natural language with AI coding agents

Local development also unlocks a powerful new workflow: using AI coding agents to write benchmark tasks through the write-kaggle-benchmarks skill. This skill comprises a set of structured instructions that teaches a coding agent how to build tasks using the kaggle-benchmarks SDK and the Kaggle CLI.

To add this skill to your agent, simply ask your agent to:

Once installed, you can describe an evaluation in plain language and get a working task on Kaggle. For example, you can tell your agent:

These powerful capabilities are driven by the new commands that we have built for Benchmarks in the Kaggle CLI.

Understand why community-driven evaluations matter

We built Kaggle Benchmarks to democratize trustworthy AI evaluations. We believe that if a capability can be measured, labs will race to improve it. By providing these clear, objective signals, our hope is to empower AI labs to drive model improvements in the areas that matter most.

For AI to truly benefit humanity, evaluations must reflect the full diversity of real-world challenges. We believe this launch is a significant step toward enabling anyone, anywhere, to build the evaluations that will shape the future of AI.

Ready to build? Try Kaggle Benchmarks today.

Get more stories from Google in your inbox.

Done. Just one step more.

Check your inbox to confirm your subscription.

You are already subscribed to our newsletter.

You can also subscribe with a