Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Kaggle is making AI benchmark creation effortless

Дата публикации: 04-06-2026 16:00:00

Today, we’re launching local development for Kaggle Benchmarks.

Основное содержимое страницы с новостью.

Your browser does not support the audio element.

Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental

As AI models evolve from simple chatbots into reasoning agents that write code, use tools and solve complex problems, traditional benchmarks are no longer enough. The community needs dynamic, rigorous evaluations — built by the people who use these models in the real-world.

That’s why we launched Kaggle Benchmarks. Since then, the global AI community has created more than 10,000 evaluation tasks, creating the trustworthy, transparent public leaderboards that help labs measure and accelerate AI progress.

Today, we are taking the next step by launching local development for Kaggle Benchmarks.

Use Kaggle Benchmarks from your local development environment

Until now, creating evaluation tasks meant working exclusively in Kaggle's web-based notebook editor, instead of developers’ preferred stack to build with.

Our new update enables developers to create, validate, push, run and download tasks directly from their local development environments like Antigravity, VSCode, Cursor and coding agents. This update is designed to meet developers where they work, making the journey from idea to evaluation faster and more intuitive.

Build evaluation tasks in natural language with AI coding agents

Local development also unlocks a powerful new workflow: using AI coding agents to write benchmark tasks through the write-kaggle-benchmarks skill. This skill comprises a set of structured instructions that teaches a coding agent how to build tasks using the kaggle-benchmarks SDK and the Kaggle CLI.

To add this skill to your agent, simply ask your agent to:

Once installed, you can describe an evaluation in plain language and get a working task on Kaggle. For example, you can tell your agent:

These powerful capabilities are driven by the new commands that we have built for Benchmarks in the Kaggle CLI.

Understand why community-driven evaluations matter

We built Kaggle Benchmarks to democratize trustworthy AI evaluations. We believe that if a capability can be measured, labs will race to improve it. By providing these clear, objective signals, our hope is to empower AI labs to drive model improvements in the areas that matter most.

For AI to truly benefit humanity, evaluations must reflect the full diversity of real-world challenges. We believe this launch is a significant step toward enabling anyone, anywhere, to build the evaluations that will shape the future of AI.

Ready to build? Try Kaggle Benchmarks today.

Get the latest news from Google in your inbox

Sign up for our newsletters with product updates, event information, special offers, and more.

Your information will be used in accordance with Google's privacy policy. You may opt out at any time.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Join the new AI Agents Vibe Coding Course from Google and Kaggle09.0227-04-2026
2Inside our 353,000-person vibe coding course015.0903-08-2026
3Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency024.3605-06-2026
4Local AI Weekly #1: It's Happening011.5711-08-2026
5The latest AI news we announced in May 2026018.705-06-2026
6Qt Toolkit To Introduce Edge AI Submodule, Initially For Vision AI With Qt04.0810-08-2026
7The latest AI news we announced in April 20260504-05-2026
8The latest AI news we announced in March 20260501-04-2026
9Turn one giant AI-generated pull request to a reviewable stack09.3404-08-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 9. Тональность: 0. Информативность: 14.65. Источник: blog.google.