Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

CommonLID Update: New Tools, Growing Impact

Дата публикации: 16-06-2026 00:00:00

CommonLID, a community-built language ID benchmark, has a new website and interactive leaderboard. Its paper was accepted to ACL 2026, with a poster session on 7 July. Source code, a PyPI package, and the dataset are now available.

Основное содержимое страницы с новостью.

Since our previous post announcing CommonLID, a new community-built language identification benchmark, the project has continued to grow. We’ve been focusing on two areas: improving usability and spreading the word.

A screenshot of the CommonLID website

A screenshot of the CommonLID Website

On the usability side, we now have a dedicated website for CommonLID: commonlid.org. You can find all the information and resources related to it here, and it’s where we’ll post any future updates. We’ve also made it easier to explore the state of the art for language identification through an interactive leaderboard. You can use it to compare performance on CommonLID by model and by language on a variety of metrics, all in the browser. If you’re looking to run your own evaluations, we’ve provided the source code and a PyPI package, which replicates the analysis in the paper and can be extended to other models and datasets.

A screenshot of the CommonLID leaderboard featuring 9 different LID systems

A screenshot of the CommonLID leaderboard featuring 9 different LID systems

As for spreading the word about CommonLID, in April we found out that the paper has been accepted to the main conference at ACL 2026 in San Diego! ACL is one of the top conferences in natural language processing, and we’re looking forward to connecting with the community there. We’ll be presenting our work on the 7th of July in poster session G. Aside from the paper, Common Crawl team members have also given multiple talks featuring CommonLID. These include presentations for EleutherAI, Cohere Labs (on YouTube), the Mozilla Data Collective and others.

For more information, check out the CommonLID CommonLID Paper and dataset, now available both on Hugging Face and through the Mozilla Data Collective. We want to keep improving this resource, so if you’d like to contribute, please raise an issue on the Hugging Face repo or get in touch via Discord.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data014.9710-02-2026
2Common Crawl Joins Project Tapestry07.2327-07-2026
3Common Crawl Foundation at LREC 2026012.1530-06-2026
413th Web-as-Corpus Workshop @ EMNLP 2026018.9529-06-2026
5Announcing the First Stable Release of CC-Downloader012.7810-08-2026
6You can now build directly on Common Crawl from the browser06.6906-05-2026
7Three papers accepted to the International Conference on Computational Linguistics01002-11-2012
8Summer Conference Success for Oxford Computational Linguistics Group09.729-05-2013
9From single pull requests to full software packages: Detecting malicious code at scale08.802-06-2026
10Papers accepted to Natural Language Processing conference01018-09-2013

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 13.35. Источник: commoncrawl.org.