Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Lin Li and Yarin Gal awarded $371k grant to improve safety of LLMs

Дата публикации: 08-05-2026 11:00:00

Research Associate Lin Li and Professor Yarin Gal from the Oxford Applied and Theoretical Machine Learning Group have won a $371,000 (£274,000) grant from US philanthropy organisation Coefficient Giving to improve the safety of LLMs.

Основное содержимое страницы с новостью.

Posted: 8th May 2026

Research Associate Lin Li and Professor Yarin Gal from the Oxford Applied and Theoretical Machine Learning Group have won a $371,000 (£274,000) grant from US philanthropy organisation Coefficient Giving to improve the safety of LLMs.

The problem

Despite recent progress, LLMs can still be manipulated by adversarial or unexpected prompts to bypass safeguards. This presents serious risks: bad actors may try to use these LLMs to carry out harmful activities that would otherwise be beyond their capabilities, such as using them to craft high-grade explosives. A promising safety guardrail is adversarial training in the model’s latent space, but models are often trained on scenarios that would not occur in practice, which can limit effectiveness and lead to reduced general performance, including overly cautious behaviour and excessive refusals.

The approach

The project advances current adversarial training methods by introducing new regularisation techniques that ensure adversarial examples in the latent space reflect realistic, achievable inputs rather than artificial internal states. Existing approaches typically restrict their search to small variations of known harmful prompts, which keeps training close to familiar examples but limits the range of scenarios the model encounters. By instead focusing on ‘reachable’ representations, the new approach enables a broader exploration of adversarial inputs that could arise in practice. This allows the model to learn from more diverse and previously unseen attack patterns, resulting in stronger safety behaviour while preserving LLM performance.

The team

Research Associate Lin Li and Professor Yarin Gal will work with Stephen Casper from MIT, who has authored pioneering work on latent adversarial training. Lin has been studying adversarial training since completing his PhD and has developed several state-of-the-art adversarial training methods. Yarin brings extensive experience in developing impactful and widely adopted and cited research, including work on semantic entropy, model collapse, and Bayesian Neural Networks.

The funding will support research staff, dedicated compute infrastructure, large-scale evaluation, and dissemination through international conferences. The project is running from March 2026 to July 2027.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Oxford researchers awarded ARIA funding to develop safety-first AI for business011.0210-04-2025
2Responsible AI UK funds work on Responsible Innovation led by University of Oxford team05.2308-11-2023
3Oxford Computer Science student Lia Yeh awarded Google PhD Fellowship07.0216-10-2023
4Department of Computer Science professor secures €1.95m European research grant for quantum project06.2323-11-2023
5New G-Research scholarships in machine learning013.7704-03-2025
6Oxford computer science researcher awarded prestigious grant with potential to advance national security06.1226-10-2023
7Oxford Professors awarded prestigious UKRI Turing AI World-Leading Fellowships021.1114-06-2023
8Research Director of Frontier AI Taskforce05.7520-09-2023
9Professor Elias Koutsoupias awarded ERC Advanced Grant09.1323-06-2026
10Tim Muller to research when to trust online reviews05.6309-11-2017

Классификация: Наука. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 7.97. Источник: www.cs.ox.ac.uk.