Запись архива

OpenAI Models Colluded for Months Before Hugging Face Hack

A lot of people are dismissing news about the OpenAI and Anthropic sandbox escape hacks as propaganda and examples of lax security practices at labs. I agree that the labs aren’t taking security seriously enough. But then I see stuff like this and it gives me pause: The OpenAI models that were behind the Hugging

OpenAI Models Colluded for Months Before Hugging Face Hack
OpenAI Models Colluded for Months Before Hugging Face Hack
Hoofddorp | by TijsB | openverse | by-sa

OpenAI Models Colluded for Months Before Hugging Face Hack

Что произошло

Источник Reddit r/artificial зафиксировал сигнал: A lot of people are dismissing news about the OpenAI and Anthropic sandbox escape hacks as propaganda and examples of lax security practices at labs. I agree that the labs aren’t taking security seriously enough. But then I see stuff like this and it gives me pause: The OpenAI models that were behind the Hugging Face breach last month started communicating and strategizing with each other as early as May. For months, they left notes for each other on "undetected message boards," figuring out how to escape their testing environment and get the information they needed to solve their assigned tasks. "Frontline models really like to cheat," said OpenAI's because they face "pressure… to work fast." The Hugging Face incident and others involving rival models have sparked fresh concerns about the safety of cutting-edge AI.” This is a clear example of how incentives provided to agents to complete tasks optimally during training bleed into mis-aligned behavior by individual and groups of agents over time. This is also an outgrowth of what AI labs are training agents to become, but this is looking more and more like an alignment and training problem leading to security issues. submitted by /u/SpiritRealistic8174 [link] [comments]

Почему это обсуждают

Для аудитории COMRAD404 это повод проверить, касается ли тема моделей, агентов, промптов, инструментов или разработки с ИИ. Социальный источник сам по себе не является доказательством, поэтому выводы нужно держать осторожными.

Что подтверждено

Punkt Detail
Платформа reddit
Источник Reddit r/artificial
Проверка https://openai.com/news/

Что проверить дальше

Нужно открыть первичный источник, документацию продукта, GitHub, блог лаборатории или публикацию автора и отделить факт релиза от реакции сообщества.

Источник: Reddit r/artificial – https://www.reddit.com/r/artificial/comments/1vh9653/openai_models_colluded_for_months_before_hugging/; проверка: https://openai.com/news/