错位悬赏：众包收集AI智能体异常行为 (Misalignment Bounty: Crowdsourcing AI Agent Misbehavior)

Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents pursuing unintended or unsafe goals. The bounty received 295 submissions, of which nine were awarded. This report explains the program's motivation and evaluation criteria, and walks through the nine winning submissions step by step.

翻译：先进AI系统有时会表现出与人类意图不符的行为。为收集清晰、可复现的实例，我们启动了'错位悬赏'项目：通过众包方式征集智能体追求非预期或不安全目标的案例。该项目共收到295份提交，其中9份获得奖励。本报告阐述了项目的设立动机与评估标准，并逐步解析了九个获奖案例。

相关内容

关注 0

人工智能杂志AI(Artificial Intelligence)是目前公认的发表该领域最新研究成果的主要国际论坛。该期刊欢迎有关AI广泛方面的论文，这些论文构成了整个领域的进步，也欢迎介绍人工智能应用的论文，但重点应该放在新的和新颖的人工智能方法如何提高应用领域的性能，而不是介绍传统人工智能方法的另一个应用。关于应用的论文应该描述一个原则性的解决方案，强调其新颖性，并对正在开发的人工智能技术进行深入的评估。官网地址：http://dblp.uni-trier.de/db/journals/ai/

【AAAI2026】无限叙事：免训练的角色一致性文生图技术

专知会员服务

8+阅读 · 11月18日

【Emtiyaz Khan】自适应人工智能的贝叶斯学习规则，The Bayesian Learning Rule for Adaptive AI

专知会员服务

26+阅读 · 2022年3月11日

KG-BERT：基于BERT的知识图谱补全，KG-BERT: BERT for Knowledge Graph Completion

专知会员服务

195+阅读 · 2020年5月31日

【Google无监督大规模视觉表示迁移】Large Scale Learning of General Visual Representations for Transfer

专知会员服务

12+阅读 · 2020年1月7日