本文へ移動研究·UK AI Security Institute Blog一次·· 7/23(木) AISI、監視モデルの手法を公開
UK AI Security Institute Blog要
AI要約
3行で- ・AISI が制御用監視モデルのレッドチーム手法を公開しました。
- ・Google DeepMind と Anthropic の監視で弱点を確認しました。
- ・進化的探索で疑いスコア 3 を達成しました。
AI要約です。詳しくは出典へ。誤りを報告
典
出典
1ソース · 一次情報
UK AI Security Institute Blog一次How our Control Red Team is stress-testing frontier monitorsaisi.gov.uk 出典を読む