本文へ移動
10/11(日) 00:28
あとで読む
出典
研究The Decoder· 3時間前

Epoch AI、エージェント研究能力を評価

The DecoderManuel Uthこの話題 1ソース →
要

AI要約

3行で
  • Epoch AI がベンチマーク InnovationEval を公開しました。
  • Claude Fable 5 と GPT-5.6 Sol を試しました。
  • 両モデルとも人間の参考手法に届きませんでした。

AI要約です。詳しくは出典へ。誤りを報告

典

出典

1ソース
The DecoderAI agents overstate their results and remain far from autonomous research, study findsthe-decoder.com
出典を読む
この話題を見る