
OpenAIは、AstraがPreparedness Frameworkの枠組みでサイバーセキュリティの臨界的能力閾値を達成した同社の初のモデルであると伝えた。モデルのリリースにはより厳格なセキュリティ対策が予定されている。
この発表は、モデルのサイバーセキュリティ機能の拡張とリリースの厳格な制限を結びつけようとする試みと関連している。ただし、出典はOpenAI Newsの要約のみで、評価の方法、試験結果、保護措置のリストなどの詳細は提供されていない。
次の重要なサインは、閾値の基準と具体的なsafeguardsの説明の公開になる。現時点で、評価がどのようなシナリオのサイバーセキュリティを対象としているか、独立した情報源で結果が検証されているかは不明。
編集部コメント
なぜ重要か
If the statement is accompanied by a detailed methodology and independent verification, it could serve as a reference for discussions on the safe release of models with advanced cyber capabilities. The closest observable signal is the publication of criteria for assessment and the list of safeguards. Significant uncertainty remains due to a single source and lack of full text.