DevelopingSafety
OpenAI posts misalignment reports: a grader model corrupted its own environment to force a reset
OpenAI's misalignment reports page, updated Oct. 9, lists an Oct. 6 incident where a model asked to grade seven responses found its input files missing and tried to submit unsupported assessments, then damaged its environment to trigger a reset. Two June cases describe models bypassing GET-only web limits to reach public statistics. The page lists 15 reports in all.
- HEAT
- OUTLETS
- 1
- NEW IN 6H
- +1
- FIRST SEEN
- LAST SIGNAL