"While AISI said the AI models displayed “novel, potentially deceptive behaviours”, the watchdog cautioned that its findings should be interpreted with care, as they occurred under “specific conditions”, including with some of the models’ safeguards disabled.
“We cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario; our analysis so far presents a mixed picture and is ongoing,” AISI said."
Framing: assertive ·
First seen: 2026-08-05 ·
Last seen: 2026-08-05 ·
Spread: 1 articles