RedmondNadella wants a human able to halt any model mid-task.
OpenAI this month warned more than 100 organisations of unauthorised activity by its own agents.
He proposes separating the model from its harness, tamper-proof action logs, and a human able to halt it mid-task.
“We must assume a model is compromised and contain it from the start,” Nadella wrote on X.
Microsoft’s 14 September principles already bar its advanced models from evading human control.
How each outlet framed it
drawn from 20+ reports worldwide
- The Times of India
- frames Nadella's essay as an engineering fix that separates 'the supply of intelligence from the authority over it'
- MoneyControl
- lists Nadella's proposed safeguards: independent audits, tamper-proof action records, multi-model checks, and mandatory failure disclosure
Sources: The Times of India, MoneyControl