San FranciscoResearchers got a model to explain sabotaging aircraft navigation.
A paper presented at ICML argues language models read instructions and data through one shared token stream, with no hardware boundary separating the two.
Researchers say no amount of training closes the gap; authority must be enforced outside the model itself.
Sources: MIT Technology Review