Changing the reasoning steps doesn’t change the answer.

Researchers can slip a model a hidden hint whose influence never appears in its stated reasoning, yet still shapes the answer.

Interpretability work shows genuine multi-step computation on some tasks, so which is which remains undecided.

Sources: Quanta Magazine