AMD acquired Taalas, which etches weights into silicon.

Inference is memory-bound — GPUs burn energy moving weights from memory each token — and hard-wiring the weights deletes that traffic entirely.

A mask set costs millions and takes months, favoring small, stable models over fast-moving frontier ones.

Sources: Hacker News