TorontoAMD acquired Taalas, which etches weights into silicon.
Inference is memory-bound — GPUs burn energy moving weights from memory each token — and hard-wiring the weights deletes that traffic entirely.
A mask set costs millions and takes months, favoring small, stable models over fast-moving frontier ones.
Sources: Hacker News