AMD Bought Taalas, Whose Chips Burn Model Weights Straight Into Transistors
AMD agreed to acquire Toronto-based Taalas, whose chips hardwire an AI model's weights permanently into silicon to eliminate the memory reads that cap GPU inference speed. Taalas claims 17,000 tokens per second on Llama 3.1 8B at a tenth of an H200's power.