Models Move Into Inference Silicon
AMD acquired Taalas, whose model-specific integrated circuits encode model weights in silicon and claim inference rates up to 17,000 tokens per second. The mechanism could create a specialized hardware tier between general accelerators and fixed application ASICs, especially for high-volume agent workloads. It remains a weak signal because the evidence is one acquisition and early demonstrations; independent benchmarks, customer deployments or additional model-specific chips would confirm the category.