AMD acquires Taalas to embed AI model weights directly into custom silicon chips. The startup's hardware achieves speeds exceeding 16,000 tokens per second for Llama 3.1-8B by hard-coding models into silicon, though it sacrifices flexibility. This acquisition signals a major shift toward application-specific hardware as chipmakers race to drastically accelerate LLM inference performance.
Opening Kapyn…