AMD Buys Taalas to Etch AI Models Into Silicon
AMD has agreed to buy Toronto's Taalas, which hardcodes a model's weights into the chip itself. Its HC1 test part runs Llama 3.1 8B at roughly 17,000 tokens per second — and can never run anything else.
AMD said on August 6 that it has agreed to acquire Taalas, a Toronto startup doing something most chip designers deliberately avoid: printing a specific model's trained weights permanently into the silicon. There is no memory to load from, because the model is the circuit. The result is very fast and completely inflexible.
Founded in 2023 and out of stealth only since February 2026, Taalas has raised roughly $219 million, including a $169 million round in February led by Quiet Capital with Fidelity and investor Pierre Lamond participating. Its CEO, Ljubisa Bajic, previously ran Tenstorrent, making this his second attempt at a non-GPU answer to inference.
What the HC1 is
The demonstrator part is not small. Taalas built it on TSMC's 6nm process at an 815 square millimetre die carrying 53 billion transistors, and it serves Meta's Llama 3.1 8B at up to roughly 17,000 tokens per second per user. The company's comparative claims are more aggressive still — it has cited figures ranging from tens of times faster than Nvidia parts to 73x an H200 at a tenth the power — but those are vendor benchmarks against a single 8B model, and no independent evaluation has been published.
- Process: TSMC 6nm
- Die size: 815 mm², 53 billion transistors
- Model: Llama 3.1 8B, hardcoded
- Throughput: up to ~17,000 tokens/sec per user
- Next part: HC2, targeting models around 20 billion parameters
The trade AMD is making
Hardwiring weights removes the memory wall — the constant shuttling between HBM and compute that dominates the power and latency budget of a conventional accelerator. What it removes in exchange is the ability to change your mind. A chip is one model, forever. Taalas softens this somewhat: it says changing weights and matrix dimensions touches only two mask layers rather than requiring a fresh design, which keeps respin costs closer to a mask change than a new tapeout. The harder constraint is scale, since larger models can spill past a single reticle and need multiple chips stitched together.
That economics only works for stable, high-volume inference — a widely deployed model with a long baseline of steady traffic, not the frontier model that gets replaced in four months. AMD's plan is to sell it as a complement rather than a replacement: the Toronto team joins AMD's Artificial Intelligence Group, and the chips are slated to sit alongside Instinct GPUs in Helios racks with Epyc CPUs, programmed through ROCm. Vamsi Boppana, AMD's SVP for AI, framed the goal as matching the right compute to each workload. Terms were not disclosed, and the deal is subject to customary closing conditions and regulatory approvals. AMD shares rose about 1.5% on the news.
The strategic timing is legible enough. AMD spent this summer buying its way into inference volume it does not get from GPUs alone, including a $5 billion, two-gigawatt arrangement with Anthropic. Google is reported to be pursuing the same hardwired-silicon idea for Gemini, and OpenAI has its own custom part in Jalapeño. The open question is whether any model stays still long enough to be worth etching — a bet on model stability as much as on silicon.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.