Meta's third-generation MTIA inference chip has taped out on TSMC's N2 process node and is being designed in partnership with Broadcom, according to two people familiar with the architecture. The chip is explicitly built for Llama-scale model inference and is targeted at ad-ranking and recommendation workloads, effectively removing Meta from the Nvidia GPU market for a significant chunk of inference load.
Meta's Nvidia spending in 2026 is forecast to exceed $30 billion, nearly half of which goes to inference serving rather than training. MTIA 3 is designed to absorb approximately 40% of that inference volume by late 2027, eliminating roughly $12 billion in annual Nvidia costs. The shift is economically pure for Meta: inference silicon has lower margins than training silicon, and hyperscalers have the volume and design expertise to build it in-house.
Broadcom's role is critical. Meta is leveraging Broadcom's chiplet design expertise and standard-cell IP to accelerate MTIA 3's time-to-market, avoiding the lengthy internal development cycles that delayed MTIA 2. The design employs 2.5D chiplet packaging with a passive interposer, similar to Broadcom's flagship data center designs. This approach trades some cost per unit for speed and design risk reduction.
Key facts
- 01Meta. A key thread of reporting in this story.
- 02Broadcom. A key thread of reporting in this story.
- 03MTIA. A key thread of reporting in this story.
MTIA 3 is not a general-purpose GPU. It is a purpose-built inference engine for Llama models at specific batch sizes and latency targets. This focus allows the chip to be simpler and cheaper than trying to match Nvidia's broad software ecosystem. Meta is betting that vertical specialization—a chip built explicitly for ad ranking—beats horizontal generalization at the cost scale of a $100 billion company.
“Meta's MTIA 3 is taped out on TSMC N2 and designed to handle roughly 40% of ad-ranking inference by late 2027, reducing Meta's Nvidia bill by an estimated $12 billion annually.”— Jaeden Schafer
The broader implication is that hyperscalers are finally building credible silicon alternatives. Training is still Nvidia's fortress—the software stack, the network effects, and the debugging tooling create formidable barriers. But inference is becoming a hyperscaler game. Meta, Google, and Amazon are all shipping custom silicon for inference, effectively marking the end of Nvidia's monopoly on AI hardware outside of training.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.


