Skip to main content
Live
Main content

Meta's next MTIA inference chip taps Broadcom and TSMC N2

The third-gen accelerator skips Nvidia entirely for Llama-scale inference serving.

Jaeden Schafer
Editor in Chief · · 5 min read

Meta's third-generation MTIA inference chip has taped out on TSMC's N2 process node and is being designed in partnership with Broadcom, according to two people familiar with the architecture. The chip is explicitly built for Llama-scale model inference and is targeted at ad-ranking and recommendation workloads, effectively removing Meta from the Nvidia GPU market for a significant chunk of inference load.

Meta's Nvidia spending in 2026 is forecast to exceed $30 billion, nearly half of which goes to inference serving rather than training. MTIA 3 is designed to absorb approximately 40% of that inference volume by late 2027, eliminating roughly $12 billion in annual Nvidia costs. The shift is economically pure for Meta: inference silicon has lower margins than training silicon, and hyperscalers have the volume and design expertise to build it in-house.

Broadcom's role is critical. Meta is leveraging Broadcom's chiplet design expertise and standard-cell IP to accelerate MTIA 3's time-to-market, avoiding the lengthy internal development cycles that delayed MTIA 2. The design employs 2.5D chiplet packaging with a passive interposer, similar to Broadcom's flagship data center designs. This approach trades some cost per unit for speed and design risk reduction.

Key facts

  • 01Meta. A key thread of reporting in this story.
  • 02Broadcom. A key thread of reporting in this story.
  • 03MTIA. A key thread of reporting in this story.

MTIA 3 is not a general-purpose GPU. It is a purpose-built inference engine for Llama models at specific batch sizes and latency targets. This focus allows the chip to be simpler and cheaper than trying to match Nvidia's broad software ecosystem. Meta is betting that vertical specialization—a chip built explicitly for ad ranking—beats horizontal generalization at the cost scale of a $100 billion company.

Meta's MTIA 3 is taped out on TSMC N2 and designed to handle roughly 40% of ad-ranking inference by late 2027, reducing Meta's Nvidia bill by an estimated $12 billion annually.
Jaeden Schafer

The broader implication is that hyperscalers are finally building credible silicon alternatives. Training is still Nvidia's fortress—the software stack, the network effects, and the debugging tooling create formidable barriers. But inference is becoming a hyperscaler game. Meta, Google, and Amazon are all shipping custom silicon for inference, effectively marking the end of Nvidia's monopoly on AI hardware outside of training.

Related · from this week
Meta starts production of MTIA AI chips in September to cut Nvidia dependence
Jaeden Schafer · 5 min read →
ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Related topics
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from News

Meta logo
Models

Meta starts production of MTIA AI chips in September to cut Nvidia dependence

The social giant is spending up to $145B this year on AI compute and wants its own silicon to blunt GPU costs.

Jaeden Schafer5 min read
Meta logo
News

Meta launches Muse, a personal AI agent built on a security-first pitch

The Superintelligence Labs product ships on iOS, Android, and WhatsApp with Stripe checkout and a $300,000 bug bounty ceiling.

Jaeden Schafer5 min read
Meta logo
News

Meta launches AI Mode on Facebook, pulling answers from public posts and Groups

Facebook's new AI Mode summarizes public posts, Groups, and Reels in response to plain-language questions — and joins a growing AI subscription push.

Jaeden Schafer4 min read