MTIA, short for Meta Training and Inference Accelerator, is the family of custom chips Meta designs to run artificial intelligence workloads in its own data centres. It is an application-specific integrated circuit (ASIC): a processor built for a narrow set of jobs rather than for general use. Its first job was inference - running already-trained models - for the ranking and recommendation systems that decide which posts, Reels and advertisements each of Meta's users sees. Meta built it because, in its own words in May 2023, graphics processing units (GPUs) "were not always optimal" for those workloads at its scale. The chips are not sold to anyone else.
How the chip is built
The 2023 design, now called MTIA 100, set the template. Manufactured by TSMC on a 7-nanometre process, it ran at 800 MHz within a 25-watt thermal design power (TDP). At its centre sat a grid of 64 processing elements (PEs) arranged eight by eight. Each PE contained two cores based on RISC-V, the open instruction set architecture, one of them with a vector extension, plus fixed-function units for matrix multiplication, accumulation, data movement and non-linear functions. Each PE had 128 KB of local memory, backed by 128 MB of shared on-chip SRAM and off-chip LPDDR5 memory. Peak throughput was 102.4 trillion operations per second (TOPS) at INT8 precision and 51.2 teraflops at FP16. The chips sat on M.2 boards connected over PCIe Gen4, with 12 accelerators per server.
The second generation, published on April 10, 2024 and since renamed MTIA 200, kept the eight-by-eight grid but moved to TSMC's 5-nanometre process, a 1.35 GHz clock and a 90-watt TDP. Dense INT8 compute rose to 354 TOPS, on-chip SRAM doubled to 256 MB and LPDDR5 capacity reached 128 GB. A rack held 72 accelerators. Meta reported three times the performance of the first chip across four key models, and six times the model-serving throughput at platform level with 1.5 times better performance per watt.
Low-power memory was a deliberate choice. Recommendation models lean on enormous embedding tables - lookups that map users and items into numerical vectors - and need capacity more than the extreme bandwidth that large language models demand. That calculation changed with generative AI. The roadmap Meta published on March 11, 2026 switched to high-bandwidth memory (HBM) and a modular chiplet design, in which separate compute, network and input-output dies can be upgraded independently and built on different process nodes. Each PE still carries two RISC-V vector cores, now alongside a dot-product engine, a special-function unit, a reduction engine and a DMA (direct memory access) engine.
Software is the other half. MTIA plugs into PyTorch, the framework Meta created, through torch.compile and a compiler backend called Triton-MTIA, and a plugin for vLLM, an open-source inference server, swaps GPU kernels such as FlashAttention for MTIA versions. According to Meta, production models can move between GPUs and MTIA without chip-specific rewrites.
Where it sits in the ad system
Every ad Meta serves passes through a funnel. A retrieval stage narrows tens of millions of candidate ads to a few thousand; ranking models then score those survivors for the likelihood of a click or a conversion. Both stages run inference on every page load, billions of times a day, under tight latency limits.
The clearest documented link is Andromeda, Meta's ad retrieval engine. When Meta unveiled it on December 2, 2024, it said the system ran on Nvidia's Grace Hopper Superchip and on MTIA, and delivered a 10,000-fold increase in model capacity, a 6% gain in recall and an 8% gain in ad quality on selected segments. Meta expected MTIA and future GPUs to support roughly 1,000 times more model complexity. In January 2026, Meta told investors Andromeda had been extended to run on Nvidia, AMD and MTIA chips, which together with model changes nearly tripled its compute efficiency.
The heavier models are a different matter. GEM, Meta's generative ads recommendation model, is trained on "thousands of GPUs", according to Meta's engineering blog of November 10, 2025, which does not mention MTIA. Its knowledge is transferred into hundreds of smaller models that serve live traffic. Meta has not published which chips serve each of those models, so the precise share of ad decisions running on MTIA is unknown.
Origin and evolution
Meta's custom silicon effort started badly. In 2022, according to Reuters, Meta scrapped the rollout of an earlier in-house chip after it missed internal targets, ordered Nvidia GPUs instead and redesigned several data centres to suit them. MTIA 100 was presented publicly on May 18, 2023.
The second generation followed in April 2024 and reached production in 16 data centre regions in less than nine months from first silicon, according to Meta. In March 2025, Reuters reported that Meta had begun testing its first chip designed for training, made by TSMC. Later that year Meta agreed to buy Rivos, a Santa Clara start-up designing RISC-V processors; terms were not disclosed, although Rivos had reportedly sought funding at a valuation above $2 billion.
The March 2026 roadmap named four new chips. MTIA 300, in production for ranking and recommendation training, draws 800 watts with 216 GB of HBM. MTIA 400, the first aimed at general generative AI work, delivers 12 petaflops at MX4 precision within 1,200 watts and links 72 chips in one scale-up domain. MTIA 450 and MTIA 500, built primarily for generative AI inference, are scheduled for mass deployment in early 2027 and later in 2027. From the 300 to the 500, Meta cited a 4.5-fold increase in HBM bandwidth and a 25-fold increase in compute. It also claimed a cadence of roughly one new chip every six months, against the one-to-two-year cycle common in the industry, and said it had deployed hundreds of thousands of MTIA chips.
Why it matters for marketers
The economics of Meta's ad business increasingly depend on compute. Meta spent $72.22 billion on capital expenditure in 2025. Its 2026 guidance began at $115 billion to $135 billion in January, was raised to $125 billion to $145 billion in April and was narrowed to $130 billion to $145 billion in July, when second-quarter capex reached $31.08 billion. Advertising produced about 97.6% of second-quarter revenue.
Custom silicon is Meta's attempt to make each ranking decision cheaper. If inference costs fall, Meta can run larger models on every auction, which it links directly to performance: Lattice and GEM improvements produced a more than 6% conversion rate gain for landing page view ads in the first quarter of 2026. Cheaper inference also underpins the automation in Advantage+, where retrieval and ranking replace manual targeting. Advertisers never select MTIA, yet its cost curve shapes how much modelling sits behind each impression they buy.
Limitations and disputes
Independent evidence is thin. Meta publishes peak specifications and relative gains against its own earlier chips, but no third-party benchmarks such as MLPerf results. Its claims against GPUs are qualitative: MTIA 400 is "competitive with leading commercial products", and at the Hot Chips conference in August 2026 Meta said MTIA 300 "achieves parity with GPUs at a competitive TCO", or total cost of ownership, without publishing dollar figures. Even specification figures vary: Meta's March 2026 table gave MTIA 400 9.2 TB/s of HBM bandwidth, while ServeTheHome's report of the Hot Chips presentation gave 9.4 TB/s.
Dependence on outside suppliers has grown, not shrunk. On February 17, 2026, Meta agreed a multiyear deal for millions of Nvidia Blackwell and Rubin GPUs. A week later AMD announced an agreement for up to 6 gigawatts of Instinct accelerators, with AMD granting Meta a warrant for up to 160 million shares. The Information reported on February 26, 2026 that Meta would also rent Google's Tensor Processing Units through Google Cloud. In the same month the Financial Times reported "technical challenges" with Meta's MTIA training chips. In June 2026, The Information reported that more than a quarter of former Rivos staff had been laid off and that work on a chip for training Meta's largest models had been paused. Meta has not publicly addressed those reports.
Not the same as
GPU. A general-purpose parallel processor, dominated by Nvidia, that handles both training and inference for almost any model. MTIA trades that flexibility for efficiency on Meta's own workloads.
Google TPU. The Tensor Processing Unit is Google's custom AI ASIC, announced in May 2016. Unlike MTIA, it is rented to outside customers through Google Cloud.
AWS Trainium and Inferentia. Amazon Web Services' chips for training and inference respectively, announced in 2020 and 2018, sold as cloud capacity.
Microsoft Maia. Microsoft's AI accelerator, first presented as Maia 100 in November 2023 for Azure workloads.
No other widely used meaning of MTIA appears in advertising or ad tech.
Recent developments
On April 14, 2026, Broadcom and Meta extended their partnership through 2029. Meta committed to more than 1 gigawatt of MTIA capacity initially, the first chip in the deal is described as an industry-first 2-nanometre design, and Broadcom chief executive Hock Tan left Meta's board to become an adviser on its silicon roadmap. Mark Zuckerberg said the work would support "personal superintelligence" for billions of people. On July 9, 2026, Reuters reported an internal memo showing Meta would start producing its latest chip, reportedly codenamed Iris, in September, with Broadcom on design and TSMC on manufacturing, while planning 7 gigawatts of compute in 2026.
PPC Land reported today that Santosh Janardhan, Meta's head of infrastructure, put 2026 infrastructure spending at "well over $100 billion" and described separate chips for ranking, large language models, training and inference.
Timeline
- 2022 - Meta scraps the rollout of an earlier in-house inference chip after it misses internal targets and turns to Nvidia GPUs, according to Reuters.
- May 18, 2023 - Meta presents MTIA v1, a 7nm, 25-watt inference chip with 64 RISC-V-based processing elements.
- April 10, 2024 - Meta presents the second-generation MTIA on TSMC 5nm, at 90 watts.
- December 2, 2024 - Meta unveils Andromeda, its ad retrieval engine running on Nvidia Grace Hopper and MTIA.
- March 11, 2025 - Reuters reports Meta has begun testing its first in-house training chip.
- Late September 2025 - Reuters and Bloomberg report that Meta will acquire RISC-V chip start-up Rivos.
- January 28, 2026 - Meta says Andromeda now runs on Nvidia, AMD and MTIA chips and guides 2026 capex to $115 billion to $135 billion.
- February 17, 2026 - Meta signs a multiyear deal for millions of Nvidia GPUs.
- February 24, 2026 - AMD and Meta announce a deal for up to 6 gigawatts of AMD accelerators.
- February 26, 2026 - The Information reports a multibillion-dollar Meta deal to rent Google TPUs.
- March 11, 2026 - Meta details MTIA 300, 400, 450 and 500 and renames earlier chips MTIA 100 and 200.
- April 14, 2026 - Broadcom and Meta extend their custom silicon partnership through 2029, starting with more than 1 gigawatt.
- April 29, 2026 - Meta raises 2026 capex guidance to $125 billion to $145 billion.
- June 2026 - The Information reports layoffs among former Rivos staff and a paused training chip.
- July 9, 2026 - Reuters reports Meta will begin producing its latest chip in September 2026.
- August 25, 2026 - Meta presents MTIA 300 and 400 at the Hot Chips conference.
Related PPC Land coverage
- Meta unveils advanced AI infrastructure with new custom chips - The April 2024 second-generation MTIA announcement.
- Meta's AI automation draws skepticism from advertisers despite performance claims - Andromeda's hardware and Meta's expectation that MTIA will support far larger models.
- Meta's ad business hits record $58.1 billion as AI drives conversion gains - Fourth quarter 2025 results, with Andromeda extended to Nvidia, AMD and MTIA chips.
- Meta Q1 2026: $56.3B revenue as AI tools double advertiser adoption - Raised capex guidance and more than 1 gigawatt of Broadcom custom silicon.
- Meta profit drops 8% to $15.8bn as legal charges hit ad gains - Second quarter 2026 results and the narrowed capex range.
- Meta's infrastructure chief puts 2026 AI spending at over $100bn - Santosh Janardhan on Meta's chip strategy and suppliers.
- Explaining Advantage+ - The automated campaign suite that depends on Andromeda and GEM.
- Meta reports 26% revenue growth amid infrastructure spending surge - Third quarter 2025 results, Andromeda's 14% ad quality gain and rising compute needs.
- Meta bets billions on AI data centers requiring massive energy and water - The Prometheus and Hyperion clusters that house Meta's compute.
Summary
Who. Meta designs MTIA with Broadcom as co-development partner and TSMC as manufacturer. Meta's infrastructure organisation operates it; advertisers use it only indirectly.
What. A family of custom AI accelerators, from MTIA 100 and 200 to the planned MTIA 300, 400, 450 and 500, built around RISC-V-based processing elements and tuned for ranking, recommendation and, increasingly, generative AI inference.
When. First presented on May 18, 2023, with a second generation in April 2024, a four-chip roadmap in March 2026 and the latest chips scheduled for mass deployment through 2027.
Where. Inside Meta's own data centres, serving systems such as the Andromeda ad retrieval engine alongside Nvidia, AMD and rented Google hardware.
Why. To cut the cost and power of running AI models billions of times a day. Meta's results against GPUs remain largely self-reported, and its growing purchases from Nvidia, AMD and Google show MTIA has not removed its reliance on external chips.
Discussion