Inference is the act of running a trained machine learning model on new input to produce an output. A language model answering a question, an image model drawing a product shot and a bidding model pricing an impression are all performing inference. The term exists to separate this from training, the earlier and far more expensive process in which a model learns its parameters from data. Training happens once, or periodically. Inference happens every time the model is used, which is why its cost now sets the economics of almost every AI feature in advertising.

How a model produces an answer

During inference the model's weights are fixed. Input is converted into numbers, passed through the network's layers in a single forward pass, and an output emerges. For a classifier, such as a model estimating predicted click-through rate (pCTR), one pass yields one probability. For a large language model (LLM), the process repeats.

LLM inference runs in two phases. Prefill processes the whole prompt at once, measured in tokens, the subword units models read and bill by. Decode then generates the reply one token at a time, each new token requiring another pass. To avoid recomputing the prompt on every step, servers store intermediate results in a key-value (KV) cache held in the accelerator's high-bandwidth memory. Memory, more than raw arithmetic, often caps how many users a chip can serve.

Operators judge inference on a handful of numbers. Time to first token measures how long a user waits before anything appears. Tokens per second measures generation speed. Throughput measures how many requests a system completes in aggregate. These pull against each other. Batching, grouping many requests on one chip, raises throughput and lowers cost per token but can lengthen individual waits. A 2023 paper by Woosuk Kwon and colleagues at UC Berkeley introduced PagedAttention, which manages KV cache memory the way an operating system manages pages; according to the authors, their vLLM system improved throughput two to four times at the same latency.

Hardware splits by workload. Graphics processing units (GPUs) from Nvidia handle most inference in data centres, alongside custom chips such as Google's tensor processing units (TPUs), Amazon's Inferentia and Trainium, and Meta's Training and Inference Accelerator (MTIA). Smaller models also run on-device, on phones and laptops, trading capability for privacy and zero serving cost to the provider.

Inference inside the ad auction

Programmatic advertising ran inference at scale long before chatbots. Every real-time bidding (RTB) request asks a demand-side platform (DSP) to score an impression and return a price before a deadline. The OpenRTB field tmax carries that limit in milliseconds, and Google's Authorized Buyers deadlines range from 80 to 1,000 milliseconds, as the PPC Land explainer on latency sets out. Index Exchange's chief executive has described real-time bidding as working within a 100-millisecond window. Network time, auction logic and the model's own forward pass all share that budget, which limits how large a bidding model can be.

Platforms work around the ceiling with cascades. Meta's Andromeda system, according to a December 2024 Meta engineering post, narrows tens of millions of candidate ads to a few thousand before heavier ranking models choose the winner. Meta reported in January 2026 that Andromeda runs on Nvidia, AMD and its own MTIA chips, nearly tripling compute efficiency. Its Adaptive Ranking Model goes further, routing each ad request to a heavier or lighter model depending on estimated conversion odds, a direct trade between inference cost and outcome that Meta credited with a 1.6% rise in conversion rates.

The sell side has started moving models closer to the auction. PubMatic launched Decision Fabric in June 2026, letting partners run their own models inside its auction with responses under 10 milliseconds. Bedrock became the first DSP to run its bidder inside an exchange, Index Exchange, in April 2026, within a container window under five milliseconds. Amazon Web Services built a dedicated network for bidding workloads in October 2025, promising single-digit millisecond latency.

Origin and evolution

The word comes from logic and statistics, where inference means drawing conclusions from evidence. Machine learning borrowed it for applying a learned model, a usage that became standard in the 2010s.

Specialised hardware made the distinction commercial. Google deployed its first TPU in its data centres in 2015 and disclosed it at Google I/O in May 2016; the chip was built for inference, not training, and a 2017 paper led by Norman Jouppi reported it ran 15 to 30 times faster than contemporary CPUs and GPUs on Google's production workloads. MLPerf, an industry benchmark, published its first inference results on November 6, 2019, covering 44 systems from 14 organisations.

Generative models changed the scale. Per-token API pricing arrived with GPT-3 in June 2020, and ChatGPT's launch in November 2022 turned inference into a mass consumer workload. In September 2024 OpenAI released o1, a model that improved answers by spending more computation at inference time, generating internal reasoning before replying. That shifted part of the industry's scaling effort from training to serving. Google unveiled Ironwood, its seventh-generation TPU and its first designed specifically for inference, on April 9, 2025. "This is what we call the 'age of inference'," wrote Amin Vahdat, Google's vice president for machine learning, systems and cloud AI. Nvidia agreed on December 24, 2025 to license inference chip technology from Groq in a deal reported at about $20 billion, hiring founder Jonathan Ross. In April 2026 Google split its eighth TPU generation into the TPU 8t for training and the TPU 8i for inference.

Why it matters for marketers

Inference is the line item behind every AI feature an advertiser touches, from automated bidding to generated creative to agentic buying tools. Prices have fallen steeply. According to Stanford's 2025 AI Index, querying a model performing at GPT-3.5's level cost $20 per million tokens in November 2022 and $0.07 by October 2024. Alphabet said Gemini serving costs fell 78% during 2025, and Anthropic priced Haiku 5.5 this month 90% below its predecessor, at $0.10 per million input tokens and $0.50 per million output.

Volume has risen faster than unit prices have fallen. Alphabet reported 3.2 quadrillion tokens a month in June 2026, and Sundar Pichai said Google is "supply-constrained" against capital expenditure that has grown from about $30 billion to roughly $180 billion. Serving costs also help explain why OpenAI introduced advertising in ChatGPT, which passed a $1 billion run rate on August 31, 2026.

Agencies feel it directly. Some now meter AI agents by the token, and a KPMG survey found 49% of leaders had cut back agent rollouts when costs outran value.

Inference as a privacy term

In data protection law the word carries a second, older weight. An inference is a conclusion drawn about a person, such as an interest, health condition or sexual orientation deduced from behaviour. The Court of Justice of the European Union (CJEU) held in case C-184/20 on August 1, 2022 that data indirectly revealing sensitive traits falls under Article 9 of the General Data Protection Regulation (GDPR), the rule covering special category data. Inferred interest segments are a basic input to targeting, so the ruling reaches into audience building. Oxford academics Sandra Wachter and Brent Mittelstadt argued in 2019 for a "right to reasonable inferences", contending that existing law protects the data going in better than the conclusions coming out. The European Commission's Digital Omnibus proposal of November 19, 2025 would narrow Article 9 for inferred data, a change the European Data Protection Board and Supervisor rejected in February 2026.

Limitations and disputes

Cost figures are hard to audit. Providers publish list prices, not margins, and reasoning models consume hidden tokens billed at output rates. One widely cited estimate holds that serving GPT-5 requires more than 200,000 GPUs running continuously, a figure attributed to a vendor investigation rather than OpenAI. Falling per-token prices can coexist with rising bills, because agents and reasoning multiply consumption per task.

Energy and concentration draw criticism. The International Energy Agency projects data centre consumption near 945 terawatt hours a year by 2030. UBS estimates OpenAI and Anthropic could account for more than 48% of Google Cloud revenue in 2027, making inference demand a concentrated bet. On-device inference raises its own disputes: Chrome was found to install a 4 GB Gemini Nano model without explicit notice.

In auctions, latency forces compromise. Smaller, faster models fit inside tmax but predict less precisely, and the models that score impressions are rarely open to buyer inspection.

Not the same as

Training builds the model by adjusting its weights against data; inference uses those weights unchanged. Fine-tuning is additional training on a narrower dataset and still changes weights. Prediction is the output of inference for many model types, such as a pCTR score, but inference also covers generation, ranking and classification. Statistical and causal inference are older uses, drawing conclusions about populations or effects from data; marketing mix models such as Google's Meridian perform Bayesian inference in that sense, which has nothing to do with serving a neural network.

Recent developments

Price competition has accelerated through 2026. Google released Gemini 3.6 Flash in July at $1.50 per million input tokens and $7.50 output, using 17% fewer output tokens than its predecessor, an efficiency gain that cuts inference bills without a list price change. Anthropic's Haiku 5.5 followed on October 7. Hardware is diverging too, with inference-specific silicon from Google, Nvidia's Groq licence and Meta's MTIA all aimed at lowering the cost per answer rather than the cost of training. In advertising, models are moving inside exchanges, closer to the impressions they score.

Timeline

  • 2015: Google deploys its first tensor processing unit, built for inference, in its data centres.
  • May 2016: Google discloses the TPU at Google I/O.
  • April 2017: Norman Jouppi and colleagues publish the TPU performance paper, reporting 15 to 30 times speed gains on inference.
  • November 6, 2019: MLPerf publishes its first inference benchmark results, covering 44 systems from 14 organisations.
  • June 2020: OpenAI opens the GPT-3 API with per-token pricing.
  • August 1, 2022: CJEU rules in C-184/20 that data indirectly revealing sensitive traits falls under GDPR Article 9.
  • November 2022: ChatGPT launches, making LLM inference a mass consumer workload.
  • September 12, 2023: Kwon and colleagues post the PagedAttention and vLLM paper.
  • September 2024: OpenAI releases o1, scaling computation at inference time.
  • December 2, 2024: Meta details its Andromeda ads retrieval engine.
  • April 9, 2025: Google unveils Ironwood, its first inference-specific TPU.
  • October 2025: AWS launches RTB Fabric for bidding workloads.
  • November 19, 2025: European Commission proposes the Digital Omnibus, narrowing Article 9 for inferred data.
  • December 24, 2025: Nvidia agrees to license Groq's inference technology in a deal reported at about $20 billion.
  • April 2026: Google announces TPU 8t for training and TPU 8i for inference; Bedrock runs its bidder inside Index Exchange.
  • June 1, 2026: PubMatic launches Decision Fabric for in-auction partner models.
  • July 21, 2026: Google releases Gemini 3.6 Flash.
  • August 31, 2026: ChatGPT advertising passes a $1 billion run rate.
  • October 7, 2026: Anthropic releases Haiku 5.5 at $0.10 per million input tokens.

Summary

Who. Model providers such as Google, OpenAI, Anthropic and Meta run inference and set its price. Chipmakers including Nvidia, Google, Amazon and Groq build the hardware. Advertisers, agencies, DSPs and exchanges pay for it through every AI feature and bidding model they use, while regulators and courts govern inferences drawn about people.

What. Inference is the stage at which a trained model is run on new input to produce an output: a bid, a ranking, a probability or generated text. Its weights do not change. In data protection law, an inference is also a conclusion drawn about a person, treated as personal data.

When. The machine learning sense took hold in the 2010s, with Google's inference-focused TPU deployed in 2015. Generative AI made it a mass workload from November 2022, and reasoning models from September 2024 shifted computation toward serving.

Where. Inference runs in hyperscale data centres on GPUs and custom accelerators, increasingly inside ad exchanges, and on consumer devices such as phones and browsers.

Why. Training is a one-off cost; inference recurs with every query, impression and agent action. Its price and latency decide which AI features are viable in advertising, how large bidding models can be within a millisecond deadline, and how much automation marketing organisations can afford to run.