Semantic describes a system that acts on what content means rather than on the characters it contains. A keyword filter asks whether the string "shot" appears on a page. A semantic system asks whether the page concerns a homicide, a basketball game or a photography workshop, and answers differently for each. That is the entire distinction: form against meaning, matching against interpretation.

The word travels across three layers of digital marketing, and the layers are routinely confused. Publishers declare meaning through markup. Verification and contextual vendors infer meaning from text, images and video, then sell the result as targeting segments. Search engines and language models compute meaning as geometry, turning passages into numeric vectors whose distances stand in for relatedness. All three get called semantic. Only the first involves anybody stating anything.

How meaning gets encoded

Declared meaning is the oldest form. Schema.org vocabulary, expressed in JSON-LD, microdata or RDFa, labels a page's entities and properties for machines that will never read the prose. Semantic HTML does a cruder version through elements such as article, nav and header, which carry structural meaning that a generic div does not. Coverage is thin outside a small core. A joint Google and Schema.org dataset published in June 2026 found 76.9% of the vocabulary's 5,545 terms present on fewer than 1,000 domains, while a handful of properties including name, description, image, url, headline and author cleared 10 million domains each.

Inferred meaning comes from classification. A crawler fetches the page, natural language processing (NLP) extracts entities, topics and sentiment, and the output is mapped onto a taxonomy. IAB Tech Lab's Content Taxonomy supplies the shared vocabulary: version 2.2 went out for comment on October 28, 2020, and migration to 3.0 is disruptive enough that an open-source mapping tool was released to automate it in February 2026, since News/Opinion classifications and entertainment genres do not translate cleanly between versions.

Computed meaning dispenses with categories altogether. An embedding model converts a passage into a vector of hundreds or thousands of numbers, and similarity becomes the cosine distance between two of them. Chrome's on-device history search, reconstructed from Chromium source code, chunks pages into passages and stores them as 1540-dimensional vectors, supporting conversational queries with no category anywhere in the pipeline.

Where it sits in the bid request

Real-time bidding carries declared and inferred meaning, not computed meaning. In OpenRTB, the Site object holds cat for the site's categories, sectioncat for the current section and pagecat for the page. The Content object carries its own cat array plus keywords or kwarray for free text. None of those codes means anything without cattax, the field naming which taxonomy they belong to; absent cattax, Content Category Taxonomy 1.0 is assumed. Version 2.6, published in April 2022, introduced cattax along with genres and gtax, a connected television pair describing programme genre separately from subject matter.

Buyers rarely rely on a seller's self-declared labels. Contextual data vendors run their own classification and push segments into the demand-side platform (DSP), which checks each incoming request against a cached verdict keyed to a URL or domain, the mechanism set out in PPC Land's explainer on pre-bid filtering. Peer39 page-level quality, safety and category data reached a DSP in Asia Pacific on August 30, 2011, and The Trade Desk by October 31, 2013.

Origin and evolution

The vocabulary arrived from computer science rather than advertising. The Resource Description Framework became a World Wide Web Consortium (W3C) Recommendation on February 22, 1999, giving the web a syntax for machine-readable statements about resources. Two years later, in Scientific American volume 284 issue 5, dated May 2001, Tim Berners-Lee, James Hendler and Ora Lassila set out the Semantic Web: an extension of the existing web in which information carries well-defined meaning. The Web Ontology Language followed as a Recommendation on February 10, 2004. By 2006 the same authors conceded the vision remained largely unrealised.

Commercial adoption took a narrower route. Schema.org launched on June 2, 2011, backed by Bing, Google and Yahoo, with Yandex joining that November, replacing open-ended ontologies with a pragmatic vocabulary tied to search results. Google's Knowledge Graph arrived in May 2012, moving entity relationships into the ranking stack.

Search followed. Hummingbird, live from late August 2013 and announced on September 26, 2013, rebuilt the core algorithm around query meaning; Amit Singhal, then senior vice president of search, called it the most significant change since 2001. RankBrain added machine-learned interpretation of unfamiliar queries in 2015. BERT reached search in October 2019, affecting roughly one query in ten in United States English at launch.

Contextual advertising picked up the word in the same period. Peer39, acquired by O3 Industries in 2021, described its approach as a holistic semantic method examining the relationships between words rather than the words alone. Integral Ad Science bought the specialist ADmantX in November 2019 and launched Context Control in April 2022, combining sentiment and emotional classification at page level. Equativ's merger with Sharethrough brought semantic targeting in-house, and a 2024 partnership with Illuma added continuously retrained contextual models. Seedtag markets its own classification of text, images and video under a neuro-contextual label.

Why the distinction matters commercially

Semantic classification analyses content, not people, which places it outside most data protection restrictions and makes it the default fallback as identifiers disappear. Contextual performance material cited by Seedtag put identifier coverage gaps at 54% of mobile impressions and 36% of desktop impressions.

Vendors report substantial performance gains. Integral Ad Science, detailing its system on March 2, 2026, cited a 300% click-through rate increase for Samsung and a 39% reduction in cost per conversion for consumer packaged goods brands, across a library of more than 380 contextual segments. At the 2022 launch the company claimed its Ozone technology was 42% more accurate than competing offerings. Those figures are self-reported and have not been independently audited.

The same shift is reshaping organic visibility. Assistants decompose a prompt into parallel sub-queries and merge the results: analysis of five million fan-out queries found ChatGPT using Reciprocal Rank Fusion to score pages appearing across several sub-queries more highly than pages answering only one, with the current year appended in 5.44% of prompts. A survey of 131 practitioners published in September 2026 ranked relevance the strongest positive ranking factor and search intent match the single strongest signal, with meta description last.

Limitations and disputes

Semantic systems fail where the meaning is not reachable. Google's Gary Illyes argued in June 2026 that content buried inside JavaScript cannot be reliably extracted by automated systems, and that pages serving semantic elements and readable paragraphs in the initial response produce cleaner machine output. No amount of classification sophistication compensates for text a parser never sees.

Availability is a second limit. Peer39 analysis from March 2026 found only 40% of connected television bid requests carrying usable programme-level content signals, leaving buyers to target an environment whose content is largely undeclared.

Precision is contested. Semantic classification was sold as the remedy for blunt keyword blocking, yet the category definitions can be broader than their labels suggest. Sensitive content categories Microsoft added in August 2026, powered by Integral Ad Science, extend a crime category to educational, informative and scientific coverage of the topic, and a kids' content category to parenting forums and travel writing. An advertiser applying those controls may exclude far more editorial inventory than intended.

Embeddings introduce opacity of a different kind. A cosine distance carries no explanation, so a page excluded or promoted by vector similarity offers nothing to appeal against. The original Semantic Web programme foundered for a related reason: publishers had little incentive to describe their own content accurately for third parties.

Not the same as

Contextual targeting is the commercial application; semantic analysis is one method of producing it. Contextual targeting can also run on keywords, URLs or seller-declared categories with no meaning extraction at all.

Keyword matching operates on strings and is the thing semantic methods were built to replace. The two coexist: OpenRTB keywords fields still carry free text alongside semantic categories.

The Semantic Web is the W3C programme of machine-readable, linked data, largely unrealised in its original form and surviving mainly as schema.org markup.

Semantic similarity in audience products describes user-level modelling rather than content analysis. Google's replacement for audience expansion in Display & Video 360 uses semantic similarity to find users resembling an advertiser's highest-value customers, which is the same mathematics applied to people, not pages.

Recent developments

Consolidation continued through 2026. Peer39 acquired the MRC-accredited Adloox business from Scope3 on June 2, 2026, extending contextual data into Google and Meta environments, with chief executive Mario Diez describing the goal as a universal signal layer feeding increasingly automated buying systems.

Classification targets are widening. Integral Ad Science opened a beta on April 2, 2026 for a tool that excludes mass-produced AI-generated content from programmatic campaigns, running on the same infrastructure as its topic and sentiment segments but scoring production quality instead of subject matter.

Analysis tooling has moved the same way. Screaming Frog's version 22.0, released in June 2025, pulled embeddings into the crawl to score content against target queries by semantic similarity rather than exact match. Research on synthetic consumer panels reported semantic similarity scoring reaching 88% correlation attainment against 65% for supervised gradient-boosted models, suggesting meaning-based comparison is spreading from targeting into measurement and research.

Timeline

  • February 22, 1999: The Resource Description Framework becomes a W3C Recommendation
  • May 2001: Berners-Lee, Hendler and Lassila publish "The Semantic Web" in Scientific American
  • February 10, 2004: The Web Ontology Language becomes a W3C Recommendation
  • 2006: The original authors describe the Semantic Web vision as largely unrealised
  • June 2, 2011: Schema.org launches with backing from Bing, Google and Yahoo
  • August 30, 2011: Peer39 page-level semantic data becomes available through a DSP in Asia Pacific
  • May 2012: Google introduces the Knowledge Graph
  • September 26, 2013: Google announces Hummingbird, rebuilt around query meaning
  • 2015: RankBrain adds machine-learned interpretation of unfamiliar queries
  • October 2019: BERT reaches Google Search, affecting about one in ten United States English queries
  • November 2019: Integral Ad Science acquires ADmantX
  • October 28, 2020: IAB Tech Lab releases Content Taxonomy 2.2 for comment
  • April 2022: OpenRTB 2.6 introduces cattax, genres and gtax; IAS launches Context Control
  • August 2024: Chrome's on-device history embeddings system is announced
  • March 2, 2026: IAS publishes performance figures for Context Control Targeting
  • June 2, 2026: Peer39 acquires Adloox from Scope3

Summary

Who. Standards bodies including W3C and IAB Tech Lab define the vocabularies; contextual and verification vendors such as Peer39, Integral Ad Science, Seedtag and Illuma sell classification; search engines, DSPs and language models consume it; publishers supply the raw content and the declared markup.

What. Semantic refers to processing content by meaning rather than by string matching, implemented three ways: declared markup such as schema.org, inferred classification into taxonomies through natural language processing, and computed similarity between numeric embeddings.

When. The vocabulary dates to RDF in 1999 and the Semantic Web proposal of May 2001, entered search with Hummingbird in September 2013 and BERT in October 2019, and entered programmatic buying through page-level contextual data from 2011 onward.

Where. In page markup and HTML structure, in the cat, cattax, keywords and genres fields of OpenRTB bid requests, in pre-bid segment libraries inside DSPs, and in the retrieval layers of search engines and AI assistants.

Why. Classification of content rather than people sits outside most privacy restrictions, which made semantic methods the default answer to identifier loss. The same techniques now decide which pages AI systems retrieve, moving the term from a targeting concern to a visibility one.