The Interactive Advertising Bureau published Measuring Visibility in the AI Era on August 3, 2026, a set of measurement guidelines covering brand and publisher visibility inside AI-generated answers. The document defines a shared metric vocabulary, splits data into two quality tiers, and lists what measurement companies have to disclose before a buyer can judge whether a number means anything.
A market has formed around a question nobody has agreed how to answer. More than 20 companies now sell tools that claim to measure how brands and publishers appear inside AI-generated responses, and those tools use different query sets, different platform coverage, different scoring rubrics and different definitions of the underlying events they count. Two of them can measure the same brand in the same category during the same week and return materially different results. Buyers have budgets. They have had no basis for comparison.
IAB released Measuring Visibility in the AI Era on August 3, 2026, from its New York headquarters, positioning the document as the industry's standardized set of measurement guidelines for tracking brand and publisher visibility in AI-powered discovery platforms. The framework runs to 36 pages and is dated August 2026. It forms part of Project Eidos, the trade body's broader measurement modernization programme.
A vocabulary problem before a measurement problem
The framework opens with a diagnosis rather than a product. According to IAB, consumers increasingly discover brands, products and content through AI platforms, and a new measurement problem has emerged alongside that shift: the tools sold to track the phenomenon produce different answers for the same subject, and without a shared way to evaluate that data, the industry cannot separate reliable signals from noise.
"Consumers are increasingly discovering and considering brands and products in AI platforms, but measurement frameworks haven't kept pace," said Caroline Giegerich, VP, AI at IAB, in the announcement. "This playbook gives the industry a common foundation for evaluating AI visibility consistently and with confidence. The brands, publishers, and agencies that embrace it now will better understand their presence in AI-powered discovery and be able to take informed action to improve their visibility."
Giegerich leads the IAB AI Visibility Measurement Framework Working Group, described in the document as a cross-industry team drawn from brands, agencies, publishers and measurement providers.
The framework identifies the absence of definitional agreement as the root condition. There is no common definition of a mention. There is no standard for what constitutes a citation. There is no shared method for evaluating whether a tool's output carries enough rigor to inform a strategy decision. According to the document, only 16% of brands systematically track AI visibility today, and the absence of a common standard is part of the reason.
That figure sits against a set of adoption numbers the framework treats as settled. ChatGPT has more than 900 million weekly active users, a level OpenAI disclosed on February 27, 2026 alongside a $110 billion funding round. Google AI Overviews, according to the framework, reach over 2.5 billion monthly users and appear in almost half of searches, more often on questions and longer queries. AI Overviews now appear on 14% of shopping queries, which the document reads as evidence that AI-generated answers increasingly shape purchase decisions rather than informational ones alone. PPC Land recorded a lower figure for AI Overviews in May 2026, when Google's own executives placed the surface above two billion monthly users, and independent measurements of AI Overview frequency have diverged sharply, with Similarweb putting the share of US searches at 43% while Adthena recorded 18%.
McKinsey estimates, cited in the framework, that brands unprepared for the shift could see traffic declines of 20% to 50% from traditional search channels.
On the publisher side the document treats the effect as already measured rather than projected. Citing Chartbeat data reported by Axios in March 2026, the framework records search referral traffic down 60% for small publishers, 47% for medium publishers and 22% for large publishers over the preceding two years. AI chatbot traffic, despite growth exceeding 200% between December 2024 and December 2025, still accounts for less than 1% of total publisher page view referrals. The gap between what search took away and what AI referrals returned remains, in the framework's phrasing, vast.
What the document does and refuses to do
The framework states its own limits early. It does not rate measurement providers. It does not prescribe tools. It does not build measurement systems. What it supplies is a shared vocabulary, a set of quality criteria and a list of disclosure requirements, on the argument that those three things are what a functioning market requires and what this one lacks.
That restraint matters commercially. A trade body ranking vendors would produce a winner list and a fight. A trade body defining what a mention is, and what a provider has to reveal about how it counted mentions, shifts the competitive terrain without naming anyone. Measurement companies, according to the document, gain a way to differentiate on methodological rigor rather than on marketing claims. Brands and agencies gain a basis for comparing competing providers. Publishers gain a common language for demonstrating citation value.
Scope is narrow and stated plainly. The framework covers organic, non-paid AI visibility measurement only. It does not address answer engine optimization or generative engine optimization tactics, agentic media buying specifications, or commerce attribution. Paid placement measurement and publisher downstream attribution are named as gaps warranting future standardization work.
The document flags one consequence of that boundary as increasingly awkward. Organic and paid visibility now co-exist on the same response surface. A brand may be cited organically while simultaneously holding a paid placement inside the same AI answer. This framework measures the organic component and does not isolate paid influence, which makes standardized paid placement measurement, in the document's assessment, an urgent adjacent priority. IAB's forthcoming attribution framework is designated to handle attribution methodology separately.
The 4 P's: metrics arranged as a causal chain
The central structure is a four-level hierarchy the framework calls the 4 P's of AI Visibility, ordered to reflect how visibility converts into business value: Presence, Prominence, Portrayal and Persuasion. Each level answers one question. Does the brand appear, or is the publisher cited? Where and how prominently? In what context, and with what accuracy? Does visibility drive action?
The hierarchy is nested rather than sequential. Presence is the outer condition; without a mention, no downstream metric has anything to describe.
Presence
Four metrics sit at this level for brands.
Mention Rate is the frequency with which a brand appears in AI-generated responses across a defined query set, calculated as responses containing a brand mention divided by total responses. The framework calls it the most fundamental visibility metric and notes the obvious sensitivity: results depend on query set composition, and different query sets produce different rates.
Citation Rate is the frequency with which a brand is cited as a source, calculated as responses containing a brand citation divided by total responses. The distinction the framework draws between the two is precise. Mention Rate captures whether a brand is spoken about. Citation Rate captures whether it is relied upon, signalling the degree to which AI platforms treat a brand's owned content or domain as an authoritative reference. A brand can post a high Mention Rate against a low Citation Rate, and the framework treats the gap between them as a meaningful signal in itself.
That gap has been measured at scale. Semrush, now an Adobe company, analysed 126 million United States AI search prompts collected between January and April 2026 and found that only 36 of more than 1,200 tracked brands appeared in the top 100 most-mentioned list on every platform in every month of the study window. The same study treated mention and citation as structurally separate metrics, documenting that a brand can be named constantly and cited almost nowhere.
Share of Voice expresses brand mentions as a proportion of all brand mentions within a defined competitive category. The framework's reasoning is that a brand's absolute Mention Rate matters less than its relative share. Co-mention patterns, meaning the competitors a brand consistently appears alongside, are treated as an analytical lens on the underlying Mention Rate and Share of Voice data rather than as a separate metric, because they reveal which competitive set a model has placed the brand into.
Two disclosures are mandatory for Share of Voice, and both are methodological choices that change the result: how the competitive category is defined, meaning which brands constitute the set and on what basis, and how the total category mention universe is sized within that set. According to the framework, both choices determine directly whether results can be compared across tools.
Visibility Momentum measures the rate of change in visibility metrics over time, calculated as the percentage change in Mention Rate or Share of Voice between defined measurement periods. The framework attaches a warning to it. The metric is vulnerable to model updates and platform changes that move baselines independently of anything the brand did, and providers are required to disclose whether trend data has been re-baselined after platform changes.
Prominence
For brands, Prominence carries a single metric: Position. The framework defines it as where the brand appears in the formatted AI response as a user sees it, rather than in the model's raw text output. That distinction is deliberate and consequential. Position captures whether a brand is the first entity mentioned, its rank when several brands are listed, whether it appears as a standalone recommendation or one option among many, how far into a response the mention occurs even when no other brands appear, and how many times it is referenced within a single response.
Measuring it in list-format responses is simple. Narrative responses have no universal method, and the framework requires providers to disclose how they operationalize Position in that case. It also requires disclosure of which of three measurement approaches is used: analysis of the entire rendered response including content a user would need to scroll to reach, processing of a screenshot of the response, or measurement limited to what is visible on screen at first view, with anything below the fold treated as not appearing. Three providers using those three approaches will report three different Position figures for identical output.
Portrayal
This level is where the framework introduces what it describes as a brand safety dimension unique to AI measurement.
Sentiment captures whether a brand is described positively, neutrally or negatively. The framework qualifies its usefulness: sentiment carries most meaning in informational queries, while recommendation and commercial contexts tend to be dominated by neutral descriptive language. Classification accuracy varies across tools and is particularly unreliable for nuanced or comparative language. Providers are required to disclose classification methodology, framing taxonomy and accuracy benchmarks.
Framing captures context rather than tone: whether a brand is positioned as a market leader, a budget alternative, an outdated option or a niche specialist, or categorically attributed to the wrong type, genre or sector. The framework's illustration is a family film classified as a violent thriller.
Hallucination Rate is the frequency of hallucinated brand mentions as a share of total mentions within a defined query set, segmented by platform. The framework's definition of a hallucinated mention is an association, attribute or characterization the model fabricates, with no basis in any underlying source, or one attributed to a real source that does not actually support the claim.
Factual Inaccuracy Rate covers a different failure. It is the frequency of brand mentions that reference the brand accurately but attach materially incorrect information: wrong product specifications, outdated pricing, misattributed features, incorrect company history. Here the model is reflecting a source faithfully, and the source is wrong.
The framework insists the two be reported separately because they trace to different remedies. Hallucinations are an AI platform quality issue. Factual inaccuracies trace to inaccurate underlying content propagating through AI responses, and the source of truth for comparison must be brand-supplied. For both metrics, providers must disclose detection methodology, report rates per platform rather than as an aggregate because rates vary significantly across models, and surface flagged mentions to clients rather than silently excluding them.
That last requirement is the sharpest operational demand in the metrics section. A provider that quietly filters hallucinated mentions out of a Mention Rate calculation produces a cleaner number and conceals a reputational exposure. The framework's position is that hallucinated mentions inflate Mention Rate, distort Share of Voice and corrupt Sentiment and Framing analysis, and that a brand falsely associated with a product, claim or context it has no connection to represents a risk providers have a responsibility to surface rather than suppress. Buyers are directed to treat the absence of a documented detection methodology as a material gap in any provider's quality claim.
Independent research has repeatedly documented the underlying error rates. NP Digital research published on February 2, 2026 found that 47.1% of marketers encounter AI inaccuracies several times each week, with more than a third acknowledging that hallucinated or incorrect AI-generated content had already been published publicly.
Persuasion
Recommendation Strength measures the degree to which a brand is actively recommended rather than passively cited. The framework distinguishes between a brand named as the best option with specific reasoning and one merely included in a generic list. Assessment considers whether the model uses specific descriptive language indicating active endorsement, the prominence of the recommendation within the overall response, and whether the recommendation is qualified or unconditional. Providers must disclose the scoring rubric, because the boundary is a judgment call that varies across tools.
Post-Citation Click-Through Rate is the percentage of AI mentions that lead users to a brand's digital property. The framework designates it a bridge metric: defined here as part of the visibility vocabulary, with full attribution methodology deferred to IAB's forthcoming attribution framework. Availability depends on platform-level data, and not all AI platforms provide click-through data.
Publisher metrics run on a parallel track
Publisher metrics occupy the same four levels with different contents, serving editorial publishers, news organizations, trade press and any content creator whose work is ingested and surfaced by AI platforms.
At Presence, publishers get Citation Rate, defined as responses containing a citation to the publisher divided by total responses, described as the publisher equivalent of Mention Rate. The framework supplies the definitional anchor the whole document rests on here. A citation is a linked reference containing a hyperlink to the source, or a named reference identifying the publication by name without a hyperlink. Implied citations, where content appears drawn from a source without explicit attribution, are excluded, because they cannot be measured consistently across platforms or providers.
That exclusion is doing significant work. It draws the boundary of what the framework claims to measure and concedes, explicitly, that publisher influence on AI systems extends past the observable. The document repeats the concession in its guidance section: citation metrics capture visible contribution, including content that is cited, paraphrased or relied on in a response, but do not include the full set of ways publisher content shapes AI systems.
Citation Decay Rate measures how Citation Rate changes over time for a specific piece of content or a publisher's portfolio, which the framework frames as the shelf life of content inside AI systems. Decay patterns may reflect model retraining schedules or index refresh cadence rather than any change in content relevance, and providers must disclose the time intervals used and whether decay curves are adjusted for known model update events.
At Prominence, publishers get Content Utilization Rate, the degree to which content is substantively drawn upon versus superficially cited. The framework separates a response that quotes, paraphrases or builds upon a publisher's reporting from one that merely lists the publisher as a source. Unlike simple citation detection, this requires comparing publisher content against the AI response itself, and the framework directs providers to use semantic, lexical or hybrid similarity methods and to disclose which.
At Portrayal, publishers get Attribution Clarity, measuring how explicitly the publisher is identified when its content is used, on a spectrum from full attribution, meaning a linked citation with publication name, article title and date, down to partial attribution consisting of publication name only. Publishers also get Hallucination Rate and Factual Inaccuracy Rate, defined against their own failure modes: a hallucinated publisher citation attributes a claim, quote, piece of reporting or source material to a publication that never produced it, while factual inaccuracy occurs when a citation is valid but the model misrepresents the substance of the cited work.
The framework treats the publisher stakes here as categorically different from the brand case. A publisher's authority rests on the accuracy of its editorial record, and being falsely cited for reporting that does not exist is described as a material harm to that reputation. Misrepresentation damages credibility even when attribution is correct.
At Persuasion, publishers get Post-Citation Click-Through Rate, indicating whether AI visibility translates into audience traffic.
The publisher metric set has an obvious destination beyond reporting. The framework states that publishers can bring citation data into licensing conversations as an input, while cautioning that visibility metrics represent one signal among several. That connects to an argument IAB Europe advanced in September 2025, when it published a framework for AI publisher compensation examining models tied to value extracted from content rather than to crawl volume.
Metrics the framework declines to standardize
Three metrics appear in current provider offerings and are held back from the core vocabulary, on the grounds that measurement infrastructure is nascent, interpretability problems remain unresolved, or validation against real business outcomes is insufficient.
Competitive Displacement Rate counts the frequency with which a brand is mentioned while a direct competitor is not. The arithmetic is simple: responses mentioning Brand A but not Brand B, divided by total responses. The framework rejects it on denominator grounds. A competitor's absence could mean the model preferred the focal brand, or it could mean the query never surfaced the category at all, and the metric cannot distinguish between those cases.
Attention Proxies are indicators approximating user attention within AI responses, including mention position, response length relative to brand mention placement, and citation click signals where available. The framework classifies them as directional only and not yet linked to real business outcomes, requiring any provider offering them to disclose how each proxy is operationalized.
Multimodal Visibility is flagged as a priority for future standardization rather than a current metric. As AI systems generate image, voice and video responses, text-based metrics alone become insufficient. The open questions the framework lists are specific: how tone and delivery affect brand perception in voice responses, how brands are visually represented in AI-generated imagery or product cards, and what the correct unit of measurement is when a single response combines text, images and interactive elements. Providers are directed to begin documenting how their tools handle non-text response formats.
Directional versus decision-grade
The second structural contribution is a two-tier classification that separates data suitable for noticing something from data suitable for spending against it.
Directional measurement answers questions of general trajectory. Is brand visibility broadly increasing or decreasing? Is the brand appearing in AI responses for its core category? The framework treats this tier as valuable for early signal detection, internal briefings and competitive awareness, and as insufficient for budget allocation, provider selection or executive strategy decisions that require precision and reproducibility.
Decision-grade measurement meets a higher standard across sample size, query volume, prompt type coverage, testing cadence, reproducibility, data validation, methodology documentation and platform coverage including multiplatform aggregation. The framework illustrates the difference by the kind of statement each tier can support. Decision-grade data can carry a conclusion such as a share of voice in the mid-range SUV category declining four points quarter over quarter, driven by a specific competitor's gains in recommendation-type queries.
Both tiers are described as legitimate. The failure the framework names is treating directional data as decision-grade without recognizing the gap.
One qualification runs against any simple tier label. Not every metric behaves the same way at the same tier. Presence metrics such as Mention Rate and Citation Rate are measurable at directional scale and suited to trend monitoring. Prominence and Portrayal metrics, including Position, Sentiment, Framing and Recommendation Strength, involve methodological judgment calls that amplify variance and demand higher reproducibility before they can be used at decision-grade. Persuasion metrics, particularly Post-Citation CTR, depend on platform-level data availability and remain directional in most implementations. The framework directs buyers to evaluate provider claims metric by metric against the criteria matrix rather than assuming a provider's overall tier classification applies uniformly across everything it reports.
The document also defines the space between the tiers. Directional data can support higher-stakes conclusions when signals converge across multiple sources and measurement periods and are corroborated by first-party analytics. It is not to be used for high-stakes decisions resting on a single measurement period, or when results conflict across platforms or providers.
The criteria matrix and its one hard number
Nine dimensions carry minimum thresholds for each tier. Most are qualitative. One is not.
On query volume, the framework treats fewer than 50 queries per measurement program as exploratory rather than directional, on the reasoning that volumes below that floor cannot meaningfully characterize a category. Directional work requires a minimum query set covering the category. Decision-grade requires a large, diverse query set with category and subcategory coverage, using the example of running shoes, trail running shoes and running shoes for flat feet, with disclosure of total volume, queries per category and variety of prompt formats.
On sample size, measured as responses collected per query, directional work requires multiple responses per query to capture typical variation. Decision-grade requires enough responses per query to establish a stable distribution, with the provider disclosing responses per query and how consistency is determined.
On prompt type coverage, the framework defines four intent types: informational, comparison, recommendation and transactional. Directional requires coverage of at least two intent types with the distribution disclosed. Decision-grade requires coverage across all four, with the distribution disclosed and results available segmented by intent type.
On testing cadence, monthly or quarterly measurement suits trend monitoring. Weekly or more frequent measurement is required for competitive response and budget allocation.
On reproducibility, directional work requires the provider to document how much variation is typical across runs. Decision-grade requires the provider to define acceptable variation ranges within a 7-day time window, report confidence levels, and maintain a documented process for verifying consistency.
On data validation, directional requires a summary-level description including how data quality is assessed and known limitations. Decision-grade requires full disclosure of what external signals the data is validated against, such as platform-native data where available, brand or publisher first-party analytics, or controlled panel data; how comparability is maintained when drawing from multiple sources; controls applied for sample integrity including outlier handling and entity disambiguation; the baseline anchoring the data and how it was established; and the known limitations of the validation approach.
On methodology documentation, directional requires a summary covering platforms included, data collection method and update frequency. Decision-grade requires full documentation of why the provider chose the queries it did, how it converts raw AI responses into numbers, how it handles each platform individually, and how results from different platforms are combined into one figure.
On platform coverage, single-platform data is acceptable at directional tier with disclosure. Decision-grade requires platforms collectively representing a substantial majority of consumer AI traffic in the target market, with weighting and market share basis disclosed.
On multi-platform aggregation, directional requires disclosure of which platforms are included in any combined figure. Decision-grade requires per-platform results reported separately, evidence of how much results vary across platforms, and documentation of how platforms are weighted in any combined score.
The reproducibility and aggregation requirements respond to a documented behaviour. Similarweb research published in November 2025 found that citation sets change by roughly 50% each month and that only 11% of citations overlap between major AI platforms, leaving 89% of citation opportunities platform-specific.
Two kinds of bias, two kinds of control
The framework separates bias by origin, because the controls differ.
Prompt-driven bias occurs when a provider's query set is unbalanced across query types, phrasings or intents, causing certain brands or content types to appear artificially stronger or weaker than they would in a representative sample. The example given is a query set heavy on recommendation-format prompts, which will favour brands that perform well in recommendation contexts and understate their performance in informational or comparison queries. The control is disclosure: providers must publish their prompt taxonomy distribution so buyers can assess whether the query set reflects how consumers actually use AI platforms.
Platform-driven bias occurs when a given platform's training data or response behaviour systematically favours or disadvantages certain brands or publishers independent of actual market position. It is detectable through cross-platform comparison using identical query sets, and the control is a requirement that providers report results per platform separately so that platform-specific patterns stay visible rather than disappearing into a combined score.
Query set construction as a quality dimension
The framework singles out prompt library construction as one of the most consequential methodological choices in AI visibility measurement, and one of the least visible to buyers. Two providers measuring the same brand in the same category can produce materially different results based solely on differences in how they construct, weight and refresh their query sets.
The disclosures required cover prompt structuring methodology; thematic clustering approach, meaning how queries are grouped into categories and whether those groupings reflect industry convention, proprietary taxonomy or client customization; distribution of queries across the four intent types; and query selection rationale, including whether query sets are grounded in real consumer search volume data or built from synthetic user data.
The framework's argument for elevating this from a disclosure item to a quality dimension is compact: a provider with strong reproducibility metrics but an opaque or poorly constructed prompt library may produce precise answers to the wrong questions.
Research published in April 2026 by Peec AI, analysing 5 million query fanouts collected across ChatGPT, Perplexity and Grok, found that AI platforms do not simply search for what a user typed, which places the composition of a synthetic prompt library at a further remove from observed behaviour.
Executive reporting and the composite score problem
AI visibility data is described as probabilistic, non-deterministic and sensitive to query composition, model version and platform behaviour. Reporting it using the conventions of traditional search metrics, the framework argues, sets expectations the data cannot meet.
Three reporting principles follow: report trends over time rather than point-in-time snapshots, always disclose the measurement tier, and contextualize results against known limitations. The document's illustration is a Share of Voice figure reported as approximately 22%, plus or minus 4 points based on current measurement precision, which it treats as more useful to a decision-maker than a bare 22% implying false precision.
Composite visibility scores receive a specific set of conditions. Providers offering an index that aggregates multiple metrics must disclose the full weighting methodology, the normalization method used, whether max-normalized, z-scored or percentile rank, the individual metric values feeding the composite, and the limitations of the aggregation approach. The framework's position is that the component metrics it defines constitute the authoritative measurement vocabulary, composite scores are a presentation layer, and the underlying data must always remain accessible.
What providers must disclose
The disclosure framework is where the document exerts the most direct pressure on vendors, and it is organized into three groups.
Under what is being measured, providers must disclose platform coverage with model version specificity; prompt library construction covering size, structure and refresh cadence, total query volume, distribution of queries across categories, how prompts are weighted within the library and whether prompts were tested across varied demographic and cultural framings; query sourcing, meaning whether query sets are synthetic, derived from real consumer search volume data, drawn from observed user behaviour or a combination, with a description of the source and selection method where real-world data informs the set; and source type classification, covering which source types the provider tracks, whether brand-owned, retailer, editorial publisher or authority and institutional, and how those sources are weighted or differentiated in outputs.
Under how it is being measured, providers must disclose data collection architecture, choosing among active query simulation where provider-constructed queries are sent to platforms and responses captured, passive behavioural observation where real user behaviour is tracked through an opted-in panel, platform-native data where first-party reporting from the AI platform itself is used, or a hybrid. Outputs from different architectures may not be aggregated or presented as equivalent without explicit disclosure.
Data collection methodology requires disclosure of how queries are issued, how responses are captured and how often collection occurs, along with the access method used, whether official APIs, licensed data feeds, browser-based observation, user-permissioned panels, synthetic sessions or automated scraping, and whether that method complies with applicable platform terms of service, rate limits and access restrictions. Providers must also disclose whether queries are issued with live web retrieval enabled, disabled or both, and results collected under different retrieval configurations are not to be combined without disclosure.
That retrieval-configuration requirement addresses a real source of divergence. Measurement of the same brand with retrieval enabled and disabled returns different answers, because one reading reflects a model's stored representation and the other reflects live web content.
Panel validity applies to any provider using passive behavioural observation, requiring disclosure of panel size, recruitment methodology, incentive structure and demographic representativeness against the target AI-using population. Panels that are small, narrowly recruited or demographically unrepresentative produce signals that cannot be generalized, and the framework directs buyers to treat undisclosed panel composition as a material gap.
Data validation and factual accuracy classification complete the group, the latter requiring providers to state whether their tools distinguish between hallucinated mentions, accurate mentions and factually incorrect mentions in which the brand is genuinely referenced but characterized with inaccurate information. The source of truth used for comparison must be brand-supplied, and each condition must be reported separately rather than aggregated into a single mention count.
One item in this section carries an apparent production error. The disclosure requirement labelled Mention Detection, Attribution, and Sentiment Logic reproduces, word for word, the text belonging to Data Collection Architecture, describing the four data collection approaches rather than any mention detection or sentiment logic requirement. The executive summary lists attribution and sentiment logic as a distinct required disclosure, and the minimum disclosure tier names it explicitly, so the intent is clear even though the table cell does not carry it.
Under how the data is maintained, historical data and versioning requires disclosure of how far back historical data extends, how model updates are handled in the historical record, and whether trend data is re-baselined after platform changes. The framework notes that trend metrics including Visibility Momentum and Citation Decay Rate are interpretable only when the historical baseline is stable or when re-baselining events are explicitly disclosed, and requires providers to document when re-baselining occurred, what triggered it, and how pre- and post-baseline data compare.
Across the whole disclosure section, one instruction governs the rest. Where a provider cannot or will not disclose against a required item, that absence is itself to be treated as a signal.
Two disclosure tiers and a certification that does not exist yet
Minimum disclosure covers the items required to place a provider's output into context: platform coverage, data collection architecture, prompt library construction at summary level, and attribution and sentiment logic at summary level. It supports directional use cases and establishes that a provider has a coherent methodology, without documenting every methodological choice in detail.
Enhanced disclosure covers every item in the required disclosures section at full documentation depth, including query sourcing methodology, full prompt library construction detail, panel validity for passive architectures, factual accuracy classification and historical data versioning. The framework expects it from any provider whose output is positioned as decision-grade or used to inform budget allocation, provider selection or executive reporting.
The framework recommends a common disclosure format that providers can complete in response to buyer requests for proposals and that buyers can apply consistently across procurement, with the stated goal of reducing the friction of cross-provider evaluation.
It then names its own destination. The disclosure framework is described as the foundation for a potential future IAB provider certification program, which would formalize the two tiers and provide independent verification of provider disclosures. The current document establishes no certification and evaluates no individual provider. It defines the vocabulary and format such a program would build on.
Certification has demonstrated pull elsewhere in IAB's orbit. At the IAB Australia retail media summit covered in late July 2026, 80% of surveyed brands and agencies said a retailer holding IAB certification would increase their willingness to work with that partner.
Non-determinism and why one reading is not a measurement
The final section of the framework addresses operating a measurement program over time, and it opens with the property that makes the whole exercise difficult.
AI platforms do not produce deterministic outputs. The same query issued twice to the same model can return different responses, with different brands surfaced, different sources cited and different framing applied. According to the framework, this is how large language models generate text and it cannot be engineered away.
Two consequences follow for measurement design.
The first is that single-response measurement is not measurement. A brand's visibility on a given query is a distribution rather than a value, and any metric derived from one response per query reflects a sample of one. The framework cites recent empirical research measuring citation patterns across Perplexity, ChatGPT and Google Gemini which found that what looks like a clear difference between two brands or publishers frequently falls within the range of normal random variation. Reliable measurement, on this reading, requires both enough responses to characterize the distribution and reporting of results as a range rather than a single number.
The second is that observed change must be distinguished from inherent variability. A four-point movement in Share of Voice between periods carries meaning only if it exceeds the variability expected from re-running identical queries against identical platforms on the same day. Without a documented variability baseline, the framework describes trend interpretation as guesswork.
Programs are required to document the variability observed across repeated identical queries within a defined time window and make that variability visible in reporting. Specific quantitative thresholds for acceptable variability remain an open working group question, since appropriate ranges differ across metrics, platforms and query types.
That open question is the framework's most candid admission. It requires variability baselines without specifying what an acceptable baseline looks like, which leaves the single most consequential parameter to individual programs for now.
Geography, time and model drift
AI platform behaviour varies by geography and over time, and the framework treats programs that fail to control for either as mixing signal with noise.
Geographic variation is produced by platform availability, model routing, regional content preferences and localization of responses. According to the framework, a query issued from the United States and the same query issued from Germany will often return different brands, different sources and different framing even when the query is identical in language. Programs covering multiple markets must disclose the geographic provenance of queries and report results per market rather than rolling geographies into a single score. Decision-grade cross-market comparison requires per-market measurement with query design adapted for local language and cultural framing.
Temporal variation is produced by platform updates, index refresh cycles and diurnal or weekly patterns in model behaviour. The framework directs programs to issue queries on a consistent cadence and at consistent times where possible, and where queries are distributed across a collection window, to document that window and hold it comparable across measurement periods.
Model drift receives the longest treatment. Platforms update models on timelines that are not always publicly announced, and the updates can shift visibility results materially and immediately. The framework's illustration is blunt: a brand's Share of Voice can change five points overnight because a model was retrained, with no change to the brand and no change to its competitive position.
Recency bias is identified as a systematic form of drift. Models weigh newer content more heavily, which means a brand or publisher's visibility can decline across a measurement period not because competitive position changed but because training or retrieval favours more recent sources. The framework directs programs to document recency bias as a baseline variable, particularly when measuring evergreen content or tracking visibility across model retraining cycles.
The distinction the section turns on is between platform-driven shifts, meaning changes caused by the AI platform itself through model updates, citation or ranking behaviour changes, retrieval configuration changes or response format changes, and market-driven shifts, meaning changes reflecting actual brand or competitive activity such as a campaign launch, a new competitor or a piece of content gaining traction. Telling them apart is described as the central task of any stable measurement program, and programs that fail will report platform change as brand performance.
Drift detection requires monitoring for abrupt shifts large enough to fall outside a program's normal variability range. Before treating a shift as platform-driven, the operator is expected to rule out brand activity, competitive activity and changes to the program's own query set. The strongest signal that a shift is platform-driven is its pattern across the category: a sudden change appearing for multiple brands on the same platform, or across a full category on that platform, almost always reflects model behaviour rather than competitive dynamics.
Model update impact assessment follows a three-step sequence once brand and competitive activity are excluded: check published platform release notes and provider communications for an announced model update; compare the affected platform against others in the program to establish whether the shift is isolated; and examine the shift across query types to see whether it concentrates in particular intents or distributes evenly.
Baseline reset protocols state that not every model update warrants a reset, since minor updates can be absorbed within normal variability ranges. A reset is warranted when a platform announces a major model version change; when citation behaviour, response format or answer-surface presentation shifts in ways that affect metric calculation, including changes to citation placement, card formats, source drawers, default collapsed or expanded states, mobile layouts or the introduction of sponsored modules even where the underlying generated text is unchanged; or when the query set itself is modified in ways that break comparability with prior periods.
When a reset occurs, the program must document the triggering event, the date and the nature of the change, re-run against the updated platform to establish a new baseline, and report pre- and post-reset data as separate series rather than stitching them into a continuous trend line.
The inclusion of presentation changes as a reset trigger is notable given the pace of change on measurement surfaces themselves. Microsoft Clarity moved its Citations feature to general availability on May 13, 2026, added competitive topic analysis on July 9, 2026 at no charge, and pushed automatic topic classification into beta on July 22, 2026.
Platform entry, cadence and the language of uncertainty
The set of platforms that matter for measurement is determined, according to the framework, by consumer share. A platform earns inclusion when it represents enough consumer AI traffic in the target market to materially affect a brand's visibility picture; below that threshold it adds noise more than signal and is better treated as a monitoring target than a core measurement surface. Consumer share also governs weighting once a platform is included, since a platform reaching 40% of consumer queries in a category contributes more than one reaching 5%, and combined scores are expected to reflect that asymmetry rather than treating platforms as equal inputs.
Consumer share is defined as the relative frequency with which consumers in the target market actually use a platform for AI-driven discovery, typically expressed as active users, query volume or session share. Programs are directed to rely on multiple sources where possible, including platform-reported metrics, third-party measurement and first-party referral analytics, and to disclose the basis on which consumer share is assessed. The framework acknowledges that the current AI usage data landscape is fragmentary and not directly comparable across platforms, and sets the standard as documented, consistent methodology across measurement periods rather than a fixed quantitative threshold.
That fragmentation is visible in the public record. Similarweb data reported in June 2026 put ChatGPT at 52.7% of worldwide generative AI website traffic, down from 76.4% a year earlier, with Claude tripling its share over the same window.
New platforms are to be added on a defined schedule rather than reactively, and additions handled as baseline events: establish a new baseline including the platform, retain the prior baseline for historical comparison, and report on both bases until enough time has passed on the new baseline to support trend analysis. Inclusion criteria, including the consumer share threshold a program uses, are to be documented in advance so that additions follow stated standards rather than ad hoc judgment.
On cadence, the framework's principle is that measurement frequency matches the decision cycle it informs and does not exceed it. Measuring weekly for a quarterly decision produces data more volatile than useful; measuring quarterly for a weekly decision produces data already stale on arrival. Monthly measurement suits trend monitoring and early signal detection. Biweekly to weekly suits competitive benchmarking and agency performance reporting. Weekly or more frequent measurement is expected for active budget allocation, provider selection or executive strategy decisions. Programs covering platforms known to update frequently may need higher frequency simply to detect drift in time, independent of the decision cycle.
The framework closes with reporting language patterns, presented as starting points rather than scripts, aimed at an audience the document assumes is accustomed to deterministic measurement. Its own characterization of the problem is that a marketing leader who has spent two decades reading Google search numbers as exact figures will read a range with a variability band as imprecise unless the report actively reframes what kind of number it is.
For introducing the data for the first time, the framework offers: "AI platforms don't return the same answer every time, even to the same question. That's why we report ranges instead of single numbers. The range shows what normal variation looks like. Changes inside it aren't meaningful, changes outside it are."
For reporting divergent platform results, where figures cannot be reconciled into a single combined story: "Visibility on [platform A] increased four points this period; visibility on [platform B] declined three points. The divergence is not reconcilable into a single trend. We are reporting per-platform figures and treating each platform's trajectory separately for this period."
Additional patterns cover a platform-driven shift coinciding with a confirmed model update, and a decline that persists after a platform stabilizes and therefore requires the narrative to shift from platform noise to performance signal.
The framework's argument for the whole reporting section rests on one distinction: the difference between reporting that visibility is down and reporting that visibility is down because the platform updated and has not yet stabilized is the difference between a strategic response and a non-response.
Who built it
The working group is led by Giegerich and lists members from across the buy side, sell side and measurement vendors: Breeze Zhao, Sr Director Measurement Science at Walmart; Daniel Flynn, Sr. Director, AI Visibility and Measurement at EMARKETER; Graham Wilkinson, EVP, Chief Innovation Officer and Global Head of AI at Acxiom; Haozhe Xu, Manager, Measurement at WPP Media; Ihab Rizk, Senior Product Manager at Microsoft Clarity; Jason Hartley, Head of Media Innovation and Trust at PMG; Justin Inman, Founder and CEO at emberos; Kristina Meinig, VP, Market Development at the Alliance for Audited Media; Paul Longo, GM, AI Advertising at Microsoft Clarity; Scott Cunningham, Fractional Chief Product Officer at the Alliance for Audited Media; Simon Poulton, EVP, Innovation and Growth at Tinuiti; Skye Yang, Marketing Science and AI Measurement Leader; and Todd Paris, Cofounder and CEO of IQRush.ai.
Two members come from Microsoft Clarity, whose Citations product reports page citations in AI-generated answers, share of authority relative to competing domains, AI referral traffic, grounding queries and cited pages, and which has released a sequence of AI visibility features through 2026 at no additional cost.
"As AI becomes a new layer of discovery and decision-making, brands need more than traffic reports. They need to know how AI systems are finding their content, when they are being cited, and where they have opportunities to improve," said Ihab Rizk, Senior Product Manager at Microsoft Clarity, in the IAB announcement. "The IAB's framework is an important step toward a common measurement language for this new era, and it aligns with our belief that AI visibility should be clear, actionable, and accessible to every business."
Measuring Visibility in the AI Era is published as part of Project Eidos, the IAB initiative named after a Greek verb meaning to see, announced on February 2, 2026 and framed around bringing clarity, consistency and confidence to measurement. Project Eidos produced Campaign Data Standards 1.0 on May 14, 2026, addressing reconciliation between planning tools, buying platforms and finance systems, and a Redefining Media Types Standard covering video classification released for comment in July 2026.
Why this matters for the marketing community
The framework arrives into a market where optimization spending has outpaced the evidence base supporting it, and where the measurement layer beneath that spending has never been standardized.
The first consequence is procurement. A brand or agency evaluating AI visibility vendors has, until now, compared marketing claims. The disclosure format converts that comparison into a document exercise with a defined failure condition: a provider that will not answer a required item has, by the framework's own instruction, supplied an answer. That shifts negotiating position without requiring the buyer to develop methodological expertise in-house. It also sets a floor that vendors will be measured against whether or not they participate, since the requirements are public and buyers can insert them into requests for proposals immediately.
The second consequence concerns what marketers currently pay for. A critical academic survey of 45 generative engine optimization studies, published on July 20, 2026, concluded that experiments provide strong evidence that a document already inside an engine's context can influence how it is cited, but establish far less often that a page will be retrieved at all, and almost never that any of it produces a durable effect on clicks or conversions. The same survey found that content rewrites confined to page bodies can reduce a page's presence in the top ten by 16%. Optimization budgets are moving against a surface where causal evidence is thin, and the framework's contribution is not to settle that argument but to specify what a credible measurement of the outcome would look like.
The third consequence is definitional and affects reporting inside organizations. The separation of Mention Rate from Citation Rate, and of Hallucination Rate from Factual Inaccuracy Rate, gives marketing teams a way to describe results that currently get compressed into a single visibility score. Adobe research published in June 2026 found that 98% of marketers lack a confident AI search strategy, with 12% describing themselves as completely lost. A survey of 481 marketers accompanying Semrush's AI Visibility Index found that 45% could not properly measure their brand's visibility in AI-generated answers, and only 9% could measure the full set of metrics that study identified as relevant. Those gaps are partly capability and partly vocabulary.
The fourth consequence sits with publishers. Citation Rate, Content Utilization Rate and Attribution Clarity give publishers standardized quantities to bring to licensing negotiations, at a moment when several large publishers are weighing whether to block Google's crawler entirely and when Reddit's annual licensing arrangement with Google is under renewal. A publisher able to state a Content Utilization Rate, with a disclosed similarity method behind it, argues from a different position than one asserting that its content is valuable. The framework is careful to describe visibility metrics as one input rather than a valuation, and the exclusion of implied citations means the measured figure will understate actual contribution.
The fifth consequence is about brand safety, which the framework treats as a Portrayal problem rather than a placement problem. Hallucination Rate and Factual Inaccuracy Rate, reported per platform and with flagged mentions surfaced rather than filtered, describe an exposure that has no equivalent in display or video buying. A brand cannot buy its way out of being described inaccurately in an answer it did not pay for.
The sixth consequence concerns the vendors themselves. A market of more than 20 tools with no shared vocabulary produces, in the framework's words, a race to the bottom on claims rather than a competition on quality. Some consolidation has already occurred. Adobe acquired Semrush in November 2025. Others exited: the founder of Lorelight shut down his AI visibility tracking tool in late 2025, concluding that generative engine optimization fits inside a comprehensive search suite rather than standing alone as a product. A public rigor standard changes the terms on which the remaining vendors compete, and the prospect of certification changes them further.
What the framework does not resolve is the paid surface. Organic and paid visibility already share the same response area, and the document measures only one of them, at a point when ChatGPT advertising has moved from a pilot in February 2026 to self-serve availability and Adthena counted 7,378 ChatGPT advertisers in a single quarter. A brand cited organically alongside its own paid placement in the same answer currently has no standardized way to separate the two contributions. The framework names that as an urgent adjacent priority and defers it.
The variability threshold question is the other unresolved item, and it is the one that determines whether decision-grade means anything operationally. Until a working group settles what magnitude of change exceeds normal noise for a given metric, platform and query type, programs will set their own baselines and comparability across programs will remain partial.
Timeline
- November 16, 2023 - First arXiv version of the foundational generative engine optimization paper is published, later reviewed in a critical survey of 45 studies
- October 14, 2025 - Adobe makes LLM Optimizer generally available, an enterprise tool for monitoring and improving visibility in generative AI interfaces
- November 2025 - Amplitude launches AI Visibility tools with competitive tracking, connecting AI mentions to downstream conversion data
- November 2, 2025 - The founder of Lorelight shuts down his AI visibility tracking tool, concluding the category belongs inside broader search suites
- November 12, 2025 - Similarweb publishes an AI citation analysis framework documenting citation sets changing 50% monthly with 11% overlap between platforms
- December 2024 to December 2025 - AI chatbot referrals grow over 200% while remaining under 1% of total publisher page views
- February 2, 2026 - IAB unveils Project Eidos alongside State of Data 2026 research finding up to 75% of marketers say advanced measurement approaches underperform
- February 2, 2026 - NP Digital research finds 47.1% of marketers encounter AI inaccuracies weekly
- February 27, 2026 - OpenAI discloses ChatGPT has reached 900 million weekly active users
- March 17, 2026 - Chartbeat data reported by Axios shows small publishers lost 60% of search referral traffic over two years, medium publishers 47% and large publishers 22%
- April 2026 - Peec AI collects 5 million query fanouts across ChatGPT, Perplexity and Grok, finding AI platforms do not search for what users typed
- May 13, 2026 - Microsoft Clarity moves Citations to general availability, reporting citations, share of authority and grounding queries
- May 14, 2026 - IAB releases Campaign Data Standards 1.0, the first output of Project Eidos
- June 1, 2026 - Adobe survey finds 98% of marketers lack a confident AI search strategy
- June 11, 2026 - Similarweb data shows ChatGPT at 52.7% of generative AI web traffic, down from 76.4% a year earlier
- June 25, 2026 - Semrush publishes the AI Visibility Index 2026, finding 36 of more than 1,200 brands held consistent visibility across every platform
- July 9, 2026 - Microsoft Clarity releases Topic Insights free to all users, adding competitive topic analysis to Citations
- July 20, 2026 - A survey of 45 studies concludes generative engine optimization rewrites can cut a page's AI retrieval by 16%
- July 21, 2026 - Several large publishers weigh blocking Google's crawler as Reddit's licensing deal enters renewal
- August 3, 2026 - IAB publishes Measuring Visibility in the AI Era, defining the 4 P's hierarchy, the directional versus decision-grade split and the provider disclosure framework
Related PPC Land coverage
- Semrush: 36 brands win AI visibility everywhere, 1,200 vanish on one - Quantifies the mention-versus-citation gap the IAB framework formalizes, across 126 million prompts and four platforms.
- Small publishers lost 60% of search traffic as AI reshapes the web - The Chartbeat dataset the framework cites for its publisher traffic decline figures.
- Microsoft Clarity Citations goes GA: now track how AI answers cite your site - Details the platform-native citation measurement approach the framework lists as one of four data collection architectures.
- Microsoft Clarity gives away AI visibility tool rivals charge for - Documents the free competitive analysis feature launched by a company with two representatives on the framework working group.
- Similarweb launches AI citation analysis framework as zero-click searches rise - Establishes the citation volatility figures that underpin the framework's reproducibility requirements.
- Survey of 45 studies finds GEO rewrites can cut a page's AI retrieval 16% - Academic review concluding that commercial claims about AI visibility optimization outrun the evidence.
- Adobe: 98% of marketers have no confident AI search strategy - Survey data on the readiness gap the framework is designed to close.
- IAB's campaign data standards want to end the reconciliation nightmare - Covers Campaign Data Standards 1.0, the first deliverable from the same Project Eidos initiative.
- AI poised to unlock $32 billion in marketing measurement value as current systems falter - Reports the State of Data 2026 research and the launch of Project Eidos.
- Nearly half of marketers encounter AI errors weekly as study exposes trust gap - NP Digital research on the hallucination rates the framework requires providers to detect and disclose.
- ChatGPT drops to 52.7% as Claude triples its AI traffic share - Similarweb platform share data relevant to the framework's consumer share weighting requirement.
- Your analytics are lying: Similarweb traces AI recommendations to real traffic - Panel-based clickstream research on the downstream traffic effects the Persuasion tier attempts to capture.
- IAB Europe unveils framework for AI publisher compensation - Earlier work on the licensing conversations the publisher metric set is designed to inform.
- Reddit and USA Today face Google exit as search traffic drops 28% - Context on the publisher bargaining position that citation measurement now feeds into.
- Marketers still call it SEO as 81% reject GEO pitch, Fractl finds - Survey evidence on how AI search work is labelled and budgeted inside organizations.
- OpenAI hits 900M users - The disclosure behind the ChatGPT scale figure the framework cites.
Summary
Who: The Interactive Advertising Bureau, through its AI Visibility Measurement Framework Working Group led by Caroline Giegerich, VP, AI at IAB, with members from Walmart, EMARKETER, Acxiom, WPP Media, Microsoft Clarity, PMG, emberos, the Alliance for Audited Media, Tinuiti and IQRush.ai. Ihab Rizk, Senior Product Manager at Microsoft Clarity, provided commentary in the announcement.
What: Measuring Visibility in the AI Era, a 36-page set of measurement guidelines defining a shared metric vocabulary for brand and publisher visibility in AI-generated answers. It introduces the 4 P's hierarchy of Presence, Prominence, Portrayal and Persuasion; a two-tier quality classification separating directional from decision-grade data across nine criteria, including a floor treating fewer than 50 queries per program as exploratory; a provider disclosure framework covering platform coverage, prompt library construction, query sourcing, data collection architecture, panel validity, factual accuracy classification and historical versioning; and operating practices for measurement stability under non-deterministic model behaviour.
When: Published August 3, 2026, dated August 2026 on the document itself.
Where: Released from IAB's New York headquarters and published on IAB.COM. The framework covers organic AI visibility measurement across AI-powered discovery platforms including ChatGPT, Google Gemini, Microsoft Copilot, Perplexity and Claude, with per-market reporting required for programs covering multiple geographies.
Why: More than 20 companies sell AI visibility measurement tools using different methodologies that produce different answers for the same brand or publisher, and only 16% of brands systematically track AI visibility. The framework supplies the shared vocabulary, quality criteria and disclosure requirements the market has lacked, without rating providers or prescribing tools, and establishes the disclosure format a potential future IAB certification program would build on.
Discussion