Twenty of 182 businesses named by a search-enabled OpenAI configuration carried at least one factual claim that the evidence contradicted, according to an audit published on 18 August 2026. The larger number was 479 claims the researchers could not resolve either way.
Most measurement of brand presence inside AI assistants counts appearances. A study released on 18 August 2026 by the research firm Empirank counted something else: whether the sentences attached to those appearances were true.
The audit, filed under the internal identifier AIV-009 and written by Empirank founder James Tandy, examined 1,257 testable factual claims spread across 182 businesses that a search-enabled OpenAI configuration recommended in response to local queries. Twenty of those businesses, or 11.0 percent, carried at least one claim that the reviewed evidence contradicted. Five carried at least one error the study classified as high severity.
The claim-level picture looks calmer. Twenty-four individual claims were judged incorrect, a rate of 1.9 percent. That gap between 1.9 percent and 11.0 percent is the finding the study builds its argument around, and it turns entirely on which denominator a reader picks.
How the sample was built
Empirank began with a pool of 3,000 sampled businesses and narrowed it to those the model actually recommended. Nine recommendations were then dropped because the underlying business identity could not be resolved safely, a step that removes ambiguity about which entity a claim refers to before any accuracy judgment is made.
The surviving 182 businesses produced 1,515 statements. Of those, 1,257 were factual claims capable of being tested against evidence. The remainder were descriptive or evaluative language that no evidence set can confirm or deny.
Each testable claim received one of four labels. Confirmed correct meant the available evidence directly supported it. Incorrect meant the evidence contradicted it. Unsupported meant the evidence neither confirmed nor contradicted it. Ambiguous meant the wording or the evidence did not allow a reliable decision. No claim was classified as outdated.
Quality control ran through independent AI agents rather than human reviewers, a methodological choice the study discloses in its own limitations section.
The distribution of results
Of the 1,257 factual claims, 749 were confirmed correct, or 59.6 percent. Another 479, or 38.1 percent, were unsupported. Twenty-four, or 1.9 percent, were incorrect. Five, or 0.4 percent, were ambiguous.
Restricting the comparison to the 773 claims where the evidence pointed decisively in one direction changes the headline sharply. Within that subset, 96.9 percent were correct. The study treats that figure as reassuring about evidence-backed statements while declining to let it stand in for the whole, because it excludes by construction the 479 claims that the evidence could not settle.
The separation is deliberate. Combining unsupported claims with incorrect ones would produce a single inaccuracy rate above 40 percent, a number the study argues would overstate model error and, more practically, would hide the distinction between a brand that needs a correction and a brand that needs better public evidence about itself.
Where the errors clustered
Not all claim categories behaved alike. Identity claims proved the most reliable, at 99.5 percent confirmed correct across 182 tested claims. Accreditation claims sat at the opposite end, with only 19.2 percent confirmed across 26 claims.
That accreditation figure carries a caveat the study states plainly: the low confirmation rate came mostly from missing evidence rather than from accreditations being disproved. A certification that no accessible page documents is not a certification the audit can call false. It is one the audit cannot call anything at all.
The pattern has a practical edge for any category where credentials do commercial work. Trade licences, professional memberships, and industry awards are exactly the claims a prospective customer weighs, and exactly the claims that tend to live on membership registries, association directories, or nowhere public at all.
What a citation did and did not prove
The study also tested whether the presence of an attributable citation predicted accuracy. Citation-supported recommendations reached a 59.9 percent confirmed-correct share across 1,199 claims and 170 businesses. Recommendations without an attributable citation reached 53.4 percent across 58 claims and 12 businesses.
Empirank describes the 6.5 percentage-point difference as descriptive rather than causal, and gives two reasons. The uncited group contained only 12 businesses, too few to carry a reliable comparison. And citation presence was measured at the recommendation level, meaning a single cited page attached to a recommendation does not establish that every sentence inside that recommendation traces back to it.
A second split looked at source type. Claims supported only by first-party material reached 75.2 percent confirmed correct across 105 claims. Claims supported only by third-party material reached 57.5 percent across 1,064 claims. The study flags the groups as differing in both size and content, and stops short of treating the result as proof that owned sources produce more accurate answers.
That restraint matters, because the opposite conclusion is the one a vendor selling answer engine optimization services would find most convenient.
The stated limits
Empirank lists its constraints rather than burying them. The response set was frozen, captured from one OpenAI configuration in August 2026, and the study notes that results may shift across models, interfaces, or time. The audit covered the 182 identity-validated businesses the model recommended, not the full 3,000-business sample. Citation presence was assessed at the recommendation level. The uncited comparison group was small. Quality control was automated.
"Visibility without factual control is an incomplete AEO result," Tandy wrote in the study's closing section.
In an email sent to PPC Land on 24 August 2026, Tandy restated the separation between the two error categories, writing that "unsupported does not mean false" and describing the 479 unresolved claims as a research limit rather than an accusation against the model.
Why the denominator argument lands where it does
Agencies do not manage claim totals. They manage named clients, and a single wrong sentence attached to one client is a problem regardless of how small it looks against a pooled percentage. That is the structural point the study presses, and it applies well beyond this dataset.
The same arithmetic recurs across published AI accuracy work. NP Digital research covered in February 2026 found that 47.1 percent of marketers encounter AI inaccuracies several times each week, with more than a third reporting that hallucinated or incorrect AI-generated content had already been published. Gracenote measurement published in June documented that AI systems fabricated roughly one in five streaming title responses when working without grounded metadata. Different domains, same shape: an aggregate rate that reads as tolerable, and a per-entity exposure that does not.
The measurement industry this lands in
Visibility tooling has grown quickly, and it has grown around presence rather than truth. An IAB framework published on 3 August 2026 found that only 16 percent of brands systematically track AI visibility while more than 20 vendors sell tools that disagree with one another about the same brand. That framework explicitly treats hallucinated mentions as a category providers have a responsibility to surface, on the reasoning that a filtered-out false mention produces a cleaner report and conceals a real reputational exposure.
Cloudflare moved in a related direction on 6 August 2026, releasing a dashboard that separates how often assistants name a brand from how often they cite it. The Empirank audit adds a third axis to that pair. A brand can be named, cited, and still described inaccurately.
Instability compounds the problem. SISTRIX measurement covered in May 2026 found ChatGPT rotating 74 percent of its citations per week, against 56 percent for Google AI Mode. Research published on 18 August 2026, the same day as the Empirank study, found that brands written into ChatGPT's own internal search query reached the final answer 68.9 percent of the time against 2.1 percent for brands merely retrieved, a gap of roughly 33 times. Personalization adds another layer, with earlier reporting documenting how location and account history reshape the same prompt before any page is fetched.
A frozen response set, in that environment, is a snapshot rather than a stable measurement. Empirank says as much.
The commercial stakes
The reason inaccuracy in a recommendation is not merely an editorial concern is that recommendations move traffic. Similarweb clickstream research published in June 2026 found that brands recommended by ChatGPT were 2.5 times more likely to receive a visit within seven days, with 55.9 percent of that AI-influenced traffic arriving through branded search rather than a referral link. Demandbase data released on 12 August 2026 recorded ChatGPT referrals to B2B sites rising 303 percent year on year.
Consumer trust in the layer doing the recommending remains thin. Yelp research with Morning Consult, published in April 2026, found that only 15 percent of United States adults trust AI search results a lot even as 65 percent had used such a tool in the previous six months, with 63 percent double-checking results against other sources. Gracenote's April survey of more than 4,000 users found three in four verifying chatbot answers about TV content before acting on them.
Verification behaviour of that scale means an incorrect claim inside a recommendation does not simply sit unread. It gets checked, and the check either confirms the brand or contradicts it.
OpenAI is also building a paid layer on the same surface. The company opened its self-serve ChatGPT Ads Manager to all United States advertisers on 5 May 2026, with cost-per-click bidding alongside a $60 CPM model. Organic recommendation accuracy and paid placement now occupy the same interface.
What the study proposes
Empirank sets out an order of operations rather than a tool. The audit begins at the recommendation level, covering the facts attached to every monitored mention rather than the mention alone. Identity and location facts come first: business names, addresses, phone numbers, opening hours, service areas, and branch details. Commercial proof follows, aligning services, credentials, awards, and commercial claims across a company website and authoritative third-party profiles. High-severity errors are escalated, and the same prompts are repeated after source corrections to observe whether the answer changes.
The prioritisation logic is stated in terms of customer behaviour. Errors involving location, availability, credentials, service scope, and branches are more likely to change what a customer does than a descriptive flourish is.
Whether that sequence produces measurable improvement is not something this study establishes. It measures a single frozen response set once, and the retest step it describes is a proposal rather than a reported result. Related work has found optimisation claims in this field resting on thin evidentiary ground, and the documented tendency of AI-generated advice to recirculate through retrieval corpora gives any published playbook in the category a short half-life.
What the audit does establish is narrower and more durable. Across one bounded evidence set, roughly six in ten factual claims attached to recommended local businesses could be confirmed, roughly four in ten could not be resolved, and about one in nine businesses carried a statement the evidence contradicted outright.
Timeline
- 4 September 2025 - OpenAI and Georgia Tech researchers publish findings on the statistical causes of hallucination in language models, proposing evaluation reforms based on explicit confidence targets
- 2 February 2026 - NP Digital research finds 47.1 percent of marketers encounter AI inaccuracies several times each week
- 8 April 2026 - Gracenote publishes survey data showing three in four viewers verify AI chatbot answers about TV content
- 14 April 2026 - Yelp and Morning Consult report that 15 percent of United States adults trust AI search results a lot
- 5 May 2026 - OpenAI opens its self-serve ChatGPT Ads Manager to all United States advertisers
- June 2026 - Similarweb clickstream research measures a 2.5 times visit uplift for ChatGPT-recommended brands
- 3 August 2026 - IAB publishes its AI visibility measurement framework, finding 16 percent of brands systematically track AI visibility
- 6 August 2026 - Cloudflare releases a dashboard separating brand mention rate from citation rate
- 12 August 2026 - Demandbase reports ChatGPT referrals to B2B sites up 303 percent year on year
- August 2026 - Empirank captures the frozen OpenAI response set underlying the audit
- 18 August 2026 - Empirank publishes study AIV-009, reporting 24 incorrect and 479 unsupported claims across 1,257 tested
- 24 August 2026 - James Tandy sends the study to PPC Land by email
Related PPC Land coverage
- Only 16% of brands track AI visibility as IAB sets measurement standard - The IAB framework requiring providers to disclose how hallucinated mentions are detected rather than filtered out.
- Cloudflare scores brand visibility inside Claude and GPT answers - A dashboard splitting mention rate from citation rate, the two metrics this audit sits beside.
- Nearly half of marketers encounter AI errors weekly as study exposes trust gap - NP Digital survey and 600-prompt accuracy test quantifying how often practitioners meet AI inaccuracy.
- Only 15% of users trust AI search results, Yelp study finds - Consumer verification behaviour that determines whether an incorrect claim gets caught.
- Your analytics are lying: Similarweb traces AI recommendations to real traffic - Panel data linking AI recommendations to site visits arriving mostly through branded search.
- Brands named in ChatGPT's own query win mentions 33x more often - Research published the same day, measuring how internal query construction decides which brands surface.
- SISTRIX April 2026: AI citation drift and the death of keyword research - Weekly citation rotation rates that complicate any frozen-snapshot measurement.
- LLM tracking tools face accuracy crisis from personalization features - How account history and location reshape identical prompts before retrieval begins.
- Gracenote data shows AI hallucinates 1 in 5 streaming titles completely - A parallel accuracy measurement in entertainment metadata, with the same grounding argument.
- The AI slop loop: how fake SEO advice is gaming search results - Documentation of how unverified AI-generated guidance re-enters retrieval corpora.
- ChatGPT referrals to B2B sites gain 303% in a year, Demandbase finds - Referral volume data establishing the commercial stakes attached to assistant recommendations.
- OpenAI opens ChatGPT Ads Manager to all US businesses with CPC bidding - The paid layer now operating on the same surface as organic recommendations.
Summary
Who: Empirank, an answer engine optimization research firm, with the study written by founder James Tandy. The audited entities were 182 identity-validated local businesses recommended by a search-enabled OpenAI configuration.
What: An audit of 1,257 testable factual claims drawn from 1,515 statements. Results: 749 claims confirmed correct (59.6 percent), 479 unsupported (38.1 percent), 24 incorrect (1.9 percent), and five ambiguous (0.4 percent). At business level, 20 of 182 businesses carried at least one incorrect claim, and five carried at least one high-severity error. Identity claims were 99.5 percent correct across 182 claims; accreditation claims were 19.2 percent confirmed across 26. Citation-supported recommendations reached 59.9 percent confirmed correct across 1,199 claims, against 53.4 percent across 58 uncited claims.
When: The response set was captured in August 2026 and frozen. The study was published on 18 August 2026 and sent to PPC Land on 24 August 2026.
Where: The audit examined local business recommendations produced by OpenAI's search-enabled configuration, drawn from an initial pool of 3,000 sampled businesses.
Why: Brand monitoring in AI assistants has been built around whether a name appears. The audit argues that presence measurement leaves the accompanying factual claims unchecked, and that a 1.9 percent claim-level error rate translates into an 11.0 percent business-level exposure that agencies managing named clients carry directly.
Discussion