A researcher enumerated every restaurant, cafe and bar in two Balinese submarkets, then asked four production AI systems 2,208 times where to eat. Most of the market never came up once.

A paper posted to arXiv on August 7, 2026 at 10:23:45 UTC reports the first audit of AI venue recommendation measured against a complete market census rather than a sampled list of prominent brands. The study, titled Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census, was written by Vladimir Pitenin of Norly Research and filed under Information Retrieval and Computers and Society. It runs to 31 pages and 10 figures, carries the identifier arXiv:2608.07069, and reports a headline that has no equivalent in the existing literature: at least 85.6 percent of the venues in the study population were never recommended by any audited system in any run.

The distinction between that figure and the ones circulating in commercial visibility reports is methodological rather than rhetorical. Prompt audits published to date evaluate outputs against curated catalogues of tens or hundreds of well-known brands. According to the paper, a catalogue without a real-world denominator can report relative prominence but cannot produce a population rate, which makes the phrase "never recommended" unmeasurable. Enumerating an entire market supplies the denominator.

A census instead of a sample

The population frame was defined operationally: every place listed on Google Places under five food-service types - restaurant, cafe, bar, coffee shop and bakery - inside fixed polygons covering greater Canggu, including Berawa, Batu Bolong and Pererenan, and greater Ubud, including Penestanan, Sayan and Pengosekan. According to the paper, Google Places was chosen because it functions simultaneously as the dominant consumer discovery layer and as a principal grounding source for the systems under audit.

Enumeration used the Places Nearby Search API across an adaptive grid. The API returns at most 20 results per query, so any cell returning exactly 20 was recursively subdivided into four half-radius cells down to a 130 metre floor. Without that step, according to the paper, dense centres truncate silently and the census undercounts precisely where venues cluster. The completed grid required 760 requests. A second pass probed roughly 220 AI-recommended names that had failed to match anything in the census against Places Text Search, recovering hotel restaurants and coworking cafes that food-type enumeration misses; 70 venues in the final registry carry that lookup provenance. The final count: 4,776 operating venues, each with a snapshot of its Google profile - rating, review count, price level, hours, website and business status - taken at collection time.

Four production assistants were queried through their public search-grounded interfaces: OpenAI on gpt-5.2-2025-12-11 through the Responses API with the web-search tool, Anthropic on claude-sonnet-5 with the server-side web-search tool, Google on gemini-3.5-flash with Google Search grounding, and Perplexity on sonar. Two configuration choices are disclosed. Claude's search budget was capped at two searches per run, a cost decision made before confirmatory collection that bound in practice - 97.6 percent of that arm's runs consumed both searches. Gemini's grounding API exposes no user-location parameter, so location context was carried in the query text for every system.

The instrument comprised eight personas - digital nomad, couple on a date, business meeting host, budget backpacker, family with children, dietary-constrained diner, specialty coffee enthusiast and late-night group - crossed with six first-person paraphrase templates and two areas, producing 96 unique queries. The confirmatory wave ran over seven calendar days with repetitions spread across days: Perplexity 10 per query, OpenAI and Gemini five, Claude three. That yields 2,208 runs, 12,439 valid venue mentions and 1,407,600 venue-level exposure opportunities, against 9,214 successes.

Most of the market is invisible

Those 2,208 runs produced 9,791 venue recommendations. Collectively they covered 689 of the 4,776 census venues. The remaining 4,087 - 85.6 percent - were never recommended once. The share never even mentioned in passing was 84.5 percent.

Narrowing the denominator barely moves the number. Among venues confirmed operational, the never-recommended share is 85.8 percent. Among the 2,173 established venues carrying at least fifty Google ratings, the denominator least open to the charge of padding with marginal businesses, it remains 72.6 percent. Widening the frame moves the figure the other way: a capture-recapture analysis against an independent Foursquare enumeration, corrected for duplicates and matcher recall, produced an estimate of roughly 9,400 food entities across the broadest defensible frame, under which the never-recommended share exceeds 92 percent. Every reported invisibility rate is therefore a floor.

What follows the entry filter is not the winner-take-all pattern the discourse often assumes. The single most-recommended venue, Seniman Coffee in Ubud, accounts for 1.9 percent of all recommendations. The pooled Gini coefficient across recommended venues is 0.668. The top five capture 7.7 percent, the top ten 13.8 percent, the top 25 28.5 percent. Being recommended at all is the scarce event; among the visible, the field is comparatively open.

Two margins, two different rulebooks

The paper's central structural claim is that visibility operates at two separate margins governed by different signals.

Entry

A pre-registered binomial generalised linear model estimated the probability that a venue is recommended in an eligible run, with engine and persona fixed effects, standard errors clustered by venue, and Benjamini-Hochberg correction across nine hypothesis terms. Four factors survived correction, and they share a theme:

  • Own website: odds ratio 1.92, confidence interval 1.40 to 2.62, adjusted p below .001. The largest effect at this margin.
  • Review volume: 1.64 per standard deviation of log review count, 1.37 to 1.97.
  • Listed price information: 1.54, 1.13 to 2.10, adjusted p of .010.
  • Web-mention volume: 1.44, 1.12 to 1.85, adjusted p of .010.

Star rating is null here. Its odds ratio is 0.89 with a confidence interval of 0.76 to 1.03 and an adjusted p of .135. Conditional on review volume and documentation, average rating shows no detectable association with whether a system surfaces a venue at all. Review recency leans positive at 1.41 but misses the corrected threshold with an adjusted p of .054 and is reported as suggestive.

One coefficient is excluded from interpretation by the author. Venues with listed opening hours appear less likely to be recommended, at 0.54 with an adjusted p of .009 - but all 70 venues added to the census through mention-driven lookup lack the hours field, because the lookup request omitted it, and those venues are recommended nearly by construction. The variable partly encodes registry provenance rather than profile completeness. The paper reports it for transparency and treats covariate contamination by the auditor's own discovery path as a general hazard.

Rank

Restrict the analysis to the 1,855 runs that recommended at least two matched venues, and the pattern inverts. A conditional logit over within-run choice sets, covering 8,450 venue-run alternatives, finds star rating significantly predicting first position at 1.17 per standard deviation, with a confidence interval of 1.08 to 1.28. Review volume holds at 1.30. Listed price sits at 1.32 and an own website at 1.33, both attenuated from their entry values. Web-mention volume goes null within choice sets at 0.92. The count of review and blog domains covering a venue, null at entry with an odds ratio of 0.90, becomes positive at 1.23 for first position.

Documentation admits; rating orders. According to the paper, that dissociation reconciles two audit traditions that appear to contradict each other. A pre-registered conjoint published in June 2026 by Baig and colleagues, randomising attributes across twelve models choosing among five hotels, found a top rating raising selection probability by 31.6 percentage points and a high price lowering it by 30.0. That design places every candidate before the model by construction, so it can only measure the within-set margin - the same margin at which this study also finds rating dominant. The census adds an earlier margin no conjoint can pose.

The Foursquare result the study was designed to find

Presence in an open point-of-interest dataset is a widely repeated assumption in local visibility work. According to the paper, part of the study was designed to test it, and the answer is a plain null. Listing on Foursquare returns an odds ratio of 0.84 at entry with an adjusted p of .252. Within the 480 Foursquare-listed venues in the case-control set, the graded ladder is uniformly non-significant: data-quality rating 1.12 with a p of .352, tip count 1.17 at .133, popularity score 1.12 at .308. At the ranking margin the coefficient is nominally negative at 0.86, which the paper declines to interpret because that margin conditions on entry.

Two readings are described as compatible with the data: assistant retrieval does not touch the dataset for venue discovery in this market, or its signal is redundant with the web presence already measured. Neither supports treating dataset inclusion as a visibility lever.

Staleness, not invention

Before variant recovery, 19.7 percent of valid wave mentions - 2,452 - lacked a census match. Every unmatched name was classified through a web-verified taxonomy. According to the paper, 38.4 percent were name variants of registered venues, recovered through 142 evidence-audited mappings; 3.8 percent were venues verified as permanently closed; 2.4 percent were real venues outside the study polygons; and 54.6 percent were unverifiable long-tail singletons or variant candidates below the evidence bar. After recovery, 1,511 mentions, or 12.1 percent of valid mentions, remain unresolved.

There is an arithmetic inconsistency worth noting between the narrative and the figure that illustrates it. The figure caption describes the same pre-recovery pool of 2,452 mentions but reports 1,181 name variants at 41.3 percent, 552 registry-gap mentions at 19.3 percent and 956 unverified at 33.4 percent; those counts, with the smaller classes, sum to 2,860 rather than 2,452. The confirmed-closed count of 93 is identical in both places.

That count is the substantive finding. Outright fabrication is close to absent: a single name, appearing 10 times, survived conservative scrutiny as likely invented, which is 0.08 percent of all valid mentions. Against that, the systems recommended permanently closed venues 93 times across 14 confirmed-closed establishments. According to the paper, "staleness, not hallucination, is the practical failure mode." The generation layer is faithful to what retrieval hands it; the sources themselves are out of date. A closed cafe's reviews, listicle entries and blog mentions persist - the same signals the entry model rewards - so the documentation trail that creates visibility sustains it after the business is gone. A sampled audit would score most of those recommendations as successes.

The commercial consequence of that mechanism has already been visible in the wild. Google's AI overview told customers Anna Mae's Bakery in Millbank, Ontario was closing on July 31, 2026, after confusing it with an Illinois bakery of the same name; the owner learned about it from a screenshot.

Answers that change every run

Identical queries repeated on the same engine return substantially different venue sets. Mean top-20 Jaccard overlap runs 0.45 for Gemini, 0.40 for Perplexity, 0.29 for ChatGPT and 0.22 for Claude. Meaning-preserving paraphrases perturb results further for every engine, sharply for Perplexity at 0.19 against 0.40 and Gemini at 0.30 against 0.45.

Yet a pre-registered test-retest holdout of 16 queries across four engines, re-run two weeks after the wave for 144 runs, found cross-period similarity comparable to the same-period rerun baseline: 0.593 against 0.449 for Gemini, 0.344 against 0.396 for Perplexity, 0.309 against 0.286 for OpenAI, 0.254 against 0.224 for Claude, pooling at 0.375. Share-of-voice rank correlation across 359 venues came in at 0.47, attenuated by design because the retest arm is roughly 2.6 times smaller than the wave arm. The churn is sampling noise rather than temporal drift, which supports treating visibility as a persistent venue property observed through a stochastic channel.

Cross-engine agreement is low. Pairwise top-20 Jaccard ranges from 0.33 between Perplexity and OpenAI to 0.54 between OpenAI and Claude. Only eight venues appear in all four engines' top-20 lists; 15 appear in exactly one. The engines also behave differently in bulk: Perplexity recommended 451 distinct venues at 6.4 mentions per run, Claude 342 at 5.8, ChatGPT 328 at 4.7, Gemini 291 at 4.9. Census match rates ran 91 percent for Gemini, 89 percent for Perplexity, 85 percent for ChatGPT and 84 percent for Claude.

The direct implication the paper draws for the generative engine optimization industry is unusually blunt: single-shot, single-engine visibility checks, which is the standard product offering, measure noise.

What the systems read

The 26,993 grounding citations attached to wave runs span 986 domains. Pooled across engines, the single most-cited domain is not a platform but one venue's own website, finnsbeachclub.com at 5.3 percent - out-citing TripAdvisor at 2.3 percent. Regional listicle and travel-blog domains follow: thehoneycombers.com at 3.9 percent and wanderlog.com at 2.2 percent, with Reddit and YouTube at 1.3 percent each.

Source diets diverge sharply by engine. ChatGPT's most-cited domain is reddit.com at 4.0 percent, followed by tripadvisor.com at 3.2 percent and corner.inc at 3.0. Claude leads with tripadvisor.com at 8.2 percent, then finnsbeachclub.com at 5.4 and wanderlog.com at 5.2. Gemini leads with wanderlog.com at 3.4 percent. Perplexity leads with finnsbeachclub.com at 8.1 percent, thehoneycombers.com at 5.5 and facebook.com at 4.0. A measurement caveat applies to Gemini, whose API exposes only grounding-redirect domains; its true sources were recovered from citation titles by deterministic mapping, with unrecoverable citations dropped.

The audit as instrument

Two of the paper's findings concern auditing rather than AI. First, factor-importance conclusions proved sensitive to entity-matching quality. A remediation pass fixing name-matching errors, with no change to the underlying data, reversed the third-party data-provider coefficient from a significant positive of 1.70 to a null of 0.84, because the original matcher's errors correlated with data-provider absence. Second, the hours-listed covariate inherited structure from the collection path, manufacturing a spurious negative association that survived multiple-comparison correction.

Both were caught by pre-registered validation gates. Mention extraction was validated by independent double annotation with human adjudication, reaching precision of 97.6 percent and recall of 99.1 percent at mention level, with strict run-level agreement of 84.0 percent rising to 91.5 percent after class-based remediation. The pre-registered accuracy gate of 95 percent left its metric level unspecified; the paper reports all three numbers and notes that mention-level metrics clear it while strict run-level does not. Entity matching resolved 87.9 percent of valid mentions, audited at roughly 98 to 99 percent population-weighted correct-entity accuracy on a 100-mention sample weighted toward low-score strata.

The funding is disclosed prominently rather than buried. The study was funded and conducted by Norly, a company selling review-management and AI-visibility tools to hospitality businesses. The insulation measures listed are pre-registration before confirmatory collection, a dataset frozen before models ran, adjudicated validation gates with failures reported, and results that include null findings against commercially convenient hypotheses - the Foursquare test among them, alongside a rating null that cuts against common industry messaging.

Why the marketing community has a stake

The economic chain the paper cites is established for the pre-AI case and now extends forward. Yelp's rounding of ratings to a displayed half-star gave Luca a causal estimate that a one-star increase raises restaurant revenue by 5 to 9 percent, concentrated entirely in independent restaurants. Anderson and Magruder found that crossing a displayed half-star threshold makes a restaurant 19 percentage points more likely to sell out evening seating. Iannelli and Ai, joining clickstream panels to the same users' assistant conversations, measured an AI recommendation lifting same-brand searches by 4.3 percentage points and brand-site visits by 2.4 among previously unengaged consumers.

The audience for these answers is already large. Sixty-five percent of Americans use an AI tool weekly, and separate Yelp research found 57 percent using AI tools for local business discovery monthly while only 43 percent of Gen Z respondents trust AI summaries for restaurant decisions. Platform investment has followed the traffic: ChatGPT added location sharing in March 2026Yelp reservations went live inside ChatGPT in August 2026, and Google Maps added restaurant ordering through Square and Toast in the same month.

The census result lands on a measurement industry that has been sizing the same problem with different instruments and reaching directionally similar conclusions. Semrush analysed 126 million United States prompts and found only 36 of more than 1,200 tracked brands holding top-100 mention status on every platform in every month. Citation concentration in local queries is steep: Yelp accumulated 512,680 citations across four AI platforms in the fourth quarter of 2025, 3.4 times the second-placed platform. What none of those datasets could produce is a population rate, because none enumerates a real market.

The evidentiary backdrop for optimisation claims is thin by academic standards. A critical survey of 45 generative engine optimization studies published in July 2026 concluded that no reviewed technique produces a stable cross-platform effect, and that body-only rewrites can cut a page's top-ten presence by 16 percent. The IAB's August 2026 measurement framework reported that only 16 percent of brands track AI visibility systematically and required vendors to disclose whether prompt libraries are synthetic or observed. The Bali audit adds a second disclosure requirement in practical terms: a visibility number generated from one run on one engine carries a Jaccard reliability between 0.22 and 0.45 against itself.

For the platform side, the finding that structured records govern entry aligns with what Google has been signalling about its own stack. Business Profile data now functions as the layer feeding Gemini, Search and Maps, and Google's own account of how AI decides which businesses get recommended rests on completeness and currency rather than keyword optimisation. The audit measures a related but distinct surface - four assistants querying the open web, not one platform querying its own database - and finds the same class of variable doing the work.

Limits

The paper is explicit about what its design cannot support. The estimates are associations under controls, not causal effects; established, professionally run venues both maintain websites and accumulate digital footprint, so unmeasured scale could contribute to the website coefficient despite controls for review volume and web mentions. The census covers the Google-listed food-service market, so unlisted micro-warungs and atypically typed outlets are underrepresented, which biases the invisibility rate downward. The audit route is the search-grounded API rather than the consumer app, and provider-side serving configurations can change without notice. The setting is two adjacent Balinese submarkets, English-language queries, and a tourist and remote-worker demand profile; replication in a structurally different market is described as required before generalising. Three pre-registered predictors - review-text similarity to query intent, cross-platform rating consistency and social presence - were not operationalised, a data-collection gap the paper discloses rather than works around.

The protocol, registry-construction method and derived data are released for replication.

Timeline

Summary

Who: Vladimir Pitenin of Norly Research, a company selling review-management and AI-visibility tools to hospitality businesses, which funded and conducted the study and discloses the competing interest.

What: A census-denominated audit of AI venue recommendation covering 4,776 restaurants, cafes and bars, evaluated against 2,208 search-grounded runs from ChatGPT, Claude, Gemini and Perplexity across 96 persona-conditioned queries. At least 85.6 percent of venues were never recommended by any system, and 72.6 percent among venues with fifty or more ratings. Entry into answers tracks documentation - own website at odds ratio 1.92, review volume 1.64, listed price 1.54, web mentions 1.44 - while star rating is null at 0.89 and predicts only first position, at 1.17. Foursquare presence shows no positive effect at either margin. Fabrication accounted for 0.08 percent of mentions; permanently closed venues were recommended 93 times across 14 establishments.

When: The confirmatory wave ran over seven days, with a test-retest holdout two weeks later. The analysis snapshot was frozen on August 3, 2026, and the paper was submitted to arXiv on August 7, 2026 at 10:23:45 UTC.

Where: Two bounded Balinese submarkets, greater Canggu and greater Ubud, defined by fixed geographic polygons, with English-language queries.

Why: Recommendation carries measured demand consequences in food and drink, where independent operators dominate and documentation is thinnest. Without a real-market denominator, no prior audit could report how many businesses never appear at all. The result reframes what visibility work is measuring: entry is governed by infrastructure a business controls, ordering by reputation, and single-run visibility checks by noise.