Fewer than half of the most popular origins on the web expose an XML sitemap where lintlab's crawler looked for one, and most of those that do carry at least one measurable flaw in their URLs, dates or markup, according to data lintlab released today. A separate pass by the same crawler found only 66 of 1,000 origins serving an llms.txt file, the plain-text summary proposed for AI systems.

In Short

A small company that sells website-checking software scanned the 1,000 most-visited websites in Google's Chrome data to see whether they publish a sitemap, the list of pages that helps search engines find content. It found one on fewer than half of the sites that answered, problems in most of the ones it found, and an AI-oriented file called llms.txt on just 66 sites. For anyone who pays for search traffic or runs a large site, the figures show that basic signals to search engines are often imperfect even at the very top of the web, though the firm behind the numbers also sells a tool that checks for them.

The study and who ran it

lintlab describes itself as a maker of tools that check web pages, sites and documents. lintlab disclosed that it sells an XML sitemap checker on the Apify marketplace, and the study page, last updated today, links to it. Those details matter for how much weight the numbers can bear: they are vendor-supplied, they have not been independently verified, and lintlab publishes aggregates only. "We publish aggregates only and name no site," the page states.

The population is unusually well defined. lintlab took the global August 2026 snapshot of the Chrome UX Report, or CrUX, from the crux-top-lists repository on GitHub, a 1,000,000-row file, and kept the 1,000 origins in its top-1k popularity bucket. It published SHA-256 digests of both the source file and the extracted list, so anyone can confirm they are measuring the same set. CrUX groups origins by popularity among Chrome users who opted in to syncing their browsing history and sharing usage statistics. Origins inside a bucket carry no order, and a site's www host and each of its subdomains count as separate entries. The data is licensed under CC BY 4.0, and lintlab is explicit that the sitemap measurements are its own: "Google has not reviewed or endorsed this study." The same dataset has been widening in other directions, too; Chrome added four experimental ad metrics to it on September 15, 2026.

The crawl itself ran from 22:15:40 to 22:47:14 UTC on October 1, a little over half an hour. Each origin was measured exactly as listed, without switching between www and the bare domain; one listed origin used plain HTTP. The crawler, identifying itself as lintlab-sitemap-study/1.0, sent one request per second per host, ran up to 20 origins in parallel, used a 15-second timeout and made at most two requests to any one URL. It rendered no pages and used no logins. In total it fetched 1,677 sitemap files containing 13,894,476 URL entries and ran 7,456 status checks. Parsing relied on saxes 6.0.0 for XML and robots-parser 3.0.1, gzip was detected from a file's first bytes rather than its headers, and lintlab reports that 36 automated tests passed.

Fewer than half, and that is a floor

Of the 1,000 origins, 916 responded as websites. lintlab found a sitemap on 400 of them, or 43.7%. Every origin was assigned exactly one outcome:

  • Sitemap, no issue found: 112 origins (12.2% of 916)
  • Sitemap with issues: 248 (27.1%)
  • Sitemap, not fully checked: 40 (4.4%)
  • No sitemap at the places checked: 256 (27.9%)
  • Blocked: 165 (18.0%)
  • Robots rules prevented the fetch: 77 (8.4%)
  • Sitemap fetch error: 18 (2.0%)

Within the group where a sitemap was found, the split is 112 clean (28.0%), 248 with at least one issue (62.0%) and 40 that could not be fully examined (10.0%).

Discovery started with robots.txt, where 417 origins (45.5% of 916) declared a sitemap through a Sitemap: line. Another 71 of the sitemaps lintlab found came from the default paths /sitemap.xml or /sitemap_index.xml rather than from a robots.txt declaration, which by subtraction means 329 were reached through robots.txt. The page does not reconcile that 329 with the 417 declarations line by line. The likely home of the difference is the blocked, disallowed and fetch-error groups, but that is an inference from the published totals, not a figure lintlab reports.

How large is the true share? The answer depends on the denominator. Among the 656 origins where the crawler was neither refused, disallowed by robots rules nor tripped by a fetch error, 400 had a sitemap, which works out at 61.0%. That calculation is PPC Land's, from lintlab's published counts. It shows how much the headline 43.7% is held down by origins lintlab simply could not see.

Fifty origins also answered at least one sitemap address with an ordinary HTML page and a 200 status, 75 such responses in all. That is a soft 404: the status code reports success while no sitemap exists at the address. The robots.txt file itself was reachable on 779 of the 916 websites (85.0%).

Why the denominator is 916

The 84 excluded origins gave lintlab neither a robots.txt file nor a page it could read. For 50 of them, robots.txt did not answer within 15 seconds. RFC 9309, the standard that formalised the robots exclusion protocol, treats an unreachable robots.txt as an instruction to assume everything is disallowed, so lintlab did not request their homepages either. Most of the rest failed with other timeouts or with connection, TLS, HTTP/2 or DNS errors. lintlab acknowledges that some of these may still serve pages to ordinary browsers.

Across all 1,000 origins, robots.txt returned an ordinary 4xx on 75, a 403 or 429 on 87 and a 5xx on 2, and could not be reached on 84. Under RFC 9309 a plain 4xx means a crawler may fetch anything. lintlab applied that rule but chose to treat 403 and 429 more cautiously, counting them as blocked and not retrying.

Blocked is not missing

"Blocked is not the same as missing," the study notes, and the 165 refused origins are the main reason the sitemap figure is a floor. An origin counted as blocked when robots.txt or a sitemap request returned 403 or 429, carried Cloudflare's cf-mitigated header, or matched a challenge signature. On 200 and 503 responses that signature had to be strict, such as a challenge page titled "Just a moment" or "Access denied". A 404 or 410 never counted as a challenge. Challenge pages served with a 200 status decided only 2 of the 165 cases; the rest came from 403, 429 or other error responses.

The first count was wrong, and the page says so. An early rule treated some 404 pages carrying a Cloudflare script as blocked. After fixing it, lintlab re-fetched the 191 origins it had first called blocked, about an hour after the crawl. Of those, 165 remained blocked, 23 had no sitemap at the checked places, 2 gave no website and 1 timed out. The re-fetch finished by 23:48 UTC on October 1. A full repeat crawl run less than two hours before the published one, with identical settings but before a separate fix to how failed status checks were counted, put 106 of the 1,000 origins into a different outcome, though no outcome's total moved by more than 8.

None of that is surprising for a third-party crawler. Large sites have spent two years hardening their defences: 35.7% of the top 1,000 websites were blocking OpenAI's GPTBot by August 2024, and Cloudflare began blocking training and agent crawlers by default on ad-carrying pages for newly onboarded domains from September 15, 2026. A study that refuses to bypass challenges will always undercount on the current web.

The XML parses; the URLs and dates do not

The study's sharpest observation is about where the faults sit. "The XML itself was rarely the problem. The URLs and dates in it were," lintlab writes. Malformed XML turned up on only 4 sites, and a wrong or missing namespace on 15.

The large sites also tend to build sitemaps in layers. Of the 400 with a sitemap, 292 (73.0%) began with a sitemap index, a file that lists other sitemap files, and 278 (69.5%) served at least one gzip-compressed file in the sample. Each site counts once per issue type, and one site can carry several:

IssueSitesShare of 400
Sampled URL answered with a redirect9223.0%
One lastmod value on every URL in a file8822.0%
URL on a different host from the listed origin7218.0%
Same URL listed more than once6917.3%
Sampled URL answered 4xx (not 403/429)4110.3%
Wrong or missing namespace153.8%
lastmod in the future143.5%
Scheme mismatch (http vs https)82.0%
Malformed XML41.0%
Sampled URL answered 5xx41.0%
Invalid lastmod30.8%
URL containing a #fragment20.5%
More than 50,000 entries in one file10.3%
More than one loc in one entry00.0%

lintlab did not score the priority or changefreq fields. Google ignores both, according to the study, and the sitemaps.org protocol treats them as hints.

Redirects in the list

The most common fault was a listed URL that answered with a redirect, on 92 sites. lintlab counted the first hop without following it. The reasoning offered on the page rests on Google's documentation: Google asks site owners to list the URLs they want shown in results and generally shows canonical URLs, so a redirecting address is not the final one.

Google's own staff have not treated redirects as an emergency. John Mueller advised against investing heavily in redirect-chain audits on February 3, 2026, arguing that problems usually show up in ordinary browsing. Yet redirects remain one of the inputs to canonical selection, and Google's troubleshooting guide lists a 3xx redirect among the ways an unexpected canonical preference can arise. Google also added clarifications on how long canonical changes take to be re-evaluated on July 10, 2026. A sitemap that points at the pre-redirect address sends a weaker signal into that process than one listing the destination.

One date for every page

The second most common fault is subtler. On 88 sites, every entry in at least one multi-URL file carried a lastmod value, and all of the values were identical. Such a date may record when the file was generated rather than when each page changed. According to lintlab's reading of Google's documentation, Google uses lastmod only when it is consistently and verifiably accurate.

The aggregate figures are large. Of the 13,894,476 URL entries in the sampled files, 6,331,629 (45.6%) carried a lastmod, and 99.3% of those parsed as valid dates. Some were valid and impossible: 53,804 dates, spread over 14 sites, lay later than the moment of the check. Only 3 sites had lastmod values lintlab rejected as invalid, and lintlab notes that its test is stricter than the W3C Datetime format, which also permits year-only, year-month and minute-precision values that lintlab would flag.

The field matters to both large search engines. Google's Gary Illyes said in May 2024 that the last-modified date in sitemaps remains a signal of site activity, though not the only one governing crawl frequency. Microsoft's Fabrice Canel and Krishna Madhavan went further on July 31, 2025, when Bing described sitemaps and accurate lastmod values as key inputs for deciding what to recrawl in AI-assisted search, or what to skip. A date that never varies tells neither engine anything.

Hosts, duplicates and size limits

On 72 sites, a listed URL sat on a different host from the origin, and lintlab treated www and the bare domain as different hosts. The sitemaps.org protocol says all URLs in a sitemap must come from a single host, and Search Console reports such entries as "URL not allowed" unless cross-site submission covers them. lintlab did not check for cross-submission through robots.txt or for verified multi-domain setups, so some of these may be deliberate. The www question has its own history: Google's site move guide was extended in June 2026 to cover www and non-www variants separately.

Duplicates affected 69 sites. A repeated entry adds no new URL but still counts towards the 50,000-entry ceiling that sitemaps.org and Google both set per file. Only one site breached that ceiling in a sampled file. Eight sites listed URLs on a different scheme from the origin, two sites used URL fragments, which Google Search generally does not treat as separate pages, and no entry in the sample carried more than one loc element.

What the sampled URLs returned

lintlab picked up to 20 listed URLs per site by hash and sent each a HEAD request over HTTP/1.1, falling back to GET only when HEAD returned 405 or 501 and stopping that GET after 64 KiB. Of the 7,456 checks:

  • 6,135 (82.3%) returned 200, on 350 sites
  • 863 (11.6%) redirected, on 92 sites
  • 196 (2.6%) returned 403 or 429, on 15 sites
  • 182 (2.4%) returned another 4xx, on 41 sites
  • 44 (0.6%) returned a 5xx, on 4 sites
  • 35 (0.5%) ended in a request error, on 4 sites
  • 1 returned another 2xx code

lintlab kept 403 and 429 apart from other client errors because they can mean the checker was refused rather than the page missing. The 35 request errors break down into 20 where the server closed the connection without replying, 13 GET fallbacks that ran past the 64 KiB cap after a 200 header had arrived, and 2 timeouts at 15 seconds. They did not change any site's outcome, and 2 of the 112 clean sites had one. While filling each sample, the crawler skipped 5,470 URLs because of robots.txt: 5 were disallowed for its user agent and 5,465 sat on hosts whose robots.txt it could not read. Thirty-four sites with a sitemap received fewer than 20 checks, and 12 received none.

The llms.txt count

The figure lintlab chose for its email subject line came from a second pass, run today. The crawler sent one request, with no retries, to /llms.txt on each of the same 1,000 origins. Sixty-six (6.6%) answered with HTTP 200 and a plain-text or Markdown content type. Another 627 (62.7%) returned 404. A further 112 returned a 200 status with some other content type, usually an HTML "not found" page, and were not counted. The study does not break down the remaining 195 responses.

The content-type filter does real work. By PPC Land's arithmetic, a scanner counting every 200 response would have reported 178 origins, or 17.8%, nearly three times the figure lintlab published. lintlab is also careful about what the 66 files prove: "Publishing the file doesn't show that any AI system reads it."

That caveat matches the rest of the record. PPC Land reported in July 2025 that adoption had stalled because no major AI provider parsed the file. A ProGEO.ai study published on March 31, 2026 found llms.txt on 37 of the Fortune 500, or 7.4%, against 92.8% with robots.txt and 76% with at least one Sitemap directive. Originality.ai counted 36,120 llms.txt instances by May 2026, up from 4,088 a year earlier, while Ahrefs server logs from 137,000 domains showed 97% of the files receiving no requests at all that month, with AI retrieval bots accounting for 1.1% of the requests that did arrive and SEO audit tools for 21.7%.

Google's position has hardened in the same period. Its AI search guide, published on May 15, 2026, said there is no requirement to produce an llms.txt file to appear in search, and on June 15, 2026 Google added a note that the files neither help nor hurt rankings. The Chrome team, meanwhile, put an llms.txt audit into Lighthouse on May 5, 2026; it flags server errors and marks a missing file as not applicable.

Is 6.6% high or low? The comparisons are imperfect. The Fortune 500 study measured corporate domains; CrUX counts origins, so a single company's www host, app host and regional subdomains appear separately, and the top-1k bucket mixes content sites, platforms and services. The two figures sit in the same range, which suggests the most prominent sites are no keener on the file than the largest companies. The gap with sitemap declarations is the clearer finding: in the same population, 417 origins declared a sitemap in robots.txt, more than six times the number serving a valid llms.txt.

The limits lintlab lists

The page is unusually candid about what it did not measure. From each sitemap index the crawler fetched at most 5 child sitemaps, chosen by hash, with no more than 8 files and a depth of 3 per origin, and it used at most 20 Sitemap: lines per robots.txt. As a result 248,513 index children went unfetched, and 1,296 Sitemap: lines on 20 origins went unused. File-level issue counts are therefore floors, while the redirect and error counts depend on which URLs the hash happened to select and could move either way.

"No sitemap" means none at the robots.txt lines and the two default paths; a site may submit a sitemap directly to search engines at another address. Files were capped at 50 MiB uncompressed, a cap no origin reached. The whole exercise is one crawl from one network location, and live responses change.

The page also records Google's own caveats, as lintlab reads them: a site of about 500 pages or fewer with good internal links may not need a sitemap, and a sitemap helps discovery without guaranteeing that every URL is crawled or indexed. It closes with an eight-point checklist built from the same checks, ranging from robots.txt declarations and namespaces to lastmod discipline and replacing redirecting URLs with their destinations.

The commercial interest

The study doubles as a demonstration of a product. lintlab's XML Sitemap Checker on Apify finds sitemaps through robots.txt, follows indexes and gzip files and returns each URL as a JSON record with any issues found. Pricing is per event: $0.001 per sitemap file parsed, $0.0003 per URL extracted and $0.0005 per URL status check, with extraction of 10,000 URLs on the free plan costing $3.00 plus $0.001 per file, according to lintlab.

Four of the study's checks are not in the product: the exact namespace URI, scheme mismatches, fragments and the single-lastmod test. The last of those identified the second most common problem in the entire study, on 88 sites. That gap cuts against any reading of the research as a simple sales funnel, but it does not remove the conflict. The figures come from a party with something to sell, about sites nobody outside lintlab can identify.

Why this matters for marketers

Sitemaps are old plumbing. Google's Martin Splitt and Gary Illyes noted in April 2025 that sitemap files, created around 2005-2006, have never been formally standardised, which helps explain the variety lintlab found in how the biggest sites implement them. They are also the lever site owners have over recrawl timing. Google's crawling overview of March 3, 2026 named sitemap files as the main way to signal new and updated pages, and described recrawl intervals ranging from every few minutes for breaking-news homepages to a month for unchanged pages. For retailers, that interval decides whether the price Googlebot last saw matches the one at checkout.

That makes the lastmod finding the one with the most commercial weight. A product feed or news section whose sitemap stamps every URL with the same date is supplying a freshness signal that, on Google's stated terms, may be discarded, and Bing has said it leans on the same field more heavily as answers are generated by AI. Redirecting entries and stray hosts are cheaper faults, but they feed into canonical selection, the step in indexing that decides which URL is shown and whose traffic appears in which Search Console property.

The llms.txt figure adds another independent count to a debate that has so far produced plenty of adoption data and almost no consumption data. Sixty-six files among a thousand of the most visited origins on the web is not a groundswell. It is also not evidence of anything happening downstream, which is the question marketers paying for AI-visibility work need answered, and which neither lintlab nor any earlier study has settled.

Timeline

Summary

Who: lintlab, a seller of web-checking tools including an XML sitemap checker on Apify, ran and published the study. It names no spokesperson, and the email carrying the results to PPC Land stated that an AI agent wrote it. The subjects are the 1,000 origins in the top popularity bucket of Google's Chrome UX Report.

What: Of the 916 origins that responded as websites, lintlab found a sitemap on 400 (43.7%), a figure it calls a floor because 165 origins refused its crawler and robots rules kept it out of 77 more. Of the 400, 248 (62.0%) had at least one measured issue, led by redirecting URLs on 92 sites and a single lastmod value repeated across a file on 88. A separate check found 66 of 1,000 origins (6.6%) serving a valid llms.txt file.

When: The sitemap crawl ran on October 1, 2026, between 22:15 and 22:47 UTC, with a re-fetch of 191 origins finished by 23:48 UTC. The llms.txt check ran today, and the study page was updated today.

Where: The source list is the August 2026 global CrUX snapshot from the crux-top-lists repository on GitHub. The crawl ran from a single network location, and the results are published as aggregates on lintlab's website.

Why: Sitemaps are the main signal site owners send search engines about new and changed pages, and both Google and Bing say they rely on accurate lastmod dates. The study suggests many of the most visited origins send that signal imperfectly, while the llms.txt count adds to evidence that the AI-oriented file remains rare and unproven. The figures are vendor-supplied and unverified, and lintlab sells a tool covering most of the checks.