A sitemap that validates, loads and sits in plain view can still go unread by Google. In a podcast episode published on October 1, 2026, John Mueller of Google's Search Relations team said a "couldn't fetch" status on such a file often has nothing to do with the XML: Google's systems may be too busy, or may see too little demand for a site's content to want it, a judgment he tied to how Google perceives the site's quality.
In Short
Websites keep a list of their pages, called a sitemap, so Google can find new and changed content, and on October 1, 2026 two Google staff explained how Google actually uses those lists. If Google says it could not read your list, the file may be fine, because Google may simply be busy or may not want more from a site it does not rate highly. So a tidy file is no guarantee of attention, and newer AI-oriented files such as LLMs.txt do not replace it.
A fetch error that is not always an error
Episode 114 of Search Off the Record, titled "Do sitemaps still matter?", went up on the Google Search Central channel on YouTube on October 1, 2026. It runs to roughly 26 minutes across 14 chapters and pairs Mueller with Martin Splitt, both members of the Search Relations team. The channel lists 807,000 subscribers; the video had drawn 3,110 views, 61 likes and eight comments when the page was captured. Splitt opened by saying he gets emails from site owners worried by what they see in their reporting, and he held the central question back for the final minutes.
That complaint concerns Search Console. Splitt asked why the tool reports "couldn't fetch" for a file that validates, is publicly reachable and is declared in robots.txt. Mueller's answer was not about XML at all. "We actually see this a lot in the forums. And there are primarily two reasons for that," he said, according to the transcript Google published alongside the video.
The first is host load, the ceiling on how much Google's systems are prepared to request from a server. "It could be the case that we don't have any time to crawl this sitemap file because we're too busy with other things," Mueller said, adding that Google reports that outcome as "couldn't fetch" too. The second is crawl demand. When Google's systems conclude there is no real need to crawl much more from a site, the sitemap goes untouched. "And the crawl demand is very often based on the perceived quality of a website. And that can have a really large impact on how much we crawl and index from a website," he said. "So it's not purely a technical thing."
The status, in other words, can be a verdict. Mueller described the sequence plainly. Systems that assume a site's overall quality is "not fantastic" spend less time crawling and indexing it, and so do not bother with its sitemap for the moment. If quality improves significantly over time, Google will return to the file, "but maybe we just don't want to." According to the episode description published by Google Search Central, when a valid, accessible file is reported as unfetchable, "the cause is often host load throttling or low crawl demand linked to perceived site quality."
Neither half of that explanation is new on its own. In July 2025, Mueller told a site owner on Bluesky that a technically sound but barely indexed site often means Google's systems have doubts about the site as a whole. Host load is the better documented of the two. Google's crawling overview of March 3, 2026 states that crawl rates drop automatically when servers slow down or return errors, and in episode 105, published on March 12, 2026, Gary Illyes and Splitt described Googlebot as one client of a shared internal crawling platform with automatic throttling. Throttling can also start on Google's side. Between August 8 and 28, 2025, crawl rates fell sharply for sites hosted on Vercel, WP Engine and Fastly, a disruption Google attributed to its own systems.
Microsoft describes a different cadence for Bing. Its July 31, 2025 guidance said Bing fetches a submitted sitemap immediately and revisits it at least once a day.
Two fields gone, one on probation
Much of the episode concerned what Google reads inside a sitemap and what it discards. Two optional fields from the original format no longer count. The first was priority. "There used to be a priority field in sitemaps. And you can imagine if you tell SEOs and ask them what is their priority for their pages, everything is number one," Mueller said. "Every URL is a maximum priority. It ended up not being very useful."
Change frequency failed for a subtler reason. Asked what SEOs entered there, Splitt guessed "Always fresh." Mueller agreed but declined to call it deception: on a server-side dynamic website, the PHP file behind each page runs fresh on every request anyway. "So theoretically it's correct, but it's not very useful." What survives is narrower. "Basically, what we primarily focus on is the URL and the date, the change date," he said.
That date, the lastmod field, is on probation. "My understanding is we want to use it, but a lot of sites get it wrong. So we have this weird love-hate relationship, I think, with it," Mueller said. Where the dates within a file look reasonable, Google takes them into account. Where they do not, it falls back on the URLs themselves and processes new ones as they appear. A site that stamps every URL with the current date simply has its dates disregarded, and the practice carries no penalty: "this is not like a spam thing where our spam systems would say this is a bad site." According to the episode description, Google "evaluates date reliability and will ignore lastmod signals if they are inaccurate or abused."
One detail on the YouTube page sits awkwardly with that account. The chapter marker at 04:46 is labelled "Deprecated Fields: Priority, Change Frequency, and lastmod," which groups lastmod with the two abandoned fields. Neither the conversation nor the written description treats lastmod as deprecated; both describe it as evaluated case by case.
Google's line on the field has been consistent. Gary Illyes said in May 2024 that the last-modified date remains a signal of site activity, though not the only factor in crawl frequency. Microsoft goes further. Its July 2025 guidance called lastmod a key signal for deciding which URLs Bing recrawls and which it leaves alone, required ISO 8601 formatting with both date and time, and still listed change frequency among the structured signals AI-assisted search relies on. That is the very field Google says it set aside.
Small sites, news sites and shops
Does every site need one? Mueller's answer was a qualified no followed by a practical yes. "I think smaller websites probably don't need a sitemap because we can just crawl them," he said, before conceding that owners struggle to place themselves on the scale. Is a 50-page site small, medium or big? Since almost every content management system, static hosting tools included, now generates the files by default unless switched off, his usual recommendation is to leave the feature on. "There's no harm in having a sitemap file if your site is small."
Larger publishers are a different case. Mueller pointed to the dedicated news sitemap format, in which, he said with some hesitation, publishers are "only supposed to list the last 1,000 pages that changed". For any site with "a non-trivial amount of content" that changes regularly, he called the file "super helpful".
E-commerce gave him his clearest example. "If you have an e-commerce site, that's pretty common, like the price changed. By the time we notice with normal crawling, maybe that's a bit late." A sitemap entry lets Google go straight to the changed product page instead of rediscovering it through a category section. Google's March 2026 crawling overview made a similar case, naming sitemaps as the main mechanism site owners have to influence recrawl timing and noting that breaking news homepages may be recrawled every few minutes while unchanged pages can wait a month. Large companies appear to have absorbed the point. A ProGEO.ai study published on March 31, 2026 found that 76% of Fortune 500 companies declare at least one sitemap in their robots.txt file.
Canonical hints and timestamped URLs
Splitt floated a workaround: adding timestamps to the URLs in a sitemap to steer Google towards the newest version. Mueller rejected it. The file, he said, is meant to contain the URLs a site wants treated as canonical, written the way the site wants them indexed rather than with internal tracking or date-stamp additions. If a static-looking URL carries a date stamp, Google will try to index the version without one as the canonical.
The sitemap also feeds canonical selection itself. "My understanding is we also use sitemap files with regards to choosing the canonical," Mueller said. Where Google discovers several similar URLs, listing one of them makes it "a little bit more likely that we will pick that as canonical."
The hedge matters. Google updated its canonicalization troubleshooting guide on July 10, 2026 to clarify how long re-evaluation takes, and the practical upshot was that canonical fixes can take around two weeks to register, with even an explicit rel="canonical" element offering no guarantee of the outcome. A sitemap entry is one input among several, not an instruction.
RSS as a short-form sitemap
The episode's most practical comparison concerned RSS feeds, which blogging systems often generate without being asked. "It's very similar. I mean, if you ask those who made the RSS standard, they will say it's very different," Mueller said. Treated purely as a list of URLs with dates, a feed carries the most recently changed pages, often capped at 10 or 20 entries, each stamped with a date and time. "So, in a sense, an RSS file can be used as a sitemap file," he said, adding that a feed can be submitted in Search Console as one.
The difference is scope. A large site may be split across 1,000 sitemap files because each is size-limited, and a system looking for changes would have to read them all first. A feed lets the same system see at a glance that, say, 50 pages changed recently and go there. Mueller's conclusion was that each format has its use for finding new and updated pages. Feeds are also easier to locate, because they are usually linked from the head of a site's HTML pages.
The endorsement comes after Google's own news product moved away from submitted feeds: Google News stopped using feeds supplied through Publisher Center in March 2025.
50,000 URLs, 50 megabytes and a cat
Splitt then tested Mueller on the format's limits. "I should look it up before telling you, but my memory is 50,000 URLs and 50 megabytes," Mueller said, specifying that the size limit applies to the uncompressed file. Sitemaps can be gzipped, but compression does not raise the ceiling. Site owners can submit any number of files in Search Console, link them from robots.txt, or group them through a sitemap index file that points to other sitemaps. "You can only nest them once, but you can submit as many sitemap index files as you want."
The episode description presents the 50,000-URL and 50MB uncompressed figures as exact limits, while Mueller offered them from memory with an explicit caveat. The numbers in the two sources agree; the certainty does not. Microsoft publishes its own arithmetic for the same structure: 50,000 URLs per file and up to 50,000 child files per index, or 2.5 billion URLs through a single index file.
Naming is flexible. A file called sitemap.xml gets guessed by many systems, but the name hardly matters to Google: martinscat.xml would do! Mueller suggested the .xml extension is probably not even required, only common. Some site owners prefer to keep their sitemaps out of public view, which Mueller described as "perfectly fine". They can use an unusual name and tell Google about it directly. The cost is that other systems cannot find the file, which means Bing probably needs its own submission.
What AI crawlers can find
That cost extends to AI companies, and here Mueller offered observation rather than documentation. "That also means if you care about AI training crawlers, if you want to make sure that your content is in these AI systems, they usually don't have a console or any setup where you can submit a sitemap file," he said. Without a submission route, discovery depends on a generic file name or on feeds. "I don't know if they document this anywhere, but I've seen that happen in my server logs where some AI crawler accesses my sitemap file and they read it, they do something with it." He reported the same pattern for his RSS files, and closed the thought with a qualification: "And that's perhaps something that the site owner wants."
The composition of AI crawling gives the remark some weight. Cloudflare figures cited in IAB Australia guidance put roughly 52% of AI crawler requests toward model training, against about 2.6% representing real-time fetches triggered by a person's question. Measurement tools have begun to surface the traffic Mueller described from his own logs. Microsoft Clarity's Bot Activity dashboard, released on January 21, 2026, reports which AI systems request a site's content and which pages receive the most automated requests, and the XML files that carry sitemaps and feeds are among the resources those crawlers use for discovery.
HTML sitemaps and LLMs.txt
Two other formats sometimes described as sitemaps got shorter shrift. An HTML sitemap, Mueller said, "is basically almost like a map of your website for users. It's not something that replaces an XML sitemap file." Crawlers can follow its links like any others, but it lacks the strict structure needed for submission. "It's basically a collection of links," and on a commerce site it typically lists categories rather than every product.
Asked whether LLMs.txt would remove the need for a structured XML file, Mueller began with "No." He described LLMs.txt as a Markdown file, sometimes carrying links to parts of a site, intended to help AI systems understand that site better. Functionally, he placed it closer to an HTML sitemap than to the XML format. "I think the hope is bigger than the reality, but maybe that will change at some point," he said. Google's systems cannot process it as a sitemap because it lacks the strict format. Search systems might one day read Markdown files and follow their links, he allowed, but "currently none of this happens."
His position on implementation was indifferent rather than hostile. "So it's fine if you want to play around, or if your CMS makes LLMs.txt files automatically, but I would not rely on it," Mueller said. "I would not treat that as a priority." Sitemaps, he added, will not lose their usefulness completely, because they are comprehensive in a way an LLMs.txt file is not meant to be.
The file has a short and contested history. Jeremy Howard proposed it on September 3, 2024, and by July 2025 analysis from Ahrefs found no major model provider parsing it. Google has since said twice in its documentation that the file does nothing for Search: its AI search guide of May 15, 2026 said none is required, and a June 15, 2026 note said such files will neither help nor hurt visibility or rankings. Yet Chrome, another Google product, added an llms.txt check to Lighthouse on May 5, 2026 under a new agentic browsing audits section. Mueller himself, in episode 113 at the end of July 2026, listed an LLMs text file alongside sitemap files and HTML sitemaps as things a well-structured site may have.
Publishing continues regardless. Originality.ai counted 36,120 llms.txt files by May 2026, 8.8 times the figure a year earlier, while Ahrefs server logs from 137,000 domains showed 97% of files receiving no requests that month. Among the Fortune 500, the ProGEO.ai study counted 37 companies, or 7.4%, with a file in place.
Extensions still in the format
Near the close, Splitt raised hreflang annotations, which can sit inside sitemaps to connect the language versions of a page. "I'm wondering actually how many people are using that," he said, and asked listeners to say in the comments whether they use sitemaps or other mechanisms. Mueller added the image and video extensions. "I don't know how important those are nowadays because it's easier for us to recognize images and videos on pages in the meantime, but it is something that is still supported."
The team covered hreflang in a July 25, 2024 episode, in which it said the annotations can be implemented through HTML tags, HTTP headers or XML sitemaps and cited Web Almanac data showing about 9% of websites using hreflang in 2022.
A semi-standard at 20
The episode opened as personal history. Sitemaps were "basically how I made my way to Google", Mueller said. The format had been shared with Microsoft and Yahoo and emerged as "this semi-standard thing that everyone agreed upon, or at least like three companies agreed upon." At around the same time, Mueller built a Windows-based generator that crawled a site like a normal crawler and wrote the URLs it found into a sitemap file. That was unusual then. Many site owners had never crawled their own sites, and PHP-based sites of the period, before WordPress and other mainstream systems produced clean URLs, threw up parameter loops and addresses that made little sense. Helping people untangle those problems drew him into the early SEO community and, eventually, to Google's attention.
Was that 500 years ago, Splitt asked? "Seriously, this must be like 20-ish years," Mueller replied.
The record broadly supports him. Search Console itself began on June 2, 2005 as Google Sitemaps, an XML submission tool, before being renamed Google Webmaster Tools in August 2006 and Search Console in May 2015. Formal status never followed. In an episode released on April 17, 2025, Splitt and Illyes noted that sitemaps, created around 2005 and 2006, have never been formally standardised, whereas robots.txt was eventually standardised by the IETF after decades as a de facto convention.
Why this matters for marketers
Sitemaps rarely feature in media plans, yet the mechanics in episode 114 reach paid and organic work alike. Advertisers who change prices, promotions or stock levels on landing pages depend on Google noticing quickly, and a lastmod field that Google has learned to distrust weakens one of the few freshness signals a site controls directly.
The quality point cuts in a less comfortable direction. If a "couldn't fetch" status can reflect weak crawl demand, then a clean, validated file offers no protection for a site Google has chosen not to prioritise, and no amount of XML repair changes that. The explanation fits a pattern Google has repeated for years, from Splitt's August 2024 account of the "Discovered - currently not indexed" status to Mueller's 2025 comments on barely indexed sites. It also leaves site owners reading a single error message that does not distinguish between their own server, a Google-side throttle and a quality judgment. For agencies reporting technical health to clients, that ambiguity is now on the record.
For publishers weighing AI exposure, Mueller's server-log anecdote is a reminder that discovery by AI crawlers runs largely through conventions, such as generic file names and feeds linked from page heads, rather than through any submission process. An obscure sitemap name chosen to keep the file from competitors has effects well beyond Google. And for teams under pressure to publish LLMs.txt files, a Google representative has said again, on the record, that Google's search systems do not use them at present.
Timeline
- May 22, 2024 - Gary Illyes says the lastmod date in sitemaps remains a signal of site activity, though not the sole driver of crawl frequency.
- July 25, 2024 - A Search Off the Record episode explains that hreflang can be implemented through HTML tags, HTTP headers or XML sitemaps.
- August 21, 2024 - Martin Splitt explains that the "Discovered - currently not indexed" status is not necessarily an error.
- September 3, 2024 - Jeremy Howard proposes the llms.txt file.
- April 17, 2025 - Splitt and Gary Illyes note that sitemaps have never been formally standardised.
- July 2, 2025 - llms.txt adoption stalls as major AI platforms decline to parse the file.
- July 15, 2025 - Mueller tells a site owner on Bluesky that barely indexed but technically sound sites usually point to doubts about the site overall.
- July 31, 2025 - Microsoft repositions sitemaps and accurate lastmod values as key inputs for AI-assisted search.
- August 8-28, 2025 - Google crawl rates fall sharply for sites on Vercel, WP Engine and Fastly, a fault Google attributes to its own systems.
- January 21, 2026 - Microsoft Clarity adds Bot Activity tracking for AI crawler traffic.
- March 3, 2026 - Google publishes a crawling overview naming sitemaps as the main lever over recrawl timing.
- March 12, 2026 - Search Off the Record episode 105 describes Googlebot as one client of a shared crawling platform.
- March 31, 2026 - ProGEO.ai finds llms.txt at 7.4% of Fortune 500 companies and a sitemap directive at 76%.
- May 5, 2026 - Google adds an llms.txt audit to Chrome Lighthouse.
- May 15, 2026 - Google's AI search guide says no llms.txt file is required.
- June 15, 2026 - Google notes that llms.txt files neither help nor hurt Search visibility.
- July 2, 2026 - Originality.ai counts 36,120 llms.txt files while Ahrefs finds 97% receive no requests.
- July 10, 2026 - Google clarifies how long canonical changes take to be re-evaluated.
- July 31, 2026 - Episode 113 of Search Off the Record covers blocking internal search pages, with Mueller listing sitemap files, HTML sitemaps and LLMs text among the elements of a well-structured site.
- October 1, 2026 - Google Search Central publishes episode 114, "Do sitemaps still matter?", in which Mueller links "couldn't fetch" reports on valid sitemaps to host load and crawl demand.
Related PPC Land coverage
- Google drops 2007 rule requiring blocked internal search pages - Episode 113, in which Mueller and Splitt discuss infinite crawl spaces, robots.txt and server load.
- Google's secret crawl logic, finally explained in one page - Google's March 2026 overview of discovery, recrawl frequency, sitemaps and automatic crawl rate adjustment.
- Googlebot is not a program - Google engineers finally explain what it really is - Episode 105 on the shared crawling platform and its built-in throttling.
- Google says poor indexing on strong hosting indicates quality issues - Mueller's July 2025 explanation that weak indexing on sound hosting often reflects site quality.
- Bing emphasizes sitemaps critical role in AI-powered search era - Microsoft's file limits, lastmod formatting rules and case for sitemaps in AI search.
- Google Search: Last-Modified Date in Sitemaps still considered a signal for activity - Gary Illyes on how Google treats the lastmod element.
- Google forces publishers to wait two weeks for canonical fixes to register - Google's canonical troubleshooting guidance and re-evaluation timing.
- llms.txt adoption rises 8.8x but 97% of files get zero AI requests - Adoption counts and server-log data showing the files go largely unread.
- Google adds llms.txt to Lighthouse as agentic web standards heat up - Chrome's llms.txt audit and the debate over what the file does.
- Only 7.4% of Fortune 500 have an llms.txt file, study finds - ProGEO.ai data on robots.txt, sitemap directives, JSON-LD and llms.txt among large companies.
- Microsoft Clarity exposes AI bot traffic with new visibility dashboard - The January 2026 tool that shows which AI systems crawl a site.
- How web standards shape the internet's governing framework - Why robots.txt became an IETF standard while sitemaps remained a de facto one.
- Google clarifies Hreflang Implementation for multilingual websites - The 2024 episode on hreflang, including implementation through XML sitemaps.
Summary
Who: John Mueller and Martin Splitt of Google's Search Relations team, speaking in episode 114 of Google Search Central's Search Off the Record podcast. The discussion concerns site owners, publishers, retailers, SEO practitioners and agencies that rely on sitemaps for discovery in Google and other systems.
What: Mueller said a "couldn't fetch" report on a valid, accessible sitemap often stems from host load or low crawl demand, which he linked to perceived site quality. He also confirmed that Google set aside the priority and change frequency fields, uses lastmod only when dates look reliable, uses sitemap listings as one input into canonical selection, treats RSS feeds as usable short-form sitemaps, and does not process HTML sitemaps or LLMs.txt files as sitemaps.
When: The episode was published on October 1, 2026.
Where: On the Google Search Central YouTube channel and podcast platforms, with a transcript published by Google.
Why: Sitemaps remain the main signal site owners send about new and changed pages. The episode shows that a technically correct file does not guarantee Google will read it, that AI crawlers appear to find sitemaps and feeds through generic conventions rather than submission, and that Google still sees no role for LLMs.txt in its systems.
Discussion