惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
DataBreaches.Net
P
Proofpoint News Feed
The Cloudflare Blog
宝玉的分享
宝玉的分享
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
月光博客
月光博客
美团技术团队
Spread Privacy
Spread Privacy
Latest news
Latest news
Cisco Talos Blog
Cisco Talos Blog
T
Threatpost
Project Zero
Project Zero
博客园 - 司徒正美
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Simon Willison's Weblog
Simon Willison's Weblog
Apple Machine Learning Research
Apple Machine Learning Research
腾讯CDC
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
F
Fortinet All Blogs
Security Latest
Security Latest
Blog — PlanetScale
Blog — PlanetScale
T
Tailwind CSS Blog
Cyberwarzone
Cyberwarzone
The Hacker News
The Hacker News
Scott Helme
Scott Helme
T
Tor Project blog
Engineering at Meta
Engineering at Meta
H
Help Net Security
Recorded Future
Recorded Future
Microsoft Azure Blog
Microsoft Azure Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
I
Intezer
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
P
Privacy & Cybersecurity Law Blog
T
The Blog of Author Tim Ferriss
I
InfoQ
C
Cybersecurity and Infrastructure Security Agency CISA
大猫的无限游戏
大猫的无限游戏
F
Full Disclosure
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Microsoft Security Blog
Microsoft Security Blog
博客园 - 三生石上(FineUI控件)
L
LINUX DO - 热门话题
V
Vulnerabilities – Threatpost
S
SegmentFault 最新的问题
人人都是产品经理
人人都是产品经理
G
GRAHAM CLULEY
A
Arctic Wolf
P
Privacy International News Feed

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
When Google Search Console Couldn’t Fetch My Sitemap, I Added an HTML Sitemap Fallback
Bob · 2026-06-03 · via DEV Community

Quick Summary: If GSC gives you a cryptic Could not fetch error on a technically valid XML sitemap, stop rewriting the XML too early. In this case, I controlled variables across Cloudflare Pages, Vercel, Cloudflare Workers/OpenNext, and domain changes, which pointed to a production-domain or GSC-state anomaly. The practical SEO fallback was to add a crawlable HTML sitemap.

I recently spent a lot of time debugging a sitemap problem that looked simple at first.

Google Search Console kept saying:

Could not fetch

The strange part was that the sitemap worked everywhere else.

It returned 200 OK. It had the correct application/xml content type. The XML was valid. robots.txt pointed to it. Browser access worked. Command-line checks worked.

But Google Search Console still refused to read it.

This post is a write-up of the debugging process: what I checked, what I ruled out, why the hosting platform was probably not the final root cause, and why I eventually added an HTML sitemap as a crawl-discovery fallback.

The Site Setup

The project is a Next.js tool site for browser-side video and image processing.

It has:

  • mostly static SEO/tool pages
  • localized routes
  • no login
  • no database
  • no payment system
  • client-side media processing
  • some heavier runtime assets served through the Cloudflare ecosystem

The production deployment was on Cloudflare Pages.

The XML sitemap was generated by the app and included localized URLs across supported languages.

At a high level, this should have been a boring sitemap setup.

It was not.

The Initial Symptoms

In Google Search Console, both the dynamic XML sitemap and a static XML sitemap showed the same failure:

Could not fetch

The discovered page count stayed at zero.

The frustrating part was that direct HTTP checks looked normal:

/sitemap.xml        -> 200 application/xml
/sitemap-static.xml -> 200 application/xml
/robots.txt         -> 200 text/plain

Both XML files passed validation.

The static sitemap was especially important. It was a plain file under public/, not a Next.js metadata route. If both the generated sitemap and a static XML file failed in GSC, then the issue was probably not limited to the Next.js sitemap route.

Checking robots.txt

The robots.txt file was simple:

User-Agent: *
Allow: /

Sitemap: /sitemap.xml

There was no disallow rule blocking the sitemap or the main pages.

So the obvious robots explanation did not fit.

Checking Cloudflare DNS and Custom Domains

Because the site was on Cloudflare Pages, I spent a lot of time checking Cloudflare configuration.

The production domain and www domain were active in Cloudflare Pages. SSL was enabled. The DNS records pointed to the Pages deployment and were proxied through Cloudflare.

I also checked old verification records, email records, custom domain status, and the basic DNS setup.

Nothing obvious looked broken.

The domain resolved. The site loaded. The sitemap returned 200. The custom domains were active.

So the basic Cloudflare Pages domain setup did not explain why GSC could not fetch the sitemap.

Checking Cloudflare Bot and Security Rules

The next suspicion was that Googlebot might be getting challenged or blocked by Cloudflare.

I checked Cloudflare AI Crawl Control and security events.

Googlebot was recognized as a search engine crawler. The relevant controls were not blocking it.

Cloudflare Security Events also showed requests to the sitemap path from verified crawlers, including Google-related user agents.

There was a custom rule for verified bots:

cf.client.bot

The rule skipped security products for verified bots.

The important part: I did not find evidence of:

  • WAF block
  • Managed Challenge
  • JS Challenge
  • Interactive Challenge
  • Googlebot being denied access

Googlebot-style requests could receive 200 responses.

That made a simple Cloudflare security-block explanation unlikely.

Checking HTTP Details

I also checked the sitemap response in several ways.

The sitemap returned:

200
application/xml

I checked normal requests, Googlebot-style user agents, HTTP/1.1 behavior, compressed responses, and XML validation.

The XML remained valid.

The dynamic sitemap had Next.js route headers, as expected. But the static sitemap also failed in GSC, which again suggested that the issue was not just the Next.js metadata route.

At this stage, the sitemap looked technically valid from the outside.

Google Search Console still disagreed.

Fixing Build and Deployment Noise

During the investigation, I also found unrelated build noise.

The Cloudflare build could fail because next/font/google tried to fetch fonts during the build. That was not directly the sitemap bug, but it made deployment verification less stable.

I removed the Google Fonts dependency and switched to a system font stack.

After that, both the normal Next build and the Cloudflare build completed successfully.

This mattered because I needed a stable deployment baseline before blaming Google, Cloudflare, or the domain.

The Vercel Diagnostic Test

At one point, I deployed the same project to Vercel as a diagnostic comparison.

The goal was not to move production to Vercel.

The goal was to answer a narrower question:

Is the sitemap XML/project output itself fundamentally broken?

The deployment itself worked.

But the key detail is this: when the production domain was used, Google Search Console still could not fetch the sitemap.

That meant simply changing the hosting platform was probably not enough.

Later, I tested the same project with a different temporary domain. Under that different domain, Google Search Console could fetch the sitemap successfully.

That changed the interpretation.

The issue was probably not just:

Cloudflare Pages vs Vercel

The stronger signal was:

The failure was likely tied to the production domain or Google/GSC state associated with that domain.

This distinction matters. Without it, it is easy to draw the wrong conclusion and think that a hosting migration alone would fix everything.

Here is the simplified control-variable table:

Deployment path Domain used GSC status What it suggested
Cloudflare Pages Production domain Failed: Could not fetch Not explained by basic Cloudflare DNS, SSL, WAF, or robots settings
Vercel diagnostic deployment Production domain Failed: Could not fetch Moving the same project to Vercel did not fix the production-domain problem
Same project on a different temporary domain Temporary domain Success The sitemap output was likely valid; the production domain or GSC state became the stronger suspect
Cloudflare Workers + OpenNext Production domain Failed: Could not fetch Swapping backend infrastructure did not fix the production-domain problem

Trying Cloudflare Workers and OpenNext

Because the old Cloudflare Pages build chain used @cloudflare/next-on-pages, and that adapter is deprecated, I also tested a Workers/OpenNext path.

This was not a casual check. I actually went through the deployment path:

  • added OpenNext/Workers configuration
  • configured wrangler
  • tested a Workers custom domain
  • confirmed the Worker was serving traffic
  • saw the x-opennext response header
  • tested the homepage
  • tested robots.txt
  • tested the sitemap
  • tested the runtime routes needed by the app

At first, the Worker test domain worked.

Then I tried switching the production domain from Pages to Workers.

That required removing the production custom domains from Pages and adding them to the Worker, because Cloudflare would not allow the same hostname to be managed by both at once.

After the switch, the production domain did serve through Workers/OpenNext.

The key routes still returned valid responses:

/              -> 200
/sitemap.xml   -> 200 application/xml
/robots.txt    -> 200 text/plain

The response headers confirmed traffic was going through OpenNext Workers.

Then I submitted the sitemap again in Google Search Console.

It still failed.

I also tried a cache-busting sitemap URL with a query string.

That URL returned valid XML outside GSC.

GSC still said it could not fetch it.

That was an important result.

It meant the original hypothesis was not supported:

This was not simply a Cloudflare Pages or next-on-pages problem.

The same production domain still had the issue even after moving the delivery path to Workers/OpenNext.

The Strongest Conclusion

After all of these tests, the most likely root area was not the XML file itself.

It was also not clearly one hosting provider.

The strongest clue was the domain test:

  • same project
  • same kind of sitemap
  • different domain
  • GSC could fetch it

That points toward a domain-level or Google-side state issue.

Possible explanations include:

  • historical crawl state for the production domain
  • Google Search Console state for the domain property
  • DNS or routing history associated with the domain
  • Google-side host classification or cache

I cannot prove exactly which one it is.

But the evidence pointed away from endlessly rewriting the sitemap XML.

The Practical Problem

Even if the XML sitemap is valid, it is not very helpful if Google Search Console refuses to process it.

The site still needs its important pages discovered.

For a multilingual tools site, that matters.

So I stopped treating the XML sitemap as the only discovery mechanism.

I added an HTML sitemap.

Why an HTML Sitemap Helps

An HTML sitemap is just a normal page with internal links.

Googlebot can crawl it like any other page.

That gives the site another discovery path:

normal page -> footer link -> HTML sitemap -> localized tool pages

This does not fix the XML sitemap failure directly.

It reduces the risk of relying on only one discovery mechanism.

That was the practical goal.

How I Designed the HTML Sitemap

I kept the page intentionally simple.

The HTML sitemap:

  • returns plain HTML
  • is linked from the footer
  • uses index, follow
  • has a canonical URL
  • groups links by language
  • lists only the core tool pages
  • excludes privacy policy and terms pages
  • avoids duplicate homepage entries

The page is not trying to be a fancy user interface.

It is a reliable crawl hub.

What the HTML Sitemap Contains

The site has multiple languages and a fixed set of core tools.

The HTML sitemap lists the localized version of each core tool page.

In the current setup:

9 languages x 11 core tools = 99 tool links

For the homepage converter, the link is represented by each language's localized tool title, instead of repeating a generic brand link.

That keeps the sitemap focused and avoids unnecessary duplicates.

Why I Did Not Just Move Hosting

Moving hosting would have been the wrong lesson.

The production domain still failed even when the project was served through a different deployment path.

A different temporary domain worked.

That means the domain/GSC state was the more important signal.

Also, the product uses browser-side media processing and heavier runtime assets that fit well with the existing Cloudflare setup.

So the better solution was not to migrate everything just because one diagnostic deployment behaved differently.

The better solution was to add a second crawl-discovery path.

What I Learned

Sitemap debugging is not always about the sitemap file.

Sometimes:

  • the XML is valid
  • the headers are correct
  • the route is public
  • bots are not blocked
  • multiple hosting paths work technically

and Google Search Console still reports a fetch failure.

At that point, adding another discovery mechanism can be more useful than continuing to tweak a valid XML file.

For this case, the final strategy was:

  • keep the XML sitemap
  • keep monitoring GSC
  • add an HTML sitemap
  • link it from the footer
  • make the important localized pages discoverable through ordinary internal links

Final Setup

The live site discussed in this post is:

The HTML sitemap is not a replacement for the XML sitemap.

It is a crawl-discovery fallback.

And in this case, that was the most practical solution.