What Crawlers and AI See of htmx Content: A Measurement
I put check marks into an htmx demo and measured what curl, a headless browser and an AI assistant get to see. What stays visible and what does not.
by Jean Pierre Kolb ·
In Part 1 of this series I built an event calendar with htmx, with Unpoly and with no library. In the browser all builds look the same. The question of this part is whether that also holds for the programs that do not look at your page through a browser: search engine crawlers, AI crawlers and AI assistants that fetch a URL on request.
For classic SEO the question is old. For GEO, visibility in AI answers, it is new and sharper. Content an AI system does not see when it fetches a page is content it cannot cite. So I did not want to read up on what crawlers supposedly can do; I wanted to measure what actually reaches them.
A note on method: All measurements are from 29 September 2026 and ran against the demo as served live. Statements about Google, Bing and the AI vendors rest on their documentation and on published studies. Where a statement only comes from employees, or is only observed, I say so.
How crawlers render today
Google runs JavaScript. According to its JavaScript SEO documentation (opens in a new tab), Google Search uses an evergreen version of Chromium. The process has two steps: Googlebot first fetches the HTML, the page then goes into a queue, and only later does a headless Chromium render it and run the JavaScript. That can take seconds, or longer. One restriction matters for this demo: for pages with noindex, Google may skip rendering altogether. The demo is noindex, so it cannot show what ends up in Google's index. It shows what a crawler can see in principle.
Google does not interact with the page while rendering. The documentation on lazy loading (opens in a new tab) says that Googlebot neither scrolls nor clicks. Content that only a click on a button loads is not seen by Google. How Google still picks up content further down is not stated there. The well-known answer comes from Google employees: in 2017 (opens in a new tab) John Mueller spoke of a very tall viewport Googlebot renders with, and in 2020 (opens in a new tab) Martin Splitt put it briefly: “It doesn't scroll.”
Google only recognises links as <a> with an href attribute. The documentation on crawlable links (opens in a new tab) says explicitly that Google cannot reliably extract URLs from other elements that act as links through script events. It does not name hx-get, but hx-get falls under the same rule. That is my conclusion, not a verbatim statement from Google.
Bing announced in October 2019 that Bingbot would gradually switch to an evergreen, Chromium-based Edge, according to the announcement of the “evergreen Bingbot” (opens in a new tab). I could not find a current, explicit recommendation from Bing on JavaScript-heavy pages.
AI crawlers are the real reason for this article. The analysis by Vercel and MERJ from December 2024, “The rise of the AI crawler” (opens in a new tab), comes to a clear result: the crawlers of OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User) and Anthropic (ClaudeBot) do fetch JavaScript files, but they do not execute them. According to the study, only Google's infrastructure, which Gemini also uses, and Applebot render. A case study by Glenn Gabe (opens in a new tab) from August 2025 confirms the picture: ChatGPT, Perplexity and Claude could not read client-side rendered content; Google's AI Overviews and Bing Copilot could. Two caveats belong here. The vendors themselves do not explicitly document that they don't render, and agentic browsers that surf on a user's behalf do run JavaScript. Those are browsers, though, not crawlers.
The setup
Every content block of the demo carries a visible check mark, a unique word such as CANARY-LISTE-HTMX-EN-R-…. The first part says which block it is; the last part is a hash that no program can guess without having seen the page. The marks of robust and fragile blocks differ. In this article I shorten them, so they do not themselves appear on an indexed page and skew the measurement.
Then I fetched the same URLs three ways:
- curl: a plain HTTP request without JavaScript, the way a crawler without rendering sees it (
curl -sL, version 8.5.0). - Headless browser: a Chromium without a window (Playwright,
HeadlessChrome/155), viewport 1280 × 720. Per URL: load, scroll to the bottom, wait three seconds, then save the finished DOM. That is a generous view: Googlebot does not scroll, but partly makes up for it with a very tall viewport (see below). - AI assistant: Claude via Claude Code's WebFetch tool, asked to return every string that starts with
CANARY-. According to the documentation (opens in a new tab), WebFetch fetches the page, converts the HTML to Markdown and lets a small model answer the question on it. That no JavaScript runs in the process is not stated there explicitly. I observed it: a pure JavaScript application came back from WebFetch as an empty shell.
A small script reads the saved responses and checks for each check mark whether it occurs.
The result
| Page · block | curl | Headless | Claude |
|---|---|---|---|
| htmx · list (in the HTML) | ✓ | ✓ | ✓ |
| htmx · upcoming events (in the HTML) | ✓ | ✓ | ✓ |
| Unpoly · list and upcoming events | ✓ | ✓ | ✓ |
| no library · list, upcoming events, quick-info dialog | ✓ | ✓ | ✓ |
htmx · tab target as its own URL (/htmx/readings/) | ✓ | ✓ | ✓ |
htmx · page 2 as its own URL (/htmx/page/2/) | ✓ | ✓ | ✓ |
| fragile · list (in the HTML) | ✓ | ✓ | ✓ |
fragile · upcoming events (hx-trigger="load") | ✗ | ✓ | ✗ |
fragile · load more (hx-trigger="revealed") | ✗ | ✓ | ✗ |
fragile · tab (button with hx-get) | ✗ | ✗ | ✗ |
fragile · quick info (button with hx-get) | ✗ | ✗ | ✗ |
The curl and headless columns come from the English demo. The Claude column I only measured on the German URLs of the same demo; curl and headless produce exactly the same pattern there.
I read three things from the table.
What is in the HTML, everyone sees. All robust blocks arrive in all three views, with htmx just as with Unpoly and with no library. The library makes no difference to visibility; what matters is whether the content is in the server response.
What is loaded later is only seen by whoever runs JavaScript. The two fragile blocks behind load and revealed are missing for curl and for Claude. Going by the studies on AI crawlers, the same applies to GPTBot, ClaudeBot and PerplexityBot. The headless browser sees them, and Googlebot probably does too, because it renders.
What sits behind a click, nobody sees. The tab and quick info of the fragile build are missing in all three views, even in the headless browser. That matches Google's statement that Googlebot does not click. And hx-get is not a link a crawler could follow.
revealed depends on the height of the window
With infinite scroll via hx-trigger="revealed" I wanted to know more. htmx loads the next batch as soon as the placeholder element at the end of the list comes into view. So I loaded the same fragile page in the headless browser without any scrolling, once with a window height of 720 pixels and once with 4000:
| Window height, no scrolling | Events | “Load more” block |
|---|---|---|
| 1280 × 720 | 6 | missing |
| 1280 × 4000 | 12 | present |
Whether a renderer sees the content loaded later is therefore decided not by htmx but by the renderer's window height. With Mueller's “very tall viewport”, Googlebot would probably see this short list in full. For a long list with many loading stages I would not be sure, and a noindex demo cannot prove it. For content that should be found, that is too shaky a foundation.
Why the robust build is robust
The robust htmx build achieves the same interaction without any content depending on JavaScript. The pattern behind it is simple: every state has its own real URL, and that URL delivers a complete page.
- The “Reading” tab is a link to
/htmx/readings/. That page exists, and a crawler follows the link like any other. htmx only takes the#katalogexcerpt from the same page (hx-select) and writes the address into the history withhx-push-url. - “Load more” is a link to
/htmx/page/2/, an ordinary page 2. htmx appends its entries instead of changing pages. - “Upcoming events” is simply in the HTML. There is no reason to load a box with three events later.
The fragile build breaks this rule in exactly the places marked ✗ in the table: its fragments have no page of their own.
One detail surprised me while measuring. If someone calls a fragment directly, for instance because a crawler fishes the URL out of the source, they get a bare piece of HTML with no <html>, no <head>, no navigation. My server delivered it with 200 and a correct charset=utf-8, but without any hint that it does not belong in an index. A fragment cannot carry a <meta name="robots">, because it has no <head>. So I added the header X-Robots-Tag: noindex, nofollow for the whole demo area in the .htaccess. For a fragment without <head>, this header is the way to send an index exclusion along. An entry in robots.txt would only forbid crawling, not the indexing of a URL found elsewhere.
When the server delivers fragments
My demo is static and therefore needs separate fragment files. In a real application the server would decide, based on the HX-Request: true header, whether to send the whole page or just a fragment. Two rules from the htmx documentation (opens in a new tab) apply then:
- The server must send
Vary: HX-Request. Otherwise a cache may serve the fragment to someone who wanted the whole page, or the other way round. - If htmx does not find a history entry in its cache, it fetches the page again and also sends
HX-Request: true, yet expects the whole page. Since htmx 2.0.5 the optionhistoryRestoreAsHxRequestcontrols this; it is on by default. If you serve fragments based onHX-Request, check theHX-History-Restore-Requestheader or turn the option off.
And a URL that ends up in the address bar via hx-push-url must deliver a complete page on reload. The reload is a perfectly normal page request without htmx.
The link to this blog
Every article here has a Markdown twin at the same address with .md at the end. On top of that there is an llms.txt that points to exactly these twins. That is the experiment of this site: serving content so that AI systems can read it without HTML noise. The measurement shows a second advantage. Markdown needs no JavaScript, so the twins contain nothing a non-rendering crawler could miss.
How well that fits shows in a detail from the research. WebFetch sends an Accept header that prefers Markdown with every request. My server answers such requests with the .md twin through content negotiation. The demo has no twins, so HTML comes back there; I checked that with a request using Accept: text/markdown. For the articles the server answers the same request with the twin. Whether WebFetch actually receives it I cannot see from the outside, though, because the tool converts HTML to Markdown anyway.
What I take away
- Content that should be found or cited belongs in the first server response. Loading later via
loadorrevealedis invisible to AI crawlers as things stand. That is true regardless of the library. - Every state needs a URL that delivers a whole page. Then htmx is just acceleration, and crawlers follow ordinary links.
- Buttons with
hx-getare dead ends for crawlers. That is fine for controls, not for content. - Fragment URLs belong out of the index via
X-Robots-Tag, because they have no<head>of their own.
So the library does not decide visibility. What is visible is what you deliver without JavaScript, and htmx makes it easy to have both. Part 3 is about another group that does not read your page with their eyes: people using a screen reader or a keyboard. There the balance for htmx is less friendly.