How this site is built for AI crawlers
This is a teardown of lumirank.ca: the HTML a crawler receives, the robots.txt rule that lets it in, and the Markdown files published beside the pages. The Canadian AI Visibility Index methodology is a different document. It describes a measurement that has not been run. This page describes a site that is already live. Every path below is one you can open.
What llms.txt is, and what Google has said about needing one, is answered on what is llms.txt. This page does not repeat that argument. It shows the files this domain serves.
The files
Five URLs account for most of what a crawler is being offered beyond the HTML pages themselves.
| URL | What it is |
|---|---|
| /robots.txt | One group: User-agent: * and Allow: /. No per-bot block list. |
| /llms.txt | Curated Markdown index, generated from the same data as the HTML pages. |
| /llms-full.txt | Full text of the case studies, concatenated into one file. |
| /work/osoba-renos.md | One case study as Markdown. Public. Canonical header points at the HTML page. |
| /sitemap.xml | URL list for classical crawlers. The empty index results URL is omitted until a run exists. |
What the first response contains
Pages are rendered on the server. The words, the headings and the FAQ answers are in the HTML of the first response. A retriever that does not run JavaScript still receives the article. The check is ordinary: request the page with curl and search the response for a sentence you can see on screen.
Case studies also ship a Markdown alternate at /work/<slug>.md. The HTML page links it with rel="alternate" and type text/markdown. The Markdown response sends a Link header with rel="canonical" back to the HTML URL, so the two representations do not compete as two pages. Both are built from the same case-study record. A hand-maintained second copy is the one that goes stale, so there isn't one.
Who is allowed to crawl
robots.txt is a single group: User-agent: * and Allow: /. The wildcard is the decision. It covers Googlebot, Bingbot, Google-Extended, GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Common Crawl's CCBot, without a per-bot list that has to be updated every time a crawler is renamed. The studio sells being read by those crawlers. A disallow rule would make the site a counter-example to its own work.
Two URLs are noindex on purpose. The index results page stays out of the index until a real run exists, and it is left out of the sitemap for the same reason. The feedback form is a form. Neither is hidden from people following a link.
Accordions that stay in the HTML
FAQ blocks use the details element. The answer is in the HTML whether or not the panel is open. Closing one is a paint decision. The words are not fetched when someone clicks.
For a short list of crawler user agents, middleware sets X-Visitor: bot, adds open to every closed details, and sets data-visitor="bot" on the html element so both faces of a flip card lay out. Everyone else gets X-Visitor: human and the same sentences. Vary: User-Agent is set because the open state differs. The words do not. Google-Extended is allowed by robots.txt and is not on that user-agent list; it still receives the answers, because they were never removed from the HTML.
What fetching does not settle
Server-rendered HTML and an open robots.txt have a mechanical effect: a crawler can fetch the page and read it. Being fetched is the precondition for being quoted. It is not the quotation. llms.txt is generated here and published at the root, and Google has written that Google Search ignores it. Treat the file as a table of contents for agents that choose to read one, and spend the real effort on pages that answer a question in the first paragraph.
The quotable pages on this site are the answers, the GEO pricing guide and the home-services page. Those URLs are listed in llms.txt for that reason. The city service pages are left to the sitemap.
Questions about the setup
Will this setup get the site cited?
No. A crawler fetching a page and an assistant naming a brand are separate events. This page documents the first. Google's own documentation says Google Search ignores llms.txt. The file is here because it is generated from the same source as the pages, so it cannot drift, and because it costs nothing to publish.
Is the bot view cloaking?
No. Recognised crawlers receive the same sentences as everyone else. Middleware adds the open attribute to details elements and sets data-visitor on the html element. The words stay the same. The Markdown case studies are ordinary public URLs, linked from the HTML page, not a copy served only when the user agent looks like a bot.
Should a site block GPTBot?
This one does not. ChatGPT Search cites pages OAI-SearchBot and GPTBot can fetch, and blocking them removes the site from that path. A domain that hosts private client files has a different decision. This domain does not host those files, and the file it publishes is a wildcard allow.
The same checks, pointed at a client site, are part of technical SEO. Crawler access is settled from the logs and from robots.txt before anyone writes a page.