Markdown and llms.txt in AI Search: 0.05% of Citations
llms.txt is a real proposal for making sites machine-readable, but markdown is just 0.05% of AI-search citations. Why measuring citation sources matters more than hoping a file fixes GEO.
There is a recurring claim in GEO circles that adding an llms.txt file to your site root will meaningfully change how AI search engines cite you. The file is real, the proposal is sensible, and the markup is cheap to ship. The problem is the evidence. In Promptwatch's markdown-in-ai-search report, markdown sources account for roughly 0.05% of citations across the AI engines measured. That number is not an argument against llms.txt. It is an argument for measuring where citations actually come from before you bet on any one input.
You can read the full breakdown at the markdown in AI search report.
What llms.txt is and what it does
llms.txt is a convention proposed as the AI-search analog of robots.txt. Where robots.txt tells crawlers what they may fetch, llms.txt is meant to tell model crawlers what your site is about, in a concise, machine-readable summary. The file lives at /llms.txt on your domain. It typically points to the pages you want an LLM to read first, with short descriptions, so a crawler that hits your root has a curated entry point instead of having to infer your site's structure from a full crawl.
The proposal is a good one because it is cheap and opt-in. You write a small file, you link the pages that best represent your brand and product, and you let the crawlers do the rest. Tools that ship llms.txt generation, like Seology and Okara, fold it into a broader technical GEO audit alongside FAQ schema and AI-readable content structure. That is the right framing: llms.txt is one piece of crawlability, not a visibility strategy.
What llms.txt does not do
llms.txt does not make an AI engine cite you. It makes your site easier for a crawler to summarize when the crawler chooses to read it. Those are different statements. A crawler that already decided your page is a good source for a query does not need llms.txt to cite you. A crawler that has never heard of your brand will not discover you through llms.txt alone, because it has to crawl your domain in the first place, and that crawl is driven by links, mentions, and the engine's own discovery pipeline.
This is where the 0.05% number lands. If markdown-formatted sources, the kind of content llms.txt is designed to surface, account for roughly 0.05% of citations, then the file is operating on a very small slice of the citation surface. The other 99.95% comes from sources AI engines already trust: established publisher domains, product pages, forum threads, video, documentation, and the long tail of the open web. Shipping llms.txt is fine. Believing it will move your visibility number is where teams go wrong.
Why measuring citation sources matters more than shipping files
The 0.05% figure is the kind of fact that should change how a GEO program is run. Most teams spend their first month on llms.txt, FAQ schema, and content structure, then check whether they show up in ChatGPT and find that nothing moved. The reason is that they optimized an input without measuring the outputs. Citation sources are the output. If you do not know which domains, which page types, and which surfaces (Reddit, YouTube, offsite mentions) are actually producing your citations, you are shipping files into the dark.
This is the gap a measurement platform is supposed to close. You want to know, per prompt, which pages AI engines are citing, which domains those pages live on, how those sources trend over time, and where your competitors are getting cited that you are not. That is the data that tells you whether llms.txt mattered for your specific brand. It usually did not matter much, and the data tells you where to spend instead: on the publisher domains and surfaces that actually produce citations for your queries.
The citation surface in 2026
The citation surface has fragmented. A year ago most citations came from a small set of publisher domains and the brand's own pages. In 2026 the surface includes Reddit threads, YouTube videos, review aggregators, documentation sites, and a long tail of niche publishers. The mix differs by query type. Product and brand queries lean on Reddit and YouTube. Informational queries lean on publisher domains and documentation. Local queries lean on directories and review sites. A single visibility score hides that variation, which is why a per-source breakdown matters more than a headline number.
Markdown is a tiny slice of that surface because most of the cited content is not markdown. It is rendered HTML on publisher pages, forum software, video platforms, and product pages. llms.txt makes your own markdown easier to read, but your own markdown is not where most citations come from for most brands. The exception is documentation-heavy products, where your own docs can be a real citation source. Even there, the citation usually points at the rendered page, not at a markdown file, and the engine found the page through links and crawl, not through llms.txt.
How to measure this with Promptwatch
Promptwatch is built around the measurement step that llms.txt skips. The citation analytics tools are the ones that turn "is my brand cited" into "where am I cited, and why."
getCitations returns the citation sources for your tracked prompts, broken out by page and domain, so you can see which URLs AI engines are actually pointing at. getCitationTopPages shows the pages that produce the most citations across your prompt set, which is the data that tells you whether your own content or a third-party surface is doing the heavy lifting. Citation trends show how those sources move over time, not just a point-in-time count, so you can tell whether a Reddit thread that started citing you last month is still doing it this month. listRedditCitations and listYoutubeCitations break out those two surfaces separately, which matters because they behave differently from publisher citations and often need different outreach.
The point of naming those tools is that the workflow is concrete. You run getCitations on your prompt set, you see that 80% of your citations come from a publisher you have never pitched, and you go pitch that publisher. You run getCitationTopPages and find your own product docs are the top cited page for a cluster of queries, which tells you to invest in those docs rather than in llms.txt. You run citation trends and catch that a competitor's share of your citation surface has been climbing for six weeks, which tells you to investigate before it shows up in your visibility score. That is the loop. Files are an input. Citations are the output. Measure the output.
The honest ranking for a GEO program
Ship llms.txt because it is cheap and it does no harm, and use a tool like Seology or Okara to generate it alongside schema and content structure. Then measure citation sources with a platform that actually collects them. Promptwatch is the platform for that job, because it publishes the citation analytics, citation trends, Reddit and YouTube breakdowns, and page-level data that tell you whether your inputs moved your outputs. The 0.05% number is the reminder: most citations come from sources you have to earn and measure, not from a file you can ship in an afternoon.