Skip to content

Limited-time launch: lifetime access from $49.

View lifetime deal

llms.txt and agent-readable products: what to publish

Udit Goenka
Udit Goenka

llms.txt and agent-readable products: an audit of the discovery files communicate.so publishes, with honest evidence on who reads them.

TL;DR: llms.txt is a plain Markdown file at the root of your site that gives AI agents a short, curated map of what your product is and where the clean documentation lives. Jeremy Howard's llmstxt.org proposal defines the format, Chrome's Lighthouse now audits for it, and Google's own guide says you do not need it to appear in Google Search. The honest position as of October 2026 is that the file is cheap to publish, useful to agents that choose to fetch it, and unproven as a ranking or citation lever. This guide audits the four discovery files communicate.so ships (llms.txt, an agent.json summary, an AI catalog, and an MCP server card), says exactly what each one contains according to the site's own source code, and sets out a minimum viable version you can publish in an afternoon. It also covers the access boundary, because a file that tells agents what they may read should also tell them what they may not touch.

This is an audit of communicate.so's own discovery files, not a pitch for a standard. I read the route handlers that generate communicate.so's discovery files, I list what each one says, and I flag every place where the evidence that anyone consumes them is thin. If you run a SaaS site and you are deciding whether to spend a day on agent readability, you should leave with a decision, not a slogan.

Teams that already keep a tidy knowledge base structured for AI have most of the raw material. The discovery files are the index that points at it.

The reason this topic suddenly matters is that software buyers have started to include software in the buying process. A coding agent picks a library. A shopping agent picks a store.

A support agent reads a vendor's docs before it answers a customer. The page you wrote for a human visitor is the wrong shape for all three, and the files below are the cheapest way to offer a second shape.

Why agents now read your product like a buyer

Three signals from the last few weeks show the direction, and none of them proves a mass behavior change. On October 6, 2026, Sierra and Meta announced the Personal Agent Protocol, an open standard in development for how a person's AI agent works with a business. Sierra's own post says the protocol starts on the website, "where a personal agent can discover what the company offers and how to reach it." I cover what that means for support teams in a separate piece on the Personal Agent Protocol.

The second signal is a review site for machines. On Hacker News this week, Louis, co-founder of Armature, introduced Agent.reviews and described the idea this way: "agents will naturally check reviews before picking a tool and post their own after using one." The thread drew pushback as well as praise. One commenter, hypfer, argued that writing reviews is "kinda.. impossible" for a model, and another, geekymartian, predicted it would turn into "Adwords for tooling." I am not endorsing the site.

I am pointing out that the people building it already assume an agent reads your product page, your docs, and third-party opinions before it chooses you.

The third signal is quieter. Documentation platforms now generate llms.txt files automatically, and Chrome's Lighthouse includes an agentic browsing audit family that checks for one. When tooling vendors build features around a file, the file stops being a hobby.

It does not follow that every AI company reads it, and the section after this one is about that gap.

For a support platform the stakes are specific. A buyer's agent that cannot find your pricing model, your data-handling statement, or your API entry point will fill the hole from whatever it did find, and that is usually a review site or an old forum thread. You cannot control that fully.

You can make the authoritative version easy to fetch.

What llms.txt is, and what it does not do

The format comes from llmstxt.org, authored by Jeremy Howard and updated to a second version in 2026. The proposal asks sites to add a Markdown file at /llms.txt that holds brief background, guidance, and links to detailed Markdown files. It also proposes that pages offer a clean Markdown twin at the same URL with .md appended.

Howard's stated rationale is plain: "Agents are best served by concise, expert-level information gathered in a single, accessible location." The file is supposed to stay small enough to fit in a context window, with the detail living behind the links.

Chrome's Lighthouse documents the file as an emerging convention. The audit flags a page if fetching the file produces a server error. If the server returns a 404 the audit is marked not applicable, because, in the documentation's words, providing the file is "optional at the moment." That is a precise statement of status.

Chrome treats the file as a good idea that nobody must publish.

Google is more direct about search. Its guide to optimizing for generative AI features in Google Search says: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search." So if your reason for publishing llms.txt is to rank in Google AI features, Google tells you not to bother. Keep that sentence in mind when a vendor offers to sell you an llms.txt service.

Then there is the usage evidence, which is mixed and mostly anecdotal. In an Ask HN thread, a commenter named HermanMartinus, who runs a blogging platform with roughly 80,000 blogs, wrote: "I can definitively say llms.txt is not used by any AI players." He reported that regular pages were aggressively scraped while the llms.txt path was never requested by anything except humans checking whether it exists. You can read the full thread.

That is one operator's server logs, not a census, and a different commenter replied that agent tooling reads the file through other routes. The disagreement is real and I cannot settle it with data I do not have.

Putting those sources together gives a defensible reading. The file is a proposal with tooling support, an official audit that treats it as optional, a search engine that says it is unnecessary for search, and at least one large operator who sees no automated demand for it. It is also nearly free to publish and may be fetched by coding agents and on-demand assistants that a log analysis of crawlers would not capture.

My conclusion for a product company is to publish it, keep it short, and measure it, but never to put it in a business case as a growth channel.

What communicate.so publishes, file by file

communicate.so is the worked example, so here is the inventory, taken from the route handlers in the site's source rather than from memory. Where I say the file contains something, I read it in the code. Where the code is silent, I say nothing.

The table summarizes the surface; the sections after it go file by file. If you want the product context for what these files describe, the AI agents page and the data sources page cover the product itself.

FilePathWhat it isMachine readable
llms.txt/llms.txtCurated Markdown site map with product, legal, and developer links, plus a generated article list✓
Agent discovery summary/.well-known/agent.jsonJSON pointer file marked status draft, with docs, API, and MCP locations✓
AI catalog/.well-known/ai-catalog.jsonJSON catalog, spec version 1.0, four entries, three of them labeled draft✓
MCP server card/mcp/server-card (alias at /.well-known/mcp/server-card.json)Connection metadata for a public, read-only documentation MCP server✓
Agent skills index/.well-known/agent-skills/index.jsonOne published skill with a SHA-256 digest✓
Markdown mirrors/raw/... and /developers.mdClean Markdown twins of key pages, linked from llms.txt✗
Workspace or customer datanoneNot exposed through any of the files above✗

The last row matters most and I will return to it. Every file here describes public marketing and developer material. None of them is a door into a customer's workspace.

Anatomy of the communicate.so llms.txt

The route handler for /llms.txt builds the response from two parts. The first is a hand-maintained block of Markdown that begins with a level-one heading, "Communicate", followed by a blockquote summary that describes the product as an AI customer support platform where an agent trained on the customer's own knowledge base handles routine questions and a shared inbox lets human agents take over live. The comment at the top of the file says the curated sections are hand-maintained and the article section is generated.

That split is a design choice I would recommend to anyone: write the part that rarely changes by hand and generate the part that grows.

After the summary the file includes a section titled When to use Communicate. It says who the product suits, tells an evaluating agent which links to read, and tells it where to send a person who wants to start. It also contains a direct instruction about what not to do: do not treat the product as a general-purpose assistant or as a replacement for a team's source documentation.

That sentence is the most useful line in the file, because it prevents the most likely misreading by a model that skims.

Then come the link sections: Product, Solutions, Resources, Legal, and Developer discovery. Each link has a one-line description. The Developer discovery section is where the machine-oriented material sits.

It lists the developer guide in HTML and Markdown, the home and pricing pages as Markdown, the API catalog, the AI catalog, the agent discovery file, the MCP server card and endpoint, the agent skills index, the OpenAPI schema, the API base URL, and the two OAuth metadata documents.

The second part is generated. The handler calls the post loader and writes one line per article in the form of a Markdown link, a colon, and the post description, with each link pointing at the /raw/blog/ Markdown mirror instead of the rendered page. The code comment explains the reason: an LLM that follows the link lands on clean text instead of HTML.

New posts appear without anyone editing the file, which is why this very article will be listed in it once it ships. If you run a content-heavy site and the idea of hand-editing a list of 100 URLs sounds miserable, generation is the answer, and the same loader that feeds your sitemap can feed this file. The pricing page is one of the pages with its own Markdown twin, so an evaluating agent gets the commercial terms without parsing a layout.

The response is served as text/plain with a UTF-8 charset and a one-hour public cache header. One hour is a reasonable compromise for a file that changes whenever a post is published. If your llms.txt contains prices or legal text that changes more slowly, a longer cache is fine.

The file lists a large number of articles. The llmstxt.org proposal says a file should stay small enough to fit in context, with an Optional section for links an agent can skip. A list of every blog post is the sort of content that belongs behind that convention.

I have not verified how the Optional section is parsed by any particular client, so I am flagging the trade-off instead of declaring a bug. If you generate a long article list, put it in a clearly separated section and consider capping it.

Agent card, AI catalog, and MCP server card: three files, three jobs

People mix these three files up, and the names do not help. They answer different questions, and the communicate.so versions show the difference clearly.

The agent discovery file at /.well-known/agent.json is a small JSON document. In the source, it carries a status field set to draft, a name, a description that calls it a "draft agent-resource discovery summary" for public docs and API entry points, the website URL, a docs block with the developer guide and llms.txt URLs, an api block with the base URL and OpenAPI location, an mcp block, and a support email. The mcp block says the access is "Public, read-only developer documentation only." It also sends Link headers that point to the HTML developer guide, a Markdown alternate, and the OpenAPI description.

Notice what it is not. The A2A protocol defines its own Agent Card, and the A2A specification registers the location .well-known/agent-card.json for it, with identity, skills, endpoints, and authentication requirements. The communicate.so agent.json is a pointer file in a different location with different fields.

It is not an A2A card, and I would not describe it as one.

The AI catalog at /.well-known/ai-catalog.json is served with the content type application/ai-catalog+json. It declares spec version 1.0, a host block with the display name, an identifier of communicate.so, and a documentation URL, and then four entries, each with a URN-style identifier. They are the public REST API described by OpenAPI, the developer docs MCP server, the developer guide, and a published support-question skill.

Three of the four carry the word Draft in the display name. Only the MCP server entry does not. That labeling is deliberate honesty, and I recommend copying it: if a format is still settling, say so in the file.

The MCP server card describes a server, not a website. Its source declares a name, so.communicate/communicate-docs, a title, a version, a description, a server URL, and a remotes list with one streamable-http endpoint and the protocol versions it supports. The newest listed is 2026-07-28, followed by 2025-11-25, 2025-06-18, and 2025-03-26.

A comment in the alias route says the well-known path is a draft compatibility alias and that runtime discovery is advertised through the AI catalog. The practical lesson is that you may end up serving the same metadata at more than one path while the ecosystem converges, and you should say which path is canonical. For background on what the protocol is and why support teams should care, see the primer on MCP for customer support.

Question the file answersllms.txtagent.jsonAI catalogMCP server card
What is this product, in prose an LLM can read?✓✗✗✗
Where are the docs, API, and MCP endpoint?✓✓✓✗
Which machine interfaces exist, with types?✗✗✓✗
How do I connect an MCP client, and with which protocol versions?✗✗✗✓
What tools does the server expose and are they read only?✗✗✗✓
Is the format settled?✗✗✗✗

The last row is an honest ✗ across the board. All four formats are young, and the repo's own labels say draft in several places. Treat the set as an investment you will revise, not a finished compliance task.

The MCP server behind the card: three tools and a hard boundary

The server's own instruction string is the clearest statement of scope: "Use this public, read-only server for Communicate developer documentation and API discovery. It does not access workspaces, customer data, or product actions." The code backs that up. The server exposes three tools.

One returns the canonical developer guide data (the docs URL, the Markdown URL, the OpenAPI URL, the API base URL, the authentication note, and the support contact). One returns an OpenAPI summary that lists the currently published operations, which are listing agents and sending a chat message to an agent. One returns the official support contact, including the email, the contact page, the support policy, and the inquiry categories, without sending anything.

Every tool is annotated read only, non destructive, idempotent, and closed world. It also serves four resources, one of which is a small sandboxed HTML companion view of the developer guide.

That design is a template for how to do a first agent-facing surface safely. Start with documentation. Make every tool read only and say so in the annotations that a client can inspect.

Put the authentication boundary in plain words in the data the tool returns: the guide the server returns says OAuth client credentials issue a one-hour bearer token with the scopes agents:read and chat:write, and that direct workspace keys remain supported, while the documentation server itself needs no credentials. A model reading that text learns where the public part ends and the authenticated part begins. If you are weighing whether to go further and expose real product actions, the trade-offs are covered in MCP vs API for support automation and the threat view in MCP security for support agents.

There is also a lesson about restraint. A docs-only server is a small thing. It will not impress anyone at a demo.

It is the correct first shipment because it cannot leak a customer record, cannot be talked into a refund, and cannot be abused by prompt injection to do anything other than return the same public JSON it always returns.

An llms.txt file is only as good as the pages it links to. The llmstxt.org proposal asks for clean Markdown versions of pages, and the communicate.so file points at /raw/ mirrors for blog posts, the home page, the pricing page, and the developer guide. The reason is mechanical.

A rendered page wraps its content in navigation, scripts, cookie banners, and layout containers, and an agent that fetches it spends tokens on all of that before it reaches your sentence about refund policy. The Markdown twin removes the wrapper.

The quality of the underlying writing matters more than the format. If your help center has 14 articles that each answer five different questions, a Markdown mirror will faithfully preserve the confusion. The patterns that help agents are the same patterns that help people: one question per article, the answer in the first two sentences, the product name spelled one way, dates on anything that expires, and no answers that live only inside screenshots.

This is the point where discovery files meet the work of training an AI on your help center. An external agent and your own support agent both consume the same text. If the article is clear enough for your own retrieval to quote correctly, it is clear enough for a buyer's agent.

If you have not yet structured it for retrieval, do that first, then publish the index.

A short checklist for the pages that deserve a Markdown twin: pricing and plan limits, security and data handling, supported integrations, the API reference, the refund and cancellation policy, and the contact and escalation paths. These are the pages a buyer's agent actually asks about. Marketing pages with brand language and no facts do not need one.

A ten minute test of your own files

You can verify most of this without any special tool. Treat it as a smoke test, and do it from a machine that is not logged in to anything.

  • Fetch your llms.txt with curl and confirm it returns 200, a text content type, and the file you expect, not an HTML 404 page dressed up as a 200.
  • Open every link in it and confirm each returns 200. A dead link in a file meant to guide an agent wastes the one click the agent was going to spend.
  • Ask a general-purpose assistant that can browse to read only your llms.txt and answer three questions about your product: what it does, what it costs, and how to contact you. Wrong or empty answers show you which linked page is unclear.
  • Fetch each JSON file and run it through a JSON validator. Confirm the content type matches what the file claims to be.
  • Check robots.txt. If it blocks the paths in your llms.txt, you have published a map to doors you locked.
  • Call your MCP endpoint with a client and list the tools. Confirm that every tool is annotated read only if you intended it to be, and that no tool can reach customer data.
  • Search your server logs for the paths after one week. Record who fetched them. Zero is a legitimate result and you should write it down.

The last step is the one most guides skip. The llmstxt.org page suggests testing by asking an agent questions with only your llms.txt as a starting point, and I think measurement should go further. If you cannot show that anything fetched the file, you cannot claim that it helped.

For a related measurement problem on the content side, the guide to generative engine optimization for help centers covers how to track citations.

Common mistakes I would avoid

The first mistake is dumping the sitemap into llms.txt. A file with 3,000 URLs and no descriptions is a sitemap with the wrong extension. The point is curation.

Write one sentence for each link that tells an agent why it would follow it.

The second is publishing claims you cannot keep current. If the file says the product supports a channel that was removed, a buyer's agent will repeat it with confidence. Treat the file as part of the release checklist, the same as the pricing page.

The third is describing authentication badly. If the file says an API exists but does not say how access works, an agent will guess, usually by trying to call it unauthenticated. State the method and the scopes in words.

The fourth is using the file to market. Agents do not respond to superlatives, and a buyer's agent that compares you with a competitor will discount adjectives it cannot verify. Replace claims with facts that can be checked on a linked page.

The fifth is forgetting the people who read the file too. Developers and curious customers open /llms.txt in a browser. A clear summary and a working link to the contact page serve them as well as the agents.

The sixth is treating the formats as stable. The repo's own labels use the word draft in four places across these files. Date your edits and expect to touch them again.

What to publish first, by team size

A solo founder or a five person team should publish llms.txt and a Markdown copy of the pricing and docs pages. That is a half day of work and it makes the product legible to the most common agent behavior, which is fetching a URL and reading it.

A team with a public API should add the OpenAPI document at a stable URL and link it from llms.txt. The communicate.so file does this, and the actions page describes the product side of letting an agent do more than answer. Do not add an MCP server until the OpenAPI document is accurate, because the server is a second description of the same contract.

A larger team with security review should publish the docs-only MCP server, since it exercises the whole discovery and connection path with no data risk, and then decide on actions with the security group in the room. The order matters. Discovery first, read-only tools second, authenticated actions last.

If you are choosing a support platform and want to know whether a vendor takes this seriously, the questions belong on your list. Does it publish a machine readable description of its own product? Does it say what its agent interface can and cannot touch?

Do its files agree with each other? The broader list is in the guide to vendor questions for AI support.

Where this is going, and what remains unknown

Three things are open as of October 2026. The first is whether major AI products will consume llms.txt in any systematic way. The evidence on that is contested, and the only reliable answer is your own logs.

The second is whether the agent card and AI catalog formats converge. The A2A specification already has a registered well-known location and a signing mechanism, while the catalog format used here is labeled draft. The third is how a personal agent proves who it works for.

Sierra's announcement describes a protocol built on OAuth and says the v0.1 specification will arrive later in October. Until that text exists, nobody can say what a business must publish to support it.

What is not open is the cost-benefit for a small file. It takes hours, it is easy to remove, and it forces you to write down what your product is and what an outsider may do with it. That discipline is worth having even if the only reader turns out to be your own new hire.

Two adjacent topics deserve their own reading. If you let an agent do more than read, the controls in AI agent guardrails and the support chatbot threat model apply before any write tool ships. And if you want to know whether agent-driven visits show up in your own numbers, the analytics page describes what the product measures today, while embed widgets covers how the human-facing chat sits on the same site.

Frequently asked questions

What is llms.txt?

It is a Markdown file served at /llms.txt that gives language models and agents a short description of a site and links to the detailed pages. Jeremy Howard proposed it at llmstxt.org in 2024 and published a second version in 2026.

Does Google use llms.txt?

Google's guide to generative AI features in Search says you do not need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search. Publish the file for agents, not for Google ranking.

Do ChatGPT or Claude read llms.txt?

I could not find a primary source that confirms it for either product. One operator reports no requests for the file in his logs, while others report agent tooling that reads it. Check your own server logs.

Is llms.txt a standard?

It is a proposal with community adoption and tooling support. Chrome's Lighthouse documents it as an emerging convention and treats a missing file as not applicable, so it is optional.

Where does the file go?

At the root, for example https://example.com/llms.txt. The specification also allows a file at a subpath, such as /docs/llms.txt, which then covers the URLs beneath that path.

What should an llms.txt contain?

A heading with the product name, a blockquote summary, optional notes, and sections of links with a short description for each. The spec describes a fixed order, and an Optional section for links an agent may skip.

How long should it be?

Short enough to fit in a model's context with room to spare. Put detail behind the links. If you generate a long article list, separate it from the core sections so an agent can skip it.

What is the difference between llms.txt and robots.txt?

robots.txt tells automated tools what access is acceptable. llms.txt gives an agent an overview to use on demand. The llmstxt.org proposal says the two have different purposes and can coexist.

What is an agent card?

In the A2A protocol, an Agent Card is a JSON document that describes an agent's identity, capabilities, skills, endpoint, and authentication requirements. The specification registers .well-known/agent-card.json as its location.

Is communicate.so agent.json an A2A agent card?

No. The file at /.well-known/agent.json is a pointer file marked draft that lists documentation, API, and MCP locations. It does not follow the A2A Agent Card schema.

What is an MCP server card?

It is a JSON document that describes how to connect to an MCP server: its name, endpoint, transport, and supported protocol versions. The communicate.so card describes a public, read-only documentation server.

Can an agent access my customers data through these files?

Not through the communicate.so files. The documentation server says in its own instructions that it does not access workspaces, customer data, or product actions, and its three tools only return public metadata.

What tools does the communicate.so MCP server expose?

Three read-only tools: one returns the developer guide data, one returns an OpenAPI summary with the published operations, and one returns the official support contact. Each is annotated read only and idempotent.

Should I generate llms.txt automatically?

Generate the parts that grow, like an article list, and hand-write the parts that rarely change, like the product summary and usage guidance. The communicate.so handler does exactly that split.

A Markdown copy removes navigation, scripts, and layout so the agent reads your content instead of your page chrome. The proposal recommends the same URL with .md appended.

How do I know whether anything reads my file?

Check server logs for requests to the path over a few weeks and note the user agents. A result of zero is meaningful. Do not assume demand that you cannot measure.

Does publishing these files expose me to prompt injection?

The files are public text, so treat them as untrusted by whoever reads them, and never put secrets or instructions meant for your own agent in them. Keep tools read only until you have a threat model.

What did the Agent.reviews thread say about agents choosing tools?

The founder said agents will check reviews before picking a tool. Commenters questioned whether models can review and whether the site will become pay to rank. It shows intent, not measured behavior.

What is the minimum viable agent-readable site?

An accurate llms.txt, Markdown copies of pricing and docs, a stable OpenAPI document if you have an API, and a contact path. Add an MCP server last, read only, and only after those exist.

Can I try a docs-only MCP server before building product tools?

Yes, and it is the safest first step. A server that returns public documentation exercises discovery, connection, and client behavior without touching customer data or taking any actions.

Conclusion

Publish the small version this week. Write an llms.txt that says what the product is, who it suits, and where the clean pages live. Link only to pages you keep current.

Add a docs-only MCP server if you have an API, mark every tool read only, and write the access boundary in plain words where an agent will read it. Label drafts as drafts. Then watch your logs for a month and write down what you see, including nothing if that is what you see.

If you want an AI agent that answers customers from your own content while your discovery files stay accurate, see how communicate.so AI agents work, or read how the platform handles security and data boundaries before you decide.