The Agent-Ready Website Checklist: 12 Checks for AI Search and AI Agents (2026)
People are asking AI assistants for recommendations, and AI agents are starting to act on websites on their behalf. On our own site, AI assistants are now the biggest source of the enquiries we can trace. So the question is no longer only “can Google find my website?” but also “can an AI read it, trust it and use it?”.
This checklist covers what a business website needs for that, in the order it matters. Each item says how to check it yourself, and what we found when we checked our own site. For the background on the newer standards, see our guides to ai-catalog.json and WebMCP.
The checklist at a glance
| # | Check | Layer |
|---|---|---|
| 1 | AI crawlers are allowed in robots.txt | Reach |
| 2 | AI crawlers are not blocked by your host | Reach |
| 3 | The content is in the HTML, not only drawn by JavaScript | Read |
| 4 | Pages answer real questions, with clear headings and tables | Read |
| 5 | Structured data is accurate and complete | Read |
| 6 | You publish an llms.txt | Read |
| 7 | Every claim on the site is true and current | Trust |
| 8 | The page structure is accessible | Use |
| 9 | Forms are labelled, and ready for WebMCP | Use |
| 10 | Public tools are listed in an AI catalog | Use |
| 11 | You test it with PageSpeed Insights | Measure |
| 12 | You can tell which enquiries came from AI | Measure |
Reach: can AI get to your site at all?
1. AI crawlers are allowed in robots.txt
Your robots.txt file tells crawlers where they may go. Check that it does not block the AI search crawlers, such as OAI-SearchBot and ChatGPT-User (OpenAI), ClaudeBot (Anthropic), PerplexityBot and Bingbot, which feeds Microsoft Copilot.
One detail catches people out. When you add a named group for a crawler, it replaces the general rules for that crawler, so any paths you block for everyone, such as checkout or account pages, must be repeated in each named group.
How to check: open yourdomain.co.za/robots.txt and read it.
2. AI crawlers are not blocked by your host
This is the check almost everyone misses, because robots.txt can look perfect while the server still turns crawlers away.
When we read our own server logs for August and September 2026, our shared hosting provider was refusing about 39% of ChatGPT-User requests, 38% of OAI-SearchBot requests and 47% of ClaudeBot requests with a “429 Too Many Requests” error, because of a rate limit applied to the whole server. Googlebot and Bingbot were not affected. Nothing on the website itself showed a problem.
We fixed it by putting the site behind Cloudflare and serving pages from its cache, so most AI requests never reach the busy server. In a follow-up test, all 36 requests from ChatGPT, GPTBot and ClaudeBot succeeded.
How to check: ask your host or developer for the server access logs, and count the responses to AI crawlers by status code. A share of 429 or 403 responses means your content is not reliably reaching AI assistants. Our step-by-step guide shows how to check your server logs for AI crawler blocks, and what fixed it for us.
Read: can AI understand what it finds?
3. The content is in the HTML
Some AI crawlers do not run JavaScript. If your text only appears after scripts run, they may see an empty page.
How to check: view the page source (right click, View page source) and search for a sentence from the page. If it is there, crawlers can read it.
4. Pages answer real questions, clearly
AI assistants quote passages that answer a question directly. Use headings that match what people ask, answer early, and put comparisons and prices in tables. A clear table is one of the easiest things for an AI answer to lift.
How to check: read your key pages as a stranger. Can you find the price, what is included and how to get in touch within seconds?
5. Structured data is accurate and complete
Structured data (schema) describes your business, services, prices and questions in a form software reads directly. Accuracy matters more than volume: no invented reviews, correct opening hours, and a business description that matches the site.
How to check: run a page through Google’s Rich Results Test or the Schema.org validator. Our schema markup guide explains what to include.
6. You publish an llms.txt
An llms.txt file is a plain summary of your business and its key pages for language models. It is cheap to add and easy to keep current if it is generated from the same data as your site. See our llms.txt guide. Keep expectations honest, though: in our September logs it was requested only occasionally by AI crawlers.
Trust: is what you say true?
7. Every claim is true and current
An AI assistant repeats what your website says, including anything out of date. Old prices, turnaround times you no longer promise, or services you no longer offer all get quoted back to potential clients. When we reviewed our own site we found and corrected old build times on ten pages and a wrong claim about monthly payment plans.
How to check: search your own site for prices, timeframes and promises, and confirm each one is still true.
Use: can an AI agent act on your site?
8. The page structure is accessible
AI agents navigate a page much as a screen reader does, using its structure: headings, landmarks, labelled buttons and fields. An accessible site is easier for agents too. PageSpeed Insights now checks this directly, under “Accessibility tree is well-formed”.
9. Forms are labelled, and ready for WebMCP
Every form field should have a proper label. The newer step is WebMCP, a proposed standard that lets a form describe itself to an AI agent in the browser with a few HTML attributes. It is in trial in Chrome, so it is preparation rather than a requirement today. Never let an agent submit a form that sends, books or pays without the person pressing the button. And check your spam protection first: in our testing, a hidden honeypot field was exposed to the agent as an ordinary input, and on most sites a filled honeypot silently throws the enquiry away. Our WebMCP guide covers the details and the safety choices.
10. Public tools are listed in an AI catalog
If your website offers something an agent can use, such as an availability check, a price list or a calculator, list it in an AI catalog at /.well-known/ai-catalog.json. We list two: our domain availability checker, and a public list of our packages and prices that an agent can read to quote our real prices. If your site has no such tools, there is nothing to list. See our ai-catalog.json guide.
Measure: can you prove it works?
11. You test it with PageSpeed Insights
Google’s PageSpeed Insights now has an “Agentic Browsing” category alongside Performance and SEO. It checks the accessibility tree, layout stability, your llms.txt, your AI catalog and WebMCP. It shows a pass count rather than a score, because the standards are still settling. Our form pages pass all six, including the three WebMCP checks, once the WebMCP origin trial token is set up correctly. A third-party token must be added by a script from your own domain, not a meta tag, as our WebMCP guide explains.
One warning from experience: a new robots.txt directive for AI catalogs, Agentmap:, is currently flagged by the same tool’s SEO audit as an unknown directive. Test after every change.
How to check: run your homepage at pagespeed.web.dev and open the Agentic Browsing section, on both mobile and desktop.
12. You can tell which enquiries came from AI
If you cannot see where enquiries come from, you cannot tell whether any of this works. Record the referring site and landing page with each enquiry. Visitors who decline cookies are invisible to analytics, so store it with the enquiry itself. That is how we know that of the 4 enquiries we could trace between 10 and 30 September 2026, 3 came from ChatGPT and landed on our domain checker.
What this checklist will not do
None of these steps guarantees that an AI assistant will recommend you, and nobody can honestly promise that. They remove the reasons it would not: a site it cannot reach, cannot read, cannot trust or cannot use. The rest comes from being genuinely good at what you do, and saying so clearly and truthfully.
Want this done for your website?
We build agent-ready websites and can check an existing one against this list, including the server log check most audits skip. Our AEO and AI search guide covers the wider strategy. Request a quote and tell us about your site.
Frequently Asked Questions
What is an agent-ready website?
A website that AI search tools can reach and read, and that AI agents can use correctly: its content is in plain HTML, its claims are accurate, its structured data is correct, its forms are clearly labelled, and any public tools are listed where agents can find them.
How do I know if AI crawlers can reach my website?
Check two things: that robots.txt allows them, and that your server is not refusing them. The second needs the server access logs. On our own shared hosting, about 39% of ChatGPT-User requests were being refused before we fixed it.
Do I need llms.txt, ai-catalog.json and WebMCP?
llms.txt is cheap and worth having. ai-catalog.json is only useful if you have a public tool to list. WebMCP is worth preparing if your site’s main actions are forms, but it is still in trial in Chrome.
Will these changes improve my Google ranking?
Some overlap with good SEO, such as readable content, accurate structured data and a fast, accessible site. The newest items, AI catalogs and WebMCP, are not ranking factors as far as any evidence shows.
How can I test my website for AI agents?
Run it through PageSpeed Insights and check the Agentic Browsing section, validate your structured data, and read your server logs for how AI crawlers are being answered.