CometWeb tools
AI visibility checker
A free check of one page for AI crawlers · robots.txt for 14 crawlers, text without JavaScript, directives and schema
This AI visibility checker reads one public page the way an AI crawler gets it: the server HTML, robots.txt and the indexing directives, skipping JavaScript. It checks whether AI search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot may fetch the page and how much text is left without JavaScript. It also checks whether noindex or nosnippet limits quoting, and which structured data and llms.txt the site has. The result shows readiness to be read. Citations in ChatGPT, Perplexity or Google stay out of reach for any outside tool, because those systems keep that data to themselves.
What to fix first
Start with robots.txt. A single Disallow: / under the wrong user agent hides the page from an AI search engine. Allow the search crawlers you care about, and decide about training crawlers separately. You can test each rule in our robots.txt tester.
Then check the text without JavaScript. Open the page source (Ctrl+U) and search for a sentence from your main content. If it is missing there, a crawler that skips scripts misses it too. Server-side rendering or prerendering fixes that for the whole site.
Directives and schema come next. A noindex or nosnippet left over from a staging setup is a common reason a page is missing from results. Structured data and an llms.txt file help machines read the page, with no promise of citations.
This is one page at one moment. To check every page and repeat the check after each release, see AI search readiness in CometWeb Insight.
Check several kinds of page
robots.txt rules and rendering often differ between templates. Check the home page, one article or product and a category page. If only one template has little text in the server HTML, the fix belongs to that template, not to the whole site.
Google Search Central · AI features and your website ↗Terms in the result
- robots.txtRobots Exclusion Protocol (RFC 9309)
- A file with rules on which pages each crawler may fetch.
- User-agentCrawler name
- The token a crawler looks for to find its rules in robots.txt.
- SSRServer-side rendering
- The server sends finished HTML with the content, without waiting for JavaScript.
- nosnippetRobots directive
- Turns off the snippet in results, and in Google also use in AI Overviews.
- JSON-LDJSON for Linked Data
- schema.org structured data written in a script tag.
Enter a URL and choose Check page. Our server fetches the page, its robots.txt and /llms.txt, and the result shows which AI crawlers may fetch it, how much text they get without JavaScript and whether a directive limits quoting. To see whether a specific AI system cites your site, ask it the questions your customers ask and look at the sources of the answer.
For AI search and answers: OAI-SearchBot and ChatGPT-User (OpenAI), Claude-SearchBot and Claude-User (Anthropic), PerplexityBot and Perplexity-User (Perplexity), plus Googlebot and Bingbot for classic results. Training crawlers, that is GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot and meta-externalagent, are a separate decision about licensing your content.
OpenAI documents GPTBot (training) and OAI-SearchBot (ChatGPT search) as independent settings, so you can block GPTBot and still allow OAI-SearchBot. Anthropic (ClaudeBot and Claude-SearchBot), Google (Google-Extended does not affect Google Search) and Apple split them the same way.
Most AI crawlers read the HTML the server sends and skip JavaScript. If content appears only after scripts run, as in many single-page apps, such a crawler gets an empty shell. Server-side rendering or prerendering puts the text in the first response.
It is optional. llms.txt is a community proposal, and Google says its AI features need no special AI text files or markup. Adding one takes a few minutes, so we report it as information. Our llms.txt generator builds one from your sitemap.
Yes. Google states that nosnippet and max-snippet:0 also prevent the content being used as direct input for AI Overviews and AI Mode. A positive max-snippet limits the length of a quote, and data-nosnippet excludes only the marked parts.
Structured data helps machines understand what a page is about, but citations and rankings depend on much more than schema. Google says its AI features need no special schema.org markup. Use the types that match the page and keep the JSON valid.
The tool checks readiness: access, readable text and directives. Citations depend on each system's index and on its answer to a given question, and that data is not public. Checks across the whole site, repeated after changes, are described on the AI search readiness page of CometWeb Insight.
What else might come in handy
One page is a sample. The SEO module in Insight checks AI crawler access in robots.txt and the other signals across the whole site, then repeats the check after a release.