AThe AI Visibility Index

Crawler

AVI-Crawler

AVI-Crawler is the automated reader behind The AI Visibility Index. It fetches public pages to measure how well a site is built to be found, read and cited by AI assistants — the same job the crawlers behind ChatGPT, Gemini and Perplexity are doing when they decide whether your business is worth naming.

How to identify it

It sends this User-Agent:

AVI-Crawler/0.1 (+https://www.theaivisibilityindex.ai/bot; AI Visibility Index readiness scan)

When a page can only be read by running its JavaScript, we may retry with a headless browser. That request carries a normal Chrome User-Agent with AVI-Crawler/0.1 appended, and the same link back to this page.

What it does

  • — Fetches your homepage, robots.txt, sitemap.xml, llms.txt, and up to four further pages (typically About, Contact, FAQ, a product page).
  • — Reads public HTML only. It never submits forms, logs in, or attempts checkout.
  • — Runs at most once a day per site, a handful of requests each time. It is not a load test and will not crawl your whole catalogue.
  • — Obeys robots.txt.

How to allow it

Nothing is required — a normal site is readable already. If you run bot protection (Cloudflare, WAF rules, rate limiting) and want your reading to be accurate, allow the User-Agent above. Many well-protected sites currently block us, and we publish no score at all rather than guess: a site we cannot read shows as unreadable rather than as low-scoring.

That is worth knowing for its own sake. If automated readers can't reach you, the assistants' crawlers are hitting the same wall — and a page they can't read is a page they can't cite.

How to block it

Add this to your robots.txt. We'll stop, and your listing will show that we have no reading.

User-agent: AVI-Crawler
Disallow: /

Contact

Questions, or a listing that looks wrong: hello@theaivisibilityindex.ai. If the business is yours you can find it in the index and claim it by proving you control the domain.