W
Reviewed by Jacob Whitmore, Whito · Fact-checked for accuracy

Last Updated on August 25, 2026

The mistake most UK businesses worry about is not the one making them invisible to AI. When Whito audited 120 UK business websites, five were blocking an AI crawler. Five. The thing everybody is anxious about affects roughly four in a hundred.

The mistakes that actually keep businesses out of AI answers are duller and close to universal. About 94 in 100 of those same sites had no structured data saying what kind of business they are. About 91 in 100 answered no questions on the page. In a separate audit of 200 sites, fewer than half published a single price, and not one plumber did.

This page lists the eight that matter, each with the measured frequency, why it costs you, and the fix. Every number comes from a Whito study you can open and check, or from the AI companies’ own documentation.

94%

Of 120 UK business sites had no correct business type structured data

91%

Had no visible FAQ content. One site in 120 had FAQ structured data

4%

Were blocking an AI crawler. The problem everyone worries about is the rarest one here

0

News outlets cited across 159 AI answers recommending UK local businesses

Why these are the mistakes that count

Google’s own documentation sets the bar plainly. To appear as a supporting link in AI Overviews or AI Mode, “a page must be indexed and eligible to be shown in Google Search with a snippet”, and “there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary”.

So there is no secret AI layer to optimise. There is only whether a machine can find your site, read it, and extract facts from it that it is willing to repeat. Every mistake below is a failure at one of those three steps.

And the machines do read your site. In Whito’s study of 159 AI Mode answers recommending UK local businesses, the business’s own website was cited in 90% of them, more often than any other source including Google itself. Your website is the primary evidence. Most of them say almost nothing a machine can use.

The eight mistakes

1. Your website does not say what kind of business you are in a way a machine can read

Measured: about 7 of 120 UK business websites carried correct business type structured data. 113 did not.

Structured data is a small block of code that states your business type, name, address, phone number, opening hours and service area as machine readable facts rather than as words in a paragraph. Whito found JSON-LD of some kind on 47 of 120 sites, but mostly generic page markup. Correct business type schema appeared on about seven.

Without it, a machine has to infer what you are from your prose. It usually can. But when it is choosing between you and a competitor who has stated it as fact, inference loses.

The fix, about 30 minutes. Add one LocalBusiness JSON-LD block to your homepage with your exact name, address, phone, hours and service area. Then check it at validator.schema.org, which is free and needs no account. The details in the code must match the ones on the page.

2. You publish no prices

Measured: 45.5% of 200 UK independent business websites published any price. 0 of 20 plumbers and 0 of 20 gardeners published a single figure.

Price is the question customers actually ask, and it is one of the few facts an AI answer can quote about you that a competitor cannot claim too. Whito’s audit of 120 sites found 63 published prices. Restaurants managed 12 of 12. Accountants and builders managed 0 of 24.

The usual objection is that every job is different. That is true and it is not a reason to publish nothing. A range, a starting price, a typical job, or an hourly rate with what it includes all beat silence.

The fix, about an hour. Publish one honest number. “Most kitchen rewires cost £900 to £1,600 depending on the number of circuits” is a fact a machine can repeat and a customer can act on. “Contact us for a quote” is neither.

3. Your site answers no questions

Measured: 11 of 120 UK business websites had visible FAQ content. Exactly 1 of 120 had FAQPage structured data.

AI answers are assembled from passages that already answer a question. A page written as a brochure gives an engine nothing to lift. A page with a plain question as a heading and a direct answer underneath gives it a ready made passage.

This is the cheapest mistake on the list to fix and the least done. One site in 120.

The fix, about an hour. Write down the six questions customers ask you on the phone. Put each one on your site as a heading, with a direct answer underneath in two or three sentences. Answer in the first sentence, then explain. Do not wind up to it.

4. Your homepage is JavaScript and serves an empty page

Measured: 5 of 120 UK business websites served effectively empty HTML on the homepage.

Some website builders and custom builds deliver a page that is blank until a browser runs the code that fills it in. A human sees a normal site. A crawler that does not run that code sees nothing at all.

It is rare, at about four in a hundred, but it is total. Every other item on this list is irrelevant if this one is true of your site, because there is no content to read.

The fix, two minutes to check. Right click your homepage and choose View Page Source. If you cannot find your own business name, your phone number or your main headline in that text, a crawler probably cannot either. That is a conversation to have with whoever built the site.

5. Your public record does not match your website

Measured: 51% of 182 businesses recommended by AI tools for electrical work could not be matched to an active registered company. In 6 of 15 towns, at least one tool recommended a business whose registered company had been dissolved or was in liquidation.

This one runs the other way from the rest. It is not only about whether you appear. It is about whether what appears about you survives being checked.

If your website trades under one name, your invoices carry another, and Companies House holds a third, then anything checking you finds three businesses and cannot confirm any of them is real. That is a bad position when the engine recommending you is also, increasingly, being checked by the customer.

The fix, about 20 minutes. Look yourself up on Companies House, which is free to search. Check the registered name, trading name, registered address and officers against what your website says. Fix whichever is wrong. Then check your filings are not overdue, because that is public too.

6. You are not listed where AI actually looks

Measured: across 159 AI Mode answers, trade directories and review platforms were cited in 85 answers (53%). Checkatrade appeared in 62, TrustATrader in 26, Which? Trusted Traders in 15. Trustpilot appeared in 2. Yelp appeared in 0.

Which directory you are on is not a matter of taste. Whito counted the sources behind 159 AI answers, and the spread is severe. Checkatrade alone appeared in 50 of the 60 trade answers. Trustpilot, which many businesses treat as the serious one, appeared twice in the whole study.

Google was cited in 128 of 159 answers, second only to businesses’ own websites. An incomplete Google Business Profile is the most expensive blank space on this list.

The fix, an afternoon. Complete every field of your Google Business Profile first, since it is free and it is cited in four answers out of five. Then work out which directories your own sector’s answers actually cite, rather than paying for the one with the best sales team.

7. You are betting on press coverage

Measured: 0 of 159 AI answers cited a news outlet. No newspaper or news site appeared at all, in any sector, in any town.

This is the most counterintuitive finding in the set, so it is worth stating precisely. In one study, of one engine, on one kind of question, press coverage did nothing. Not a little. Nothing.

Coverage may well be worth having for reasons that have nothing to do with AI: credibility, links, customers who read it. But if a PR agency is selling you coverage as the route into AI recommendations for local work, ask them for the evidence, because ours points the other way.

The fix, free. Stop treating press as an AI visibility tactic for local search. Spend the effort on the directory listings and the Google profile that the same study shows are cited in half and four fifths of answers respectively.

8. You installed an llms.txt file and thought that was the job

Measured: 21 of 120 sites had an llms.txt file. 20 of those were generated automatically by Wix, Shopify or an SEO plugin. Exactly 1 in 120 had been written by a human.

An llms.txt file is a short plain text page telling AI systems what a site is about. It is being sold across the UK as an AI visibility service. Here is what the primary sources say.

The specification is a proposal published by one person, Jeremy Howard, dated 3 September 2024 and modified 10 August 2026, describing itself as open for community input. It is not a standards body document and no AI company issues it. The crawler documentation of OpenAI, Anthropic and Perplexity all covers robots.txt in detail and none of them documents llms.txt as something their crawlers read. Google states directly that creating such files “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them”.

The fix, ten minutes or nothing. Write one if you want. It is cheap and it may matter later. Do not pay anyone for it, and do not let it stand in for the seven items above, all of which are measurable today.

The mistake you are probably worrying about instead

Blocking the AI crawlers is the fear that gets the attention, and the data says it is the rarest problem here. Whito found 5 of 120 sites blocking an AI crawler, all of them through a platform toggle. A separate audit of 49 sites found 1 blocking, and found that not one small or independent business had made a deliberate, owner led decision either way.

So the risk is not that everyone is blocking. It is that the few who are did not choose to, and that anyone who does choose to usually blocks the wrong thing.

If you block thisWhat actually happens
GPTBotYour content is not used to train OpenAI’s models. You still appear in ChatGPT search. This is the training crawler, not the search one
OAI-SearchBotOpenAI states that sites opted out “will not be shown in ChatGPT search answers, though can still appear as navigational links”. This is the one that costs you visibility
Google-ExtendedYour content does not train or ground Gemini. Google states it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search”
ClaudeBotAnthropic describes this as the crawler that collects content that may contribute to training
PerplexityBotPerplexity describes this as the crawler that surfaces and links websites in its search results, and states it “is not used to crawl content for AI foundation models”

Source: each company’s own crawler documentation, read 25 August 2026. OpenAI states the settings are independent: “a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot.”

The distinction worth getting right

Opting out of AI training and being invisible to AI search are two different decisions. Blanket advice to block anything with an AI sounding name makes the second one while trying to make the first.

If you want your content kept out of model training but still want to appear in AI answers, that is a supported position and the vendors document exactly how.

All eight, ranked by how common they are

MistakeHow commonFix timeCost
No correct business type structured data113 of 12030 minutesFree
No FAQ content on the site109 of 1201 hourFree
No prices published109 of 2001 hourFree
llms.txt treated as the answer20 of 21 auto-generated10 minutes or skipFree
Public record does not match the site51% unmatched in one sector study20 minutesFree to check
Not on the directories AI citesVaries by sectorAn afternoonFree to start
JavaScript only homepage5 of 1202 minutes to checkFree to check
Blocking an AI crawler5 of 1202 minutes to checkFree

Frequencies come from three separate Whito studies with different samples, so they are not directly comparable with one another. Each row states its own base. Sources below.

Every fix on this list is free. Not one of them requires a product with AI in the name.

The order to fix them in

Eight is too many for one sitting. If you have an afternoon, do these four.

  • View your page source and check your own name is in it. Two minutes. If it is not, nothing else on this list matters yet.
  • Complete every field of your Google Business Profile. Cited in 128 of 159 answers. The highest return of anything here.
  • Publish one honest price. An hour, and it is the fact most likely to be quoted about you.
  • Write six questions and answer them on the page. An hour, and only one site in 120 has bothered.

Structured data, Companies House and directories come next. llms.txt comes last, if at all.

Methodology and sources

Every frequency on this page comes from a published Whito study, and every vendor statement was read at its original source on 25 August 2026. The three Whito studies use different samples and different methods, so a percentage from one is not comparable with a percentage from another. Each figure states its base where it appears.

Whito studies used

  • AI readiness of UK business websites. 120 websites across 10 sectors, checked 2 July 2026. Source of the structured data, FAQ, llms.txt, pricing, crawler policy and empty HTML figures.
  • The AI door test. 49 UK business websites, June 2026. Source of the finding that no small or independent business had made a deliberate, owner led crawler decision.
  • What AI actually cites when it recommends a UK local business. 159 Google AI Mode answers to the prompt “Recommend a trustworthy [trade] in [town]”. Source of every citation share, the directory counts and the zero news finding.
  • UK price transparency study. 200 independent businesses across 10 trades and 10 cities, checked 10 August 2026. Source of the 45.5% figure and the plumbers and gardeners findings.
  • Which electricians does AI recommend. 182 businesses across 15 UK towns. Source of the unmatched company and dissolved company findings.

Vendor and official documentation

  • Google, AI features and your website, last updated 10 December 2025. Eligibility for AI Overviews and AI Mode. developers.google.com
  • Google, AI optimization guide, last updated 10 July 2026. The statement that Google Search ignores llms.txt and similar files. developers.google.com
  • Google, Google crawlers and user-triggered fetchers, last updated 14 July 2026. Google-Extended. developers.google.com
  • OpenAI, Overview of OpenAI Crawlers. The four bots and the independence of the settings. developers.openai.com
  • Anthropic, Does Anthropic crawl data from the web, 7 April 2026. ClaudeBot, Claude-User and Claude-SearchBot. support.claude.com
  • Perplexity, PerplexityBot documentation. docs.perplexity.ai
  • llms.txt specification, Jeremy Howard, published 3 September 2024, modified 10 August 2026. llmstxt.org
  • Companies House company search, free to search and view. gov.uk

Known limits

  • The three Whito samples are different sizes and were collected on different dates. 120 sites in July, 49 in June, 200 in August, 159 answers in August. Do not add the percentages together or treat them as one dataset.
  • The citation study covers one engine and one question shape. Google AI Mode, asked “Recommend a trustworthy [trade] in [town]”. The zero news finding is specific to that. It does not mean press coverage never influences any AI answer about anything.
  • The 51% unmatched figure is from one sector. Electricians, 182 businesses. It is not a claim about UK businesses generally.
  • Frequency is not the same as impact. A JavaScript only homepage affects 4% of sites and is fatal to all of them. Missing structured data affects 94% and is a disadvantage rather than a wall. The ranking table sorts by how common, not by how much it costs you.
  • None of these studies proves causation. They measure what is present on sites and what engines cite. They do not demonstrate that fixing any single item causes an engine to recommend you.
  • Vendor crawler documentation changes without notice. Everything here is dated 25 August 2026. Check the source links before acting on a detail.
  • Our own free scorecard names GPTBot in its guidance on AI crawler blocks and does not yet mention OAI-SearchBot. On the evidence above, that wording understates the distinction between training and search opt outs. We are correcting it, and are recording that here rather than changing it quietly.

Questions people ask

What is the most common reason a UK business is invisible to AI?

Missing structured data. In Whito’s audit of 120 UK business websites, about 7 carried correct business type structured data and 113 did not. The second most common is having no FAQ content, which 109 of 120 lacked.

Does blocking GPTBot make my business invisible to AI?

No. GPTBot is OpenAI’s training crawler. Blocking it stops your content being used to train models but does not remove you from ChatGPT search. The crawler that controls that is OAI-SearchBot, and OpenAI states that sites opted out of it “will not be shown in ChatGPT search answers”.

Do I need to publish prices to appear in AI answers?

You do not need to, but price is one of the few facts about you that an engine can quote and a competitor cannot claim. Whito found 45.5% of 200 UK independent business websites published any price, and 0 of 20 plumbers published a single figure. A range or a typical job beats “contact us for a quote”.

Does press coverage help my business appear in AI recommendations?

Not in the study we ran. Across 159 Google AI Mode answers recommending UK local businesses, no news outlet or newspaper was cited at all, in any sector, in any town. That is one engine and one type of question, but it is a clear result.

Is llms.txt worth adding to my website?

It costs ten minutes, so add one if you want. But no major AI company has publicly committed to reading it, the specification is one person’s proposal rather than a standard, and Google states that Google Search ignores such files and that creating one will neither help nor harm your visibility. Do not pay for it.

How do I check if AI can read my website at all?

Right click your homepage, choose View Page Source, and look for your business name and phone number in the text. If they are not there, a crawler probably cannot see them either. Then visit yourdomain.co.uk/robots.txt and look for lines blocking named AI crawlers.

The sharp takeaway

Businesses invisible to AI are not usually being blocked, penalised or outspent. They are being skipped, because their website states almost nothing a machine can lift and repeat with confidence.

Ninety four sites in a hundred do not say what kind of business they are in machine readable terms. Ninety one in a hundred answer no questions. More than half publish no price. Those three gaps cost nothing to close and almost nobody closes them.

Stop looking for the thing that is blocking you. In four cases out of a hundred there is one. The rest of the time the problem is that your website has not given anything worth quoting.

Find out which of these applies to you

The free Whito scorecard runs 17 checks covering most of this list and tells you which one to fix first. About two minutes, no sign up, no card.

Check your business, free Get the free guide

Related reading

author avatar
Whito
Whito exists to stop businesses scaling the wrong way. We focus on structure, leverage, and measurable growth, not noise, not vanity metrics.