The Industrial AI Visibility Benchmark 2026

An anonymised measurement of what AI assistants can actually retrieve from 984 manufacturer, distributor and industrial-commerce catalogues.

Published 2026-08-11Fieldwork 2026-08-10Edition 2026.1Scores only · no site is identified

The finding

The industrial web is not ready to be read by machines, and the gap is not confined to small suppliers.

Across 984 catalogues the median score was 50 out of 100. 671 sites — 68% — scored below 55, the point at which an assistant is reconstructing a product rather than reading it. None reached the top band; the best result recorded anywhere in the cohort was 80.

984

catalogues

manufacturer, distributor and industrial-commerce sites measured on what AI assistants can actually retrieve from them.

50

median score

Half the market scores at or below 50 out of 100. The mean and the median sit within a point of each other: this is not a few laggards dragging an otherwise healthy field down.

0

sites graded A

Not one site in 984 cleared the top band. The best result recorded was 80.

68%

mostly opaque or worse

671 of 984 sites scored below 55, the point at which an assistant is reconstructing the product rather than reading it.

Why this was measured

Business buyers have moved. Forrester puts the share of B2B buyers using generative AI somewhere in their purchase at 94%, drawn from a survey of close to 18,000 buyers, and reports that buyers now name AI and conversational search their most meaningful research source[2][3]. The question that follows is a supply-side one, and it has had far less attention: when an assistant goes looking for a part, what does it actually find?

There is good reason to think the answer is “less than the vendor imagines”. Analysis of roughly 1.3 billion AI-crawler fetches in a single month found that the major assistants’ crawlers fetch JavaScript but do not execute it — GPTBot requested JavaScript in 11.5% of requests, Claude in 23.8%, and neither ran it[1]. A catalogue whose products are assembled in the browser is, to those crawlers, an empty shell. That study dates from December 2024 and the crawler population has grown since; it establishes the mechanism rather than today’s exact numbers.

What has been missing is a measurement of the other side of the exchange — not how assistants behave, but how ready the industrial catalogue actually is to be read. That is what this benchmark is.

ABCD25–29: 1 catalogues30–34: 261 catalogues35–39: 53 catalogues40–44: 61 catalogues45–49: 99 catalogues50–54: 196 catalogues55–59: 117 catalogues60–64: 90 catalogues65–69: 67 catalogues70–74: 27 catalogues75–79: 11 catalogues80–84: 1 cataloguesmedian 50253545556575sites
Fig. 1 — score distribution across the cohort, in bands of five. Grade thresholds marked; the copper rule is the median.

The distribution is tight and low. The middle half of the market falls between 34 and 57, and the 95th percentile — the top one site in twenty — reaches only 69. There is no long tail of excellence: the failure is systemic rather than distributional, and a site that did the obvious things well would sit clear of the entire cohort.

How the cohort graded

GradeScoreMeaningSitesShare
A85–100Agent-ready00%
B70–84Largely readable394%
C55–69Partially readable27428%
D40–54Mostly opaque35636%
F0–39Effectively invisible31532%

What was measured

Each catalogue was scored on five dimensions. They are stated here as the questions they ask, which is what a reader needs in order to judge whether the finding means anything. The weightings and the probe design are not published — see method and limitations.

01

Crawler policy

Are the assistants' own crawlers permitted, and is that permission stated deliberately rather than inherited from a default?

02

Reachability

When a crawler asks for a product page, does it get one — without a challenge page, a redirect loop, or a timeout?

03

Readability

Is the product's substance present in the served HTML, or assembled in the browser after load where a non-rendering crawler will never see it?

04

Catalogue coverage

Can the catalogue be enumerated and traversed, and do the product pages carry structured product data rather than prose alone?

05

Agent surfaces

Beyond the human website, is anything published for machines — a stated AI policy, a plain-text mirror, a declared tool endpoint?

What almost nobody publishes for machines

The scores describe how well the human website survives being read by a machine. This is the separate question of whether anything was built for the machine deliberately.

State an explicit AI-crawler policy193 / 984 · 20%

Naming the assistants' crawlers, rather than leaving them to a wildcard rule written years before those crawlers existed.

Publish an llms.txt90 / 984 · 9%

The proposed plain-text index for language models.

Serve a plain-text or Markdown mirror35 / 984 · 4%

A machine-readable rendering of a page that is otherwise delivered as an application.

Advertise a tool endpoint for agents0 / 984 · 0%

Zero. Not a rounding artefact — no site in the cohort published a discoverable agent tool interface.

The last line is the one worth sitting with. robots.txt has been a standard since 2022 [4], llms.txt is a live proposal [5], and there is now a published protocol for exposing tools to agents directly [6]. Of 984 industrial catalogues, 0 had adopted the last of these. Whatever else that is, it is not a crowded field.

The catalogue that cannot be reached

389 sites — 40% — served no product page we could read at all. Not a thin page, not a page missing structured data: nothing retrievable. Where a catalogue could be enumerated, the cohort exposed 894,308 product URLs directly, against an estimated ~54,983,275 catalogue URLs in total once sitemap indexes are scaled out.

The gap between those two numbers is the shape of the problem. Most of the industrial catalogue exists. Comparatively little of it is reachable, enumerable and marked up as a product [7] at the moment an assistant asks.

This matters beyond search. The classification models this sector already files against [9], and the Digital Product Passport regime now in force across the EU single market [8], both assume product data that is structured, addressable and machine-readable. The same work that makes a catalogue legible to an assistant is largely the work those regimes require anyway.

Method, and what it cannot tell you

Fieldwork was carried out on 2026-08-10 against live public sites, from multiple network vantage points, using only routes a member of the public could take. Nothing was logged in, no authentication was attempted, and no rate limit was deliberately exceeded. The cohort combines hand-curated majors across electronics distribution, electrical wholesale, MRO, building products and component manufacturing with a broader filtered sample of industrial commerce, deduplicated by company.

The scoring formula, the dimension weightings, the probe design and the sampling frame are not published. They are the commercial method. This is a real limit on independent replication, and it is stated plainly rather than dressed up: what follows is what the measurement cannot tell you.

  • A blocked probe is a floor, not a reading

    285 of 984 results — 29% — are low-confidence because the site challenged, throttled or refused the request. Those scores are the lowest the site could be scoring, not a measurement of it. A defensive edge configuration is indistinguishable from an unreadable catalogue when viewed from outside, and we do not pretend otherwise.

  • Absence and silence are different facts

    A 404 is evidence a thing is not there. A timeout, an empty 202 or a challenge page is evidence of nothing at all. The two are recorded separately and never collapsed into a single “not found”.

  • One day, one set of network positions

    Edge configurations change, and some sites treat traffic differently by origin, by hour and by season. This is a snapshot from a small number of vantage points, not a longitudinal study.

  • It measures retrievability, not commercial outcome

    A high score means an assistant can read the catalogue. It does not mean the assistant will recommend the part, that the data is correct, or that the buyer converts. Those are different questions and this is not evidence about them.

  • The sourced half carries a false-positive rate

    A hand-checked sample of the automatically-sourced portion found a minority of domains that were not genuinely industrial commerce. The headline figures are reported over the whole cohort without excluding them, which makes the result marginally conservative rather than flattering.

How to cite this

The figures are free to quote, in press or in research, with attribution. The distribution is published as a CSV so a claim made about it can be checked against it.

Suggested citation

Partsgraph (2026). The Industrial AI Visibility Benchmark 2026: an anonymised measurement of 984 manufacturer and distributor catalogues. Edition 2026.1, fieldwork 2026-08-10. https://partsgraph.ai/research/ai-visibility-benchmark-2026

References

  1. 01Vercel — “The Rise of the AI Crawler”. Analysis of ~1.3 billion AI-crawler fetches over one month, including 569M by GPTBot and 370M by Claude. Published 17 December 2024.
  2. 02Forrester (B. Winters) — “How GenAI And Trust Are Reshaping B2B Buying In 2026”, Forbes, 2 February 2026.
  3. 03Forrester — “The State Of Business Buying, 2026”, drawing on a Buyers’ Journey Survey of close to 18,000 business buyers. Press release, 21 January 2026.
  4. 04IETF RFC 9309 — Robots Exclusion Protocol. The 2022 standardisation of robots.txt.
  5. 05The /llms.txt proposal.
  6. 06Model Context Protocol — specification and documentation.
  7. 07schema.org — the Product type.
  8. 08Regulation (EU) 2024/1781 — Ecodesign for Sustainable Products Regulation, the framework introducing the Digital Product Passport. In force 18 July 2024.
  9. 09ETIM International — the classification model for technical products.
Where does your catalogue sit?

The grader runs the same measurement against a single site and places it in this distribution. It takes a domain, no signup, and returns the percentile alongside the score.

Grade my catalogue