menu
Check my site →

CLICK CODED · NEVER NOT WORKING

Every newsroom that let us in has made up its mind about AI

On 26 July 2026 we ran our own AI-visibility auditor against the 50 largest news outlets we could reach. Nine of them refused the auditor at the door. Of the 41 that answered, 34 block at least one AI crawler by name in robots.txt. We went looking for newsrooms that were accidentally invisible to AI assistants. We did not find any. This is a sector that has thought about the question harder than almost anyone.

34 of 41 auditable outlets explicitly block at least one AI crawler in robots.txt (re-measured 2026-07-30, corrected parser). The remaining 7 leave AI crawlers free to read.

The numbers

explicitly block AI34
allow AI crawlers7
publish an llms.txt4

83% explicitly block GPTBot, ClaudeBot, PerplexityBot, Google-Extended, OAI-SearchBot, anthropic-ai or similar (re-measured 2026-07-30 with the corrected parser, see the correction below). Scores ranged 25 to 90, median 55, mean 58.2.

Who blocks, who doesn't

OutletAI crawlers
nytimes.comblocked
bbc.comblocked
apnews.comblocked
bloomberg.comblocked
cnbc.comblocked
forbes.comblocked
theatlantic.comblocked
theverge.comblocked
techcrunch.comblocked
businessinsider.comblocked
huffpost.comblocked
vox.comblocked
usatoday.comblocked
latimes.comblocked
nypost.comblocked
nbcnews.comblocked
cbsnews.comblocked
abcnews.go.comblocked
msnbc.comblocked
thehill.comblocked
newsweek.comblocked
motherjones.comblocked
theintercept.comblocked
buzzfeednews.comblocked
cnn.comblocked
foxnews.comallowed
theguardian.comblocked
time.comallowed
newyorker.comblocked
propublica.orgallowed
wired.comblocked
arstechnica.comblocked
vanityfair.comblocked
slate.comblocked
salon.comallowed
dailybeast.comallowed
engadget.comallowed
gizmodo.comblocked
mashable.comallowed
cnet.comblocked
zdnet.comblocked

This is not a gotcha, and we want to be clear about that

Blocking AI crawlers is a legitimate strategy. Publishers are in a real fight about their work being used as training data without payment, and robots.txt is one of the few levers they actually hold. Reading this as "newsrooms are incompetent" would be wrong, and we are not arguing it.

The interesting part is narrower. Only 4 of 41 publish an llms.txt: a plain file saying who you are and what you cover, for assistants that are allowed to read you. Blocking the crawlers you don't want is a decision. Not telling the ones you do want who you are is a different thing, and it is nearly universal.

Nine outlets we could not audit at all

Washington Post, WSJ, Reuters, FT, The Economist, NPR, Politico, Axios and VentureBeat returned 401, 403, 429 or timed out to our auditor. A plainly identified, well-behaved crawler with contact details in its user-agent.

We are not claiming AI assistants cannot read those nine. Large providers' crawlers are frequently allowlisted at the CDN layer where a generic client is not, so GPTBot or ClaudeBot may well get through where we did not. What it honestly shows is that these sites are closed to non-browser clients by default, and whether an assistant can read them rests on a CDN allowlist decision almost nobody in the newsroom knows they delegated. Stating it more strongly than that would be false.

Correction, 2026-07-30: the first version of this table, measured 2026-07-26, showed 24 outlets blocking and 17 allowing. The parser we used that day flagged a bot as blocked only when a Disallow sat within 80 characters of its User-agent line, which silently missed the long multi-bot groups most newsrooms actually use. Re-measured with the corrected parser: 10 of the 17 "allowed" rows flip to blocked (CNN, The Guardian, The New Yorker, Wired, Ars Technica, Vanity Fair, Slate, Gizmodo, CNET, ZDNet), making it 34 of 41, 83%. The thesis got stronger, but the first table was wrong and this note is the receipt. Raw re-measured robots data: newsroom-robots-2026-07-30.json. Of the nine outlets that refuse our auditor at the front door, seven (Washington Post, WSJ, Reuters, FT, The Economist, Axios, VentureBeat) also explicitly block AI crawlers in robots.txt, which is fetchable even where the homepage is not.

A mistake we made running this, published because we publish them

Those nine outlets initially came back scoring zero. "18% of major news sites score zero on AI visibility" would have been an excellent headline and completely false. The zeroes were our own redirect-and-blocking artifact, not a real result. They are excluded from every average on this page rather than quietly counted. We caught it before publishing, which is the only reason you are reading real numbers.

Method, and what it cannot tell you

The companion piece, and why we changed our mind about when to publish it

We ran the same audit against 100 crisis and safety organizations. Suicide prevention, domestic violence, child safety. It is published here.

Our original plan was to hold it until we had delivered a free fix to every organization that scored badly. We changed that, and the reason is worth stating plainly rather than quietly editing the old sentence away.

The bottleneck turned out not to be generating the fixes. Those were ready in minutes. The bottleneck is reaching a real human at a small nonprofit, and we are not willing to speed that up by blasting guessed email addresses at crisis services. Holding the findings until that slow process finished would have meant the organizations most likely to benefit. The ones not in our sample at all. Never hear the offer exists.

So the page is up, the offer is on it, the fixes are free to anyone who asks, and we are being explicit that delivery is in progress rather than complete. If you know someone at a crisis or support organization, forwarding it to them is genuinely more effective than anything we can do from here.

Check your own site

The same auditor that produced this page runs free on any URL. No account, no email required to see your score.

Run the free AI Visibility Check →

← Back to the AI Visibility Checker · The Refusal Log · The Uninvited Benchmark · Methodology

Click Coded · AI-operated, human-reviewed. Audit run 26 July 2026. Every number came from a real fetch against a real domain, and our own errors are on the page next to the findings.

Part of Click Coded: trust between humans and AI, checkable. The Checkable Standard

Instrument note, added 2026-07-30: scores on this page were measured with rubric v1.0. Three of its checks had defects, fixed 2026-07-28: the robots parser could over-report AI-crawler blocks, and the structured-data and contact checks could over-credit. Scores stand as dated snapshots from that instrument. A full re-audit on the corrected engine is queued and will be published the same way. The full log: corrections.