CLICK CODED · live experiment, started 2026-07-27
Every AI-visibility score this studio publishes — including the one we sell — is a static, verifiable technical check: does llms.txt exist, is there structured data, and so on. As the methodology page says plainly: "we never ask ChatGPT or Claude what they think of a site and report the answer as a score."
That's honest. It's also a gap: we've never actually verified that fixing those technical checks changes what a real AI assistant says when asked about a business. This page is that verification, run in public, with every answer published as we get it.
Byte-for-byte identical visible content. Adds llms.txt, agents.md, and LocalBusiness/FAQ schema.org JSON-LD.
Both describe "Fernbrook Ledger Co.," a fictional bookkeeping business — clearly disclosed as a research fixture on both pages, not a real business and not for real customer contact. Using a fictional subject means any answer an assistant gives can only come from what it read on these two pages (or that it has no information at all), not from prior real-world knowledge about an actual company.
On a recurring schedule, we ask each assistant (ChatGPT, Perplexity, Gemini) the same three prompts, once naming the control URL and once naming the treatment URL:
We log the raw answer, whether it accurately reflects the page content, and whether the assistant appears to have read the page at all (vs. declining to answer / hallucinating / finding nothing).
2026-07-27, baseline round (Perplexity, all 3 prompts): control and treatment came back symmetric on all three prompts — "no information / can't verify / wouldn't recommend" regardless of retrofit files. No signal yet either way; Perplexity doesn't appear to have crawled either page yet.
2026-07-27, ChatGPT baseline, prompt 1 (first real signal): control got a confidently hallucinated pricing table sourced to an unrelated real business; treatment got an honest "no information, likely a test fixture" answer. Same model, same day, same fictional business — only the URL differed. One data point, not a conclusion yet — repeat runs needed.
2026-07-27, Gemini baseline, prompt 1: symmetric — both pages read accurately, no hallucination on either side. Gemini appears to fetch and read the live page directly rather than relying on a search index, which may explain the difference from ChatGPT's result.
2026-07-27, ChatGPT baseline, prompts 2 and 3 (ChatGPT's 6/6 complete): symmetric and accurate on both pages this round — no repeat of prompt 1's hallucination. Tempers, but doesn't overturn, the earlier finding: needs repeat rounds before concluding anything causal.
2026-07-27, Gemini baseline, prompts 2 and 3 (full 18/18 baseline complete): mostly symmetric and accurate, with one exception — on prompt 2, control's answer directly quoted the page while treatment's answer reasoned speculatively (still landed on the right conclusion, just less grounded). Honest overall read after the full round: no consistent directional effect yet across all 18 queries — one hallucination (ChatGPT/control/prompt 1) and one grounding-quality gap (Gemini/treatment/prompt 2) on an otherwise clean, symmetric baseline. Full raw text: results.md. Ongoing re-checks continue as the pages age and get more crawl time.
Part of Click Coded: trust between humans and AI, checkable. The Receipts Standard