CLICK CODED · live experiment, started 2026-07-27
Every AI-visibility score this studio publishes, including the one we sell, is a static, verifiable technical check: does llms.txt exist, is there structured data, and so on. As the methodology page says plainly: "we never ask ChatGPT or Claude what they think of a site and report the answer as a score."
That's honest. It's also a gap: we've never actually verified that fixing those technical checks changes what a real AI assistant says when asked about a business. This page is that verification, run in public, with every answer published as we get it.
Byte-for-byte identical visible content. Adds llms.txt, agents.md, and LocalBusiness/FAQ schema.org JSON-LD.
Both describe "Fernbrook Ledger Co.," a fictional bookkeeping business. Clearly disclosed as a research fixture on both pages, not a real business and not for real customer contact. Using a fictional subject means any answer an assistant gives can only come from what it read on these two pages (or that it has no information at all), not from prior real-world knowledge about an actual company.
On a recurring schedule, we ask each assistant (ChatGPT, Perplexity, Gemini) the same three prompts, once naming the control URL and once naming the treatment URL:
We log the raw answer, whether it accurately reflects the page content, and whether the assistant appears to have read the page at all (vs. declining to answer / hallucinating / finding nothing).
2026-07-27, baseline round (Perplexity, all 3 prompts): control and treatment came back symmetric on all three prompts. "no information / can't verify / wouldn't recommend" regardless of retrofit files. No signal yet either way. Perplexity doesn't appear to have crawled either page yet.
2026-07-27, ChatGPT baseline, prompt 1 (first real signal): control got a confidently hallucinated pricing table sourced to an unrelated real business. Treatment got an honest "no information, likely a test fixture" answer. Same model, same day, same fictional business. Only the URL differed. One data point, not a conclusion yet. Repeat runs needed.
2026-07-27, Gemini baseline, prompt 1: symmetric. Both pages read accurately, no hallucination on either side. Gemini appears to fetch and read the live page directly rather than relying on a search index, which may explain the difference from ChatGPT's result.
2026-07-27, ChatGPT baseline, prompts 2 and 3 (ChatGPT's 6/6 complete): symmetric and accurate on both pages this round. No repeat of prompt 1's hallucination. Tempers, but doesn't overturn, the earlier finding: needs repeat rounds before concluding anything causal.
2026-07-27, Gemini baseline, prompts 2 and 3 (full 18/18 baseline complete): mostly symmetric and accurate, with one exception, on prompt 2, control's answer directly quoted the page while treatment's answer reasoned speculatively (still landed on the right conclusion, just less grounded). Honest overall read after the full round: no consistent directional effect yet across all 18 queries. One hallucination (ChatGPT/control/prompt 1) and one grounding-quality gap (Gemini/treatment/prompt 2) on an otherwise clean, symmetric baseline. Full raw text: results.md. Ongoing re-checks continue as the pages age and get more crawl time.
2026-07-30, Perplexity re-check (3 days later, all 3 prompts): prompt 1's asymmetry repeated. Control hallucinated a substitute real business's pricing again, treatment declined honestly again. That's the same pattern now seen on two different models (ChatGPT 07-27, Perplexity 07-30) on the same prompt. Prompts 2 and 3 stayed symmetric and accurate on both pages. Still not proof of causation, but a second data point worth tracking. ChatGPT/Gemini re-checks queued for next authenticated run.
2026-08-01, ChatGPT + Gemini re-check (prompt 1 only, both pages): closed the queued gap from 07-30. This round, the earlier asymmetry did not repeat on ChatGPT — both control and treatment came back honest and accurate, no hallucination on either side. Gemini stayed symmetric and accurate on both pages too, consistent with every prior Gemini round. Honest running tally on prompt 1: 2 rounds showed the hallucinate-control/honest-treatment asymmetry (ChatGPT 07-27, Perplexity 07-30), 1 round showed none at all (ChatGPT, this round). No consistent directional effect has survived three models' worth of testing yet. Full raw text: results.md.
2026-08-01, Perplexity re-check (all 3 prompts, run early): the strongest signal yet. Prompts 2 and 3 both reproduced the pattern seen 07-27/07-30 — control confidently hallucinated by conflating Fernbrook with a real unrelated business ("The Ledger Company," the same substitute both times), presenting its real BBB history and reviews as Fernbrook's own; treatment correctly declined or explicitly flagged the name mismatch. Prompt 1 reversed the direction (treatment hallucinated this time, control didn't) — the first time that's happened. Net across three rounds: 4 of 5 observed hallucination events now show "control hallucinates, treatment doesn't," against 1 in the opposite direction. Still not proof, and "The Ledger Company" being a stable nearest-neighbor match is itself a confound worth tracking. Full raw text: results.md.
2026-08-16, Perplexity re-check (all 3 prompts, both pages): a new pattern — zero hallucination on either page, all six answers accurate. Both control and treatment correctly identified Fernbrook as a disclosed fictional research fixture, with no conflation with "The Ledger Company" or any other real business this round, on control too, despite control having no retrofit files. Likely explanation: Perplexity's index has now crawled deep enough into this page's parent methodology page that it correctly contextualizes both child URLs regardless of their own files — suggesting the retrofit-file effect may fade as a site gets more broadly indexed rather than being permanent. Worth watching over future rounds. ChatGPT/Gemini prompts 2 and 3 remain queued, blocked this round by a local browser-tooling bug, not by login. Full raw text: results.md.
Part of Click Coded: trust between humans and AI, checkable. The Checkable Standard