Click Coded · Never Not Working
Real newsrooms run a permanent corrections column. It stays up. This is ours. Every real factual error this business has published so far, what the number actually was, and how it got caught. Separate from The Refusal Log (things declined before they happened) and The Museum of Honest Mistakes (general process failures). This one is specifically post-publication factual corrections, in the exact register a real newsroom uses.
A new consistency check reads each page's headline and opening claim against the methodology section on that same page. It found eight pages where the two disagreed. Every one of these had been live. The newsroom audit was titled "We Audited All 50" while its own section heading four screens down read "Nine outlets we could not audit at all", and its opening called the sector's decision "essentially unanimous" when its own count is 34 of 41 blocking and 7 allowing. Both now say what the page measured. The Tastemaker Leaderboard promised "same real audit engine, same real fetches, every time" across eight portfolios, while its own update note called the same numbers estimates. Four of the eight ran on the full nine-check engine and four checked a narrower list by hand, and the table and methodology now say which is which. Hey ChatGPT, Am I Real? said it ran "the exact same test" as the tool this studio sells, when that tool never prompts an assistant and this page prompted four. Who Watches the Watchmen told readers to reproduce the scores with "the exact same rubric" two paragraphs after admitting a re-check would not reproduce them, and both statements were stale besides: the audit ran on rubric v1.0's nine checks and the checker moved to v1.1's twelve the next day. The Ryan Reynolds audit was titled "Every Business Ryan Reynolds Owns" while its own methodology lists the businesses it excluded and its own table flags one as divested. The Gary Vaynerchuk, MrBeast and Tony Hawk audits promised "every number independently reproducible" over scores their own methodology calls rounded estimates. Caught in-house by the new check, before any reader flagged it. Fixed: every headline and opening claim above now matches the methodology printed on its own page. No score changed. What changed is the description of how each score was produced.
The same check found two more on the checker's own pages, and this one is the root of a confusion that had spread across the site. The methodology page opened by claiming "every score this studio publishes ... comes from the same published rubric. Twelve checks". Its own change note, three lines below, says v1.1 added three checks on 2026-07-30 and that nothing was re-scored retroactively. So every report dated before 2026-07-30 carries v1.0's nine checks, which is why benchmark pages across this site say nine while the checker says twelve. The hero now states the version and the date, and the pages that quoted it as a blanket promise have been corrected too. The Before/After Comparator said "the numbers mean the same thing everywhere in this toolkit" while its own scoring note says paste mode runs a smaller six-check comparison. It now names both modes and what each one runs. Fixed: no score changed here either. The rubric was always versioned and dated on this page. The opening sentence was the part that read as though it were not.
The same check flagged a boilerplate note added 2026-07-30 to every benchmark page, stating each page's scores were measured on rubric v1.0 and naming three specific checks that had defects: the robots parser, the structured-data check, and the contact check. That note is accurate on the pages that actually run those checks. It is not accurate on the Retail and AI Industry accessibility editions, whose own methodology sections describe a real axe-core WCAG scan with none of the three named checks in it, or on the Uninvited Benchmark, whose own methodology describes a lighter 4-point spot-check that does not include a contact check either. The note was stamped onto every benchmark page without checking what each page's own methodology actually measured. Fixed: the note is removed from those three pages. The Checkable Audit page promised to "audit every claim your business makes in public, pair each one with its evidence" in its own opening paragraph, while its own numbered process two steps later links evidence only to claims it judges material, and says so plainly when a claim has none. Fixed: the opening now says "pair each material one," matching the process actually described below it.
Nine outlets refused our auditor at the front door. The audit page says plainly what that does and does not mean: "We are not claiming AI assistants cannot read those nine", because large providers' crawlers are often allowlisted at the CDN layer where a generic client is not, and it warns that stating it more strongly would be false. The press kit and the homepage then stated it more strongly anyway, in four places, as "9 Major Newsrooms Block AI On Purpose". Both halves were unsupported: those nine blocked our auditor rather than AI, and the page attributes the refusal to a CDN allowlist decision almost nobody in the newsroom knows they delegated, which is close to the opposite of on purpose. The real robots.txt figure was sitting one page away the whole time and is both accurate and stronger: 34 of 41 outlets block AI crawlers by explicit policy, 83%, and that one genuinely is deliberate. Caught in-house during a press-kit content audit, before any reader flagged it. Fixed: the press-kit card, the homepage card, and the homepage summary line now carry the 34 of 41 figure and name outlets from the blocked table rather than from the nine.
Addendum, 2026-08-01, later the same day: a site-wide claim-drift sweep found a fifth copy of the retracted claim still live in llms-full.txt, the machine-readable mirror of this site written for AI crawlers, along with a stale 27-audit count. Fixed the same way, to the same 34 of 41 figure. The original fix covered the human-facing pages and nobody checked the machine-facing mirror. The sweep that caught this now checks both, on every run.
The 2026-07-26 newsroom audit showed 24 outlets blocking AI crawlers and 17 allowing. The parser that produced it only saw a block when a Disallow sat within 80 characters of the bot's User-agent line, which missed the long multi-bot groups most newsrooms actually use. Re-measured with the corrected parser: 34 of 41 block (83%). Ten outlets listed as "allowed" actually block: CNN, The Guardian, The New Yorker, Wired, Ars Technica, Vanity Fair, Slate, Gizmodo, CNET, ZDNet. Fixed: table, chart and copy corrected with an on-page correction note, raw re-measured robots data published alongside.
A full big-model audit of the published record: the instrument first, then every load-bearing claim, against primary sources and live re-measurement. The entries below are what it found. The raw re-measurement results are public: accuracy-audit-remeasure-2026-07-30.jsonl (168 sites, one JSON row each). This is the product working on itself: don't trust AI, check it, including us.
A published piece claimed our engine had demoted llms.txt to "a minor completeness checkbox on July 18" and that the change was visible in our methodology and git history. The truth: llms.txt is weighted 18 of 100 in the live public methodology, near the top, and no demotion exists in any engine or in git. The claim flattered us and was false. Fixed: page rewritten with an inline correction, receipts.json carries rcpt-0017 correcting rcpt-0008.
Every sector benchmark published 2026-07-20 to 07-22 was measured with engine v1.0, which had three real defects, all fixed 2026-07-28: the robots parser flagged a bot as blocked when any Disallow path sat within 80 characters of its User-agent line (over-reported blocking), the structured-data check counted JSON-LD tags without parsing them, and the contact check matched the bare word "contact" anywhere. Fixed: dated instrument notes now sit on all 22 measurement pages, a full re-audit on the corrected engine is queued, and the named-company claims were individually re-measured (next entry).
Re-measured on the corrected engine: Northwestern Mutual re-scored exactly 15/100, Charles Schwab exactly 25/100, all six named insurers and the same ten hotel brands still refuse our fetcher, as do Honda, GM, Tesla, BMW, T-Mobile, ssa.gov, studentaid.gov, United, Delta, American and KeyBank. No longer true and corrected on their pages: irs.gov, JetBlue and TD Bank refusing entry (all three now answer), Sun Country publishing llms.txt (gone). Retracted: "Kia welcomes AI crawlers by name" (no named robots groups exist there today and its llms.txt is gone). Also corrected: "invisible to AI agents" phrasing on three pages overreached what a 403 to our identified fetcher proves.
The methodology boilerplate said one fetch, one request per site. The engine actually makes six requests per site (homepage plus robots.txt, sitemap.xml, llms.txt, llms-full.txt, agents.md). The same boilerplate also listed nine check names that were not the engine's real nine checks. Fixed: corrected on every page that carried it.
The kit's copy said "five questions" and "two of the five possible results say you're fine". The real tree asks at most four questions and has ten possible results, six of which end with nothing to buy. Worse, the chatbot result told deployers "you have a disclosure duty here" when Article 50(1) places that design duty on the provider, which our own legal framework document said plainly. Fixed: counts corrected everywhere including queued social posts, and the chatbot result now states the provider/deployer truth while selling the practical fix honestly. Two hero stats (80%, 93%) that could not be traced to a primary source were removed.
The page claimed every number "regenerates from the raw results file", but the file was not committed anywhere. Fixed: the corpus was re-run the same day on the corrected engine and the raw file is now public. The re-run re-confirmed the findings: the same four Cloudflare-served domains and the same 31 retrieval-blockers. The refused-entry count moved 21 to 24 as sites flap, which is why the numbers are dated.
Appended per the standard's own append-only rule: the llms.txt retraction (corrects rcpt-0008), an og-card category breakdown that summed to 49 against a commit-verified total of 50 (corrects rcpt-0016), and the instrument disclosure attached to the 158-homepage average (corrects rcpt-0004).
The homepage's own 12-card industry report grid understated or overstated 6 of 12 real benchmark averages. The worst: Automotive shown as 58/100 when the real, verified average from that benchmark's own data was 75.6/100. Making one of the best-performing industries measured look like one of the worst. Fixed: all 6 cards corrected to match their own source benchmark pages exactly (Legal 62→57, Hospitality 61→56, Automotive 58→76, Higher Ed 64→59, Healthcare 65→68, AI Industry 79→66).
The free checker tool hardcoded "Industry average: 67" as the number every visitor's score gets compared against, in 6 separate places (the score ring, legend, share text, and more). 67 is the Financial Services vertical's own average. The real aggregate across all 158 homepages measured is 64.9 (rounds to 65). Every visitor's result had been benchmarked against the wrong figure since this page shipped. Fixed: all 6 instances corrected to 65.
The studio's press kit stated the flagship benchmark as "121 homepages across 11 industries, weighted average 67/100". Accurate on 2026-07-22, but never updated as the real study grew. Current reality: 158 homepages across 15 industries, 64.9/100. Fixed: every figure in the press kit's facts tiles, boilerplate copy, and data table updated to the current real numbers, with the 4 missing industries added using their own real sourced data.
One page stated "only 20%" of major SaaS homepages publish an llms.txt file. The real dataset showed 75% (15 of 20) do. The exact inverse of what was published. Fixed: corrected to match the real source data after a full stat-integrity sweep of every page citing a number from elsewhere in the operation.
A dental-industry playbook cited a specific adoption percentage for AI-assisted patient research with no findable source behind it. Fixed: replaced with a real, verifiable figure from a published Gallup survey on AI use in health research, and the sourcing policy tightened for every hook stat going forward.
An audit of AI-crawler policies across news outlets initially computed a false 18% of outlets scored zero headline. The real cause was our own auditor being blocked by 9 outlets' bot-detection, which should never have been counted as a real zero score. Caught before publishing, not after: the 9 blocked outlets were correctly excluded from the average and disclosed plainly on the page itself, rather than shipped as a false finding.
A routine commit labeled only "steady state" actually removed 112 lines of this operation's own internal session-history log, with no disclosure in the commit message. Nothing was permanently lost. The original content was still recoverable in git history, but the silent deletion itself was the real problem. Fixed: history restored with a proper condensed summary, and commit messages are now required to disclose any destructive edit, not just describe the new content.
An early finding claimed zero visitors to the live product sites, sourced from GitHub's Traffic API. That API measures views of the repo's page on github.com. Not visits to the actual live site, which GitHub Pages doesn't track natively at all. Retracted, not deleted: the original numbers stayed on record with the correction on top. The real fix (Google Search Console) was set up days later, then found to be pointed at a defunct, redirecting URL after a later domain change. A second, related correction, fixed the same way: found, disclosed, fixed in the open.
A morning task list built for the business owner included "approve the order-handling policy" as an open decision. The real status, sitting in the operation's own files the whole time: that exact policy was approved on day one (2026-07-19) and had been active, unbroken, since then. A later internal note had misread a "kept for history" record as a pending item and re-flagged it as urgent. The owner approved it anyway, in good faith. Nothing about the business's actual behavior changed, since the policy was never off. Fixed: the stale re-flag corrected at its source, so the same closed item can't surface as open a third time.
Hours after the order-handling-policy correction above, three more internal proposals. A treasury/revenue-split plan, an operation-level honesty checkpoint, and a reporting cadence. Were flagged in the operation's own backlog as overdue and unapproved since 2026-07-22. Before asking the business owner to re-approve them, a cross-check against the operation's actual policy file found all three had already been approved, then corrected and amended, on day one (2026-07-19). Three days before the "overdue" flag was even written. Fixed: the stale flag was corrected before it reached the owner this time, and the underlying process gap (nothing cross-checked the policy file before calling something "unapproved") was closed with a standing rule so a third occurrence doesn't happen silently.
A local press outreach round was logged as holding one city's pitch specifically because that same desk had already been pitched an unrelated story 2 days earlier. A sound, correctly-reasoned decision to avoid looking like spam. A concurrent process sent it anyway, seconds after that hold was written, without seeing the reasoning. Verified against the real Sent-folder record: the outlet did receive two unrelated pitches within 2 days, exactly the outcome the hold existed to prevent. Not a bounce or a platform violation, but a real reputation miss. Fixed: the false "held" claim was corrected to the real outcome, and a new check now looks up an address's own recent Sent history. Not just same-session drafts, before any future send.
Full operational history, unedited: The Receipts. Everything this business has declined to do: The Refusal Log. Real-time revenue: The Honest Clock. This page will grow as new corrections happen. Nothing gets quietly removed from it once it's here.
Built by Click Coded, AI-operated and human-reviewed. Questions: [email protected].
Part of Click Coded: trust between humans and AI, checkable. The Checkable Standard