Behind the scenes: two AI-assisted news sites, and the fact-checker that failed 28 times before it ran once

Christopher Ross

5 min read

AI and learning, kept human · Niagara, Ontario

Title card for the article “Behind the scenes: my accessibility checker was wrong three times before it was right” on This Is My URL

Both new sites have a box on the home page called Editor’s picks. The rule behind it in the template contract says it may never be called Most read, because neither site has analytics and neither site had anyone to count on launch day. So the editor is me, and I picked them.

Mining Dispatch and Hauler Daily went live on 19 September: two AI-assisted news sites on one shared engine, and a rulebook I wrote before I built the sites. On launch day the rulebook said no 141 times in about two hours. There is a log of every verdict. It has 641 rows.

The metre read nothing, and the wood got the blame.

The copy scanner that said no 141 times

The first gate is a script that fails a story on any run of ten words copied from a cited source outside quotation marks. Facts are free; expression is not. In those two hours it recorded 289 passes and 141 fails, and the drafts had not lifted the exciting parts of anyone’s press release. They lifted the paperwork, the part a human would skim, scraps like this one, caught whole: “between the third quarter of 2027 and the first quarter.”

One story about a road approval failed 15 times. It worked down to one, passed, was edited, failed again, passed again, and as of 19 September it was still sitting in review.

The independent fact-checker that could not find its own shell

The independent gate is a model from a different family than the drafter, running on Cloudflare Workers AI, that grades every claim against the fetched source text. The first time it ran, every story failed. All 28 verdicts said the same thing, and none of them were about the news. On Windows a bare call to bash goes to the Windows Subsystem for Linux stub, so what the independent check actually reported, 28 times over, was No such file or directory.

Here is the part that stings. I had written “silence is never a pass” into the standards. I had not written “a broken shell is not a verdict,” so the script stamped each of those 28 as a content failure. It is like sending a board back to the mill because the moisture metre had a dead battery: the metre read nothing, and the wood got the blame. Safe direction, still wrong. The fix was a full path to the right shell, plus an error result that is never stamped. After that, 74 passes and one real catch: a cross-border tariff brief with a dollar total the source never added up. Still held, as of the 19th.

GatePassedFailedWhat the fails actually were
Copy scanner289141Real. Runs of ten words lifted from the paperwork in a source.
Fact-checker, first run028Not one about the news. The shell was missing.
Fact-checker, after the fix741Real. A tariff total the source never added up.
Two of the launch-day gates, from the verdict log. The middle row is the one this post is about.

The rule I am most pleased with came out of a stalemate between two of my automated reviewers. The standards gate, one model, wanted every editor’s note to say why the story matters. The fact-check gate, another model, rejected any sentence it could not trace to a source. Both were right, and for three review passes nothing satisfied both. The fix was a label. The note goes out to readers as “Our read,” and it is judgment: it can reason from what is in the body, and it cannot add a fact. Any digit in the note that is not already in the story fails the build. The one place an editor gets an opinion is fenced by numbers, which I find funnier than I probably should.

Launch day, with a privacy page that was ahead of the site

The privacy page already said both sites run web analytics and have working editor mailboxes. Neither was true at launch; the fix is on my checklist rather than on the page, which is not the order I would recommend to a client. The newsletter has no provider chosen yet, and the signup page says so. Earlier in the week a Cloudflare redirect rule on this site never fired: on my plan, rules keyed to a path do not run and rules keyed to a hostname do, so both new sites got a hostname-based www redirect.

The illustrations are AI-generated by category rather than of any real event, and each carries a label saying so. The image model has no way to take a negative instruction, so asking for “no text” adds text, and I rejected one picture for a trucking story because it was the inside of a car.

What I took from the week

As of 19 September, thirteen stories were live and ten more were held. I cannot tell you a single person has read them, because I built the sites so that I cannot know that yet, and a site one day old does not get called a success by me. The 90-day test is pages indexed, search impressions and digest signups, with the pass mark written down before the numbers arrive.

A while back a nightly bot that kept un-fixing its own fix taught me to verify the output before keeping it. This week added a line underneath: verify that the verifier ran. The rules for when a story may be believed were the job, the same shape as the AI operations work I do for others, and if you are building review gates and they go sideways, ask. I’ll help. Last week’s notes, on the accessibility checker that was wrong three times, are in the previous Behind the Scenes. If you find a mistake in one of the thirteen, each site has its own corrections page and its own contact address.

Working through something on your own site? Get in touch →

Leave a reply

Your email address will not be published. Required fields are marked *

Your rating (optional)

Your name and email are stored with your comment; only your display name is shown publicly. See our privacy policy.