Skip to content
Houtini.

AI Fact Checking: How I Check My Content Reliably With Agents

At the end of September I had AI agents fact-check all 64 articles on this site, one agent to an article, and their reports came to 155 fixes. Google's own guidance, updated on 1 October, now calls it critical to fact-check AI-generated content before you publish it. In today's post we're taking a closer look at how the agents were set up, what they found, what the recent research says about AI fact-checking, and how I got the fixes onto live pages without breaking them.

Richard Baxter Richard Baxter AI Ops & Marketing Engineer
Published 17 min read
On this page
  1. Why one wrong fact is enough
  2. How I set up the fact-checking agents
  3. What 64 articles turned up
  4. The cache bug that was my own agents
  5. Is AI a good source for fact-checking?
  6. What recent research says about AI fact-checking
  7. Fixing a live page without breaking it
  8. What I'd tell anyone fact-checking with AI
  9. Where to start on your own site
Diagram of the fact-check pipeline: read-only QA agents, five at a time, produce 64 reports covering 1,617 claims, compiled into one work order of 155 fixes; each fix is settled against a primary source and patched into the live page with its structure checked, and 49 articles were corrected live.

I strongly feel consensus in your content, with evidence-based facts or citations, is key. A single factual error on your site and you're toast, particularly with Google AIO results.
That's why I have a fact check agent.

At the end of September 2026 I had every live article on houtini.com fact-checked by AI agents. There are 64 of them, and each one got its own agent and its own report. On 2 October the reports were compiled into 155 fixes, and by the end of that day 49 articles were corrected on the live site.

Some facts had gone stale, some had always been wrong, and a few were things AI had put on my pages in the first place, such as a quote and a block of terminal output. The work was fixing all of that on live pages without breaking anything else.

Why one wrong fact is enough

For at least six months I've felt that if you want to be in Google's AI Overviews, the data you play back that isn't your own has to be perfectly accurate within the corpus of Google's knowledge. I call this consensus, and it's a massive ranking factor in Google search. (Always add some net new observation, so the document feels worth including in that corpus.)

Google uses the word too. Its guidance on creating helpful, reliable, people-first content says that for topics that could "significantly impact people's lives or well-being", which Google calls YMYL topics, the content "must be highly accurate and consistent with established expert consensus". It doesn't call consensus a ranking factor. Its guidance on generative AI content, updated on 1 October 2026 to say "It is critical to manually factcheck and review all AI-generated content for accuracy and trustworthiness before publishing". And Google has said it is piloting a new way to partner with websites whose content "meaningfully contributes to the freshness and factuality of generative AI responses".

In June 2026, on my sim racing site, a feature claim an LLM had made up became the thesis of an article. It shipped. The tool's developer then publicly called it "AI non-sense".

So I fact-checked that site end to end, all 316 articles, and 66 of them were corrected. The rule I took away from it is that the LLM is a detector, not a source, which means an answer from a model that has searched the web is only a hypothesis until a second method confirms it against the original page.

On that run the agents only staged drafts and couldn't publish anything. A 19-way parallel run lost 8 agents to rate limits and a Gemini spend cap, while batches of five ran clean.

Why Houtini goes stale fast

Most of what I write about on Houtini moves within weeks. Opus 5.5 shipped on 22 September and Sonnet 5.5 on 28 September, so model names that were current in August were stale by October. Prices and version numbers move, and so do facts you'd think were settled: Saturn's moon count went from 146 to 285 when the IAU Minor Planet Center updated it on 26 March 2026, and NASA had it at 293 by August.

How I set up the fact-checking agents

I kept the setup deliberately narrow: one agent per article, and QA only, so no edits and no publishing. The 64 articles went out in 13 batches of five, oldest published first.

Each agent pulled the live page and built an inventory of every checkable claim on it: versions, release dates, model IDs, prices, environment variable names, install commands, hardware numbers, "X supports Y" claims and URLs. Every sentence an agent flagged was quoted character for character, so a finding could be matched back to the page later.

Anything about my own tools was checked against my local repos first and never asked of an LLM. Everything else went to Gemini 3.1 Pro (gemini-3.1-pro-preview) with grounding switched on, at temperature 0.2, three to five claims per call, asking for the primary source URL for each. Grounding is Gemini searching the web while it answers. A primary source is the vendor's own docs, pricing page, repo or the person's own post - not a summary of it. Steps 4 and 5 of the agent brief are below.

# Fact-check sweep agent brief - houtini.com, September 2026 (QA ONLY: no edits, no publishing)
...
4. Everything else: `mcp__gemini__gemini_chat` with EXACTLY `model: "gemini-3.1-pro-preview"`, `grounding: true`, `temperature: 0.2`, never `max_tokens`; narrow asks of 3-5 claims per call; ask for the primary source URL for each. Gemini is a DETECTOR, not a source.
5. Double-check every candidate correction with `mcp__brave-search__brave_web_search` (and `mcp__firecrawl-mcp__firecrawl_scrape` on the primary page where needed). A correction needs a named, linkable PRIMARY source (vendor docs, repo, pricing page, official post) confirmed by a second method. `vertexaisearch.cloud.google.com` redirects are not sources.

Each finding got a severity: HIGH for wrong, MED for stale, LOW for cosmetic. Each report had four sections: to correct; uncertain (flag it, don't fix it); structural; and checked and correct.

If an agent couldn't settle a claim, it went under uncertain and was left alone. A claim that was right when I wrote it but had since moved counted as stale, a MED, with the date it moved. No agent was allowed to invent a correction to fill a gap, either. One blind spot was that the agents read the rendered page, and only four of the 64 reports checked the raw HTML for the underscore problem in the next section.

What 64 articles turned up

By 2 October there were 64 reports. When I added up their header lines, the agents had checked 1,617 claims, with 153 rows marked to correct, 191 uncertain and 33 structural.

I had a cheaper model, Sonnet, compile them into one work order. Its brief said "You compile; you do not re-verify." The work order came to 155 fixes across 49 articles: 64 HIGH, 72 MED and 19 LOW, with 33 items that needed me and 33 that were structural. It was sorted by the number of HIGH fixes first, then by harm to the reader, then by age. On harm, a command that breaks when you paste it outranks a wrong price or date, which outranks a stale model name. Top of the list was a day-one model benchmark article with eight HIGH fixes.

# Houtini fact-check fixes - work order (compiled 2026-10-02)
Articles: 64 · with fixes: 49 · fixes: 155 (HIGH 64 / MED 72 / LOW 19) · needs Richard: 33 · structural (route to update-article.md): 33

The work order grouped the errors that repeated across articles into patterns. Model names had gone stale after September's releases, and the version numbers of my own packages had moved on. Some of what the pages said about how Claude Code works didn't match Anthropic's docs, a few vLLM version pins were out of date, and one page still gave Saturn's 146 moons as the right answer. And underscores in identifiers had been eaten as italics (the site's markdown reads text between two underscores as italic), which is a markup problem rather than a factual one, but it still put the wrong text in front of a reader.

The row below is a typical Claude Code docs correction. The page said CLAUDE.local.md is gitignored automatically, and the docs say you add it to .gitignore yourself.

# claude-code-project-setup-anatomy - fact-check (QA only) - 2026-10-02
Claims checked: 34 · To correct: 3 · Uncertain: 1 · Structural: 0

| # | Section | Current text (exact quote) | Correct text | Primary source (named + URL) | 2nd method | Severity |
|---|---|---|---|---|---|---|
| 2 | CLAUDE.local.md - Your Personal Overrides | but it's automatically gitignored so your personal tweaks never land in the shared repo. | but you need to add it to .gitignore yourself so your personal tweaks never land in the shared repo. | Claude Code docs, memory page https://code.claude.com/docs/en/memory ("Add `CLAUDE.local.md` to your `.gitignore` so it isn't committed"; auto only for settings.local.json per https://code.claude.com/docs/en/settings) | Docs .md source plus Brave hit | HIGH |

Four of the finds were quotes, figures and output on my own pages that the original source never said or printed. The Desktop Commander article had Anthropic saying something they never said. It had started as a paraphrase in my research notes and turned into a quote somewhere between the notes and the page. Anthropic's real words, reported by Infosecurity Magazine on 9 February 2026, were that the security flaw the article covers "falls outside our current threat model". The DGX Spark article misquoted Timothy Carambat of AnythingLLM from one of his videos. Another article cited Salesforce's State of Marketing for 87%, when Salesforce's own page headlines 75% and the 87% traced back to aggregator sites.

The fourth was on the Claude Code system requirements article: a block of terminal output presented as what Claude Code prints after you sign in. It doesn't print that, and the credentials path in it was missing a dot. An update on 2 September had fixed the prose around it and left the block standing as if it were a capture. My call was "drop the code block and publish", and the script's record of the cut is below.

- pair 2.2-cut DELETE content[76] (code): {"_type": "code", "_key": "73b66fb74695", "language": "text", "code": "Authentication successful.\nLogged in as: you@yourcompany.com (Pro plan)\nToken stored at: ~/.claude/credentials.json"}

The lead-in where the block used to be now tells you to run claude auth status --text and read what your own install says.

The cache bug that was my own agents

On the morning of 2 October, several agents reported fetching one article's URL and getting a different article back. One report showed a CF-Cache-Status: HIT header with the canonical (the address a page declares as its own) pointing at the DGX Spark article, and another said its first fetch "returned a body belonging to another article". The work order logged it as Cloudflare edge cache cross-wiring, a site-level bug.

That looked serious, so I went after the site. I tried to reproduce the fault with 48 concurrent uncached renders, and every one came back clean. At 08:40 I put a guard into the site's Worker anyway: never store a page whose canonical URL isn't the URL requested.

The cause was my own agents, which were all writing the pages they fetched to the same scratch filename, and overwriting each other's. One agent would fetch its page, another would fetch a different page into the same file, and the first would read the second's article as its own. Their own reports finally named the shared file. Diagnosing the site first cost about an hour and a code change I didn't need, and I reverted the guard at 18:23.

97a3abd Edge HTML cache: never store a page whose canonical URL isn't the URL requested
c5cf546 Revert "Edge HTML cache: never store a page whose canonical URL isn't the URL requested"

The fix went into the agents' brief. Every agent now pulls the page with a parameter (_fc) that bypasses the page cache, saves what it fetches under a name that includes its article's slug, and checks the page's canonical before using anything. A later pass over 24 pages found every canonical matched its own URL.

The shared file explains the cross-wiring, but I can't fully explain that one HIT header, and that's where I've left it.

Is AI a good source for fact-checking?

So, is AI a good source for fact-checking? From this run, no. The agents found 153 things to correct in 1,617 claims, but when one of them asked a model where a fact came from, the answer was only ever somewhere to start looking.

The clearest example was on the Desktop Commander article, which quoted a "9.52/10 user rating". The agent asked Gemini where the figure came from, and Gemini's grounded search pointed at the Houtini article itself. The report's note is below.

- "9.52/10 user rating" (DC vs Cursor vs Claude Code section) - no independent primary source found. Gemini's own grounded search, when asked where this figure comes from, circularly cited the Houtini article itself as the origin of the number - meaning it cannot currently be traced to an MCP directory (Smithery, Glama, MCP.directory) or any other named source. Leave as-is or drop the figure; do not treat as verified.

The figure was on Desktop Commander's own homepage, as "9.52/10 User Rating", so it stayed on my page, attributed to them.

The 33 items the work order had marked for me were settled the same way, from my own files, repos, logs and primary sources: 23 resolved, 10 by a standing rule, and none left needing me. During the day I'd asked, "are you sure you can't infer my reasoning, interest and motivation from the original?" Settling them turned up four decisions that were mine to make, and each came to me with a recommendation and was answered in a line. The DGX Spark article got a price refresh, a benchmark dataset on the site was republished because the text was right but the download was stale, an article about a retired product came down behind a 301 redirect, and one article was reframed.

The DGX Spark price moved on the day itself. That morning's notes said the ASUS GX10 at $5,999 was now dearer than NVIDIA's DGX Spark, so the "cheapest NVIDIA route" framing was dead. The same day, NVIDIA raised the 128GB Spark to $6,950. It was $3,999 at launch and $4,699 earlier in 2026, so that's "an increase of nearly 75 percent", as The Register put it. The GX10 undercut it again, by $951, and the refresh was written from that day's primary pages.

What recent research says about AI fact-checking

Once the sweep was done, I wanted to know whether what it had shown me was particular to my setup. So I had the recent research on AI fact-checking pulled together from the primary pages. There are seven pieces in the table below, from February 2025 to August 2026. They range from peer-reviewed papers and a preprint (a paper that hasn't been peer-reviewed yet) to Full Fact, the UK fact-checker, testing chatbots on claims it had already checked.

SourcePublishedWhat it found
Tai et al., ACM Web Science Conference 2025February 2025 (arXiv)LLM credibility ratings showed only moderate agreement with human coders, and even correct ratings leaned heavily on linguistic features
Tow Center for Digital Journalism, in Columbia Journalism ReviewMarch 2025Eight AI search tools gave incorrect answers to more than 60% of 1,600 queries asking them to identify the news article an excerpt came from, from 37% wrong (Perplexity) to 94% (Grok 3)
Poynter, on Full Fact's AI toolsNovember 2025Full Fact's tools read more than 300,000 sentences a day for over 40 fact-checking organisations; fact-checkers decide what to investigate and publish
Qazi et al., Scientific ReportsFebruary 2026On 5,000 claims already checked by 174 fact-checking organisations in 47 languages, smaller models were confident but less accurate, larger ones more accurate but less sure
Reuters Institute, AI and the Future of NewsMarch 2026Maldita and Full Fact use large language models to detect and classify claims across millions of sentences
Li and Bakker, arXiv preprintApril 2026Raters scored 1,614 LLM-written Community Notes (the crowd-written context notes X attaches to posts) more helpful than human notes, but the model sometimes missed whether a correction was needed at all
Full FactAugust 2026Five months of testing chatbots found 39 major errors and 67 incorrect responses; Google and OpenAI said the API-based setup didn't represent their consumer apps

Where the research agrees with what I saw

The first thing the research agrees on is that language models are good at finding claims. That's how the professional fact-checkers use them, according to Poynter and the Reuters Institute: Full Fact's and Maldita's systems find and sort the claims at scale, and in Poynter's account the fact-checkers decide what gets investigated and published. My agents did the same job on my own pages, pulling 1,617 checkable claims out of 64 articles and sorting each one into a section of its report, with every flagged sentence quoted exactly.

The second is that a model makes a poor source. The Tow Center for Digital Journalism asked eight AI search tools to name the original article behind a pasted excerpt, and more than half of the responses from Gemini and Grok 3 cited fabricated or broken URLs. In Full Fact's own trial, some chatbots corrected themselves days later, often by citing Full Fact's own fact-checks. That's the same loop I hit when Gemini, asked where a rating came from, pointed back at my article.

They also agree that a person or a second check has to sit between the model and publication. Full Fact says chatbots "cannot replace the role of fact checkers or human judgement in general". Qazi and colleagues recommend confidence thresholds or human-in-the-loop mechanisms, and Tai and colleagues caution against relying solely on the models. In my sweep no agent was allowed to publish, and every fix that went live was settled on a primary source. Tai's paper found that even when the models rated content correctly, their reasoning leaned heavily on linguistic features.

Where the research still disagrees

The first disagreement is how bad the sourcing problem is right now. On the Full Fact page, Google says the study "accessed out-of-date Gemini models through a developer channel" and that a retest on the Gemini website found the claims "consistently debunked". OpenAI says testing through the API didn't reflect the consumer ChatGPT experience. The models in that trial included Gemini 3.1 Pro, which is the model my own QA agents ran on. Whether a bigger model fixes it is disputed too. Qazi's team found larger models more accurate, if less sure of themselves, while Tow found the premium tools had higher error rates than the free ones.

The research also splits on whether AI fact-checks help readers. In Li and Bakker's preprint, raters scored LLM-written Community Notes as more helpful than human notes, across different political viewpoints. Full Fact and Tow measured whether the answers were right, and found plenty that weren't, while Li and Bakker measured how helpful raters found a note. So it isn't a straight contradiction, but a note that raters find helpful can still be wrong. I never measured what readers made of my corrections.

Fixing a live page without breaking it

With the work order sorted, the obvious route for the edits was to flatten each live article to markdown, edit the text, rebuild it and upload it. I tried that on a pilot first, the Firecrawl guide, which runs to 119 blocks. The pilot stopped before a single fix went in. The article had two break blocks, which render as dividers, and after the round trip both came back as comments reading "unknown block type: break". An image caption disappeared too.

The words were all there, so any check that compares text would have passed the page with a caption and two dividers missing, which is why I edited the pages the way they're stored instead. The site's CMS, EmDash, stores each article as Portable Text. Portable Text is a list of blocks, such as paragraphs, images, code and breaks, and each block holds its text.

Two routes for applying a fact-check fix: flattening the live page to markdown and rebuilding it lost two dividers and an image caption on the pilot article, while swapping the exact text in the stored page and checking its structure changed nothing else.

I ran the edits through four small scripts. The pull script takes a fresh copy of the live post. The apply script makes exact string swaps. Each old string has to occur exactly once across the page's text. Unless a fix declares a merge or a cut I'd approved, the script refuses any change to the block count, the block types and keys, the number of spans in a block or the hero image. Spans are the runs of text inside a block, each carrying its own formatting. It also refuses an em dash or an arrow in new text, and if anything fails it writes nothing. The publish script checks the live post hasn't changed since the pull, publishes, and reads back the hero, the SEO fields and the block count. Then it fetches the public page, checks the canonical and confirms every new string renders. The fourth script dumps a block span by span, for markup fixes.

Each fix went in as a pair (one exact old-text / new-text swap). The old text was quoted from the stored page, never from the QA report, because the reports quote rendered text. Underneath the text, bold, links and inline code sit in separate spans, and an old string can't cross from one span into the next.

The underscore problem was the one kind of fix that had to cross spans. On the Docker MCP gateway article, MCP_DOCKER had rendered as MCPDOCKER, with the text between the underscores in italics. The fix was a merge pair that rejoins the three spans into one.

{
 "id": "4",
 "merge": {
  "block": 27,
  "from": 9,
  "to": 11
 },
 "text": " and look for MCP_DOCKER marked Connected. For Claude Desktop, restart it and look for MCP_DOCKER in the Search and tools menu.",
 "marks": []
}

The apply report for that pair is below.

- pair 4 MERGE content[27].children[9..11]
  - before: ' and look for MCPDOCKER marked Connected. For Claude Desktop, restart it and look for MCPDOCKER in the Search and tools menu.'
  - after:  ' and look for MCP_DOCKER marked Connected. For Claude Desktop, restart it and look for MCP_DOCKER in the Search and tools menu.' marks=[]
...
blocks: 52 (structure unchanged: True; merged blocks: [6, 27, 51])

Block 51 wasn't in the work order at all, and the apply step only found it by reading the spans. When I re-ran the Firecrawl pilot this way, its 5 pairs went in on 119 blocks with the structure unchanged, and both dividers and the caption were still there on the live page.

A fix that needed a whole sentence rewritten, or that would have cut my own first-hand experience, couldn't go in as a swap, so it went to me, or to a separate writing model that worked only from the facts and a brief, with a fresh reader checking the result. There were three publish rounds that day. The first put 46 articles live, the second 14 and a rewrite round 8 more entries, and because some articles went out in more than one round, that came to 49 different articles corrected live.

What I'd tell anyone fact-checking with AI

If you're about to point agents at your own site, these are the four things I'd set up before the first run.

Give every agent its own scratch files

Parallel agents that share a scratch path will look like a bug in the system they're testing. Mine looked like a Cloudflare cache fault. Name every scratch file by the article's slug, something like fc-<slug>.html, never a generic /tmp/page.html. Have the agent check the page's canonical URL before it uses anything it fetched. And if the results look like the site is serving the wrong pages, suspect the shared path before you suspect the site.

Let the model find errors, never settle them

Let the model inventory the claims and flag the doubtful ones, then settle each one against a primary page and confirm it by a second method. Don't accept the redirect links a grounded model returns as citations. If two checks disagree, the claim is unverified and it stays flagged.

Saturn's moon count caught me out on this. The 2 October sweep changed it to 285 in my RAG article, the IAU figure from 26 March 2026, and checked that against NASA's search snippet, which still said 274. NASA's page body already said 293 "as of August 2026", and so did JPL's table of Saturn's satellites. So the fix went live already out of date, and the run on my hallucination article caught it that same evening. When an agent checks a figure, have it read the primary page itself, not the snippet. That article was corrected to 293 on 3 October.

Edit the page the way it's stored

A round trip through a convenient format loses whatever that format can't express. On my pilot it was two dividers and an image caption. Edit the stored format directly. Make each change an exact swap of a string that occurs once, quote the old text from the stored page rather than from a report, and have the script assert the structure is unchanged and refuse to write anything if it isn't.

Keep the publish button away from the agents

My QA agents couldn't edit anything, and the agents applying fixes stopped before publishing. One script with checks did the publishing, after a review of what it was about to change. In June the agents physically couldn't publish: no publish call, no Cloudflare token. I don't take an agent's report on its own work at face value, which is why the gate sits outside the agents.

Where to start on your own site

Start with one article. Run one agent over it, QA only, with the four report sections: to correct, uncertain, structural, and checked and correct. Put the primary-source rule in its brief, the way steps 4 and 5 are written further up. Check every flag against the vendor's own page before you change a word. Then apply the fix as an exact swap with a structure check, and look at the live page afterwards.

If you want the background on why models make things up in the first place, how to prevent AI hallucinations is the next read.

Continue reading.