The Complete Guide to Checking AI's Work Before Your Company Ships It
As of June 2026, tracking projects counted 1,598 court cases worldwide involving AI-fabricated legal citations. A year earlier that figure sat around 200. Roughly 5 or 6 new cases now get documented every day.
The largest sanction in that set came to about $109,700. A federal judge in Mississippi suspended two lawyers from practicing in his district for two years after both sides filed fake citations.
Lawyers spend years training, earn a license, and carry personal liability, and they still filed invented facts because the output looked right.
Now put that failure inside your company. At $5M to $100M in revenue, AI output stops being something you personally paste into an email. Your marketing team drafts campaigns with it. Your support team answers customers with it. Your ops people summarize supplier contracts with it. One invented fact travels through three sets of hands, picks up your logo on the way, and lands in front of thousands of customers.
A lawyer files one bad brief and one judge catches it. Your team ships one bad claim and it prints ten thousand times before anyone reads it closely.
I run AI and growth at NuVision Auto Glass, a $48M US auto glass company, and my growth company builds AI workflows for US mid-market brands. The founders I talk to split into two camps. One camp trusts the output and ships it. The other camp makes everyone verify everything, which kills the time saving and quietly teaches the team to stop using the tools. Both camps lose, and the fix is one owner-level decision you make once.
Why a model invents things
You can't set policy for a failure you don't understand, so start with the mechanism. It takes two minutes and it removes the mystery.
A language model predicts the next likely word based on patterns it learned from enormous amounts of text. Then it does that again, and again, until it has built an answer. That's the entire process. It has no separate step where it checks a fact, and no internal signal that says "I'm unsure here" unless someone connected it to a tool that can look things up.
So it produces the most plausible-looking answer, always. A plausible warranty term. A plausible case name in the correct citation format. A plausible statistic with a plausible source attached.
Fluent and correct are two different skills, and nothing trained the model to tell your team which one they're getting.
Paid, purpose-built tools carry the same flaw. A peer-reviewed Stanford study found dedicated legal research AI tools hallucinated 17 to 34% of the time, and general chatbots produced error rates of 69 to 88% on legal questions. Hallucination means the model states something false as confidently as something true. Those numbers came from tools built specifically for accuracy in that field.
The mechanism gives you the policy for free. Errors concentrate where the model supplies facts you didn't give it. That single sentence drives everything below.
Sort every task by what a mistake costs
The standard advice says check all AI output. Nobody's team does that, because it turns a 5-minute task into a 20-minute one and the tools stop paying for themselves.
The workable system sorts tasks into 4 tiers by the cost of a mistake, then attaches one check to each tier. Your job as the owner runs one level up from the checking itself. You define the tiers, you name who checks each one, and you decide what runs with no check at all.
Tier 1: runs unchecked, on purpose.
Work where a mistake costs nothing and the person using it spots problems instantly. First drafts someone will rewrite. Brainstorm lists. Rough outlines. Internal summaries of documents the reader opens anyway.
Example: your marketing lead generates 20 campaign angles before a planning meeting. Half are mediocre. Nobody cares, because the meeting exists to kill the weak ones. Adding a verification step here burns payroll on nothing.
Most of your company's AI use should live in this tier, and saying that out loud matters. When you tell the team "this category ships unchecked," they stop treating every AI task as a compliance exercise and the speed benefit survives.
Tier 2: a named person reads it end to end.
Work that reaches a real customer but carries no money, no legal weight, and no binding promise. Support replies. Social posts. Job descriptions. Website copy.
The check: one named person reads the full text before it leaves, hunting for two things. Facts nobody supplied, and promises nobody authorized. Models add both constantly, because businesses in their training data usually offer that discount, that response time, that guarantee.
The rule that catches most of it: anything in the output that nobody put in the input is unverified by definition. The checker flags new specifics.
A real example. A landscaping company owner asked for a reply to a customer about turf renovation timing. The draft came back warm and clear, and it included the line "we offer a 12-month guarantee on all new turf." He offers 6 months. He never mentioned a guarantee in the prompt at all. He caught it in a 40-second read-through because he sends his own email.
At your scale that safety net is gone. The same invented sentence lands in a support macro, and every rep pastes it into every reply until a customer shows up holding the promise. That's why the checker has a name. "Someone reads customer-facing output" fails. "Priya reads customer-facing output before it ships" works, because a named person owns the miss.
Tier 3: every number and every proper noun gets verified at the source.
Work with money, dates, quantities, people, or companies in it. Quotes and proposals. Reports going to your board or your bank. Anything with a statistic. Any summary of a contract. Anything naming a real person or business.
The check runs on two rules:
1. The source rule. If the output cites something, the checker opens the source. A citation proves nothing about whether the source exists or says that. Fabricated citations look identical to real ones, which is exactly how 1,598 court filings went wrong.
2. The numbers rule. The checker confirms every figure at its origin, never against another AI answer. Models fail most on percentages, dates, and multi-step arithmetic, and they present all three with total confidence.
Contract summaries need one more habit, because the worst failure leaves no trace. Ask for a summary of a 20-page supplier agreement and you get eight clean bullet points, every one accurate. The auto-renewal clause with a 90-day cancellation window sits on page 14 and appears in none of them, because the model judged it less central than the pricing terms. Nothing in the output tells you a clause is missing. For anything that binds the company, someone reads the source document and uses the summary as a map.
Tier 4: a qualified human owns it, AI only prepares.
Work where being wrong causes real harm. Legal filings. Tax and regulatory submissions. Safety claims. Financial statements. Insurance matters. Warranty terms written into customer contracts.
The check: a qualified human does the work and uses AI to prepare material only. Draft, organize, summarize, suggest questions. The human makes every judgment and signs every fact.
Companies your size hit this tier more often than the org chart admits. An HR manager asks a chatbot about notice periods across the four states you employ people in. A product manager lets AI write the safety language on a spec sheet. An ops lead drafts the incident statement for your insurer. None of it looks dramatic while someone types, and all of it binds the company afterwards. The lawyers in those 1,598 cases ran a Tier 4 task with a Tier 1 check, and that mismatch is the entire failure.
The delegation layer only you can build
The tiers describe what to check. At your size the harder question is who, and that part never delegates upward from the team. Someone with authority over every department has to set it, and that's you.
Write one page. It needs 4 things:
- The tier list, with real examples pulled from your own business, so nobody debates which tier a task sits in.
2. A named checker for each recurring output type in each team. Support macros, proposals, board reports, published content. Names, never roles.
3. One line added to your existing review step: who checked this, and against what. That question moves accountability from "the AI" back onto a person, and it kills most of the risk by itself.
- The standing rule that nobody sends a number they didn't personally confirm at its origin.
Then train the checkers on 3 habits that make their job fast:
1. Ask the model what it's unsure about. After any factual answer, the checker asks "which parts are you least confident in, and what should I verify independently?" It flags its weakest claims surprisingly often. Treat that list as a starting point.
2. Ask the same question twice, in a fresh chat, worded differently. Where the answers disagree, the checker found the soft ground. Where they agree, confidence rises slightly and proof still stands at zero.
3. Feed the model the source instead of asking it to recall one. A model summarizing a pasted document beats a model recalling that document from memory by a wide margin, and this is the single biggest accuracy upgrade a team can adopt. Stop asking what the regulation says. Paste the regulation and ask what it means.
What the system costs, and what skipping it costs
Tier 1 and 2 work stays as fast as it was. Tier 3 adds minutes per task, mostly spent opening sources. Tier 4 saves drafting time and spends no new checking time. Sorting a task into its tier takes about 4 seconds once the one-pager exists.
That's the full overhead: one page, some names, 4 seconds per task.
Against that, weigh what the unchecked path costs at your scale. A wrong guarantee replicated through a support team. A fabricated figure in a board deck. An auto-renewal nobody surfaced until the cancellation window closed. The professionals in those court cases had licenses, training, and personal liability, and volume still beat them.
Your company runs more AI output through more hands than any law firm in that dataset. Build the one-pager this week, name the checkers, and let everything in Tier 1 run free.
If you want AI workflows built for your team with the checking step designed in rather than bolted on afterwards, that's the way we build them at NuroSparx. Tell us what your team is running and we'll show you where the risk sits. Examples of what that looks like in practice are in our case studies. The one-pager takes an afternoon, so write it this week, or book a call and we'll build the tiers with you. Once the checks hold, point that confidence at the automations they protect.
