A quality-assurance system for AI work, for people who are not software engineers. From "My $4K Cursor Hallucination" at yourscalex.com/blog. Copy any prompt straight into ChatGPT, Claude, Gemini, Cursor, or any AI chat.
You do not need all of this on day one. Each rung catches what the one below it misses. Climb as the stakes go up.
Receipts, not narration. Never accept an AI's description of its own check. Demand the raw output, the actual link, the actual screenshot. Costs nothing, catches the most.
A second model. Paste the work into a fresh chat with a different company's AI and ask it to find what is wrong. Different training, different blind spots.
External AI. If your AI works inside your files (Cursor, Claude Code, a company repo), the reviewer must live outside them: a plain chat with no access to your system, reviewing an exported packet. A sub-agent spawned by the same tool is NOT external. Same house, same kool-aid.
Two blind reviewers. Two external AIs from different companies, identical instructions, neither sees the other's answer. Where they agree, trust it. Where they diverge, that is exactly where you look. Divergence is the signal.
External humans. The only reviewer that cannot hallucinate about your business: a real customer, a real colleague, a real engineer with the production keys. Anything customer-facing gets at least one human check before it ships.
Give this to the AI doing the work. It scopes the job and bans self-grading up front.
You have ONE job in this session: [describe the single task]. Rules: 1. Stay inside this scope. If you notice other problems, list them at the end under "PARKED" and do not act on them. 2. You may not describe your own work as "done," "complete," "verified," or "tested." Instead, state exactly what you did and attach proof. 3. For every claim, attach a receipt: the exact output, the exact text, a screenshot, or a link I can check myself. Any claim without a receipt must be labeled UNVERIFIED. 4. Build on facts, not on earlier summaries. If you are relying on something a previous session claimed, label it INHERITED and flag it for re-verification. 5. When the task is finished or blocked, stop and summarize. Do not start new work.
Run this the moment any AI (or vendor) says "done." It separates what was actually verified from what is being repeated or assumed.
You just told me this is done/verified. Before I accept that: 1. Show me the raw evidence, not a summary: the actual output, the actual result, the actual thing a customer or I would see. 2. For each item, tell me: did you personally verify this in THIS session, or are you repeating it from an earlier note or assumption? Label each one VERIFIED-NOW, REPEATED, or ASSUMED. 3. What did you NOT check that could still be broken? 4. If I clicked/opened/tested this myself right now, what exactly would I see? 5. Which of your claims depend on OTHER claims being true? If the first domino is wrong, which reports fall with it?
Paste into two different fresh chats, ideally different companies (one in ChatGPT, one in Claude or Gemini). The reviewers must have no access to your working tool and must not see each other's answers.
You are an independent reviewer. You cannot fix anything; your only job is to find what is wrong. Below is a report from another AI claiming certain work is complete. Audit it: A. COMPLIANCE: Did the work stay inside its stated scope and rules? B. EVIDENCE: For every claim of "done" or "verified," is there raw proof, or only a confident description? List every claim that has no receipt. C. FOUNDATIONS: Which claims are built on top of earlier claims rather than on evidence? Trace each conclusion back to a receipt or flag the chain as unsupported. D. ACCURACY: Do the claims match my actual instructions and decisions (quoted below)? Flag anything attributed to me that I did not say. E. GAPS: What is missing, unverified, or likely to cause problems later? F. VERDICT: GO or NO-GO, with your top 3 concerns ranked. Never fill an evidence gap with a guess. If the material does not contain something, say "the material does not contain that" and ask. My original instructions/decisions, in my own words: [paste them]
For anything a customer will see or that touches money, an AI verdict is not enough. This is the checklist, not a prompt.
HUMAN CHECK — [date] — [thing being shipped] 1. Who is the one real human (not me, not an AI) who will look at this before customers do? [name] 2. What will they actually do: click the link / run the report / read the email as a customer would? 3. Production keys: who can actually push this live? (If the answer is "the AI" or "me at 11pm," stop.) 4. What real-world feedback loop will tell us within a week if this is wrong? (customer reply, usage data, a complaint) 5. Their verdict, in their words: [paste]
Big decisions get made outside the tool that executes them, and written down in your own words. If a plan later quotes you saying something not in this log, treat it as fiction.
DECISION LOG — [date] Decision made: [in your OWN words, verbatim, typos and all] Made where: [separate chat / voice memo / on paper — NOT inside the working tool] What must be true before anyone (human or AI) acts on it: [the receipt you require] Who said GO: [you, plus both external reviewers if it matters] Rule of use: an AI may restate this decision, but the restatement must link back to this log entry. If a plan quotes you saying something not in this log, treat it as fiction.
Brief (1) → work happens → Receipt Check (2) → send the packet to two external chats with (3) → where they disagree, look there first → human check for anything customer-facing (4) → record your ruling (5) → only then let anything irreversible happen. Two extra chats and ten minutes per job.