Agent Dispute Kit: How to Challenge an AI Work Claim
When an AI agent says a consequential task is complete, ask for evidence that is independent of the summary. The dispute loop is simple: identify the claim, inspect the execution receipt, and check the system that owns the result.
Try the Zambo demo · Explore Zambo
Use the public receipt verifier to inspect a receipt when you have one. A receipt is evidence about a recorded execution, not a guarantee that an outside system accepted the result or that a human decision was correct.
The three-step loop
Step one is claim definition. Write down exactly what the agent says happened, including the action, target, time, and expected result. Step two is execution evidence. Request the receipt or execution record for the action, inspect its tool, inputs or commitments, observed output, timestamp, and integrity status. Step three is outcome confirmation. Open the file, query the record, check the sent folder, or ask the system that owns the state. Stop at the boundary of what can be confirmed. A receipt is stronger than a summary, but it is not a universal outcome guarantee.
Template: request the evidence
Use a neutral request: Please provide the receipt identifier for each material action, the recorded tool and scope, the timestamp, the observed result, and the public verification path. Please separate actions that ran from actions that were planned, retried, or inferred. Please identify any step for which no receipt exists. This request does not accuse the agent of lying. It creates a common evidence vocabulary and gives an operator a chance to correct a missing link before the disagreement grows.
Template: record the finding
Use three columns in the dispute record: claim, evidence, and conclusion. In claim, quote the narrow statement being checked. In evidence, link the receipt and the independent system check. In conclusion, choose supported, partially supported, unsupported, or unresolved. Add a reason and the reviewer timestamp. Do not use verified as shorthand for correct. Verified can describe the integrity or procedure that was actually run, while correctness may still depend on an outside system or human review.
What to ask for in a receipt
Ask for a stable identifier, created-at time, tool name and version, caller scope, provenance class, canonical bytes or an output commitment, and verification status. The exact fields depend on the receipt format. Do not request private prompts, credentials, or customer data when a redacted projection is enough. If a receipt links to a public page, resolve it and compare the identifier with the action under dispute. If the record is unavailable, mark the evidence unresolved and preserve the retrieval attempt.
The critical path test
You do not need to redo every low-risk action. Find the one or two claims on which the task's value depends. A finance task may hinge on the final ledger write. A deployment task may hinge on the target version and health check. A research task may hinge on the source retrieval and quoted evidence. Require receipts for the remaining steps and independently check the critical path. This approach gives a reviewer useful confidence without pretending that a whole workflow has been recreated.
Common weak arguments
The agent's confident summary is not independent evidence. A screenshot without a stable source can be incomplete. A timestamp alone does not prove execution. A payment record does not prove a useful tool result. A valid hash does not prove the underlying output was correct. A log entry can be edited or incomplete unless its integrity and retention are understood. Name the weakness plainly, then ask what evidence would close it. The purpose of a dispute kit is to improve the decision, not to win an argument through jargon.
When evidence conflicts
Preserve both records. Compare identifiers, times, tool versions, scope, and exact bytes. Check whether one record describes a request while the other describes an observed result. Ask the system owner which state is authoritative. If the conflict remains, record it as unresolved and escalate according to the organization's process. Do not merge records by hand or rewrite an earlier receipt. A clean dispute record is itself useful evidence of how the disagreement was handled.
Privacy and fair handling
A dispute can expose more information than the original task. Share only what the reviewer needs, redact secrets, and restrict personal or customer data. Do not publish an agent's private context merely because a claim is disputed. Keep the receipt reference and the conclusion separate from payloads that require access controls. The party making a consequential claim should be able to explain the evidence, while the reviewer should avoid turning verification into unnecessary disclosure.
A concise policy clause
An organization can state: consequential agent actions must have an inspectable execution record; summaries do not constitute proof; reviewers may request the receipt, recorded result, and relevant system-of-record confirmation; and an execution receipt proves only the recorded execution and observations it covers. This clause is practical because it sets an evidence expectation without promising that automation is infallible. Link the public verifier at /verify where the workflow permits.
Close the dispute carefully
A dispute closes when the narrow claim is supported, corrected, or explicitly left unresolved. Record who reviewed it, which receipt was checked, which outside system was consulted, and what remains unknown. If the agent made an unsupported claim, correct the task summary and improve the evidence path. If the receipt was valid but the outside outcome failed, fix the workflow or system owner rather than blaming the receipt. Accountability improves when the record says exactly what happened and exactly what it cannot establish.
Review notes
Good evidence work starts with a bounded question. Name the action, the system that owns the result, the time window, and the person or service responsible for reviewing it. Then separate three states that are often collapsed into one: an action was requested, an execution layer recorded an action, and an outside system confirmed an outcome. The first state is intent. The second is execution evidence. The third is outcome evidence. A page that keeps these states separate is easier to use during an incident and harder to misuse in a status summary.
When evidence is incomplete, record the gap instead of filling it with confidence. A missing record may be a collection failure, a transport problem, a retention issue, or an action that never ran. Those possibilities have different owners and different remedies. Preserve the original record, the verification attempt, and the reviewer conclusion. If a later retry succeeds, keep the earlier attempt visible. This makes the evidence useful to someone who was not present when the work happened and gives the next operator a concrete place to start.
Reviewers should also write down assumptions. State which timestamp is authoritative, whether a result was observed directly, how a hash was computed, and which system owns the final state. If the answer depends on an external service, link the external evidence or mark it unavailable. If the action involved private data, describe the field class without copying the data into a public page. These habits turn a receipt from a decorative status badge into a durable review artifact. They also make disagreements productive: two reviewers can compare assumptions instead of arguing over a summary.
Finally, keep the evidence proportional to the decision. A low-risk read can use a light check, while a money movement, data mutation, deployment, or external communication deserves a stronger receipt and an independent state check. Do not use a larger word count or a more elaborate format to imply certainty. The relevant question is always what was recorded, what was verified, and what remains outside the record.
Frequently asked questions
What is the first step in an agent dispute?
Write the exact consequential claim being checked, including the action, target, time, and expected result.
Is a receipt proof that the result is correct?
No. It supports the recorded execution and observation. Correctness may require an independent system or human check.
What if there is no receipt?
Mark the claim unsupported or unresolved, ask for the missing evidence, and do not substitute the agent's summary.
Should a dispute expose private prompts?
No. Request the minimum evidence needed and protect secrets and personal data.
Where is the public verification path?
Use https://zambo.dev/verify when the disputed action has a public receipt.