Your AI Agent Said "Done": How to Verify the Work
The most dangerous sentence an AI agent can send is often the shortest:
Try the Zambo demo · Explore Zambo
“Done - I published the post, merged the fix, and updated the research brief.”
That sentence may be accurate. It may also be a summary of an intention, a tool call that never left the queue, a tool call that failed halfway through, or a model filling in the ending it expected to see.
The problem is not that models can be wrong. The problem is treating a confident natural-language message as an audit trail.
If an agent’s work matters to a client, a deployment, a decision, or a deadline, separate three questions:
- What did the agent intend to do?
- What execution did a system actually record?
- What changed in the outside world?
Those are different states. A good verification workflow keeps them different until evidence connects them.
Zambo’s execution-receipt model is useful here precisely because it draws a hard line: a receipt can establish the integrity of a recorded execution, but it cannot by itself prove a result in a system Zambo did not observe.1 That limitation is not a defect. It is how you avoid turning “the agent said so” into a false fact.
“Done” is a claim. An execution record is evidence.
A chat transcript is useful context. It is not a reliable source of truth for whether work occurred.
Consider these two artifacts:
Agent message
─────────────
Done. I changed the rate limiter and deployed the service.
Execution record
────────────────
request_id: req_8e2…
tool: gitlab.create_merge_request
input_digest: sha256:…
started_at: 2026-09-28T17:06:13Z
finished_at: 2026-09-28T17:06:15Z
status: succeeded
result: { merge_request_url: "https://gitlab.com/acme/api/-/merge_requests/482" }
receipt_digest: sha256:…
The first tells you what the model claims. The second gives you something inspectable: which tool ran, when it ran, what request was recorded, what result the tool returned, and a way to check whether the record has been altered.
But the second still does not prove production received the rate limiter. It proves a pull request was created - assuming the receipt verifies and accurately represents the observed tool interaction. The deployment needs a different record and, for high-consequence work, an independent read-back from the production environment.
This is the discipline behind AI agent proof of completion: do not promote a claim to a fact just because it is fluent.
Use four states, not one vague “done” state
Most agent workflows become easier to reason about when they expose state explicitly.
| State | What it means | Useful evidence | What it does not establish |
|---|---|---|---|
| Planned | The agent proposed an action. | Plan, requested parameters, approval decision. | That a tool was invoked or any system changed. |
| Attempted | The agent submitted a request to a tool or service. | Request ID, queued job ID, timestamp, authenticated caller. | That the tool completed successfully. |
| Executed | The tool or execution layer recorded a completed attempt and result. | Tool response, status, timestamps, execution receipt, output digest. | That a downstream or external system reflects the intended outcome. |
| Confirmed | An authoritative system was read after the action and returned the expected state. | Read-back response, version/revision, URL, deployment health, independent observer. | That the state will persist forever, or that every side effect was desirable. |
Call the last state confirmed, not “guaranteed.” A read-back can go stale. An API can report success while an asynchronous job later fails. A human can revert a change five minutes after you check it. Verification is a claim about a particular object, from a particular observer, at a particular time.
The practical rule is simple:
An agent may describe a result at any time. Only evidence should advance its status.
For consequential writes, keep planned, attempted, executed, and confirmed as distinct fields in the job record. Do not make an LLM-generated summary the field that determines success.
The verification checklist: what to ask after an agent says “done”
Use this checklist for a handoff, a review queue, or a human-in-the-loop approval screen.
1. Identify the exact operation
Start with a noun and a verb - not a paragraph:
- Publish: which CMS item, URL, locale, and revision?
- Change code: which repository, branch, commit, pull request, and environment?
- Research: which sources were retrieved, on what date, and what claims came from each?
If the answer is “I updated the website,” the work is not yet reviewable. Ask for the object identifiers.
2. Inspect the recorded execution
Find the execution record or receipt. Verify at least:
- tool or integration name and version;
- request and response timestamps;
- status, including partial failures and retries;
- identifiers returned by the target system;
- input/output or their digests, with the expected redactions;
- who or what initiated the action; and
- integrity/provenance fields needed by your verifier.
A receipt is strongest when another person or system can verify it without relying on the agent’s narration. Zambo describes receipts as portable records designed to make recorded execution inspectable and verifies their stored integrity; see its verification guidance.1
3. Check the result against the intended change
A 200 OK is not the same as the requested result.
Match the recorded result to the planned operation:
- Did the returned content ID equal the draft the editor approved?
- Did the PR target
main, not a similarly named branch? - Did the research job retrieve the specified primary source, not a cached article repeating it?
This catches a mundane but expensive class of agent failure: the tool worked, but the agent used the wrong target, parameter, account, or version.
4. Read the outside world back
For a write, query the authoritative system after the action. This is the step people skip when they accept a reassuring chat message.
Examples:
| Job | Execution evidence | Outside-world read-back |
|---|---|---|
| Publish an article | CMS write response and receipt with content ID | Fetch the canonical URL; verify status, revision, publication time, and rendered content. |
| Merge a code fix | Merge response, commit SHA, CI result | Query the protected branch; deploy or inspect the target environment; run a focused health check. |
| Send a research brief | Retrieval/tool records and stored source artifacts | Open the cited sources; confirm each central claim supports the wording and that dates are current. |
Use a distinct credential or observer where practical. A deployment tool that reports its own success is useful, but a production health endpoint, an artifact registry, or a downstream log is better confirmation. For sensitive systems, the reader should have least-privilege access and should not be controlled by the same agent that performed the write.
5. Preserve the link between intent, execution, and confirmation
A receipt alone is easy to detach from its context. Store the chain:
job_id
├─ approved plan / request digest
├─ execution receipt(s)
├─ returned object IDs and versions
├─ outside-world read-back(s)
└─ reviewer decision and timestamp
The chain makes later questions answerable: Which request produced this document? Which execution changed this branch? What did production show immediately afterward?
6. Record uncertainty instead of laundering it away
“Could not confirm” is a valid outcome. So are “partial success,” “write accepted; async processing pending,” and “source inaccessible.”
Do not have the agent convert those into a polished “completed” summary. Make them visible in the final status. An agent that can honestly say “executed, confirmation pending” is more useful than one that always sounds finished.
Three failure patterns this catches
Publishing: the draft was saved, not published
An editorial agent calls a CMS endpoint. The tool returns a content ID and an updated status. The agent says, “The article is live.”
That may be a category error: many CMS APIs use the same object for draft, scheduled, published, or rejected content. The receipt can show the agent called the endpoint and received updated. It cannot turn that state into “live.”
Verify it: retrieve the public canonical URL without an editor session. Check the visible title, published state, revision, timestamp, and the required sections. If the page is expected to be indexed, that is a later operational check - not evidence that the original publish action succeeded.
Code changes: the pull request exists, the fix does not
A coding agent creates a PR and receives a URL. It reports, “The bug is fixed.”
Here, the PR URL and execution record can prove a PR was created. They do not prove the diff addresses the bug, tests passed, reviewers approved it, the merge happened, or production changed.
Verify it: inspect the diff against the acceptance criteria; verify CI against the commit SHA; check merge status and the protected branch’s revision; then run a focused check in the deployed environment. The important unit of evidence is not “code was written.” It is the desired behavior was observed on the intended version.
Research: sources were found, the conclusion was not supported
A research agent returns ten links and a confident synthesis. A retrieval log or receipt can demonstrate that it fetched or searched for sources. It cannot establish that the conclusion accurately represents those sources.
Verify it: open the primary sources for material claims; check publication dates, authorship, scope, and whether the claim is stated or inferred. Save quotations or structured notes alongside the URLs. For claims likely to change, record the access date and make the brief say what it cannot establish.
This is why “research complete” should mean a reviewable evidence packet exists, not merely “the model has a conclusion.”
What receipts can establish - and what they cannot
Receipts are valuable because they are narrower than an agent’s prose. They can support statements such as:
- a named tool invocation was recorded at a time;
- the recorded input and output have not changed since a receipt was created, if integrity verification succeeds;
- a tool returned a particular status or object ID; and
- an execution record has provenance fields that a reviewer can inspect.
That is enough to make a handoff materially better. It enables review, reconciliation, and debugging without asking someone to trust a chat bubble.
But a receipt should not be used to claim more than its observer saw. It generally cannot, by itself, establish:
- that a third-party system committed a write after an asynchronous handoff;
- that a public page is visible to the intended audience;
- that a deployment is healthy under real traffic;
- that an external business outcome occurred; or
- that the work was correct, safe, authorized, or useful.
Zambo’s public documentation makes this boundary explicit: its receipt verifies the stored execution record and integrity, not an outside-world outcome it did not observe.1 Its public verifications log is also a useful example of keeping source-linked observations separate from broader conclusions.2
The correct response to this limitation is not “receipts are useless.” It is to add the evidence source appropriate to the claim: a CMS read-back for publishing, a production probe for deployment, a bank or ledger record for a payment, or primary-source review for research.
Build completion as an evidence pipeline
The architecture does not need to be elaborate. It needs to stop allowing the model’s final sentence to overwrite reality.
- Plan: emit a structured proposed action and required acceptance checks.
- Approve: bind approval to the specific request, target, and material parameters for consequential actions.
- Execute: capture a tool-level record; preserve failures, retries, and partial results.
- Receipt: make the execution record independently inspectable and integrity-checkable.
- Read back: query the authoritative outside system after the action.
- Decide: set the final state from the evidence policy, not the model’s wording.
Planning is not execution, and approval is not enforcement. Zambo’s guidance on approvals makes the same distinction: whether a proposed action is enforced depends on the tool and integration, not on how confidently an agent described its plan.3
For a low-risk read-only task, the pipeline may end at a verified execution record. For a financial, administrative, credentialed, or production write, require explicit approval and an outside-world read-back. The MCP security guidance similarly emphasizes least privilege, user control, and safeguards around sensitive actions.4
A better final message for agents
Do not ban agents from saying “done.” Make the word carry a defined evidence level.
Here is a useful completion template:
Status: Confirmed
Requested action: Publish “Your AI Agent Said Done” to /blog/agent-verification
Executed: CMS publish request recorded at 17:06:15 UTC
Receipt: [verify execution record](https://example.com/receipt/abc)
Read-back: Public URL returned 200 at 17:06:20 UTC; revision 418 matches approved digest
Limits: Search indexing and downstream cache propagation not checked
If the read-back has not happened, say Status: Executed - confirmation pending. If the tool returned an error after creating an object, say Status: Partial success and give the object ID. These are not inferior answers. They are answers a reviewer can act on.
Make the next agent handoff reviewable
The fastest way to improve trust in an agent workflow is not another prompt telling it to “be accurate.” It is a small rule in the product:
No consequential task is complete until its intended result has the right evidence.
Start with a safe, read-only job and inspect the path from request to result to receipt. Zambo’s installation guide shows how to connect an AI client, and its execution receipt overview explains what to verify. When your workflow needs a real outside-world effect, add the read-back that makes “done” worth believing.
Frequently asked questions
1. Can an AI agent hallucinate that a task is complete even if it called a tool?
Yes. The model can summarize a failed, partial, queued, or mis-targeted operation as successful. A tool call is stronger evidence than a chat message, but you still need to inspect the status and returned identifiers. For a write, read the authoritative system back afterward.
2. What is proof of completion for an AI agent?
There is no single artifact that proves every kind of completion. A defensible proof chain usually combines the intended request, an observed execution record or receipt, and - when the claim concerns an external system - an authoritative post-action read-back. The evidence should match the claim you need to make.
3. Does an AI agent audit trail prove that the work was correct?
No. An audit trail can show what was proposed, attempted, recorded, and confirmed. It does not automatically establish that the change met requirements, was authorized, or produced a desirable business outcome. Correctness still needs tests, review, acceptance criteria, or domain-specific validation.
4. What should an AI execution receipt contain?
At minimum: a stable execution identifier, tool/integration identity, timestamps, outcome status, the returned object IDs or outputs (or safe digests), and integrity/provenance data a verifier can check. It should also make redaction and its evidence boundary clear. See Zambo’s receipt documentation for its model.1
5. When is outside-world read-back necessary?
Use it whenever the final claim is about a system beyond the execution recorder: a published page, a merged branch, a production deployment, a sent payment, or an externally visible record. For low-risk read-only retrieval, a verified execution record may be enough. Increase independence and rigor with the consequence of getting it wrong.
References
- Zambo, “Receipts as execution evidence”.
- Zambo, “Independent Verifications Log”.
- Zambo, “Planning is not execution”.
- Model Context Protocol, “Security Best Practices”. See also National Security Agency, “Model Context Protocol (MCP): Security Design Considerations”.