What Is an AI Agent Execution Receipt? What It Proves and What It Doesn't

An agent saying “done” is not the same thing as having something you can review.

Try the Zambo demo · Explore Zambo

That distinction matters as agents move beyond drafting text. A coding agent may say it deployed a fix. A research agent may claim it checked a source. An agency may deliver an AI-assisted audit to a client. In each case, a reviewer needs a clearer answer than a polished summary: what operation ran, what did it return, and can I inspect the record without trusting the agent’s narration?

Quick answer: An AI agent execution receipt is a durable, inspectable record of a particular tool call or execution step. At its best, it records the operation, relevant inputs or context, time, observed result, and an integrity commitment - such as a hash over canonical bytes - so another person or system can check the stored record. It is evidence about the recorded execution, not automatic proof that every claim in the output is true or that an outside-world action occurred.

This guide explains the practical boundary: what a receipt is for, how to review one, and where it fits beside logs, traces, evaluations, and human judgment.

The problem: agent work is easy to summarize and hard to inspect

Most agent interfaces are built around a conversation. That is convenient, but a chat transcript is a weak handoff artifact when someone else needs to review the work later.

Consider a few familiar situations:

Tool use broadens the agent’s reach. Model Context Protocol (MCP), for example, is an open-source standard for connecting AI applications to external systems, including tools, data sources, and workflows.1 But a connection to a tool does not itself create a reviewable record of a specific call. A receipt is meant to fill that handoff and review gap.

The core question is not “Did the model sound confident?” It is “What evidence is available for this specific execution?”

A precise definition of an AI agent execution receipt

An AI agent execution receipt is a bounded evidence artifact for one execution. The boundary is important: it should identify the operation that ran and preserve what the system actually recorded, rather than turn a broad agent session into an unverifiable success claim.

A useful receipt normally makes it possible to inspect five things:

  1. Identity - which receipt and execution are under review.
  2. Operation - the tool or capability called, plus the relevant arguments or request context.
  3. Time and state - when the record was created and whether the call succeeded, failed, was blocked, or is awaiting confirmation.
  4. Observed result - what the system received or produced within its execution boundary, including upstream evidence when it was actually captured.
  5. Integrity check - a stable byte representation and/or hash that a verifier can recompute or compare.

Zambo describes its receipt as a public, persistent, independently checkable record of one AI agent tool call, including what ran, when it ran, the observed result, canonical bytes, and an output hash.2 The company’s own evidence guidance puts the same idea plainly: a receipt packages a tool call into a stable record that can be inspected beside existing telemetry - not in place of it.3

Why canonical bytes and a hash matter

A hash can only be checked reliably if both sides hash the same bytes. Merely serializing the “same JSON” is not always enough: whitespace, key order, and number serialization can differ. RFC 8785 describes why hashing and signing need an invariant representation and defines a JSON canonicalization scheme with deterministic property sorting.4

That does not make a hash a truth machine. It does something narrower and still valuable: it lets a reviewer test whether the stored byte sequence matches the stated commitment. The reviewer must separately assess where the result came from, whether the tool was appropriate, and whether the claim being made goes beyond the evidence.

What a receipt can prove - and what it cannot

This is the line that keeps “verifiable” from becoming a marketing adjective.

A receipt can support a claim that… A receipt alone cannot establish that…
a named tool call was recorded by the receipt system the tool’s answer was factually correct
the stored canonical bytes match the stated hash, if verification succeeds a model did not hallucinate or misinterpret the result
a result, error, or blocked state was recorded at a stated time an external business outcome happened outside the recorded boundary
the receipt includes a source or external-evidence reference, where one was actually captured a payment, post, trade, delivery, or permission change completed merely because an HTTP request returned successfully
the recorded request and result are available for a later reviewer to inspect the requester was authorized, an approval was valid, or the action was appropriate

In other words, a receipt can provide integrity and provenance evidence for a record. It does not collapse correctness, authorization, outcome confirmation, and accountability into one check.

Zambo’s receipt documentation makes this boundary explicit: a matching output hash shows that stored bytes match the commitment; it does not make an external value true or show that an action outside the execution boundary happened.2 Its documentation also distinguishes an asynchronous request being accepted from a second-phase observation such as a callback, delivery event, settlement event, or status read.2

That restraint is useful in practice. If a receipt says a tool returned “deployment accepted,” a reviewer can accurately repeat that. They should not translate it to “the deployment is live” unless the record contains a meaningful read-back, callback, or other independent confirmation.

The anatomy of a receipt

Receipt formats vary, and a production schema should be versioned. Still, the parts below are a solid review checklist.

Field or concept Why it is there Review question
Receipt ID and public or shareable URL Locates the exact record being discussed. Am I reviewing the same execution the agent cited?
Schema version Lets a verifier interpret fields consistently as the format evolves. Does my verifier understand this version?
Tool / operation Defines the capability boundary. Was this the tool I expected to run?
Arguments or request context Shows what the tool was asked to do, subject to redaction. Do the inputs support the conclusion being claimed?
Status and timestamps Preserves whether it succeeded, failed, paused, or was unavailable, and helps place it in sequence. Is this a completion, a failure record, or only an accepted request?
Observed result / preview Gives the reviewer the recorded result. A preview may not be the full upstream dataset. What did the system actually observe?
Canonical bytes, algorithm, and output hash Makes the integrity check repeatable. Does the recomputed or service-provided verification match?
Evidence and provenance fields Separates source-backed observation from a model explanation or a logged assertion. Is there an upstream source, read-back, callback, or no external evidence?
Redaction and access notes Prevents a shareable artifact from becoming a leak. What was omitted, and is that omission acceptable for this review?

Here is an illustrative, deliberately small receipt shape - not a universal schema and not a substitute for inspecting the full record:

{
  "id": "receipt_123",
  "tool": "website_audit",
  "created_at": "2026-09-28T16:10:00Z",
  "status": "completed",
  "output_hash": "sha256:…",
  "evidence_url": "https://example.com/page-audited"
}

A mature receipt can carry more: byte length, content type, source URL, provenance class, verifier checks, a chain or parent reference, and a list of external evidence. More fields are not automatically better. They should improve reviewability without publishing credentials, sensitive prompts, private records, or personal information.

The receipt lifecycle: from a tool call to a reviewable record

A receipt is most useful when it is created as part of the execution flow, rather than reconstructed after a dispute. A sensible lifecycle looks like this:

  1. Define the work boundary. The agent selects a tool and submits a request with the least context needed to do the job.
  2. Execute and capture. The tool runs. Record the operation, relevant inputs, timestamps, status, and observed output. Preserve failure states; do not rewrite a failed call as success.
  3. Normalize the record. Produce the defined canonical representation for the bytes that will be committed. This step must be specified well enough for independent checking.
  4. Commit integrity data. Calculate and store the declared hash and algorithm. If there is an evidence artifact, record its own digest and metadata where applicable.
  5. Attach provenance carefully. Distinguish among an action the system executed, information it observed from an upstream system, and an assertion it merely logged. Attach external confirmation only when it was actually observed.
  6. Publish or deliver an access-controlled record. A public link is convenient for a client handoff; a private or redacted record may be necessary for sensitive work. “Verifiable” does not require exposing secrets.
  7. Verify and review. A person or system checks identity, state, integrity, source evidence, and the exact claim being repeated.
  8. Preserve context for later. Link the receipt to a trace, ticket, commit, invoice line, approval, or follow-up read-back when that relationship matters.

For an API or workflow designer, the critical design decision is step five. Treating every 2xx response as proof of completion is a category error. For external side effects, record the request acceptance separately from a provider callback or a later read-back.

Three practical examples

1. A website audit delivered by an agency

An agency asks an agent to scan a public marketing site for missing metadata and broken canonical links. The client receives a report plus a receipt URL for the audit call.

A reviewer can check the target URL, tool, time, observed findings, and integrity commitment. That is useful proof of the recorded audit execution.

What it does not prove: that every recommendation is strategically correct, that a page will rank better after changes, or that the client implemented the work. Those require separate editorial judgment, search evidence, and potentially a later read-back.

2. A coding agent flags a dependency issue

A coding agent uses a repository audit tool and returns a receipt tied to the repository revision or other captured evidence. The maintainer reviews the receipt alongside the pull request and test results.

The receipt helps answer: “Which analysis did the agent actually run?” The pull request, CI run, and human review answer different questions: “Was the change correct, tested, and approved?”

3. An external action is initiated

An agent sends a request to create a record in a third-party system. The first receipt shows the request and immediate response. A second receipt or linked evidence record shows an observed webhook callback or a read-back of the created object.

The first record may prove request accepted. The second can support a stronger, but still bounded, claim that the system observed a corresponding downstream state. If no read-back or callback exists, say so. Do not promote the initial receipt into proof of completion.

How to verify an execution receipt

Verification does not need to be ceremonial. It needs to be repeatable and appropriately skeptical.

A five-step review

  1. Confirm identity. Open the receipt URL or API record and make sure its ID, tool, and time match the execution under discussion.
  2. Check the operation and state. Read the tool name, relevant arguments, status, and result. Look for a visible failure, pending state, or incomplete scope.
  3. Verify integrity. Use the published verifier or independently recompute the declared digest over the declared canonical bytes. A match supports the claim that the stored record has not changed relative to that commitment.
  4. Inspect provenance and external evidence. Review source URLs, captured artifacts, read-backs, transaction references, callbacks, or an explicit absence of evidence. Assess their relevance and freshness.
  5. Match language to evidence. State only what the receipt supports. “The tool returned X at time Y” is different from “X is correct” or “the external action completed.”

Zambo’s documented sequence follows this pattern: inspect the receipt ID, tool, timestamp, and result; request the verifier result; compare the hash; then inspect any upstream evidence separately from the record’s integrity.2 Its receipt verifier is the product path for checking a Zambo receipt.

A reviewer’s stop signs

Pause the review - or downgrade the conclusion - when you see any of the following:

A good verification outcome can be “unavailable,” “not enough evidence,” or “request accepted but outcome unconfirmed.” That is not a failure of the receipt model; it is honest evidence handling.

Receipts vs. logs, traces, and evaluations

These are complementary tools. Trying to force one artifact to answer every question usually creates a misleading audit trail.

Artifact Primary job Best question it answers What it does not replace
Execution receipt Shareable evidence for one bounded agent tool call “What was recorded for this particular execution, and can I check its integrity?” System-wide debugging, quality testing, or external outcome proof
Logs Event-level operational detail “What did the system record while it ran?” A stable, reviewer-oriented handoff artifact or a correctness guarantee
Trace Request path and timing across components “How did this request move through services and dependencies?” A focused, portable statement of one result
Evaluation (eval) Measured quality against defined cases or criteria “How well does this agent or workflow perform across a test set?” Proof of one production execution
Approval / authorization record Human or policy decision to permit an action “Who authorized this operation under which conditions?” Proof that the action succeeded
External confirmation Evidence from the affected outside system “Did the external system report the outcome?” Evidence of every prior step or the quality of the original decision

OpenTelemetry describes a trace as the path of a request through an application, composed of correlated spans that represent units of work.5 That makes traces excellent for diagnosing latency, dependencies, and failures. Logs and traces should usually remain in the engineering stack. A receipt is the artifact you add when one completed (or failed) AI step needs a stable link for a handoff, review, invoice reconciliation, or later check.

Evals solve another problem entirely. A regression suite might tell you whether an agent generally follows a tool-use policy. A receipt tells you what was recorded for this tool call. Use both if the work matters.

Designing receipts that hold up in real reviews

If you are building an agent product, do not start with a badge that says “verified.” Start with the reviewer’s workflow.

Make the record legible

Make provenance explicit

Use language that reveals the source of a claim: executed, observed, reported by provider, read back, or not confirmed. That avoids the common slide from “the agent sent a request” to “the business outcome happened.”

Protect sensitive work

Receipts often need redaction, access control, retention rules, and safe sharing defaults. An integrity commitment to a record is not a reason to disclose credentials, personal data, proprietary prompts, or customer information. Consider whether a reviewer needs the raw input, a minimized representation, or a secure reference to a private artifact.

Link, do not flatten

Link a receipt to the relevant trace, commit, change request, approval, test run, or external confirmation. Each record should keep its own meaning. A clean chain of distinct artifacts is more defensible than one overloaded “success” object.

Where Zambo fits

Zambo is designed around verifiable execution receipts for AI agent tool calls. Through its hosted MCP connection or direct API path, Zambo says that each successful call returns a receipt URL, allowing a user to inspect the recorded result and verify the associated record.6

The practical workflow is straightforward:

  1. Connect an MCP-capable AI client to Zambo or make a direct call.
  2. Run a concrete job, such as a public website audit, prompt-injection scan, live data lookup, or code-oriented check.
  3. Open the resulting receipt at the supplied zambo.dev/run/... link.
  4. Inspect the tool, result, time, evidence, and integrity fields; then use the Zambo verifier.
  5. Share the receipt with a client, teammate, or reviewer with a claim limited to what the record supports.

Zambo does not position its receipts as a replacement for observability tooling. Its documentation explicitly recommends using traces, logs, and metrics for system health, debugging, and alerting, while using a receipt when an AI step needs a portable evidence link.3 That is the right division of labor for teams that need both reliable operations and reviewable deliverables.

Try the evidence loop: run a read-only Zambo demo, open the receipt, and verify it before deciding whether the artifact is useful in your workflow. For client configuration and direct API examples, use the Zambo MCP installation guide.

Suggested internal links

Use these real Zambo pages to move readers from definition to a concrete proof loop:

FAQs

What is an AI agent execution receipt?

An AI agent execution receipt is an inspectable record of a specific agent tool call or execution step. It should identify what ran, when it ran, the recorded result, and a way to check the record’s integrity. It is stronger than a chat summary because another reviewer can examine the same record.

Does an execution receipt prove that an agent’s answer is true?

No. A receipt can support that a system recorded a tool call and result, and that the committed record matches its stated hash. It does not by itself prove the answer is correct, complete, current, or appropriately interpreted. Review the source evidence and the claim separately.

Does an execution receipt prove that an external action completed?

Not necessarily. A receipt for a successful request may prove that the request was recorded or accepted. To claim an external outcome - such as a payment settled, a record was created, or a post was published - you need observed confirmation from the affected system, such as a callback, transaction reference, or read-back.

How is an execution receipt different from an agent trace or log?

A trace helps engineers follow a request across components; a log captures operational events. A receipt is a portable, reviewer-oriented record for one specific agent execution. They work together: link a receipt to traces and logs when a reviewer needs both the high-level evidence artifact and diagnostic detail.

How do I verify a Zambo execution receipt?

Open the receipt URL, confirm its ID, tool, status, time, and result match the work under review, then use Zambo’s verifier to check the record and output hash. Inspect any listed external evidence independently, and limit your conclusion to what that evidence actually establishes.2

References

  1. Model Context Protocol, “What is the Model Context Protocol (MCP)?” https://modelcontextprotocol.io/docs/getting-started/intro
  2. Zambo, “AI Agent Execution Receipt - Definition & Verification.” https://zambo.dev/execution-receipt/
  3. Zambo, “Receipts as execution evidence.” https://zambo.dev/docs/receipts-as-execution-evidence/
  4. RFC 8785, “JSON Canonicalization Scheme (JCS).” https://www.rfc-editor.org/rfc/rfc8785
  5. OpenTelemetry, “Traces.” https://opentelemetry.io/docs/concepts/signals/traces/
  6. Zambo, “Install Zambo MCP.” https://zambo.dev/install/