AI Agent Receipts for A2A Agents
Last updated: 2026-09-28
Agent-to-agent work needs a handoff record that another agent or reviewer can inspect. An execution receipt can record the tool, time, scope, observed result, and integrity evidence for one step, while the A2A message remains responsible for task state and delivery.
Try the Zambo demo · Explore Zambo
Use the public receipt verifier to inspect a receipt when you have one. A receipt is evidence about a recorded execution, not a guarantee that an outside system accepted the result or that a human decision was correct.
Why an agent handoff needs evidence
An agent-to-agent exchange is not automatically proof of work. One agent can say that it searched, wrote, approved, or delivered something, while the receiving agent sees only a message. That message may be useful, but it is still a claim from the sender. A receipt creates a separate inspection point for the execution layer. It can identify the recorded action, preserve the relevant bytes or an output commitment, and show what the recorder observed at a particular time. The distinction matters when an agent hands work to another agent, when a task crosses a trust boundary, or when a later reviewer needs to reconstruct why a downstream decision was made.
Mapping a lifecycle to receipt fields
A practical A2A lifecycle has discovery, authorization, delegation, execution, response, and review phases. A receipt fits most directly around execution. The stable receipt identifier lets a receiver refer to one event without relying on a conversational position. The created-at timestamp gives temporal context. The tool object names the operation, its version, and the caller scope. The provenance class says whether the record was executed by the recorder, observed through a gateway, or logged from an agent report. Canonical bytes and an output hash let a verifier check that the committed representation has not silently changed. Verification status records the result of the stated verification procedure.
Discovery and authorization are separate
A receipt should not be used as a substitute for agent discovery metadata or authorization policy. Discovery tells a participant what an agent claims to support. Authorization determines whether a requested action is allowed. The receipt starts after the execution layer has a record to preserve. A receiving agent should validate the sender identity and permission context using the protocol that governs the exchange, then use the receipt to inspect the claimed execution. Keeping these responsibilities separate avoids a common mistake: treating a well-formed record as permission to act, or treating a permission decision as evidence that the action already ran.
Delegation and scope
When one agent delegates a step, the receipt should make the caller scope understandable without exposing credentials or private prompts. A useful scope is the permission context that the execution layer actually evaluated. It is not a request to publish every internal policy rule. The receiving agent can then ask whether the recorded scope was sufficient for the delegated action. If the receipt does not contain enough context, the correct result is an explicit evidence gap. It is better to say that scope was not independently established than to infer authorization from a successful-looking response.
Execution and observed output
The execution receipt describes what the recorder observed. If the tool returned a structured response, the receipt may commit to the canonical bytes or a digest of them. If the result was an external side effect, the receipt should state whether the side effect was observed directly, reported by another agent, or merely requested. An A2A response can carry the receipt reference alongside its task result. The downstream agent should preserve that distinction when it summarizes the handoff. A response saying delivered is not equivalent to a receipt saying the execution layer observed an accepted delivery.
Task-completion receipts are broader
A task-completion receipt usually summarizes whether a multi-step objective reached a claimed state. An execution receipt is narrower and easier to audit: it records one operation and its observed result. A task-level record can link several execution receipts, but it should not erase them. If a task says that three agents searched, compared, and published, the reviewer should be able to inspect three corresponding events or see which step has no evidence. The task summary can remain convenient for routing, while the step receipts preserve the trail needed for disagreement.
What a receiver should verify
A receiving agent can perform a small, repeatable check. First, resolve the receipt reference and confirm that its identifier matches the handoff. Next, inspect the timestamp, tool, scope, provenance class, and verification status. Then check the canonical bytes and output hash using the published verification path. Finally, compare the result with the task's acceptance condition and, when possible, inspect the system that owns the claimed outside state. This sequence verifies the record before it relies on the record. It also makes a partial result visible instead of allowing a confident summary to hide it.
Privacy and replay boundaries
A2A messages can contain sensitive instructions, identities, or customer data. A receipt should expose the minimum evidence needed for the stated review. Hashes do not make private data safe to publish if the original can be guessed or recovered. Receivers should also consider replay: an old valid receipt may describe a real execution but still be irrelevant to a new request. Check timestamps, task identifiers, freshness requirements, and any authorization binding supplied by the governing protocol. A receipt proves what it covers. It does not automatically prove that a later message is fresh or authorized.
A conservative integration pattern
A useful integration stores a receipt reference in the agent-to-agent task record, retains the exact response that the sender presented, and records the receiver's verification outcome separately. The receiver can mark the handoff as evidence-backed, evidence-partial, or unverified. Those labels describe the review state, not the quality of the underlying task. A failed fetch is a transport problem until the record can be retrieved again. A valid receipt with an incorrect tool output remains a valid record of an incorrect observation. That honest boundary is what makes the pattern useful in incident review.
Start with one real step
Do not begin by designing a universal agent registry. Choose one high-value handoff, define the claim that needs evidence, and require a receipt reference for the execution step that supports it. Test the receiver's verification path with a valid record and with a missing or changed record. Link the reviewer to the public verifier at /verify. Then document what remains outside the receipt, such as an external system's eventual state or a human approval. Small, explicit boundaries are easier to operate than a protocol that claims to prove every part of an agent conversation.
Review notes
Good evidence work starts with a bounded question. Name the action, the system that owns the result, the time window, and the person or service responsible for reviewing it. Then separate three states that are often collapsed into one: an action was requested, an execution layer recorded an action, and an outside system confirmed an outcome. The first state is intent. The second is execution evidence. The third is outcome evidence. A page that keeps these states separate is easier to use during an incident and harder to misuse in a status summary.
When evidence is incomplete, record the gap instead of filling it with confidence. A missing record may be a collection failure, a transport problem, a retention issue, or an action that never ran. Those possibilities have different owners and different remedies. Preserve the original record, the verification attempt, and the reviewer conclusion. If a later retry succeeds, keep the earlier attempt visible. This makes the evidence useful to someone who was not present when the work happened and gives the next operator a concrete place to start.
Reviewers should also write down assumptions. State which timestamp is authoritative, whether a result was observed directly, how a hash was computed, and which system owns the final state. If the answer depends on an external service, link the external evidence or mark it unavailable. If the action involved private data, describe the field class without copying the data into a public page. These habits turn a receipt from a decorative status badge into a durable review artifact. They also make disagreements productive: two reviewers can compare assumptions instead of arguing over a summary.
Finally, keep the evidence proportional to the decision. A low-risk read can use a light check, while a money movement, data mutation, deployment, or external communication deserves a stronger receipt and an independent state check. Do not use a larger word count or a more elaborate format to imply certainty. The relevant question is always what was recorded, what was verified, and what remains outside the record.
Frequently asked questions
Does an A2A receipt prove the other agent completed its task?
No. It records an execution event and its observed result. Task completion and outside effects require their own evidence.
Which fields matter most in an agent handoff?
The stable identifier, timestamp, tool and scope, provenance class, canonical bytes, output hash, and verification status provide the core inspection trail.
Can a receipt replace authorization?
No. Authorization belongs to the protocol and policy layer. A receipt helps review what the execution layer recorded after that decision.
How should a receiver handle missing evidence?
Mark the handoff unverified or partial, request the missing record, and do not turn the sender's summary into proof.
Where can a reviewer inspect a Zambo receipt?
Use the public verifier at https://zambo.dev/verify when a receipt reference is available.