How to Add Verifiable Execution Receipts to MCP Tools
An MCP tool result is useful to the calling agent. A receipt makes the execution reviewable by someone - or something - outside that conversation.
Try the Zambo demo · Explore Zambo
The goal is not a longer log. It is a portable, bounded record that answers: what tool was requested, under what authority, what the server observed, what bytes are committed, and what remains unproven? That boundary matters especially for tools that write data, spend money, change access, or publish externally.
Zambo’s execution-receipt overview and public verification log illustrate the useful distinction: integrity of a recorded execution is not the same as proof that an outside-world outcome occurred. This guide shows how to build that distinction into an MCP server from the start.
First: know what a receipt is - and is not
These terms solve different problems. Do not use them interchangeably in a UI or API.
| Artifact | What it is | What it can support | What it does not establish by itself |
|---|---|---|---|
| Execution receipt | A structured record of an attempted or completed tool lifecycle, with commitments to inputs, outputs, provenance, and status | That a defined execution record was created; with integrity checks, that its committed fields have not changed | That every field is true, that a third party accepted a request, or that the business objective was achieved |
| Attestation | A claim made by an issuer or witness, such as “the provider returned HTTP 202” | What that issuer says it observed, scoped to its identity, time, and method | Independent truth beyond the issuer’s observation |
| Digital signature | A cryptographic mechanism over bytes or a digest | Integrity and, when the verifier trusts the public key and key binding, attribution to the signing key | Correctness of the signed statement or an external outcome |
| Proof of business outcome | Evidence from the system that owns the relevant state, evaluated against an explicit acceptance condition | For example, “the invoice is marked paid in the ledger at this time” | A permanent guarantee; an external state can later change or be reversed |
A receipt may contain an attestation, be signed, and link to outcome evidence. It is still important to label those as separate layers. JSON Web Signature (JWS), for example, provides integrity protection for content using signatures or MACs; it does not certify the truth of the content.jws
Use precise status language. “The server recorded an execution attempt” is a receipt claim. “The destination system displayed the requested state at 14:03 UTC” is an external read-back claim. “The customer received the message” needs evidence from a system that can observe delivery or receipt.
1. Classify the tool before you design the receipt
Receipt detail, approval, and retry behavior should follow the harm of a wrong or repeated action - not the tool name. Start with a small, enforceable taxonomy.
| Risk class | Typical tools | Default control | Receipt expectation |
|---|---|---|---|
| R0 - read-only, bounded | Fetch public price, inspect a repository, calculate a value | Allow within scoped access and rate limits | Request/result commitments, source/provenance, normal failure record |
| R1 - reversible or draft write | Create a draft, add a label, stage a change | Identity plus policy; allow undo window where possible | Target, undo/compensation reference, idempotency key, observed result |
| R2 - durable external write | Send a message, update a CRM record, deploy a configuration | Human or policy approval bound to the exact request; least-privilege scope | Approval, dispatch attempt, external response, read-back state or explicit absence |
| R3 - financial, access, public, or destructive | Transfer funds, grant admin, delete records, publish content | Named principal, explicit confirmation, strong authorization, request-bound approval, no blind retry | All R2 fields plus policy decision, approver identity, exact scope, outcome evidence and review state |
Assess three dimensions together:
- Side effect: none, ephemeral, reversible, durable, destructive, or irreversible.
- Authority: public data, user data, tenant data, credentials, administrator, or financial authority.
- Blast radius: one object, many objects, a public audience, a regulated record, or an account boundary.
Treat unknown behavior as higher risk until tested. A tool that looks like a read can still trigger a provider-side cache warm, paid query, export, or notification.
Declare side effects in the tool contract
MCP tool annotations are useful discovery metadata, not an enforcement boundary. The MCP tools specification supports annotations and warns clients to treat them as untrusted unless they come from a trusted server; it also recommends that a human can deny calls and that clients make invocations visible.mcp-tools Keep the standard hints accurate - readOnlyHint, destructiveHint, idempotentHint, and openWorldHint - then add a server-owned contract that your policy engine actually enforces.
{
"name": "crm.create_contact",
"annotations": {
"readOnlyHint": false,
"destructiveHint": false,
"idempotentHint": true,
"openWorldHint": true
},
"x-receipt-policy": {
"risk_class": "R2",
"side_effect": "durable_write",
"approval": "request_bound",
"idempotency": "required",
"confirmation": "provider_readback_required",
"retention_class": "customer_record"
}
}
x-receipt-policy is an example extension, not an MCP-defined field. Version and document it. The server must derive its decision from the authenticated principal, target tenant, arguments, and live policy - not from an agent-provided risk label.
For R2 and R3 calls, show the user a compact preflight: tool, target, affected objects or amount, credentials/scopes, irreversible effects, and what confirmation will be collected. This supports the MCP security guidance to make consent specific to the client and scopes, rather than treating a broad prior consent as permission for every action.mcp-security
2. Make the lifecycle explicit
A single success: true cannot describe a consequential tool call. Use separate lifecycle states and retain them as events or linked receipts.
proposed ──> approved ──> executed ──> confirmed
│ │ │ │
└─rejected └─expired ├─failed └─contradicted
├─partial
└─unknown
- Proposed: The server normalized the intended call and evaluated policy. No side effect has been dispatched.
- Approved: A person or policy approved a specific request fingerprint for a limited time. This is not execution.
- Executed: The server attempted the operation and recorded what it observed. Use this even when the result is a provider error; do not equate it with a successful outcome.
- Confirmed: An independent read-back or evidence capture met a defined confirmation rule. Confirmation says what the evidence source showed at capture time, not that an inferred business objective was achieved forever.
A confirmed receipt should name the acceptance condition. For a message, it might be “destination API returns an immutable message ID and a subsequent GET retrieves that ID.” For a transfer, it might be “ledger reports accepted” rather than “settled,” unless settlement evidence is actually present.
Bind approval to the material request
An approval that covers “send the report” is too easy to replay or reinterpret. Bind the approval to a request hash over at least:
- tool name and immutable tool-contract version;
- normalized arguments and target environment/tenant;
- requested authority/scopes and authenticated principal;
- risk class and side-effect declaration;
- approval policy version, expiry, and a one-time approval nonce.
If any material field changes, invalidate approval and generate a new proposed receipt. Approval is authorization evidence, not a signature on a vague intent.
proposal = canonicalize({
tool, tool_contract_digest, tenant, principal,
normalized_arguments, requested_scopes, risk_class, policy_version
})
request_hash = sha256(proposal)
if approval.request_hash != request_hash or approval.expires_at < now():
emitReceipt(state="rejected", reason="approval_mismatch_or_expired")
stop
executeOnce(idempotencyScope(tenant, tool, idempotency_key), proposal)
3. Use a receipt payload and a separate integrity envelope
Do not hash a JSON object that includes its own hash. Create an immutable payload, canonicalize it, then attach an integrity envelope.
{
"payload": {
"receipt_schema": "https://example.com/schemas/mcp-receipt/1.0",
"receipt_id": "rcpt_01J...",
"event_id": "evt_01J...",
"parent_receipt_id": "rcpt_01I...",
"state": "executed",
"created_at": "2026-09-28T18:01:23.456Z",
"request": {
"tool": "crm.create_contact",
"tool_contract_digest": "sha256:...",
"request_hash": "sha256:...",
"idempotency_key_digest": "sha256:...",
"arguments": {"email": "[redacted]", "list_id": "customers"},
"redaction_profile": "public-v1"
},
"authorization": {
"principal_ref": "user:opaque-7f...",
"approval_ref": "apr_01J...",
"approved_request_hash": "sha256:...",
"scopes": ["crm.contacts.write"]
},
"execution": {
"attempt": 1,
"started_at": "2026-09-28T18:01:24.001Z",
"ended_at": "2026-09-28T18:01:24.842Z",
"outcome": "provider_accepted",
"result_artifact_digest": "sha256:...",
"external_reference": "crm:contact/opaque-91..."
},
"provenance": {
"server_instance": "mcp-prod-eu-2/opaque",
"server_build": "git:4b6c...",
"tool_implementation_version": "3.4.1",
"upstream": {"provider": "ExampleCRM", "api_version": "2026-06"}
},
"claims": [
{"claim": "provider accepted create request", "evidence": "execution.result_artifact_digest"}
]
},
"integrity": {
"canonicalization": "RFC8785",
"hash_algorithm": "SHA-256",
"payload_digest": "sha256:...",
"signature": {"format": "JWS", "kid": "receipt-2026-q3", "value": "..."}
}
}
The example uses references and digests rather than live credentials, personal data, or raw provider bodies. It also separates an observed provider acceptance from a verified business outcome.
Canonicalization and hashes
Different JSON serializers can change whitespace, object-key order, number rendering, and escaping without changing the apparent data. A signature or hash over arbitrary serialization will therefore fail across implementations.
For JSON receipts, adopt a named canonicalization algorithm and publish fixtures. RFC 8785, the JSON Canonicalization Scheme (JCS) defines a deterministic JSON representation for hash and signature use, including deterministic property sorting and constraints around JSON values.jcs
A practical verification sequence is:
- Validate the receipt against its declared schema.
- Extract only
payload; reject duplicate keys, invalid Unicode, non-finite numbers, and undocumented transforms before canonicalization. - Canonicalize
payloadusing the declared algorithm; encode the result as UTF-8 bytes. - Recompute the declared digest and compare it to
integrity.payload_digestusing a constant-time comparison in security-sensitive code. - If a signature is present, verify it over the documented bytes or digest, using a trusted key identified by
kid. - Report integrity, signature/key trust, schema support, and evidence availability as separate results.
A hash makes a changed payload detectable only when the verifier has an expected digest from a trusted receipt or log. A signature adds a key-bound integrity claim, but key distribution and rotation are part of the design. JWS explicitly notes that a verifier must authenticate the origin of a public key; otherwise it does not know who signed the message.jws
Provenance is more than a server name
Record enough provenance to answer “which system observed this?” without turning the receipt into a credential dump:
- authenticated actor/principal reference and tenant, using opaque IDs when appropriate;
- MCP client or agent identity where available, transport/session correlation ID, and request time;
- server deployment/build, tool implementation and policy versions;
- upstream provider identity, endpoint or region class, response reference, and capture times;
- artifact locations, retention policy, and the digest of each referenced artifact.
Do not imply that an unauthenticated user-agent string identifies an actor. Label confidence or source, such as client_asserted, authenticated_session, or server_observed.
4. Collect external evidence without overstating it
A tool server sees its own request and the response it receives. It may not see downstream processing, settlement, recipient delivery, or later reversal. For an R2/R3 call, create a follow-up confirmation event that records a source-system read-back.
{
"state": "confirmed",
"confirmation": {
"rule": "GET /contacts/{external_reference} returns matching canonical email",
"captured_at": "2026-09-28T18:01:26.101Z",
"evidence_source": "ExampleCRM API v2026-06",
"evidence_artifact_digest": "sha256:...",
"result": "supported",
"limitations": ["Provider read-back only; no proof of customer engagement"]
}
}
Preserve the original executed receipt and link the confirmation with parent_receipt_id or a signed event chain. Never overwrite executed with confirmed. If the read-back is absent, delayed, inconsistent, or access is lost, record pending, contradicted, or unobserved; do not manufacture a green outcome.
This is a good place to adopt the review pattern behind Zambo’s receipt verification guidance: inspect the receipt, verify its committed bytes, then read back the relevant state from the system that owns it.
5. Redact for the audience, not for convenience
Receipts often contain exactly the data an auditor wants and an attacker should not receive: access tokens, prompts, personal data, account numbers, internal URLs, or provider response bodies. Build redaction into the model, not as a UI afterthought.
Recommended pattern
- Create a restricted payload under access controls when policy permits it.
- Create a public or shareable payload with explicit field-level status:
present,redacted,omitted, orderived. - Sign the shareable payload and include the digest of the restricted artifact as an opaque reference.
- Record the redaction profile and version, paths removed, and reason category - not secret values.
- Use independent artifact access policy for the restricted body; a receipt URL should not be a bearer token for the body.
Avoid publishing a plain SHA-256 hash of an email address, known account number, or other low-entropy secret: a third party may be able to guess values and compare hashes. If an authorized verifier must validate a hidden field, use an access-controlled artifact, or a documented salted commitment whose verification material is disclosed only to that verifier. Be clear that a public verifier cannot validate the value of a field it cannot see.
6. Design retries as evidence, not as invisible transport behavior
Network timeouts create the most dangerous receipt ambiguity: the server may not know whether the downstream system performed the write. Make idempotency part of the tool contract.
- Require a client idempotency key for R2/R3 writes; scope its uniqueness to the tenant, principal, tool, and target system.
- Store the canonical request digest alongside the key and the first terminal or in-progress receipt.
- On a repeat with the same key and same digest, return the original receipt or continue its tracked confirmation - not a second write.
- On a repeat with the same key and a different digest, reject it and emit an
idempotency_conflictreceipt. - Pass an idempotency key to the downstream provider when it supports one. If it does not, prefer a read-before-write or a manual reconciliation state over an automatic retry after an ambiguous timeout.
A retry is an event worth retaining. Include attempt, retry_of, idempotency_key_digest, dispatch_status, and the reason for retry. The idempotency key is not an approval token; approval binding, authorization, and replay prevention are separate controls.
7. Emit failure and partial receipts
A missing record makes a clean success narrative easy and post-incident debugging difficult. Emit a receipt for every material terminal path, including policy rejection, expired approval, validation failure, provider denial, timeout, partial write, compensation, and verifier failure.
Use a result vocabulary that preserves uncertainty:
| Observed condition | Receipt outcome | Safe statement |
|---|---|---|
| Policy stopped the request before dispatch | rejected |
“No dispatch was attempted by this server.” |
| Input/authorization failed before side effect | failed_pre_dispatch |
“The server did not send the provider request.” |
| Provider returned an error response | failed_provider_reported |
“The provider reported failure; final external state may require read-back.” |
| Connection timed out after dispatch | unknown_after_dispatch |
“The server cannot determine from this attempt whether the provider acted.” |
| Some targets changed | partial |
“Named targets completed; named targets remain failed or unknown.” |
| Undo action succeeded or failed | compensated / compensation_failed |
“A compensating action was attempted; see linked evidence.” |
Sanitize error text, retain a stable internal error code, and commit a digest of the restricted diagnostic artifact. A failed receipt should not expose stack traces or secrets - and an unknown_after_dispatch receipt must not be silently retried as if it were a clean failure.
8. Version the receipt as a verifier contract
Receipts are portable only if a verifier knows what it is verifying. Publish, alongside the server:
- a schema URI and exact version in every payload;
- the canonicalization identifier and hash/signature algorithms;
- a JSON Schema, field semantics, state-transition rules, and redaction profiles;
- canonicalization and verification fixtures, including invalid and failure examples;
- public keys or a trusted key-discovery method,
kidlifecycle, revocation/retirement policy, and test vectors; - a changelog that distinguishes additive metadata from digest-affecting changes.
A workable compatibility rule is:
- Patch: wording or documentation correction; no payload semantic change.
- Minor: optional fields or new optional evidence types; older verifiers may show them as unknown.
- Major: a change to canonicalization, required-field meaning, state semantics, digest coverage, or signature input. Use a new schema URI/version and keep the old verifier available for retained receipts.
Never “helpfully” reinterpret a receipt under an unknown major schema. Return UNKNOWN_SCHEMA and link to the exact verifier needed. Likewise, do not change an existing receipt’s redaction or artifact reference in place; publish a linked correction or redacted derivative that identifies what changed.
9. Build a verifier people can actually use
A correct command-line verifier is necessary but not sufficient. The reviewer needs to make a decision quickly and see uncertainty at the same level of prominence as success.
A good verifier should show, in this order:
- Status and scope: proposed, approved, executed, confirmed, failed, partial, or unknown - plus tool, risk class, tenant/environment, and timestamp.
- What the receipt proves / does not prove: a one- or two-sentence evidence boundary specific to this event.
- Integrity: schema result, canonicalization, payload digest match, signature result, signing key identity/trust state, and key time validity.
- Authority and provenance: actor source, approval link/expiry, policy and tool version, server build, and upstream provider.
- External evidence: capture time, source, rule, result, artifacts, and limitations. A link to issuer-authored text is not independent evidence by itself.
- Privacy: redacted/omitted fields, redaction profile, artifact availability, and retention deadline.
- Reconciliation: idempotency key digest, retry/parent/child receipt links, partial targets, and recommended next action.
For machines, return typed outcomes rather than one overloaded boolean:
{
"schema": "SUPPORTED",
"payload_digest": "VALID",
"signature": "VALID",
"signing_key": "TRUSTED",
"artifacts": "PARTIALLY_REDACTED",
"external_evidence": "PENDING",
"overall": "INTEGRITY_VALID_OUTCOME_UNCONFIRMED"
}
Do not use a single green check for an integrity-valid receipt with missing outcome evidence. Make it possible to download the exact payload bytes, copy the digest, view a readable diff, and open linked evidence under the appropriate authorization. A short live demo followed by receipt verification is a useful way to test whether a receipt is understandable outside your development environment.
Implementation checklist
Before calling a receipt-aware MCP tool production-ready, verify that it can:
- [ ] classify every tool and default unclassified tools to a restrictive policy;
- [ ] expose accurate MCP annotations and an enforceable server-side side-effect contract;
- [ ] create request-bound, expiring approvals for consequential calls;
- [ ] record proposed, approved, executed, and confirmed events without rewriting history;
- [ ] canonicalize a documented payload and verify its digest independently;
- [ ] sign receipts where issuer attribution is required, with documented key trust and rotation;
- [ ] distinguish server observation, provider attestation, source-system read-back, and business outcome;
- [ ] redact secrets safely and state what public verification cannot check;
- [ ] enforce idempotency, reveal ambiguous outcomes, and retain failure/partial receipts;
- [ ] publish schema versions, fixtures, verifier code, and a human-readable verification view.
If you are evaluating an MCP workflow end to end, start with a safe read-only tool from the Zambo MCP tools catalog, then compare the record with the MCP install and verification flow. The same discipline scales to higher-risk tools: make the state, evidence source, and remaining uncertainty visible before you need them in an incident review.
FAQs
1. Does every MCP tool need a digital signature?
No. Every material execution should have an integrity model, but signing is most useful when a receipt leaves the issuer’s trust boundary or a third party needs to attribute it to a signing key. For internal, access-controlled logs, an append-only store with authenticated access may be sufficient. If you sign, publish how verifiers obtain and trust keys; a signature is not meaningful if the key origin is ambiguous.jws
2. Is a provider’s “200 OK” proof that the business action happened?
No. It is evidence that the provider returned that response to the server. Record it as a provider attestation, then define and collect the source-system read-back needed for the particular action. A delivery, settlement, publication, or compliance conclusion may require a different system and a later time.
3. Can we hash raw inputs and outputs instead of storing them?
Yes, when storage or sharing would expose sensitive data. Record which bytes were hashed, the canonicalization method, artifact retention/access policy, and any redactions. A digest alone cannot help a reviewer understand the event, and a raw hash of predictable personal data can leak through guessing. Combine commitments with scoped artifacts and a readable public summary.
4. How should a verifier handle an old receipt after the schema changes?
Select the verifier by the receipt’s declared schema URI/version. Preserve old verification code and fixtures for the retention period. An unsupported major version should produce an explicit UNKNOWN_SCHEMA result, not a best-effort green result under the newest schema.
5. What should we do after a write call times out?
Emit unknown_after_dispatch, preserve the idempotency key and request digest, and reconcile with the target system before retrying. If the provider supports idempotency, query or reuse the provider’s key. If it does not, use a documented read-back or a manual review path; an automatic repeat can turn uncertainty into a duplicate side effect.
References
- Zambo, Verifiable receipts for AI agents.
- Zambo, Independent Verifications Log.
- Zambo, Execution receipt guide.
- Model Context Protocol, Tools specification. Covers tool discovery/calls, optional annotations, the warning that annotations are untrusted unless from a trusted server, and human-in-the-loop interaction guidance.
- Model Context Protocol, Security Best Practices. Covers per-client consent, exact redirect URI validation, least-privilege scope selection, and security controls for MCP connections.
- A. Rundgren et al., RFC 8785: JSON Canonicalization Scheme (JCS). Defines a deterministic JSON representation intended for repeatable cryptographic operations.
- M. Jones et al., RFC 7515: JSON Web Signature (JWS). Defines signatures/MACs over JSON-based content and discusses key-origin authentication and the limits of MACs versus signatures.