How to Add Verifiable Execution Receipts to MCP Tools

An MCP tool result is useful to the calling agent. A receipt makes the execution reviewable by someone - or something - outside that conversation.

Try the Zambo demo · Explore Zambo

The goal is not a longer log. It is a portable, bounded record that answers: what tool was requested, under what authority, what the server observed, what bytes are committed, and what remains unproven? That boundary matters especially for tools that write data, spend money, change access, or publish externally.

Zambo’s execution-receipt overview and public verification log illustrate the useful distinction: integrity of a recorded execution is not the same as proof that an outside-world outcome occurred. This guide shows how to build that distinction into an MCP server from the start.

First: know what a receipt is - and is not

These terms solve different problems. Do not use them interchangeably in a UI or API.

Artifact What it is What it can support What it does not establish by itself
Execution receipt A structured record of an attempted or completed tool lifecycle, with commitments to inputs, outputs, provenance, and status That a defined execution record was created; with integrity checks, that its committed fields have not changed That every field is true, that a third party accepted a request, or that the business objective was achieved
Attestation A claim made by an issuer or witness, such as “the provider returned HTTP 202” What that issuer says it observed, scoped to its identity, time, and method Independent truth beyond the issuer’s observation
Digital signature A cryptographic mechanism over bytes or a digest Integrity and, when the verifier trusts the public key and key binding, attribution to the signing key Correctness of the signed statement or an external outcome
Proof of business outcome Evidence from the system that owns the relevant state, evaluated against an explicit acceptance condition For example, “the invoice is marked paid in the ledger at this time” A permanent guarantee; an external state can later change or be reversed

A receipt may contain an attestation, be signed, and link to outcome evidence. It is still important to label those as separate layers. JSON Web Signature (JWS), for example, provides integrity protection for content using signatures or MACs; it does not certify the truth of the content.jws

Use precise status language. “The server recorded an execution attempt” is a receipt claim. “The destination system displayed the requested state at 14:03 UTC” is an external read-back claim. “The customer received the message” needs evidence from a system that can observe delivery or receipt.

1. Classify the tool before you design the receipt

Receipt detail, approval, and retry behavior should follow the harm of a wrong or repeated action - not the tool name. Start with a small, enforceable taxonomy.

Risk class Typical tools Default control Receipt expectation
R0 - read-only, bounded Fetch public price, inspect a repository, calculate a value Allow within scoped access and rate limits Request/result commitments, source/provenance, normal failure record
R1 - reversible or draft write Create a draft, add a label, stage a change Identity plus policy; allow undo window where possible Target, undo/compensation reference, idempotency key, observed result
R2 - durable external write Send a message, update a CRM record, deploy a configuration Human or policy approval bound to the exact request; least-privilege scope Approval, dispatch attempt, external response, read-back state or explicit absence
R3 - financial, access, public, or destructive Transfer funds, grant admin, delete records, publish content Named principal, explicit confirmation, strong authorization, request-bound approval, no blind retry All R2 fields plus policy decision, approver identity, exact scope, outcome evidence and review state

Assess three dimensions together:

  1. Side effect: none, ephemeral, reversible, durable, destructive, or irreversible.
  2. Authority: public data, user data, tenant data, credentials, administrator, or financial authority.
  3. Blast radius: one object, many objects, a public audience, a regulated record, or an account boundary.

Treat unknown behavior as higher risk until tested. A tool that looks like a read can still trigger a provider-side cache warm, paid query, export, or notification.

Declare side effects in the tool contract

MCP tool annotations are useful discovery metadata, not an enforcement boundary. The MCP tools specification supports annotations and warns clients to treat them as untrusted unless they come from a trusted server; it also recommends that a human can deny calls and that clients make invocations visible.mcp-tools Keep the standard hints accurate - readOnlyHint, destructiveHint, idempotentHint, and openWorldHint - then add a server-owned contract that your policy engine actually enforces.

{
  "name": "crm.create_contact",
  "annotations": {
  "readOnlyHint": false,
  "destructiveHint": false,
  "idempotentHint": true,
  "openWorldHint": true
  },
  "x-receipt-policy": {
  "risk_class": "R2",
  "side_effect": "durable_write",
  "approval": "request_bound",
  "idempotency": "required",
  "confirmation": "provider_readback_required",
  "retention_class": "customer_record"
  }
}

x-receipt-policy is an example extension, not an MCP-defined field. Version and document it. The server must derive its decision from the authenticated principal, target tenant, arguments, and live policy - not from an agent-provided risk label.

For R2 and R3 calls, show the user a compact preflight: tool, target, affected objects or amount, credentials/scopes, irreversible effects, and what confirmation will be collected. This supports the MCP security guidance to make consent specific to the client and scopes, rather than treating a broad prior consent as permission for every action.mcp-security

2. Make the lifecycle explicit

A single success: true cannot describe a consequential tool call. Use separate lifecycle states and retain them as events or linked receipts.

proposed ──> approved ──> executed ──> confirmed
  │  │  │  │
  └─rejected  └─expired  ├─failed  └─contradicted
  ├─partial
  └─unknown

A confirmed receipt should name the acceptance condition. For a message, it might be “destination API returns an immutable message ID and a subsequent GET retrieves that ID.” For a transfer, it might be “ledger reports accepted” rather than “settled,” unless settlement evidence is actually present.

Bind approval to the material request

An approval that covers “send the report” is too easy to replay or reinterpret. Bind the approval to a request hash over at least:

If any material field changes, invalidate approval and generate a new proposed receipt. Approval is authorization evidence, not a signature on a vague intent.

proposal = canonicalize({
  tool, tool_contract_digest, tenant, principal,
  normalized_arguments, requested_scopes, risk_class, policy_version
})
request_hash = sha256(proposal)

if approval.request_hash != request_hash or approval.expires_at < now():
  emitReceipt(state="rejected", reason="approval_mismatch_or_expired")
  stop

executeOnce(idempotencyScope(tenant, tool, idempotency_key), proposal)

3. Use a receipt payload and a separate integrity envelope

Do not hash a JSON object that includes its own hash. Create an immutable payload, canonicalize it, then attach an integrity envelope.

{
  "payload": {
  "receipt_schema": "https://example.com/schemas/mcp-receipt/1.0",
  "receipt_id": "rcpt_01J...",
  "event_id": "evt_01J...",
  "parent_receipt_id": "rcpt_01I...",
  "state": "executed",
  "created_at": "2026-09-28T18:01:23.456Z",
  "request": {
  "tool": "crm.create_contact",
  "tool_contract_digest": "sha256:...",
  "request_hash": "sha256:...",
  "idempotency_key_digest": "sha256:...",
  "arguments": {"email": "[redacted]", "list_id": "customers"},
  "redaction_profile": "public-v1"
  },
  "authorization": {
  "principal_ref": "user:opaque-7f...",
  "approval_ref": "apr_01J...",
  "approved_request_hash": "sha256:...",
  "scopes": ["crm.contacts.write"]
  },
  "execution": {
  "attempt": 1,
  "started_at": "2026-09-28T18:01:24.001Z",
  "ended_at": "2026-09-28T18:01:24.842Z",
  "outcome": "provider_accepted",
  "result_artifact_digest": "sha256:...",
  "external_reference": "crm:contact/opaque-91..."
  },
  "provenance": {
  "server_instance": "mcp-prod-eu-2/opaque",
  "server_build": "git:4b6c...",
  "tool_implementation_version": "3.4.1",
  "upstream": {"provider": "ExampleCRM", "api_version": "2026-06"}
  },
  "claims": [
  {"claim": "provider accepted create request", "evidence": "execution.result_artifact_digest"}
  ]
  },
  "integrity": {
  "canonicalization": "RFC8785",
  "hash_algorithm": "SHA-256",
  "payload_digest": "sha256:...",
  "signature": {"format": "JWS", "kid": "receipt-2026-q3", "value": "..."}
  }
}

The example uses references and digests rather than live credentials, personal data, or raw provider bodies. It also separates an observed provider acceptance from a verified business outcome.

Canonicalization and hashes

Different JSON serializers can change whitespace, object-key order, number rendering, and escaping without changing the apparent data. A signature or hash over arbitrary serialization will therefore fail across implementations.

For JSON receipts, adopt a named canonicalization algorithm and publish fixtures. RFC 8785, the JSON Canonicalization Scheme (JCS) defines a deterministic JSON representation for hash and signature use, including deterministic property sorting and constraints around JSON values.jcs

A practical verification sequence is:

  1. Validate the receipt against its declared schema.
  2. Extract only payload; reject duplicate keys, invalid Unicode, non-finite numbers, and undocumented transforms before canonicalization.
  3. Canonicalize payload using the declared algorithm; encode the result as UTF-8 bytes.
  4. Recompute the declared digest and compare it to integrity.payload_digest using a constant-time comparison in security-sensitive code.
  5. If a signature is present, verify it over the documented bytes or digest, using a trusted key identified by kid.
  6. Report integrity, signature/key trust, schema support, and evidence availability as separate results.

A hash makes a changed payload detectable only when the verifier has an expected digest from a trusted receipt or log. A signature adds a key-bound integrity claim, but key distribution and rotation are part of the design. JWS explicitly notes that a verifier must authenticate the origin of a public key; otherwise it does not know who signed the message.jws

Provenance is more than a server name

Record enough provenance to answer “which system observed this?” without turning the receipt into a credential dump:

Do not imply that an unauthenticated user-agent string identifies an actor. Label confidence or source, such as client_asserted, authenticated_session, or server_observed.

4. Collect external evidence without overstating it

A tool server sees its own request and the response it receives. It may not see downstream processing, settlement, recipient delivery, or later reversal. For an R2/R3 call, create a follow-up confirmation event that records a source-system read-back.

{
  "state": "confirmed",
  "confirmation": {
  "rule": "GET /contacts/{external_reference} returns matching canonical email",
  "captured_at": "2026-09-28T18:01:26.101Z",
  "evidence_source": "ExampleCRM API v2026-06",
  "evidence_artifact_digest": "sha256:...",
  "result": "supported",
  "limitations": ["Provider read-back only; no proof of customer engagement"]
  }
}

Preserve the original executed receipt and link the confirmation with parent_receipt_id or a signed event chain. Never overwrite executed with confirmed. If the read-back is absent, delayed, inconsistent, or access is lost, record pending, contradicted, or unobserved; do not manufacture a green outcome.

This is a good place to adopt the review pattern behind Zambo’s receipt verification guidance: inspect the receipt, verify its committed bytes, then read back the relevant state from the system that owns it.

5. Redact for the audience, not for convenience

Receipts often contain exactly the data an auditor wants and an attacker should not receive: access tokens, prompts, personal data, account numbers, internal URLs, or provider response bodies. Build redaction into the model, not as a UI afterthought.

Recommended pattern

Avoid publishing a plain SHA-256 hash of an email address, known account number, or other low-entropy secret: a third party may be able to guess values and compare hashes. If an authorized verifier must validate a hidden field, use an access-controlled artifact, or a documented salted commitment whose verification material is disclosed only to that verifier. Be clear that a public verifier cannot validate the value of a field it cannot see.

6. Design retries as evidence, not as invisible transport behavior

Network timeouts create the most dangerous receipt ambiguity: the server may not know whether the downstream system performed the write. Make idempotency part of the tool contract.

  1. Require a client idempotency key for R2/R3 writes; scope its uniqueness to the tenant, principal, tool, and target system.
  2. Store the canonical request digest alongside the key and the first terminal or in-progress receipt.
  3. On a repeat with the same key and same digest, return the original receipt or continue its tracked confirmation - not a second write.
  4. On a repeat with the same key and a different digest, reject it and emit an idempotency_conflict receipt.
  5. Pass an idempotency key to the downstream provider when it supports one. If it does not, prefer a read-before-write or a manual reconciliation state over an automatic retry after an ambiguous timeout.

A retry is an event worth retaining. Include attempt, retry_of, idempotency_key_digest, dispatch_status, and the reason for retry. The idempotency key is not an approval token; approval binding, authorization, and replay prevention are separate controls.

7. Emit failure and partial receipts

A missing record makes a clean success narrative easy and post-incident debugging difficult. Emit a receipt for every material terminal path, including policy rejection, expired approval, validation failure, provider denial, timeout, partial write, compensation, and verifier failure.

Use a result vocabulary that preserves uncertainty:

Observed condition Receipt outcome Safe statement
Policy stopped the request before dispatch rejected “No dispatch was attempted by this server.”
Input/authorization failed before side effect failed_pre_dispatch “The server did not send the provider request.”
Provider returned an error response failed_provider_reported “The provider reported failure; final external state may require read-back.”
Connection timed out after dispatch unknown_after_dispatch “The server cannot determine from this attempt whether the provider acted.”
Some targets changed partial “Named targets completed; named targets remain failed or unknown.”
Undo action succeeded or failed compensated / compensation_failed “A compensating action was attempted; see linked evidence.”

Sanitize error text, retain a stable internal error code, and commit a digest of the restricted diagnostic artifact. A failed receipt should not expose stack traces or secrets - and an unknown_after_dispatch receipt must not be silently retried as if it were a clean failure.

8. Version the receipt as a verifier contract

Receipts are portable only if a verifier knows what it is verifying. Publish, alongside the server:

A workable compatibility rule is:

Never “helpfully” reinterpret a receipt under an unknown major schema. Return UNKNOWN_SCHEMA and link to the exact verifier needed. Likewise, do not change an existing receipt’s redaction or artifact reference in place; publish a linked correction or redacted derivative that identifies what changed.

9. Build a verifier people can actually use

A correct command-line verifier is necessary but not sufficient. The reviewer needs to make a decision quickly and see uncertainty at the same level of prominence as success.

A good verifier should show, in this order:

  1. Status and scope: proposed, approved, executed, confirmed, failed, partial, or unknown - plus tool, risk class, tenant/environment, and timestamp.
  2. What the receipt proves / does not prove: a one- or two-sentence evidence boundary specific to this event.
  3. Integrity: schema result, canonicalization, payload digest match, signature result, signing key identity/trust state, and key time validity.
  4. Authority and provenance: actor source, approval link/expiry, policy and tool version, server build, and upstream provider.
  5. External evidence: capture time, source, rule, result, artifacts, and limitations. A link to issuer-authored text is not independent evidence by itself.
  6. Privacy: redacted/omitted fields, redaction profile, artifact availability, and retention deadline.
  7. Reconciliation: idempotency key digest, retry/parent/child receipt links, partial targets, and recommended next action.

For machines, return typed outcomes rather than one overloaded boolean:

{
  "schema": "SUPPORTED",
  "payload_digest": "VALID",
  "signature": "VALID",
  "signing_key": "TRUSTED",
  "artifacts": "PARTIALLY_REDACTED",
  "external_evidence": "PENDING",
  "overall": "INTEGRITY_VALID_OUTCOME_UNCONFIRMED"
}

Do not use a single green check for an integrity-valid receipt with missing outcome evidence. Make it possible to download the exact payload bytes, copy the digest, view a readable diff, and open linked evidence under the appropriate authorization. A short live demo followed by receipt verification is a useful way to test whether a receipt is understandable outside your development environment.

Implementation checklist

Before calling a receipt-aware MCP tool production-ready, verify that it can:

If you are evaluating an MCP workflow end to end, start with a safe read-only tool from the Zambo MCP tools catalog, then compare the record with the MCP install and verification flow. The same discipline scales to higher-risk tools: make the state, evidence source, and remaining uncertainty visible before you need them in an incident review.

FAQs

1. Does every MCP tool need a digital signature?

No. Every material execution should have an integrity model, but signing is most useful when a receipt leaves the issuer’s trust boundary or a third party needs to attribute it to a signing key. For internal, access-controlled logs, an append-only store with authenticated access may be sufficient. If you sign, publish how verifiers obtain and trust keys; a signature is not meaningful if the key origin is ambiguous.jws

2. Is a provider’s “200 OK” proof that the business action happened?

No. It is evidence that the provider returned that response to the server. Record it as a provider attestation, then define and collect the source-system read-back needed for the particular action. A delivery, settlement, publication, or compliance conclusion may require a different system and a later time.

3. Can we hash raw inputs and outputs instead of storing them?

Yes, when storage or sharing would expose sensitive data. Record which bytes were hashed, the canonicalization method, artifact retention/access policy, and any redactions. A digest alone cannot help a reviewer understand the event, and a raw hash of predictable personal data can leak through guessing. Combine commitments with scoped artifacts and a readable public summary.

4. How should a verifier handle an old receipt after the schema changes?

Select the verifier by the receipt’s declared schema URI/version. Preserve old verification code and fixtures for the retention period. An unsupported major version should produce an explicit UNKNOWN_SCHEMA result, not a best-effort green result under the newest schema.

5. What should we do after a write call times out?

Emit unknown_after_dispatch, preserve the idempotency key and request digest, and reconcile with the target system before retrying. If the provider supports idempotency, query or reuse the provider’s key. If it does not, use a documented read-back or a manual review path; an automatic repeat can turn uncertainty into a duplicate side effect.

References

  1. Model Context Protocol, Tools specification. Covers tool discovery/calls, optional annotations, the warning that annotations are untrusted unless from a trusted server, and human-in-the-loop interaction guidance.
  2. Model Context Protocol, Security Best Practices. Covers per-client consent, exact redirect URI validation, least-privilege scope selection, and security controls for MCP connections.
  3. A. Rundgren et al., RFC 8785: JSON Canonicalization Scheme (JCS). Defines a deterministic JSON representation intended for repeatable cryptographic operations.
  4. M. Jones et al., RFC 7515: JSON Web Signature (JWS). Defines signatures/MACs over JSON-based content and discusses key-origin authentication and the limits of MACs versus signatures.