Agent Execution Receipts for LangChain, CrewAI, and OpenAI Agents
A handoff fails when the only proof of an agent’s work is trapped in the agent that produced it: a callback log, a task object, a trace viewer, or a pasted summary. The next developer, reviewer, or client needs a durable answer to simpler questions:
Try the Zambo demo · Explore Zambo
- What execution step was recorded?
- What did it return, and when?
- Which workflow, task, or tool attempt produced it?
- Can I inspect and verify the record without logging into the original agent UI?
An execution receipt is the portable evidence record for that handoff. It complements traces, logs, and metrics; it does not replace them. A receipt packages one execution step with its request context, observed result, timestamp, and an integrity commitment so someone else can inspect the recorded event later.receipt-evidence
Evidence boundary: a receipt can verify the record a service stored and observed. It does not prove an outside-world outcome the service did not observe - such as whether an email was received, a payment settled, or a downstream database accepted a write. Record a provider confirmation or a read-back when that distinction matters.
This guide is framework-neutral in its design rules and specific about how the current Zambo integrations fit those rules. Start with the execution-receipt overview if your team needs a shared definition, or go to receipt verification when you already have a receipt URL.
The portable receipt contract
Treat the receipt as an artifact with two identities:
- An internal, stable record in your system of record.
- A shareable reference that a reviewer can resolve independently, such as a public receipt URL and its receipt ID.
The smallest useful linkage is:
application_run_id → framework_run_or_task_id → work_attempt_id → receipt_id + receipt_url
Do not keep only the URL in ephemeral process memory or a chat transcript. Persist it alongside the workflow record before returning a final “complete” status to a human-facing surface.
Metadata to expose - and metadata to keep private
A receipt should be easy to reconcile with a workflow without becoming a data dump. Use an explicit allowlist.
| Field | Why a reviewer needs it | Exposure guidance |
|---|---|---|
receipt_id, receipt_url |
Durable lookup and independent review | Public only if the underlying record is safe to disclose. |
application_run_id, workflow_id |
Reconcile the receipt to your product’s job | Usually safe as opaque IDs; avoid customer-identifying IDs. |
framework, framework_event, integration_version |
Explain which event boundary produced the record | Expose. This is critical when a callback can produce several receipts. |
tool_name or task_name |
Establish the claimed operation | Expose a human-readable, non-sensitive label. |
attempt_number, parent_attempt_id |
Distinguish a retry from an original execution | Expose. Never overwrite the original attempt. |
started_at, completed_at, recorded_at |
Support ordering and freshness checks | Use UTC and label which clock each timestamp represents. |
status |
Separate work completion from receipt publication | Use separate values, such as work_succeeded and receipt_published. |
| Input summary and result summary | Give the reviewer enough context to assess the event | Redact secrets, personal data, credentials, and confidential business content before receipt emission. |
operation_id / idempotency key |
Connect a side effect to a retry-safe business operation | Keep the raw secret private; expose a non-sensitive reference or hash if useful. |
| External confirmation / read-back | State whether an outside effect was observed | Expose the observation and its source; do not imply it is guaranteed by the receipt itself. |
| Canonical bytes, hash algorithm, digest, verifier status | Enable independent integrity checking | Expose when the receipt is intended to be independently verified. |
The public record should state both what was observed and what was not observed. That is more useful than a generic “verified” badge.
Where to emit a receipt
The right emission point is the boundary a reviewer needs to evaluate. In practice, there are three common choices.
| Boundary | Emit when | Best for | Main risk | Mitigation |
|---|---|---|---|---|
| Lifecycle event | A tool or model starts or ends | Fine-grained debugging and tool-level evidence | Several receipts can represent one logical run | Store event type, phase, parent run, and sequence number. |
| Task completion | A bounded task has a reported result | Reviewable work packages and client handoffs | It can hide intermediate tool actions | Link the task receipt to child tool receipts or the trace. |
| Side-effect confirmation | A write has a provider response or read-back | Consequential actions such as publishing, updating, or sending | “Tool returned success” may not mean the outside effect occurred | Capture provider confirmation and, where possible, a read-back in the evidence. |
A practical default is to emit:
- a start record only when intent and timing matter;
- an end record for the observed result of each meaningful tool call; and
- a task- or run-completion record for the handoff a human is likely to read.
That produces a useful hierarchy: detailed evidence for engineering, plus a concise, reviewable artifact for the recipient. Do not emit a receipt for every token or internal thought. More events are not automatically more evidence.
Three integration patterns at a glance
| Framework | Natural emission boundary | Current Zambo pattern | What to persist | Retry / duplication consideration | Best handoff use |
|---|---|---|---|---|---|
| LangChain | Lifecycle callbacks for model, tool, or chain events | ZamboCallbackHandler records supported lifecycle events; a tool execution may yield a start receipt and an end receipt.langchain |
All receipt URLs, event type, LangChain run ID, parent run ID, and sequence | last_receipt_url is only the latest success; retain the URL list for batch or nested runs. |
Explain a specific tool result and retain a path back to a broader trace. |
| CrewAI | Explicit task-completion callback | ZamboCrewListener publishes a completed task event when your completion hook calls it.crewai |
Task ID, task description/result summary, attempt number, listener URL list, and receipt errors | The listener does not promise retry deduplication; repeated callback calls can create separate records. | Hand off one completed task with a narrow, readable result. |
| OpenAI Agents SDK | Run hooks around tool start/end, plus the tool wrapper’s returned URLs | ZamboRunHooks records tool_start and tool_end; zambo_function_tool() tracks returned receipt URLs.openai-agents |
Agent/run ID, tool-call ID, hook phase, tool wrapper URLs, and receipt errors | A single agent run can legitimately have multiple tool-event receipts; preserve order and parentage. | Review tool use inside a multi-tool agent run without relying on the original run UI. |
The framework determines where events are available. Your application determines the evidence model: which events count, what safe metadata they contain, which parent-child links exist, and what a reviewer can conclude.
Pattern 1: LangChain lifecycle receipts
LangChain is well suited to an event-oriented receipt model. Its current Zambo callback supports model, tool, and chain lifecycle events, and a completed event can return a public receipt URL.langchain That is valuable when the handoff asks, “What did this tool return?” rather than “Did the whole conversation look reasonable?”
Use it deliberately
- Attach the callback to the runnable or tool that actually executes - not merely to an object created elsewhere.
- Decide whether tool-start events are necessary. A start record establishes invocation context, but it cannot contain the eventual result.
- On tool end, persist every returned receipt URL with the framework run ID and event phase.
- On chain end, optionally produce a higher-level receipt that summarizes the handoff result, with links to relevant child receipts or traces.
- If the application retries a tool, preserve a new attempt record rather than replacing the prior one.
The published integration notes that a single tool invocation can create separate start and end receipts, and a chain may create another completion record.langchain That is why a single latest_receipt_url field is not an adequate audit model. It is a convenience for a simple UI, not a stable ledger for a run tree.
Good review surface: “Website-audit tool ended at 14:31 UTC; attempt 2 returned 12 findings; receipt R-…; parent workflow W-….”
Weak review surface: “Agent run complete” with one unlabelled link.
Read the LangChain receipt integration for the current adapter details and the receipt evidence guide for the distinction between an event record and a trace.
Pattern 2: CrewAI task-completion receipts
CrewAI’s task callback is a natural handoff boundary. The current Zambo integration is intentionally explicit: its listener publishes when the application calls the completion method from a CrewAI completion hook; constructing the listener alone does not subscribe it to every event.crewai
That explicitness is a feature. It forces you to choose the exact task whose output a reviewer should read.
Use it deliberately
- Create one completion callback per task that warrants a receipt.
- Build a receipt payload from the completed task’s safe description, result, and workflow correlation fields - not the entire crew transcript.
- Persist the receipt URL and task attempt in the same transaction or durable update that marks the task terminal.
- If publishing the receipt fails, keep
task_status = completedand setreceipt_status = pendingorfailed; do not report a successful receipt that does not exist. - When the task retries, create a new attempt. Decide which attempt is the authoritative handoff only after the task’s retry policy finishes.
The CrewAI adapter catches receipt-publishing exceptions so a receipt network problem does not replace the completed task result; it also records receipt errors separately.crewai Treat that behavior as a reminder that work success and evidence-publication success are separate states.
For a client-facing task, make the human-facing summary include the receipt URL, the task version, the completion time, and an evidence boundary. Example:
“Task
release-summary, attempt 1, completed at 14:31 UTC. The receipt records the submitted description and returned summary. It does not establish that a release was approved or sent.”
See the CrewAI receipt integration for the current listener behavior.
Pattern 3: OpenAI Agents SDK run-hook and tool-wrapper receipts
In the OpenAI Agents SDK, a run may call several tools. The right pattern is usually tool-level receipts with run-level correlation, not one undifferentiated receipt for the final prose answer.
Zambo’s published OpenAI Agents SDK integration combines ZamboRunHooks, which records tool_start and tool_end, with zambo_function_tool(), which tracks returned receipt URLs.openai-agents The division maps cleanly to a general design:
- Use hooks to capture the framework lifecycle boundary.
- Use the tool layer to preserve the receipt returned for a concrete operation.
- Use your application store to join both to the agent run, tool-call ID, and retry attempt.
Use it deliberately
- Assign an application-owned
workflow_idbefore starting the agent run. - At each tool start/end, record the agent run ID, tool-call ID, phase, and attempt number.
- Preserve each returned receipt URL as an append-only child of the run; do not collapse all tool calls into one string field.
- When a final response is shown to a human, list the receipt links that substantiate the key tool-derived claims.
- If no relevant tool completed, say so. A fluent final answer is not a substitute for an execution record.
This pattern is especially useful for a multi-tool research or delivery agent. A reviewer can follow a concise final handoff, then open the evidence for the specific tool call that matters, without re-running the agent or entering the original application.
For package and hook details, use the verified OpenAI Agents SDK receipt integration.
Preserve the receipt URL and ID as first-class run data
A receipt must survive process restarts, framework upgrades, queue redelivery, and a handoff to another system. The following is conceptual pseudocode, not an API for any framework:
on_observable_work_event(event):
# Establish a stable app-owned identity before publishing evidence.
attempt = ledger.get_or_create_attempt(
workflow_id = event.workflow_id,
logical_step_id = event.logical_step_id,
attempt_number = event.attempt_number,
event_phase = event.phase
)
safe_evidence = redact_and_allowlist({
framework: event.framework,
framework_run_id: event.run_id,
parent_run_id: event.parent_run_id,
tool_or_task: event.name,
input_summary: event.safe_input_summary,
observed_result: event.safe_result,
timestamps: event.timestamps,
operation_ref: event.nonsecret_operation_ref
})
# Persist the payload before a network call so receipt publication can retry
# without rerunning the underlying business work.
ledger.store_pending_evidence(attempt.id, safe_evidence)
receipt = publish_receipt(safe_evidence)
if receipt.created:
ledger.append_receipt(attempt.id, {
receipt_id: receipt.id,
receipt_url: receipt.url,
verifier_url: receipt.verifier_url,
integrity: receipt.integrity,
published_at: now_utc()
})
ledger.set_receipt_status(attempt.id, "published")
else:
ledger.set_receipt_status(attempt.id, "pending")
queue_receipt_publish_retry(attempt.id)
There are two important design choices here:
- Persist before publish. If your receipt provider is temporarily unavailable after work completes, you can retry publication from the saved, redacted evidence payload.
- Append, do not overwrite. A logical task may have a start receipt, an end receipt, an external-confirmation receipt, and a final handoff receipt. Each needs its own ID and role.
If your database has a receipt table, enforce a uniqueness constraint appropriate to your policy, such as (attempt_id, receipt_role, provider) - not merely (workflow_id). One workflow legitimately produces many receipts.
Retries and partial failure: model two independent outcomes
Receipt systems become misleading when they treat “the tool finished” and “a receipt is available” as the same boolean. Use a state model that makes partial failure visible.
| Work status | Receipt status | Meaning | Recommended behavior |
|---|---|---|---|
succeeded |
published |
Work completed and an evidence link is available | Show the link with the relevant evidence boundary. |
succeeded |
pending |
Work completed; receipt publication is retrying | Show “work complete; receipt pending.” Retry only publication. |
succeeded |
failed |
Work completed; receipt could not be created after policy limit | Preserve the error and internal evidence; do not fabricate a link. Escalate if receipt is a delivery requirement. |
failed |
published |
The failed attempt was recorded | Keep it. A failure receipt is useful evidence for debugging and reconciliation. |
unknown / timeout |
published |
The endpoint observed something, but final work outcome is not established | Mark the uncertainty; use a read-back or provider query before asserting success. |
Retry rules that prevent duplicate evidence and duplicate side effects
- Give the business operation its own idempotency key. A receipt ID is not a substitute for an idempotency key on a payment, email, database write, or other side effect.
- Separate “retry work” from “retry receipt publication.” Once the work result is durably known, a receipt retry should reuse the stored evidence payload rather than invoke the tool again.
- Record the attempt chain. Include
attempt_number,prior_attempt_id, retry reason, and the policy outcome. A reviewer should be able to see whether attempt 2 superseded attempt 1. - Do not silently deduplicate distinct event phases. A start event and an end event may have the same logical operation but are different records.
- Require a new confirmation after an uncertain write. A timeout after a provider call is neither a confirmed failure nor a confirmed success. Query the provider or read the resource before retrying a non-idempotent action.
For high-consequence workflows, bind approval to the exact operation or request hash and invalidate it if material inputs change. The MCP security best practices are a useful companion for authorization and tool-risk design; a receipt should document the observed step, not stand in for an approval control.
Let a human verify without the original agent UI
A shareable receipt is useful only if a reviewer can verify it from an ordinary browser and, where needed, reproduce the integrity check independently.
For a Zambo receipt, the published framework integrations describe a public /run/{UUID} receipt page and a verifier response at /api/receipt/{UUID}/verify. The verifier provides the hash algorithm, output hash, canonical bytes, byte length, status, and checks; the canonical bytes are Base64-encoded in the current response and should be decoded before computing SHA-256.langchaincrewai
Human verification checklist
- Open the saved receipt URL in a clean browser session. The reviewer should not need the framework dashboard, your agent’s memory, or an account in the original app.
- Match identity. Confirm the receipt ID in the URL is the ID saved in the handoff record.
- Match context. Check the tool/task label, event phase, safe input summary, observed result, and UTC timestamp against the expected workflow and attempt.
- Inspect the boundary. Determine whether the record shows a tool return, a task completion, or an external provider confirmation. Do not infer more.
- Check integrity. Use the linked verification response. Decode the supplied canonical bytes, calculate the stated digest, and compare it with the stored output hash.
- Record the review. Save the reviewer, review time, conclusion, and any exception in your own workflow system. The receipt verifies its record; your review log records the decision made from it.
A minimal independent verifier looks like this:
# Conceptual verifier logic; use the provider's documented response shape.
verification = GET(saved_receipt.verifier_url)
canonical_bytes = base64_decode(verification.canonical_bytes)
computed_hash = sha256(canonical_bytes)
assert computed_hash == verification.output_hash
assert verification.status indicates a valid record
Hash equality establishes that the canonical record returned by the verifier matches its stated digest. It does not establish the truth of a model’s assertion or an external effect not contained in the evidence. Zambo’s receipt documentation makes the same distinction between a portable execution record and broader observability or outcome verification.
A handoff-ready delivery template
Put this short block in the project record, ticket, client update, or operator console - not only in a framework trace:
Work item: <human-readable task/tool>
Workflow / attempt: <workflow_id> / <attempt_number>
Observed status: <succeeded | failed | uncertain>
Receipt status: <published | pending | failed>
Receipt: <receipt_url>
Recorded at: <UTC timestamp>
What it shows: <specific observed tool/task result>
What it does not show: <unobserved external outcome, if any>
Follow-up: <read-back, approval, or escalation required>
That template makes the receipt usable when ownership changes. The receiving person can decide whether the record is sufficient, whether an outside confirmation is still required, and where to find the evidence without reconstructing the original agent execution.
Implementation checklist
Before calling a framework integration complete, verify that your application can answer all of these:
- [ ] Which event boundary creates the receipt, and why is it the right review boundary?
- [ ] Are receipt URLs and IDs persisted with application-owned workflow, step, and attempt IDs?
- [ ] Can one workflow retain multiple ordered receipts without overwriting earlier ones?
- [ ] Are work completion and receipt-publication status represented independently?
- [ ] Does a receipt-publication retry avoid rerunning a completed side effect?
- [ ] Are public payloads allowlisted and redacted before emission?
- [ ] Can a reviewer open and integrity-check the receipt without the original framework UI?
- [ ] Does the handoff state what the receipt proves - and what it does not prove?
To start with a working integration, choose the LangChain, CrewAI, or OpenAI Agents SDK guide. For a broader connection path, see Install Zambo MCP and the framework integrations index.
FAQs
1. Should every agent event get a receipt?
No. Emit receipts for boundaries another person must inspect: a meaningful tool result, a completed task, a consequential side effect with confirmation, or a final handoff. Keep traces and logs for high-volume diagnostic events. A receipt should make a specific claim reviewable, not create unreadable evidence noise.
2. Is a receipt URL enough to correlate work across systems?
No. Save the URL and receipt ID with your application-owned workflow ID, logical step ID, framework run/task ID, event phase, and attempt number. The URL lets a human inspect the evidence; your internal IDs make it reconcilable after queues, retries, and framework changes.
3. What should happen if the task succeeds but receipt publication fails?
Report two statuses: the task succeeded, and the receipt is pending or failed. Persist the redacted evidence payload first, retry only the receipt publication, and do not rerun a completed side effect just to obtain a receipt. If evidence delivery is contractually required, escalate rather than presenting the task as fully handoff-ready.
4. Does a valid SHA-256 check prove that an email, payment, or database write happened?
No. It verifies the integrity of the recorded canonical receipt. To support an outside effect, include an observed provider response, transaction reference, or read-back in the evidence, then state precisely what that observation establishes. The receipt still does not prove facts the endpoint did not observe.receipt-evidence
5. How can a non-technical reviewer verify a receipt without our agent dashboard?
Give them the receipt URL, the expected receipt ID, the work-item context, and the evidence boundary. They can open the public record, compare its event/result/timestamp to the handoff, and follow the verifier link. For a higher-assurance review, a developer can independently recompute the digest from the documented canonical bytes - without access to the original LangChain, CrewAI, or OpenAI Agents SDK run.
Sources and integration references
- Zambo, “Receipts as execution evidence”. Defines receipts as stable records for an execution step; distinguishes them from traces, logs, metrics, and unobserved outside outcomes.
- Zambo, “LangChain verifiable receipts”. Documents the callback lifecycle events, URL handling, public receipt page, and verification response described in this guide.
- Zambo, “CrewAI verifiable receipts”. Documents explicit task-completion publishing, retry/duplicate considerations, separate receipt errors, and independent verification.
- Zambo, “OpenAI Agents SDK verifiable receipts”. Documents
ZamboRunHooks, tool start/end recording, the function-tool wrapper, and receipt URL tracking.