AI Agent Receipts for CI/CD Pipelines
Last updated: 2026-09-28
A pipeline receipt should answer a narrow question for each step: what operation ran, under which scope, at what time, and what result did the execution layer observe? Linking those records helps a reviewer investigate a build or deploy without treating an agent summary as proof.
Try the Zambo demo · Explore Zambo
Use the public receipt verifier to inspect a receipt when you have one. A receipt is evidence about a recorded execution, not a guarantee that an outside system accepted the result or that a human decision was correct.
Receipts per pipeline step
A pipeline is a sequence, not one indivisible action. Checkout, dependency inspection, test execution, artifact creation, approval, deployment, and post-deploy checks can each have different permissions and failure modes. Record a receipt at the boundary where the agent actually invokes a tool or observes a result. Keep the pipeline run identifier outside the receipt if that is the system's existing convention, then store the receipt id as a reference. This lets the pipeline UI remain familiar while giving an incident reviewer a stable evidence trail.
Build evidence
For a build step, the receipt can identify the tool, version, caller scope, timestamp, and observed status. An output hash can commit to a result or canonical response without publishing secrets. The build system still owns source revision, dependency lock state, runner identity, and artifact storage. A receipt should link or reference those records rather than pretending to contain them. If a build failed before a receipt was created, record the missing event as a collection gap. Do not create a success receipt from the agent's planned command.
Tests and interpretation
A test command's receipt can show that the command executed and what output the recorder observed. It does not prove that the tests were complete, that the assertions were well designed, or that the environment matched production. A reviewer should compare the receipt with the test selection, source revision, dependency state, and runner configuration. If an agent summarizes all tests passed while the receipt covers only a subset, the summary is too broad. The evidence should narrow the claim, not expand it.
Deploy verification
Deployment has at least two distinct claims. One is that a deployment operation ran. The other is that the target environment served the intended version and remained healthy. An execution receipt can support the first claim and may support the second if the execution layer directly observed a health check. It cannot prove an outside system remained healthy after the observation window. Pair the receipt with a target version, health check record, rollback decision, and monitoring link. A failed or delayed check should remain visible.
Failed deploy forensics
During an incident, preserve failed receipts instead of keeping only the final successful story. The sequence of timestamps, tool scopes, observed errors, and output commitments can show where a pipeline diverged. Redact credentials and confidential payloads, but retain enough context to distinguish a timeout, rejected permission, invalid artifact, and target-side failure. If a transport error truncated the result, label it as a transport uncertainty and retry the retrieval. Do not convert an unavailable record into a fabricated failure reason.
Incident-response handoff
An incident handoff benefits from a compact index of receipt ids, pipeline run identifiers, source revisions, target environments, and reviewer notes. The index is a navigation aid, not a replacement for the individual records. The incoming responder should verify the highest-impact steps first: the version that was deployed, the permission used, the first failing check, and the rollback action. A receipt can establish that a recorded call occurred. It cannot establish that the incident was fully contained without system evidence.
Permissions and least privilege
A receipt's caller scope makes a useful review prompt. Ask whether the pipeline step had the minimum permission needed and whether the permission was expected for that stage. Keep authorization decisions in the policy system. The receipt should not expose bearer tokens, private keys, or complete secret values. A successful call is not evidence that the permission was appropriate. Review scope separately from execution status, especially when an agent can choose tools dynamically.
Reproducibility
Receipts make a run more inspectable, but they do not make it automatically reproducible. Reproduction depends on source revision, dependencies, inputs, environment, clocks, network responses, and external state. Preserve references to those controlling systems. When a step is cheap and safe to rerun, do so and record the new observation separately. Never rewrite an old receipt to make it match a later rerun. Historical evidence should remain historical.
A minimal pipeline pattern
At pipeline start, create a run record. For each agent action, store the receipt reference and the pipeline step name. After each critical step, perform an independent check against the system that owns the result. At completion, generate a summary that links every material claim to one or more receipt ids and marks unobserved claims. During incident review, open the receipts through /verify, then compare the records with deployment and monitoring evidence.
Adoption without disruption
Start with one high-risk action such as production deployment or permission-changing configuration. Instrument the invocation boundary, not every internal function. Define the claim the receipt supports, the retention period chosen by the pipeline owner, and the reviewer who handles gaps. Once the workflow is useful, add build and test steps. The goal is not more logs for their own sake. The goal is a short path from an agent's claim to evidence a reviewer can inspect.
Review notes
Good evidence work starts with a bounded question. Name the action, the system that owns the result, the time window, and the person or service responsible for reviewing it. Then separate three states that are often collapsed into one: an action was requested, an execution layer recorded an action, and an outside system confirmed an outcome. The first state is intent. The second is execution evidence. The third is outcome evidence. A page that keeps these states separate is easier to use during an incident and harder to misuse in a status summary.
When evidence is incomplete, record the gap instead of filling it with confidence. A missing record may be a collection failure, a transport problem, a retention issue, or an action that never ran. Those possibilities have different owners and different remedies. Preserve the original record, the verification attempt, and the reviewer conclusion. If a later retry succeeds, keep the earlier attempt visible. This makes the evidence useful to someone who was not present when the work happened and gives the next operator a concrete place to start.
Reviewers should also write down assumptions. State which timestamp is authoritative, whether a result was observed directly, how a hash was computed, and which system owns the final state. If the answer depends on an external service, link the external evidence or mark it unavailable. If the action involved private data, describe the field class without copying the data into a public page. These habits turn a receipt from a decorative status badge into a durable review artifact. They also make disagreements productive: two reviewers can compare assumptions instead of arguing over a summary.
Finally, keep the evidence proportional to the decision. A low-risk read can use a light check, while a money movement, data mutation, deployment, or external communication deserves a stronger receipt and an independent state check. Do not use a larger word count or a more elaborate format to imply certainty. The relevant question is always what was recorded, what was verified, and what remains outside the record.
Frequently asked questions
Does a pipeline receipt prove a deploy succeeded?
It can record that a deploy call ran and what it observed. Target state and continued health need independent checks.
Should every internal function create a receipt?
Usually no. Start at meaningful tool and system boundaries so the evidence remains reviewable.
What belongs in an incident handoff?
Receipt ids, pipeline and source references, timestamps, observed failures, target environment, and explicit evidence gaps.
Can receipts replace pipeline logs?
No. Receipts provide a focused verification unit. Detailed logs remain useful for debugging and context.
Where can the verifier be opened?
Use https://zambo.dev/verify for a public receipt reference.