Measured product lead custody receipt.
What The Model Saw, Not Just What It Said
A perception-receipt research lane for AI video pipelines · ZPE-Video · GPD multi-hypothesis exploration · github.com/Zer0pa/ZPE-Video
Live experiment. Not a release. Ambition headlined. Claims bounded.
Video AI often tells you what it decided, not what it saw.
ZPE-Video is testing whether a model's visual decision can leave a useful receipt: the object, frame, state, and context that made the decision possible.
The work is still research. Cross-runtime receipts give us a starting signal, but the hard question is ahead: what has to be preserved so a person can audit a machine's view of a moving scene?

An AI processes video — but the right evidence object is still being tested.
What AI saw in the video is becoming testable.
Perception traces from AI video pipelines scatter across Parquet, JSON, pickle, and MCAP containers. The open question is whether a smaller receipt profile can preserve enough state for bounded audit disputes.
zpe-video tests whether detector/tracker state can become a re-derivable evidence object. The current lane uses fixed receipts, manifests, reject vectors, baselines, and GPD branch scoring before any public product claim.
Receipt evidence is tested through bytes, baselines, branches, and blockers.
Cross-runtime bytes are promising, but not public-gate closed.
Deterministic now means a research target: the same detector/tracker input plus the same wire-format spec should produce a byte-identical perception receipt across runtimes. Local Python, Node, and Rust evidence is strong; public macOS/Linux CI is the remaining blocker. The scope is the record, not computer vision truth.
The receipt carries detector and tracker state — boxes, labels, CRCs, manifest binding. It does not prove detector correctness, reconstruct pixels, or establish legal chain of custody. Compression is not the wedge. The next work is public CI, branch scoring, and stronger baseline pressure.
WHAT THE MODEL SAW, as a research question.
The ambition is not to compete with video codecs. It is to test whether detector/tracker state can become a reproducible receipt: enough to answer bounded perception disputes without pretending to reconstruct the whole video.