What shipped, dated.
The real build history — for a tool about knowing what changed and when, keeping one seems like the least we could do.
L3: drift timeline report + eval-gap suggestions
Drift across many snapshots, and diff-driven coverage nudges: "your change touched the search tool — no eval case exercises it."
Packaging & onboarding
pip install agentsnap, an agentsnap init scaffold, and a documented config schema. Target: first snapshot in under 30 minutes, timed.
Open-core release
The snapshot + diff CLI layer goes public on GitHub. Drift attribution and reporting stay commercial.
v0.2 — probe & drift, demo video
Probe-at-snapshot-time (--with-probe): eval sets run against the live agent, results stored in the manifest. Drift scoring with per-case noise floors — trajectory edit distance, regex assertions, output similarity, all deterministic. One-page HTML drift report with a coverage statement. Plus the One-minute demo, recorded from real output.
Trust kit: secrets redaction at capture
Key-like fields and credential-shaped strings (sk-…, AKIA…, ghp_…, bearer tokens) redacted before hashing, before disk — stored as <redacted>:sha256-… so rotations still diff. Round-trip covered by tests.
v0.1 — snapshot + diff core
Content-hashed config manifests covering model, temperature, system prompt, tool schemas, retrieval settings, and a KB fingerprint. Typed, severity-ranked diffs. CI-friendly exit codes. 23 tests on the load-bearing edge cases.
Want the next entry to be about your stack?
The first adapter gets chosen by the first design partner's infrastructure. That could be a very specific kind of leverage.