# Preservation Audit Audit date: 2026-08-25. ## Preservation Scope This audit covers the local repository metadata and working tree, the configured Git remote, the Spark checkout inspected through `ssh spark`, local top-level benchmark/profiler artifacts plus the packaged SageAttention wheel, Spark's selected physical research artifacts, and SHA-256 reconciliation of top-level benchmark JSON filenames. It does not claim to preserve model checkpoints, mounted parity tensors, generated media, container layers, external package repositories, or every file under the broader Spark output tree. No commit, tag, branch, archive, upload, or copy was created by this audit. The audit records are this document, `research/REPRODUCIBILITY.md`, and the physical file inventory in `research/artifact_manifest.json`. ## Git Topology And Commit Context - Current local branch: `fl2va-direct-runtime`. - Local `HEAD`: `bee998d8403aea18f9636b3773a8e7719ff3817c` (`Record RTX Pro 6000 multi-GPU scaling`, committed 2026-08-22 16:42:09 +0700). - `origin/fl2va-direct-runtime` and `origin/master` both resolve to `bee998d8403aea18f9636b3773a8e7719ff3817c`; `git ls-remote` confirms the same two server heads. - Local `master` remains at `ea4f4a2dadbd9754bcadfcfbe9a789f773cf838a`, which is the merge base with `fl2va-direct-runtime`; the feature branch is 76 commits ahead and zero behind that local branch. - `fl2va-direct-runtime` was created from `8730920634c3f4572353222e4ad307f01862fa16` on 2026-08-19 and then advanced linearly to `bee998d`. - The other advertised head is `vae-decode-optimization` at `8730920634c3f4572353222e4ad307f01862fa16`. - There are no tags locally or on the configured remote. - Local reflogs contain the visible branch history but no deleted branch tip. Local `git fsck --full --no-reflogs --unreachable` finds only the empty blob `e69de29bb2d1d6434b8b29ae775ad8c2e48c5391`; it finds no useful dangling or unreachable commit. There is therefore no recoverable deleted branch history in the audited local object database. - The Spark checkout is older: branch `fl2va-direct-runtime` and its tracking ref are at `9f62f6ea838150967422fe648ee25b8b2491356d`. Spark also has a local preservation branch `preserve/spark-parity-diagnostics-2026-08-13` at `f1ddc416f732c3de3a9f9de92eec3724d7f65ebc`, based on `a8df9ef1730b9cf30f9341fd1298214a3b406dce`. It is historical parity context, not a preservation point for the current optimization tree. The exact former P1 dispatch source and some other removed experimental source cannot be recovered from Git. They were never present in the reachable commit history, reflogs do not expose a deleted tip containing them, and `fsck` exposes no useful dangling commit. Current summaries can preserve decisions and measured results, but they cannot reconstruct those exact source states. ## Dirty-Worktree Risk The local tree has extensive tracked modifications and untracked source, documentation, tests, tools, benchmark JSONs, and profiler reports relative to `bee998d`. The post-commit research narrative and implementation are therefore not protected by the remote heads. A clone of the remote reproduces only the committed baseline. The Spark checkout is independently dirty and differs from both local `HEAD` and the local working tree. It contains modified runtime source and many untracked research artifacts while remaining at `9f62f6e`. Neither dirty tree is a backup of the other. Editing, cleaning, container rebuilds, output rotation, checkout replacement, disk loss, or an accidental `git clean` can permanently remove unique evidence. Git cannot restore untracked profiler binaries or source that was never committed. The risk is immediate rather than theoretical: the expected Spark production container exists, but container state and Spark's mutable output directories must not be treated as durable preservation. Final validation temporarily started that container, confirmed `/ready` with Sage2 as both the initial and current backend, and then stopped it again to leave the research environment quiescent. ## Local Inventory The local inventory is byte-exact as observed: | Class | Location | Count | Bytes | | --- | --- | ---: | ---: | | Benchmark JSON | top-level `benchmarks/*.json` | 148 | 1,175,484 | | Nsight Compute reports | top-level `benchmarks/*.ncu-rep` | 14 | 199,677,223 | | Nsight Systems reports | top-level `benchmarks/*.nsys-rep` | 2 | 7,686,256 | | Nsight CSV exports | top-level `benchmarks/*.csv` | 2 | 169,373 | | SageAttention wheel | `wheels/sageattn3-1.0.0-cp312-cp312-linux_aarch64.whl` | 1 | 1,680,352 | | Selected local physical artifacts | all rows above | 167 | 210,388,688 | All 167 selected files have one path, byte size, and SHA-256 record in `research/artifact_manifest.json`; the manifest totals reproduce this table. The two CSV exports and wheel explain why the selected total is larger than the JSON plus `.ncu-rep` plus `.nsys-rep` subtotal. The unrelated `benchmarks/rtxpro6000-server-2gpu-topology.txt` and research source files are outside this physical-artifact total, although they remain within the broader preservation scope. ## Spark Inventory The audit originally retained a selected Spark aggregate of 279 physical artifacts totaling 665,950,155 bytes. Within it, the top-level output JSON set contains exactly 215 files totaling 2,259,709 bytes at `/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/benchmarks`. The selected 279-file figure is a scoped preservation aggregate, not an unqualified recursive count of the output directory. Re-enumeration found 280 current top-level files totaling 547,708,583 bytes, including the exact 215-file JSON set, while the recursive directory has 291 JSONs because older/nested result collections are also present. The original 279 paths were not retained, so the 665,950,155-byte aggregate cannot be reproduced or honestly assigned to physical files. `research/artifact_manifest.json` therefore inventories the complete reproducible 280-file top-level set and records the discrepancy explicitly: one additional observed file and 118,241,572 fewer bytes than the retained aggregate. It must not be presented as a reconstruction of the lost 279-path selection. Spark also has the same 1,680,352-byte SageAttention ARM64 wheel in the checkout. The current top-level output includes large `.ncu-rep`, `.nsys-rep`, `.sqlite`, `.qdstrm`, CSV, tensor, media, log, and report files. Presence on Spark is not a second durable copy unless the artifact is also independently stored and checksum-verified. ## JSON Reconciliation Reconciliation uses basename identity and SHA-256 over local top-level `benchmarks/*.json` and Spark top-level output `*.json`: | Result | Count | | --- | ---: | | Identical filename and SHA-256 | 106 | | Same filename, different SHA-256 | 2 | | Local only | 40 | | Spark only | 107 | | Local total | 148 | | Spark total | 215 | The two hash mismatches are: - `gb10-cute-p0-attn-qkv-128rows.json` - `gb10-fully-fused-one-step-fresh.json` The arithmetic reconciles exactly: 106 identical plus 2 mismatches plus 40 local-only equals 148 local files; 106 plus 2 plus 107 Spark-only equals 215 Spark files. Same-name mismatches must be preserved as two versions, not resolved by timestamp or overwritten. Local-only and Spark-only files are unique evidence until copied into an immutable, checksummed collection. ## Large Artifact Policy Git is appropriate for source, patches, small JSON summaries, exact commands, manifests, and checksums. Raw `.ncu-rep`, `.nsys-rep`, `.sqlite`, `.qdstrm`, tensor captures, media, and wheels should not be added casually to ordinary Git history. They are large, tool/version-sensitive, and can make every clone carry research payload indefinitely. For each large-artifact preservation set: 1. Freeze an explicit relative-path manifest with byte size and SHA-256 for every file, including both versions of hash-mismatched JSONs. 2. Record producer revision or working-tree archive hash, exact command and environment where known, host/GPU facts, profiler version, and linked summary JSON. 3. Store the payload in immutable or versioned object storage, an archival share, or a dedicated artifact release. Use Git LFS only if its remote retention and clone policy are intentionally managed. 4. Keep at least two independently administered copies and verify restoration by checksum. A mutable Spark output directory and its checkout count as one failure domain, not two. 5. Do not delete raw profiler databases after exporting CSV until the export and summary are verified and the raw file is archived. 6. Never replace same-name mismatches in place. Preserve origin-qualified paths such as `local/` and `spark/` in the archive. ## Recoverability Conclusions - `bee998d` and all reachable committed ancestors are recoverable from the configured remote. - The present local optimization source, tests, documents, and many artifacts are not recoverable from that commit or remote because they are dirty or untracked. - Spark provides additional unique evidence but is older, dirty, mutable, and not an exact working-tree replica. - The 106 identical JSONs have two observed copies; the 40 local-only and 107 Spark-only JSONs do not. The two mismatches represent four distinct payloads. - Raw local profiler evidence is concentrated in 16 reports totaling 207,363,479 bytes; Git cannot recreate it. - No tag identifies the research state, no deleted branch tip or useful dangling commit was found, and the exact former P1 dispatch source plus other removed source cannot be recovered from Git. - The preserved P3 patch and launcher are a positive exception: they retain an exact external source commit, patch, and orchestration path. - The production image still exists locally on Spark, but the container is stopped and an image present in one Docker store is not archival recovery. ## Missing Evidence - The original path list and class-by-class composition behind the retained 279-file, 665,950,155-byte Spark aggregate; the complete current 280-file top-level inventory does not resolve that historical discrepancy. - Exact commands and complete environments for most historical benchmarks and profiler captures. - A content-addressed archive of each dirty working tree. - Base-image digest, Python patch version, full installed-package lock, firmware, clock/power state, and model/checkpoint hashes. - Checksums and retention status for mounted parity captures, model files, generated media, build logs, and P3 run worktrees. - A Git-recoverable copy of the former P1 dispatch and other removed experimental source. - A currently passing `/ready` response and a fresh complete test-suite result. Until those gaps are closed, the evidence supports the recorded conclusions but does not provide complete independent replay of every historical experiment.