diff --git a/CURRENT_STATE.md b/CURRENT_STATE.md index a15c8f5..4e00834 100644 --- a/CURRENT_STATE.md +++ b/CURRENT_STATE.md @@ -1,6 +1,6 @@ # H3 Runtime Current State -Status date: 2026-08-25 +Status date: 2026-08-26 This document is the canonical snapshot of implemented scope and remaining work. Historical handoffs in `PLAN.md` and `PARITY.md` may describe older states. @@ -13,7 +13,9 @@ Historical handoffs in `PLAN.md` and `PARITY.md` may describe older states. - Qwen text and vision conditioning, token refinement, video VAE encoding, H3 packed denoising, beta/RES sampling, video/audio decoding, and final MP4 mux. - Resident HTTP runtime with warmup, readiness reporting, request-level backend - selection, timing stages, optional latent saving, and diagnostic intermediates. + selection, timing stages, canonical-benchmark-only per-step CUDA timings, + peak sampling memory and latent checksums, FC2 dispatch deltas, optional + latent saving, and diagnostics. - SageAttention2 as the default quality backend. - SDPA, forced cuDNN SDPA, FlashAttention-4, Sage3, Comfy Kitchen INT8, KJ Sage, head-sliced, and Sol-Attn experimental backends. @@ -70,7 +72,10 @@ Historical handoffs in `PLAN.md` and `PARITY.md` may describe older states. - H3-specific attention backend optimized for real GB10 tensor shapes. - Blackwell-native CUTLASS/CuTe or cuBLASLt NVFP4 GEMMs. - CUDA graph capture and shape buckets. -- A four-GEMM NVFP4 roofline study for QKV, attention output, FC1, and FC2. + +The four-GEMM roofline study and canonical FC2 cuBLASLt integration are complete. +Remaining owned-GEMM work concerns QKV, attention output, and FC1 only and must +be justified against the post-FC2 profile. The active NVFP4 fusion profile is now measured on one canonical GB10 block: 32 of 53 launches belong to the four scale/pack/GEMM paths. Native packed data, @@ -140,7 +145,7 @@ trajectory validation or enable the backend until launch fusion or a different persistent scheduler passes this block gate. See `benchmarks/gb10-cute-qkv-runtime-block-gate-summary.json`. -The fresh post-optimization canonical run is `288.93 s` with unchanged video +The historical pre-FC2 post-optimization canonical run was `288.93 s` with unchanged video and audio checksums. Block 24 is `468.22 ms` median, of which production NHD Sage2 attention consumes `258.47 ms`. Internal attribution places `238.81 ms` in the SM89 attention mainloop, versus `7.68 ms` Q/K quantization and `10.65 ms` @@ -176,8 +181,8 @@ latents. `H3_NVFP4_FC2_LT_SPLITK1=1` enables only the validated `M=37,810` descriptor; nearby row counts can differ by two BF16 elements and therefore retain the existing Comfy fallback. The extension and measured runtime ABI are prepared during H3 model loading rather than on the first canonical request. -The Spark hot service is intentionally -stopped until performance work is complete. See +The resident service has now passed post-FC2 deployment and repeated canonical +baseline validation and is active on Spark at port `8001`. See `research/fc2_nvfp4_scheduling/RESULTS.md`. The Spark hot runtime was rebuilt and recreated with image @@ -203,15 +208,17 @@ rejected. Active LoRA retains the exact materialized fallback instead of using either producer fusion. See `benchmarks/gb10-nvfp4-lora-producer-fusion-turbo4-isolated.json`. -The active Spark image is now -`sha256:1d340e14cb6fc45ccfdbe63dde8db2a2b3aea94b493702c8a08e1f8d5b4f7b83` -with both accepted NVFP4 producer fusions enabled for eligible base requests. -Health, startup warmup, three deployed CUDA parity tests, and a resident real -generation smoke pass. The final smoke completed in `2.321 s` with `0.217 s` -sampling and produced a valid 22-frame 320x192 H.264 file. See -`benchmarks/gb10-post-lora-revert-deployment-smoke.json`. +The authoritative profile image is +`sha256:a15d0c09dd8cc82aaf2b564d3da760ea5ab8dec974f73075f7d30ac3a504815c`. +The active production overlay is +`sha256:6d880d628334c981c3d155bf5244e65e26e22cc9273c80145f646eee3c3698c2`; +it inherits the profiled binary/ABI layers and makes telemetry +canonical-benchmark-only. Production flags enable both accepted producers and +`H3_NVFP4_FC2_LT_SPLITK1=1`; ordinary requests do not create step events or +copy latents for checksums. `/ready`, startup warmup, repeated exact canonical +generation, profiler captures, and 58 tests pass. -A fresh fully fused Nsight recapture now supersedes the old approximately +A previous pre-FC2 fully fused Nsight recapture superseded the old approximately 515 ms block profile. Block 24 is `458.78 ms` median uninstrumented and `465.78 ms` across the Nsight GPU span, with 41 kernels and only `0.084 ms` of inter-kernel idle time. Sage2 is `57.60%` of kernel time, the four NVFP4 GEMMs @@ -221,6 +228,23 @@ the same distribution. Hardware counters attribute `56.25%` of the warm-cache off-chip request proxy to the NVFP4 GEMMs even though Sage2 remains the time bottleneck. See `benchmarks/gb10-fully-fused-fresh-nsight-summary.json`. +The authoritative post-FC2 resident baseline is now the median of three warmed, +unprofiled canonical runs: `256.464`, `255.447`, and `255.135 s`, giving +`255.447 s`. Every run produced the established video SHA-256 +`c62d23a42972eab907ba42f93c50247ff17a9c454b4a53fe93d2e34f9fefe578` and audio +SHA-256 `852005383770480a6503504e1ffec86dd1fb63a69c6400f92da18e39e0986de2`, +with 600/600 FC2 dispatches and zero fallback. Peak sampling allocation and +reservation were `44,445,830,144` and `48,708,452,352` bytes. + +The same-script generic block-24 decomposition control is `427.410 ms`; it +bypasses guarded FC2 and is not production-FC2 timing. A complete warmed step +has a `20.907 s` Nsight GPU span, `20.895 s` kernel time, 2,744 kernels, +`11.519 ms` total launch gaps, and `99.98%` launch-API/GPU overlap. The new time +ranking is Sage2 `62.36%`, NVFP4 GEMMs `19.95%`, packing `9.70%`, norm/RoPE +`5.19%`, and gate/add `2.63%`. Sage2 is therefore the next-ranked investigation, +but no new optimization has begun. See +`benchmarks/gb10-post-fc2-production-profile-summary-20260826.json`. + The real block-24 Sage2 scheduler study and exact SM89 P0 retune are complete. Manual preparation plus the unchanged prequantized SM89 mainloop is byte-exact against public SageAttention @@ -291,15 +315,24 @@ Production remains unchanged. See `benchmarks/gb10-sage2-p3-temporal-pair-analysis.json` and `research/sage2_temporal_pair/`. -The follow-on four-GEMM NVFP4 roofline study selects MLP FC2 as the next exact -kernel target. QKV, attention output, and FC1 sustain `315-321` dense-equivalent -TFLOP/s and sit above the practical GB10 compute/memory ridge point. FC2 instead -uses the stream-K kernel, sustains only `106.02` TFLOP/s, and generates a -`12.16 GB` warm-cache off-chip request proxy, `16.11x` its ideal unique bytes. -Its measured arithmetic intensity is `479` FLOP/byte versus a practical ridge -near `1,177` FLOP/byte. The next experiment should reduce FC2 K-split rereads and -reuse fused SwiGLU producer tiles; the other three GEMMs are not first targets. -See `benchmarks/gb10-four-gemm-nvfp4-roofline.json`. +Vortex Exact Attention is initialized as an isolated clean-sheet research +project under `research/vortex_exact_attention/`. Phase 0 imports and verifies +the retained SageAttention 2.2.0 exactness contract at commit +`d1a57a546c3d395b1ffcbeecc66d81db76f3b4b5`; it does not repeat P0-P3 or alter +dispatch. Ten retained oracle artifacts match their recorded SHA-256 values. +Phase 1 selects the VEA-B Q128 paired-owner architecture: four QK/softmax warps +own `RS/RS_f8/m/d`, four separate PV warps own `RO`, and two producer warps own +K/V staging. Its `180-207 ms` mainloop range is a heuristic screen, not achieved. VEA-A +is the fallback and two-CTA DSM VEA-C is conditional on capability probes. No +kernel exists. Missing self-contained canonical Q/K/V fixtures, checkpoint +hash, and deployed Sage2 binary hash remain blockers. Production remains Sage2. + +The four-GEMM roofline selected FC2, and the guarded public split-K-1 schedule +closed that target. Fresh NCU values at block 24 are `26.38 ms` QKV, `8.48 ms` +attention output, `36.17 ms` FC1, and `16.20 ms` FC2. Sage2 remains much larger: +its NCU-replayed mainloop is `259.00 ms`, with 255 registers/thread, `16.65%` +achieved occupancy, `98.85%` L2 hit rate, and `161.50 GB` L2 requests. No new +kernel work starts until this ranking is accepted. ## Quality Work Remaining @@ -341,7 +374,8 @@ See `benchmarks/gb10-four-gemm-nvfp4-roofline.json`. ## Recommended Execution Order -1. Run the four-GEMM NVFP4 roofline study on QKV, attention output, FC1, and FC2. +1. Treat the post-FC2 profile as the GB10 baseline; investigate Sage2 only under + a separately approved experiment with exactness and absolute latency gates. 2. Capture matched SM120 and SM100 component profiles and package Sage2 on SM120. 3. Resume NVFP4 GEMM/epilogue work only with a design that preserves the accepted producer fusions and exact BF16 boundaries. diff --git a/PERFORMANCE_ROADMAP.md b/PERFORMANCE_ROADMAP.md index 5afd8d4..cfdc230 100644 --- a/PERFORMANCE_ROADMAP.md +++ b/PERFORMANCE_ROADMAP.md @@ -22,13 +22,11 @@ This is the best first engineering target because: 5. It provides a low-risk test of the profiling, fallback, parity, and deployment machinery required by later custom kernels. -The next component is NVFP4 scale, activation packing, GEMM, and epilogue -fusion. Component 1 established the numerical contracts and validation -infrastructure required to attempt it safely. A custom attention target follows -this component and should preserve SageAttention 2.2.0's validated numerical -policy while adopting zero-copy strided QKV input and SM-specific fused -execution. Sol sparsity must remain experimental until full latent, audio, and -visual quality gates pass. +The accepted NVFP4 producer fusions and canonical FC2 public split-K-1 schedule +are complete. The authoritative post-FC2 profile ranks Sage2 first at `62.36%` +of warmed-step kernel time, followed by NVFP4 GEMMs at `19.95%` and packing at +`9.70%`. No follow-on optimization begins from this document alone; Sage2 is the +next investigation only after a separate exactness and absolute-latency plan. Completed result: blocks 0, 24, and 49 are bit-exact and `3.6-3.9%` faster; the canonical 12-step sampling run is checksum-identical and improves from @@ -49,7 +47,7 @@ Internal closeout summary: This is measured engineering evidence, not a projected optimization. -## Measured GB10 Baseline +## Historical GB10 Baseline The representative workload is 1344x768, 124 frames, 100 synthetic refined-text tokens, and 37,810 packed tokens. Measurements use block 24 unless noted. @@ -99,6 +97,42 @@ Sol residual-INT8 Q/K internals at tau 1.3 with exact conditioning KV: INT8 P-by-V is not useful on GB10: it increases total Sol attention time to `138.65 ms`. +## Authoritative Post-FC2 GB10 Baseline + +The exact resident workload is 1344x768, 124 frames, 12 steps, seed `440420`, +Sage2, 100 synthetic refined-text tokens, and 37,810 packed tokens. After one +unmeasured canonical warmup, three unprofiled resident runs took `256.464`, +`255.447`, and `255.135 s`; the authoritative median is `255.447 s`. All three +runs dispatched FC2 600/600 times with zero fallback and matched the retained +video/audio SHA-256 values exactly. Peak allocated/reserved sampling memory was +`44.446/48.708 GB`. + +Fresh block-24 and warmed-step evidence: + +| Measurement | Post-FC2 result | Previous pre-FC2 result | Change | +| --- | ---: | ---: | ---: | +| Generic block decomposition control p50 | `427.410 ms` | `458.775 ms` | `-6.84%` | +| Warmed-step GPU span | `20.907 s` | `23.707 s` | `-11.81%` | +| Warmed-step kernel time | `20.895 s` | `23.694 s` | `-11.81%` | +| Kernel count | `2,744` | `2,694` | `+50` | +| Total positive launch gaps | `11.519 ms` | `12.995 ms` | `-11.36%` | + +Post-FC2 warmed-step kernel ranking: + +| Rank | Component | Time | Share | +| ---: | --- | ---: | ---: | +| 1 | Sage2 preparation and mainloop | `13.029 s` | `62.36%` | +| 2 | Four NVFP4 GEMMs | `4.168 s` | `19.95%` | +| 3 | NVFP4 scale and packing | `2.027 s` | `9.70%` | +| 4 | Norm and RoPE | `1.084 s` | `5.19%` | +| 5 | Residual gate/add | `0.549 s` | `2.63%` | + +The profile is GPU-bound: summed kernel time is `99.94%` of the GPU span and +kernel-launch API intervals overlap GPU kernels by `99.98%`. CPU launch work is +not the next bottleneck. The block decomposition control calls generic NVFP4 +linears and bypasses guarded FC2; production FC2 evidence comes from resident +NSYS and targeted NCU, not its `mlp_fc2` attribution. + ## Quality Boundary Current Sol routing is a performance bound, not an accepted Sage2 replacement. @@ -122,12 +156,12 @@ The recoverable ranges below overlap and must not be added directly. | Priority | Target | Measured cost | Plausible saving | Risk | Portability | | ---: | --- | ---: | ---: | --- | --- | | Complete | Fuse modulation and residual gates | `39.1 ms/block` projected scope | `3.6-3.9%` measured block gain | Low | Validated on SM121 | -| 1 | Fuse NVFP4 scale, pack, GEMM, and epilogue | `151.1 ms/block` projections | `30-75 ms/block` | Medium | Architecture-specific | -| 2 | Retain Vortex scale discovery | up to `21 ms/block` versus Comfy scale path | `3-9%` block gain already measured | Low | SM100/120/121, retune launch | -| 4 | Sage-compatible zero-copy dense attention | `259.9 ms/block` attention plus `20.3 ms` layout | `70-140 ms/block` | High | Separate SM100/120/121 policy | -| 5 | Fuse SwiGLU with FC1/FC2 boundaries | `25.8 ms/block` plus projection traffic | `5-15 ms/block` | Medium | SM100/120/121 | -| 6 | CUDA graphs or persistent denoiser execution | Not isolated yet | likely `1-5%` sampling | Medium | Shape-bucket-specific | -| 7 | Quality-gated sparse scheduling | up to `140 ms/block` attention bound | Unknown accepted saving | Very high | Backend-specific | +| Closed | Canonical FC2 public split-K-1 | `53.618 ms` baseline FC2 | `15.636 ms` accepted | Low | GB10 canonical shape only | +| 1 | Sage2 preparation and mainloop | `13.029 s/step`, `62.36%` | Not assigned | High | Separate SM100/120/121 policy | +| 2 | Remaining four NVFP4 GEMMs | `4.168 s/step`, `19.95%` | Not assigned | Medium | Architecture-specific | +| 3 | NVFP4 scale and packing | `2.027 s/step`, `9.70%` | Not assigned | Medium | SM100/120/121, retune launch | +| 4 | Norm and RoPE | `1.084 s/step`, `5.19%` | Not assigned | Medium | Shape-specific | +| 5 | CUDA graphs or persistent execution | `11.519 ms` measured launch gaps | Small | Medium | Shape-bucket-specific | Native Sage NHD was tested and rejected as a standalone optimization. It was numerically exact but did not materially improve median path time and had worse @@ -349,7 +383,7 @@ identical video and audio tensors. Enable `H3_NVFP4_SWIGLU_FUSION=1` for Spark single-GPU deployments. Evidence is in `benchmarks/gb10-nvfp4-swiglu-fusion-summary.json`. -The mandatory post-fusion profile supersedes the earlier approximately 515 ms +The previous pre-FC2 post-fusion profile superseded the earlier approximately 515 ms block distribution. With all production flags enabled, fresh block-24 timing is `458.78 ms` median without Nsight and `465.78 ms` across the Nsight GPU span. The 41 kernels have only `0.084 ms` total inter-kernel idle time, so CPU launch @@ -404,10 +438,21 @@ closed. Guarded production integration is complete behind improves by `8.19%` with 600 successful dispatches and zero fallback. Shape testing found two-BF16 differences at several nearby row counts, so only `M=37,810` is eligible and -all variable shapes retain Comfy. The hot service remains stopped until the -performance program is complete. See `research/fc2_nvfp4_scheduling/RESULTS.md` and +all variable shapes retain Comfy. The post-FC2 baseline and profiling are now +complete, and a telemetry-opt-in overlay of the profiled binary image is +resident again. See +`research/fc2_nvfp4_scheduling/RESULTS.md` and `benchmarks/gb10-fc2-nvfp4-library-sweep-20260825.json`. +The mandatory post-FC2 recapture now establishes the production baseline. +Three exact resident runs have a `255.447 s` median. Block 24 falls from the +previous `458.775 ms` to `427.410 ms`, while warmed-step GPU span falls from +`23.707 s` to `20.907 s`. Aggregate NVFP4 GEMM time falls `33.52%`; Sage2 now +occupies `62.36%` of kernel time. Targeted NCU confirms FC2 at `16.20 ms`, +`91.06%` L2 hit rate, `26.07 GB` requested L2 traffic, and `77.65%` memory +throughput. Sage2 remains the dominant exact target, but no implementation has +started. See `benchmarks/gb10-post-fc2-production-profile-summary-20260826.json`. + The follow-up real block-24 Sage2 decomposition now selects the next exact kernel experiment. Manual preparation plus the existing prequantized mainloop is byte-exact against public SageAttention 2.2.0. Uninstrumented median phase @@ -466,6 +511,16 @@ variant crossed 3%. Evidence is in `benchmarks/gb10-sage2-p0-register-scheduler-analysis.json`, the per-variant latency JSON files, and the P0 NCU reports. +Vortex Exact Attention now owns the next exact-attention design investigation. +Phase 0 imports and verifies the retained Sage2 evidence rather than repeating +P0-P3. The exact contract, fixture hashes, proof limits, and prior-work exclusion +matrix are under `research/vortex_exact_attention/`. Phase 1 recommends VEA-B, +a Q128 ten-warp pipeline with permanent disjoint QK/softmax and PV/output owners. +Its `180-207 ms` mainloop and `205-235 ms` complete-attention ranges are +heuristic screening models, not measured results. Register allocation, shared capacity, named +barriers, and INT8/FP8 overlap must pass capability probes before kernel work. +Production dispatch remains unchanged. + Materializing the optional BF16 LoRA input while these fused producers pack NVFP4 was also tested against the official Turbo-4 adapter. The isolated 1344x768, 124-frame trajectory was bit-exact but regressed from `131.11 s` to @@ -535,7 +590,7 @@ Only after the dense backend passes: ### GB10 / SM121 -The current two-step profile implies roughly `305 s` for a warmed 12-step +The authoritative exact resident median is `255.447 s` for a warmed 12-step 1344x768/124-frame sampling run. | Target | Estimated sampling | Speedup | @@ -579,6 +634,10 @@ The current B200 result measures this runtime path, not B200's hardware ceiling. - `benchmarks/gb10-37810-token-optimization-profile-seed440420.json` - `benchmarks/rtxpro6000-server-1v2-sdpa-scaling-seed440420.json` - `benchmarks/b200-vs-rtxpro6000-server-sdpa-1344x768-seed440420.json` +- `benchmarks/gb10-post-fc2-resident-baseline-20260826.json` +- `benchmarks/gb10-post-fc2-production-profile-summary-20260826.json` +- `benchmarks/gb10-post-fc2-warmed-step-20260826.nsys-rep` +- `benchmarks/gb10-post-fc2-block24-targeted-20260826.ncu-rep` - `tools/profile_h3_block.py` - `tools/profile_attention_components.py` - `tools/profile_hybrid_attention.py` diff --git a/SAGE2_BLACKWELL_DESIGN.md b/SAGE2_BLACKWELL_DESIGN.md index 03ca6b6..9a16045 100644 --- a/SAGE2_BLACKWELL_DESIGN.md +++ b/SAGE2_BLACKWELL_DESIGN.md @@ -150,3 +150,14 @@ This closes temporal warp separation as an exact Sage2 optimization on the current SM89 instruction stream. A multi-percent attention gain now requires a substantially new mainloop or a quality-gated Blackwell-native attention algorithm rather than another exact schedule rearrangement. + +## Successor Research + +Vortex Exact Attention supersedes further source-identical Sage2 retuning. Its +Phase 0 contract imports this document and retained P0-P3 evidence without +claiming unrecoverable P0 source. Phase 1 evaluates only architectures that move +or partition `RS/RS_f8`, `m/d`, and `RO` ownership. The recommended VEA-B design +uses separate QK/softmax and PV/output warps around the still-supported exact +INT8 `mma.sync` arithmetic on SM121. No kernel or production dispatch exists; +projected latency ranges are not achieved results. See +`research/vortex_exact_attention/`. diff --git a/research/EXPERIMENT_REGISTRY.md b/research/EXPERIMENT_REGISTRY.md index d482709..538c139 100644 --- a/research/EXPERIMENT_REGISTRY.md +++ b/research/EXPERIMENT_REGISTRY.md @@ -1,7 +1,7 @@ # Experiment Registry Generated from the repository working tree, documentation, benchmark artifacts, -and git history available on 2026-08-25. The machine-readable registry is +and git history available on 2026-08-26. The machine-readable registry is `research/experiment_registry.json`; it is authoritative for record fields and artifact existence flags. @@ -91,6 +91,9 @@ missing rather than inferred. | `sage2-p1-entry-fusion-final` | Finalized Sage2 P1 entry fusion | `performance_rejected` | Exact, but complete-block gate failed; callable wrapper/export and validator remain without production integration. | `sage2_entry.py`; entry validator; P1 analysis/parity/NCU | Trajectory intentionally skipped | | `sage2-p2-direct-v-preparation-final` | Finalized Sage2 P2 direct V preparation | `performance_rejected` | Exact isolated gain missed absolute go threshold; callable wrapper/export and validator remain without production integration. | `sage2_entry.py`; V-prep validator; P2 analysis/randomized/repeat artifacts | Block/trajectory intentionally skipped | | `sage2-p3-temporal-pair-final` | Finalized Sage2 P3 temporal pair | `performance_rejected` | Sanitizer-clean/exact but effectively neutral. | Retained patch/runner and P3 artifacts | NCU/block/trajectory intentionally skipped | +| `gb10-post-fc2-production-profile-20260826` | Authoritative GB10 post-FC2 production baseline and profile | `production_accepted` | Exact `255.447 s` resident median and fresh NSYS/NCU ranking establish Sage2 as the dominant next investigation. | Post-FC2 baseline/profile summaries and raw reports | Matched SM120/SM100 profiles | +| `vortex-exact-phase0-contract-import` | Vortex Exact Attention Phase 0 contract import | `research_retained` | Imported and verified retained exactness evidence without repeating closed Sage work; no kernel or dispatch exists. | `research/vortex_exact_attention/NUMERICAL_CONTRACT.md`; reference inventory | Self-contained canonical inputs, checkpoint and deployed binary hashes | +| `vortex-exact-phase1-architecture-decision` | Vortex Exact Attention Phase 1 architecture decision | `research_retained` | Three structural ownership designs modeled; VEA-B recommended subject to capability probes. No speedup is claimed. | `research/vortex_exact_attention/DESIGN.md`; architecture model | Compiled resources, barrier/overlap probes, executable parity and timing | ## Interpretation Notes @@ -110,11 +113,11 @@ missing rather than inferred. | Status | Count | | --- | ---: | -| `production_accepted` | 4 | -| `research_retained` | 6 | +| `production_accepted` | 6 | +| `research_retained` | 8 | | `architecture_rejected` | 7 | | `performance_rejected` | 15 | | `quality_rejected` | 2 | | `temporarily_blocked` | 3 | | `incomplete` | 25 | -| **Total** | **62** | +| **Total** | **66** | diff --git a/research/experiment_registry.json b/research/experiment_registry.json index b3eab07..720176e 100644 --- a/research/experiment_registry.json +++ b/research/experiment_registry.json @@ -1,7 +1,7 @@ { "schema": "h3-blackwell-runtime-experiment-registry", "version": "1.0.0", - "generated_at": "2026-08-25", + "generated_at": "2026-08-26", "allowed_statuses": [ "production_accepted", "research_retained", @@ -817,6 +817,83 @@ "reproducer_commands": [], "timestamp": null, "evidence_missing": ["Candidate source"], "production_behavior": "Not integrated.", "source_recovery": "Latency artifact retained." }, + { + "id": "vortex-exact-phase0-contract-import", + "name": "Vortex Exact Attention Phase 0 contract import", + "family": "vortex-exact-attention", + "status": "research_retained", + "hypothesis": "The retained Sage2 inventory is sufficient to import and verify the exact observable contract without repeating closed P0-P3 experiments.", + "implementation_strategy": "Verify the pinned SageAttention commit and retained artifact hashes, consolidate every numerical boundary and proof limit, and keep the project isolated from runtime dispatch.", + "source_locations": ["research/vortex_exact_attention/README.md", "research/vortex_exact_attention/NUMERICAL_CONTRACT.md", "research/vortex_exact_attention/EXPERIMENT_LOG.md", "research/vortex_exact_attention/benchmarks/reference_inventory.json"], + "active_source_location": null, + "commit_hash": "c502842", + "benchmark_artifacts": [{"path": "research/vortex_exact_attention/benchmarks/reference_inventory.json", "exists": true}], + "profiler_artifacts": [{"path": "benchmarks/gb10-sage2-block24-mainloop-full.ncu-rep", "exists": true}, {"path": "benchmarks/gb10-post-fc2-block24-targeted-20260826.ncu-rep", "exists": true}], + "environment": {"gpu": "NVIDIA GB10", "cuda": "13.0", "driver": "580.95.05", "pytorch": "2.9.1+cu130", "triton": null, "container_image": null, "commit_context": "documentation and retained-artifact verification only"}, + "metrics": {"verified_artifact_count": 10, "reference_commit": "d1a57a546c3d395b1ffcbeecc66d81db76f3b4b5", "canonical_tokens": 37810, "mainloop_baseline_ms": 237.089, "no_eligible_cycles_percent": 63.53}, + "correctness_evidence": ["All ten retained artifact SHA-256 values match the imported inventory.", "The pinned public SageAttention commit exists.", "Existing short-shape, blocks 0/24/49, and trajectory outputs are retained as oracles."], + "decision_rationale": "Phase 0 is complete as an imported specification. Missing self-contained canonical inputs and binary/checkpoint hashes remain explicit blockers before executable prototype acceptance.", + "reproducer_commands": [], + "timestamp": "2026-08-26", + "evidence_missing": ["Self-contained canonical Q/K/V tensors", "Checkpoint SHA-256", "Deployed Sage2 extension binary hash", "Exceptional-value fixtures"], + "production_behavior": "No kernel and no dispatch; production Sage2 remains unchanged.", + "source_recovery": "P3 is fully recoverable; P1/P2 are partially reusable in the shared experimental CUDA worktree; exact P0 candidate source is unrecoverable and only timing/JSON/NCU evidence is claimed." + }, + { + "id": "vortex-exact-phase1-architecture-decision", + "name": "Vortex Exact Attention Phase 1 architecture decision", + "family": "vortex-exact-attention", + "status": "research_retained", + "hypothesis": "Separating QK/softmax ownership from PV/output ownership can reduce simultaneous per-warp live state and dependency stalls while preserving Sage2's exact tile and accumulation order.", + "implementation_strategy": "Model three clean-sheet SM121 producer-consumer architectures, reject designs requiring material score tensors or reordered reductions, and select a prototype only after resource and synchronization gates.", + "source_locations": ["research/vortex_exact_attention/DESIGN.md", "research/vortex_exact_attention/PERFORMANCE_MODEL.md", "research/vortex_exact_attention/HARDWARE_CAPABILITIES.md", "research/vortex_exact_attention/benchmarks/architecture_model.json"], + "active_source_location": null, + "commit_hash": "b842bad", + "benchmark_artifacts": [{"path": "research/vortex_exact_attention/benchmarks/architecture_model.json", "exists": true}], + "profiler_artifacts": [{"path": "benchmarks/gb10-sage2-block24-mainloop-full.ncu-rep", "exists": true}, {"path": "benchmarks/gb10-post-fc2-block24-targeted-summary-20260826.json", "exists": true}], + "environment": {"gpu": "NVIDIA GB10", "cuda": "13.0", "driver": "580.95.05", "pytorch": "2.9.1+cu130", "triton": null, "container_image": null, "commit_context": "design-only; no kernel execution"}, + "metrics": {"baseline_mainloop_ms": 237.089, "model_gate_ms": 220, "preferred_model_ms": 190, "candidate_count": 3, "recommended_candidate": "VEA-B", "recommended_mainloop_projected_ms": [180, 207], "achieved_speedup": null}, + "correctness_evidence": ["Every candidate retains strict KV tile order and the exact two-level PV accumulation contract.", "Every candidate assigns QK/softmax and output accumulation to different owners.", "The model validator rejects same-warp ownership, fewer than three candidates, and measured-result labels."], + "decision_rationale": "VEA-B has a credible modeled path below 190 ms without depending on unverified DSM. VEA-A is fallback; VEA-C remains conditional. All latency ranges are projections.", + "reproducer_commands": ["python research/vortex_exact_attention/tools/validate_models.py", "python -m unittest discover -s research/vortex_exact_attention/tests -p test_*.py"], + "timestamp": "2026-08-26", + "evidence_missing": ["Compiled role-specific register allocation", "Named-barrier handoff cost", "Concurrent INT8 QK and FP8 PV issue evidence", "DSM feasibility for VEA-C", "Executable exactness and latency"], + "production_behavior": "No kernel and no dispatch; production Sage2 remains unchanged.", + "source_recovery": "Design documents, machine-readable model, validator, and test are retained under research/vortex_exact_attention/." + }, + { + "id": "gb10-post-fc2-production-profile-20260826", + "name": "Authoritative GB10 post-FC2 production baseline and profile", + "family": "production-baseline", + "status": "production_accepted", + "hypothesis": "The guarded FC2 schedule materially improves the exact resident workload and shifts the next bottleneck away from FC2.", + "implementation_strategy": "Measure one unreported resident warmup plus three exact canonical runs, recapture one complete warmed denoising step with Nsight Systems, and collect ordered block-24 kernel sections and L2 traffic with Nsight Compute.", + "source_locations": ["CURRENT_STATE.md", "PERFORMANCE_ROADMAP.md", "tools/benchmark_hot_runtime.py", "tools/profile_sampling_stages.py", "tools/summarize_nsys_profile.py", "tools/summarize_ncu_profile.py", "tools/build_post_fc2_profile_summary.py"], + "active_source_location": "src/h3_blackwell_runtime/fc2_lt.py", + "commit_hash": "a29b8960b0f887c20e74dafa16a24c37d6508b4e", + "benchmark_artifacts": [ + {"path": "benchmarks/gb10-post-fc2-resident-baseline-20260826.json", "exists": true}, + {"path": "benchmarks/gb10-post-fc2-block24-profile-20260826.json", "exists": true}, + {"path": "benchmarks/gb10-post-fc2-production-profile-summary-20260826.json", "exists": true} + ], + "profiler_artifacts": [ + {"path": "benchmarks/gb10-post-fc2-warmed-step-20260826.nsys-rep", "exists": true}, + {"path": "benchmarks/gb10-post-fc2-warmed-step-20260826.sqlite", "exists": true}, + {"path": "benchmarks/gb10-post-fc2-warmed-step-nsys-summary-20260826.json", "exists": true}, + {"path": "benchmarks/gb10-post-fc2-block24-targeted-20260826.ncu-rep", "exists": true}, + {"path": "benchmarks/gb10-post-fc2-block24-targeted-traffic-20260826.ncu-rep", "exists": true}, + {"path": "benchmarks/gb10-post-fc2-block24-targeted-summary-20260826.json", "exists": true} + ], + "environment": {"gpu": "NVIDIA GB10", "cuda": "13.0", "driver": "580.95.05", "pytorch": "2.9.1+cu130", "triton": null, "container_image": "sha256:a15d0c09dd8cc82aaf2b564d3da760ea5ab8dec974f73075f7d30ac3a504815c", "commit_context": "isolated a29b896 post-FC2 build plus response-only telemetry"}, + "metrics": {"resident_sampling_median_seconds": 255.44699439899978, "resident_sampling_runs_seconds": [256.46369375299946, 255.44699439899978, 255.13481797700024], "block24_median_ms": 427.40975750000004, "warmed_step_gpu_span_seconds": 20.90717264, "warmed_step_kernel_seconds": 20.895065376, "kernel_count": 2744, "fc2_dispatches_per_run": 600, "fc2_fallbacks_per_run": 0, "component_ranking": ["sage2", "nvfp4_gemms", "nvfp4_packing", "norm_rope", "gate_add"]}, + "correctness_evidence": ["All three resident runs match the established canonical video and audio SHA-256 values.", "Each run records 600 FC2 attempts, 600 successes, and zero fallbacks."], + "decision_rationale": "The exact resident median and fresh profiler evidence establish the post-FC2 production baseline. Sage2 is the dominant next-ranked investigation at 62.36% of warmed-step kernel time; no new optimization is authorized by this entry.", + "reproducer_commands": ["python tools/benchmark_hot_runtime.py --server http://127.0.0.1:8001 --runs 3 --image sha256:a15d0c09dd8cc82aaf2b564d3da760ea5ab8dec974f73075f7d30ac3a504815c --output benchmarks/gb10-post-fc2-resident-baseline-20260826.json", "python tools/profile_sampling_stages.py --attention sage2 --steps 1 --warmup-runs 1 --uninstrumented --cuda-profiler-capture --output /output/h3-blackwell-runtime/benchmarks/gb10-post-fc2-warmed-step-capture-20260826.json"], + "timestamp": "2026-08-26", + "evidence_missing": ["Matched SM120 and SM100 component profiles"], + "production_behavior": "Guarded canonical FC2 dispatch remains enabled; ordinary prompt conditioning and all noncanonical shapes retain their prior behavior and fallback.", + "source_recovery": "Summary JSON, raw NSYS/NCU reports, exports, capture metadata, and deterministic parser tools are retained in benchmarks/ and tools/." + }, { "id": "sage-p0-padded-v-smem", "name": "Sage P0 padded-V shared-memory layout", diff --git a/research/vortex_exact_attention/EXPERIMENT_LOG.md b/research/vortex_exact_attention/EXPERIMENT_LOG.md index 772d3e4..af13c9e 100644 --- a/research/vortex_exact_attention/EXPERIMENT_LOG.md +++ b/research/vortex_exact_attention/EXPERIMENT_LOG.md @@ -63,3 +63,8 @@ on cluster residency and DSM latency. These are projections only. No kernel was implemented, no inference was run, and production dispatch and services were not modified. + +Phase commits: + +- Phase 0 contract import: `c502842` +- Phase 1 architecture decision: `b842bad`