h3-blackwell-runtime/research/vortex_exact_attention/EXPERIMENT_LOG.md
2026-08-26 13:39:20 +07:00

3.7 KiB

Experiment Log

This is append-only. Projected ranges are design estimates, never achieved results.

2026-08-26: Phase 0 Import And Verification

Status: research_retained

Hypothesis: the retained Sage inventory is sufficient to import the exact observable contract without repeating prior experiments.

Actions:

  • Verified public SageAttention commit d1a57a546c3d395b1ffcbeecc66d81db76f3b4b5 exists.
  • Rehashed ten retained contract, parity, timing, and NCU artifacts; all matched benchmarks/reference_inventory.json.
  • Imported Q/K preparation, K smoothing, INT8 quantization, QK traversal, online-softmax, FP8 V/P, two-level PV accumulation, final conversion, and ragged-tail contracts into NUMERICAL_CONTRACT.md.
  • Recorded missing production binary/checkpoint hashes and missing standalone canonical input tensors as blockers rather than fabricating provenance.
  • Did not execute inference, modify dispatch, or touch the production service.

Decision: Phase 0 is verified as an imported specification. No kernel work is authorized until Phase 1 selects a structural state-ownership architecture.

Closed Prior Work

Work Result Preservation / reopening rule
P0 register caps Three-CTA threshold at 168 registers caused catastrophic spills; no gain Timing/JSON/NCU retained; exact candidate source missing and must not be claimed reproducible
P0 scope/lifetime narrowing Fewer spills, approximately 0.06% gain Reopen only with structural ownership change
P0 score aliasing Approximately 0.01%, noise Do not repeat unchanged
P0 early K prefetch Approximately 0.30% slower Reopen only with a different producer/consumer pipeline
P0 softmax interleaving Neutral/slower Do not repeat unchanged
P0 padded V shared memory Same excessive wavefronts, no gain Reopen only if staging instructions/dataflow change
P1 entry fusion 20.9-23.4% isolated; only 0.535-0.863% complete-block Reuse partial implementation only inside an owned attention pipeline
P2 direct V preparation 39.48%, 4.17 ms; only 0.909% of block Reuse partial implementation only inside an owned attention pipeline
P3 temporal pairing Exact/sanitizer-clean; approximately 0.10%; private state remained live Complete patch and launcher retained; new work must transfer/partition ownership
Hopper WGMMA retarget CUDA 13 rejects WGMMA for sm_121a Reopen only with changed ISA/toolchain support
SM121 UMMA INT8 QK Public SM120/121 Blackwell MMA path does not expose a Sage-compatible UMMA INT8 operation Retain supported INT8 mma.sync unless capability changes

P1/P2 source is partial and mixed in research/shared_cuda_experimental_worktree/. P3 source and launcher are complete under research/sage2_temporal_pair/. Exact P0 variant patches are unrecoverable; historical commands and environments must not be invented.

2026-08-26: Phase 1 State-Ownership Decision

Status: research_retained

Evaluated three architectures that permanently separate QK/softmax state from PV/output state: VEA-A Q64 split-D intra-CTA, VEA-B Q128 paired owners, and VEA-C two-CTA DSM ownership split. All retain strict tile order, exact FP8 score bytes, per-tile RO rescaling, two-level PV accumulation, and final conversion.

Decision: recommend VEA-B after register/shared/barrier/dual-pipe capability probes. Its projected 180-207 ms mainloop screening range has a credible sub-190 path without depending on unverified DSM. VEA-A is the fallback; VEA-C is conditional on cluster residency and DSM latency. These are projections only.

No kernel was implemented, no inference was run, and production dispatch and services were not modified.