4.9 KiB
Phase 2A VEA-B Capability Report
Date: 2026-08-26
Status: advance_to_isolated_aligned_prototype
This phase measures only the VEA-B ownership, resource, handoff, and concurrent issue capabilities. It does not implement attention, measure attention latency, or authorize production dispatch.
Decision
VEA-B passes the Phase 2A capability gates with CUDA block-scope mbarriers as
the two-slot handoff primitive. Proceed to one isolated aligned-shape attention
prototype. Do not add H3_ATTENTION=vortex_exact or alter the Sage2 fallback.
The inline-PTX named-barrier primitive is rejected for the prototype. It is faster in the synthetic barrier-only probe, but Compute Sanitizer racecheck reports five shared-memory hazards because that path does not provide a sanitizer-recognized happens-before edge. The mbarrier path reports zero hazards and remains far below the synchronization budget.
Authoritative Probe
Environment: NVIDIA GB10, SM121, CUDA 13.0, PyTorch 2.9.1+cu130, production
image sha256:6d880d628334c981c3d155bf5244e65e26e22cc9273c80145f646eee3c3698c2.
The run uses 48 blocks, 320 threads, 591 epochs, 20 warmups, and 100 measured
iterations.
| Gate | Measurement | Result |
|---|---|---|
| Selected handoff registers | 96/thread | pass, limit 200 |
| QK role registers | 54/thread | pass |
| PV role registers | 138/thread | pass |
| Combined ownership registers | 139/thread | pass |
| Local memory / ptxas spills | 0 / 0 | pass |
| Selected dynamic shared memory | 51,200 bytes | pass |
| Resident ten-warp CTAs/SM | 1 | pass |
| mbarrier-only p50 / p95 | 0.076800 / 0.078880 ms |
pass, budget 11.85445 ms |
| Synthetic payload p50 / p95 | 1.874944 / 1.879474 ms |
pass |
| 591-epoch payload checksum | 4,671,090 |
pass, zero publication errors |
| Deterministic repeat | identical | pass |
| INT8/FP8 positive clock overlap | 48/48 blocks | pass |
| INT8/FP8 overlap p50 ratio | 0.9999929 |
capability pass |
| Combined issue p50 | 0.127072 ms |
below 0.224848 ms serial p50 sum |
The overlap probe establishes that distinct VEA-B warp roles can make progress
through INT8 and FP8 mma.sync loops concurrently. It is not a throughput model
for the complete attention mainloop.
Sanitizer And NCU
- Selected mbarrier payload: memcheck 0 errors; racecheck 0 hazards.
- Tensor issue probe: memcheck 0 errors; racecheck 0 hazards.
- Rejected inline named barrier: memcheck 0 errors; racecheck 5 hazards.
- NCU selected handoff: 96 registers/thread, 51.2 KiB dynamic shared memory,
shared-memory block limit 1,
20.91%achieved occupancy, and zero local spill requests. - NCU tensor issue: 26 registers/thread and zero local spill requests.
The synthetic handoff payload has severe shared-load bank conflicts in the single-block NCU capture. Its deliberately simple byte access pattern is not an attention layout result, but the aligned prototype must choose and profile a bank-aware score/K/V layout rather than copy this access pattern unchanged.
Canonical Fixtures
Self-contained BF16 NHD Q, K, V, and Sage2 output tensors were captured at
[1, 37810, 56, 128], saved, reloaded, and compared byte-for-byte. The output
tensor SHA-256 is
4c666c20f5f8f651158a2ced33ccff08f3bada07665c595b99008d171db30574,
matching the locked public Sage2 oracle. Checkpoint and deployed Sage binary
hashes also match the numerical contract.
Invalid And Rejected Runs
The first full probe is invalid for payload conclusions. K/V allocation reserved two 8 KiB slots but indexed one 16 KiB region, allowing the next epoch to overwrite consumer data. Deterministic repetition exposed the defect; the corrected probe uses slot-relative 8 KiB indexing and passes deterministic, checksum, memcheck, and racecheck validation.
The original 16.631 ms inline-barrier number from that invalid run is not
retained as a hardware measurement. The corrected authoritative inline path is
0.056128 ms p50, but remains rejected on the sanitizer gate.
Durable Evidence
/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/probes/vea-b-capability-20260826-authoritative.json/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/build/vea-b-build-20260826-authoritative.log/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/fixtures/canonical-20260826/manifest.json/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/sanitizer//home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/ncu/
Next Gate
Implement only one aligned short-shape VEA-B prototype with the selected
mbarrier handoff and exact D=128 arithmetic. It must compare against the captured
prepared-byte fixtures before any canonical or integrated timing. Phase 2A does
not validate the heuristic 180-207 ms mainloop screen.