h3-blackwell-runtime/research/vortex_exact_attention/HARDWARE_CAPABILITIES.md
2026-08-26 14:44:28 +07:00

4.2 KiB

Hardware Capabilities

Interpretation Rule

Blackwell-native means using target-specific scheduling, state ownership, or data movement supported by the target. Compiling the SM89 algorithm for sm_121a is not sufficient. SM100, SM120, and SM121 are separate schedules.

Capability Matrix

Capability SM100 / B200 SM120 / RTX PRO 6000 SM121 / GB10
Compute capability 10.0 12.0 12.1
TMA global/shared tensor copies Documented Documented Documented
Native matrix programming model tcgen05.mma with TMEM warp-level mma.sync SM120-family warp-level mma.sync
Programmer-visible TMEM Yes Not exposed Not exposed
Native UMMA/tcgen05 INT8 Exposed Not exposed Not exposed
Sage-compatible INT8 path tcgen05.kind::i8 can be investigated supported classic INT8 mma.sync supported classic INT8 mma.sync
FP8 PV path tcgen05 FP8 forms warp-level FP8 mma.sync warp-level FP8 mma.sync
Hopper WGMMA Not a portable target path Not a target path CUDA 13 rejects for sm_121a
Registers per SM / max per thread 64K / 255 64K / 255 64K / 255
Shared memory per SM / block 228 / 227 KiB 100 / 99 KiB 100 / 99 KiB
Maximum resident warps 64 48 48
Maximum resident threads 2048 1536 1536

The public SM120/121 Blackwell MMA interface does not expose a Sage2-compatible UMMA INT8 QK operation. Exact SM121 work therefore retains the supported INT8 mma.sync arithmetic while making TMA staging, role scheduling, ownership, and handoff Blackwell-specific. FP8/FP4 QK would define vortex_fast, not vortex_exact.

SM121 Facts From Retained Evidence

  • The current 128-thread kernel uses 255 registers/thread and 32 KiB dynamic shared memory.
  • Two CTAs provide eight resident warps, approximately 16.67% of the 48-warp ceiling.
  • A 168-register cap reaches three CTAs but spills catastrophically; compiler lifetime cleanup is not enough.
  • CUDA 13 rejects wgmma.fence, wgmma.mma_async, wgmma.commit_group, and wgmma.wait_group for sm_121a.
  • CUTLASS 4.6 executes SM120-family block-scaled narrow MMA on GB10, but that is not an exact replacement for Sage2 INT8 QK.
  • Phase 2A measures VEA-B role resources, one-CTA residency, bounded mbarrier handoff, and concurrent INT8/FP8 issue. TMA integration remains a complete prototype concern rather than a measured capability claim.

Capability Probe Status

  1. Complete: 320-thread QK, PV, combined-ownership, handoff, and tensor-issue resource probes with zero local spills.
  2. Complete: one 51.2 KiB ten-warp CTA/SM, confirmed by occupancy API and NCU.
  3. Complete: corrected two-slot 591-epoch mbarrier handoff, deterministic and sanitizer-clean, with 48-block p50/p95 measurements.
  4. Complete: simultaneous INT8 and FP8 warp progress in 48/48 blocks, plus an NCU tensor-issue capture.
  5. Remaining for the aligned prototype: integrated TMA staging, bank-aware payload layout, complete-kernel resources, sustained clocks, and exact output.
  6. Remaining before VEA-C: cluster/DSM support, placement, simultaneous residency, and remote shared-memory latency on GB10.

Portability Policy

  • SM121 is first and must compile specifically for sm_121a.
  • SM120 may share API and contract code, but requires its own geometry, resource model, profiler evidence, and acceptance gates.
  • SM100 should use a separate tcgen05/TMEM design. Porting the SM121 register accumulator schedule unchanged would discard the principal SM100 capability.

Sources

Documented API support does not establish throughput, latency, queue depth, or bit identity. Those remain target-specific measurements.