101 lines
4.9 KiB
Markdown
101 lines
4.9 KiB
Markdown
|
|
# Phase 2A VEA-B Capability Report
|
||
|
|
|
||
|
|
Date: 2026-08-26
|
||
|
|
|
||
|
|
Status: `advance_to_isolated_aligned_prototype`
|
||
|
|
|
||
|
|
This phase measures only the VEA-B ownership, resource, handoff, and concurrent
|
||
|
|
issue capabilities. It does not implement attention, measure attention latency,
|
||
|
|
or authorize production dispatch.
|
||
|
|
|
||
|
|
## Decision
|
||
|
|
|
||
|
|
VEA-B passes the Phase 2A capability gates with CUDA block-scope mbarriers as
|
||
|
|
the two-slot handoff primitive. Proceed to one isolated aligned-shape attention
|
||
|
|
prototype. Do not add `H3_ATTENTION=vortex_exact` or alter the Sage2 fallback.
|
||
|
|
|
||
|
|
The inline-PTX named-barrier primitive is rejected for the prototype. It is
|
||
|
|
faster in the synthetic barrier-only probe, but Compute Sanitizer racecheck
|
||
|
|
reports five shared-memory hazards because that path does not provide a
|
||
|
|
sanitizer-recognized happens-before edge. The mbarrier path reports zero hazards
|
||
|
|
and remains far below the synchronization budget.
|
||
|
|
|
||
|
|
## Authoritative Probe
|
||
|
|
|
||
|
|
Environment: NVIDIA GB10, SM121, CUDA 13.0, PyTorch 2.9.1+cu130, production
|
||
|
|
image `sha256:6d880d628334c981c3d155bf5244e65e26e22cc9273c80145f646eee3c3698c2`.
|
||
|
|
The run uses 48 blocks, 320 threads, 591 epochs, 20 warmups, and 100 measured
|
||
|
|
iterations.
|
||
|
|
|
||
|
|
| Gate | Measurement | Result |
|
||
|
|
| --- | ---: | --- |
|
||
|
|
| Selected handoff registers | 96/thread | pass, limit 200 |
|
||
|
|
| QK role registers | 54/thread | pass |
|
||
|
|
| PV role registers | 138/thread | pass |
|
||
|
|
| Combined ownership registers | 139/thread | pass |
|
||
|
|
| Local memory / ptxas spills | 0 / 0 | pass |
|
||
|
|
| Selected dynamic shared memory | 51,200 bytes | pass |
|
||
|
|
| Resident ten-warp CTAs/SM | 1 | pass |
|
||
|
|
| mbarrier-only p50 / p95 | `0.076800 / 0.078880 ms` | pass, budget `11.85445 ms` |
|
||
|
|
| Synthetic payload p50 / p95 | `1.874944 / 1.879474 ms` | pass |
|
||
|
|
| 591-epoch payload checksum | `4,671,090` | pass, zero publication errors |
|
||
|
|
| Deterministic repeat | identical | pass |
|
||
|
|
| INT8/FP8 positive clock overlap | 48/48 blocks | pass |
|
||
|
|
| INT8/FP8 overlap p50 ratio | `0.9999929` | capability pass |
|
||
|
|
| Combined issue p50 | `0.127072 ms` | below `0.224848 ms` serial p50 sum |
|
||
|
|
|
||
|
|
The overlap probe establishes that distinct VEA-B warp roles can make progress
|
||
|
|
through INT8 and FP8 `mma.sync` loops concurrently. It is not a throughput model
|
||
|
|
for the complete attention mainloop.
|
||
|
|
|
||
|
|
## Sanitizer And NCU
|
||
|
|
|
||
|
|
- Selected mbarrier payload: memcheck 0 errors; racecheck 0 hazards.
|
||
|
|
- Tensor issue probe: memcheck 0 errors; racecheck 0 hazards.
|
||
|
|
- Rejected inline named barrier: memcheck 0 errors; racecheck 5 hazards.
|
||
|
|
- NCU selected handoff: 96 registers/thread, 51.2 KiB dynamic shared memory,
|
||
|
|
shared-memory block limit 1, `20.91%` achieved occupancy, and zero local spill
|
||
|
|
requests.
|
||
|
|
- NCU tensor issue: 26 registers/thread and zero local spill requests.
|
||
|
|
|
||
|
|
The synthetic handoff payload has severe shared-load bank conflicts in the
|
||
|
|
single-block NCU capture. Its deliberately simple byte access pattern is not an
|
||
|
|
attention layout result, but the aligned prototype must choose and profile a
|
||
|
|
bank-aware score/K/V layout rather than copy this access pattern unchanged.
|
||
|
|
|
||
|
|
## Canonical Fixtures
|
||
|
|
|
||
|
|
Self-contained BF16 NHD Q, K, V, and Sage2 output tensors were captured at
|
||
|
|
`[1, 37810, 56, 128]`, saved, reloaded, and compared byte-for-byte. The output
|
||
|
|
tensor SHA-256 is
|
||
|
|
`4c666c20f5f8f651158a2ced33ccff08f3bada07665c595b99008d171db30574`,
|
||
|
|
matching the locked public Sage2 oracle. Checkpoint and deployed Sage binary
|
||
|
|
hashes also match the numerical contract.
|
||
|
|
|
||
|
|
## Invalid And Rejected Runs
|
||
|
|
|
||
|
|
The first full probe is invalid for payload conclusions. K/V allocation reserved
|
||
|
|
two 8 KiB slots but indexed one 16 KiB region, allowing the next epoch to
|
||
|
|
overwrite consumer data. Deterministic repetition exposed the defect; the
|
||
|
|
corrected probe uses slot-relative 8 KiB indexing and passes deterministic,
|
||
|
|
checksum, memcheck, and racecheck validation.
|
||
|
|
|
||
|
|
The original `16.631 ms` inline-barrier number from that invalid run is not
|
||
|
|
retained as a hardware measurement. The corrected authoritative inline path is
|
||
|
|
`0.056128 ms` p50, but remains rejected on the sanitizer gate.
|
||
|
|
|
||
|
|
## Durable Evidence
|
||
|
|
|
||
|
|
- `/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/probes/vea-b-capability-20260826-authoritative.json`
|
||
|
|
- `/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/build/vea-b-build-20260826-authoritative.log`
|
||
|
|
- `/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/fixtures/canonical-20260826/manifest.json`
|
||
|
|
- `/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/sanitizer/`
|
||
|
|
- `/home/daniel/StoryStudioAssets/H3-output/h3-blackwell-runtime/research/vortex_exact_attention/phase2a/ncu/`
|
||
|
|
|
||
|
|
## Next Gate
|
||
|
|
|
||
|
|
Implement only one aligned short-shape VEA-B prototype with the selected
|
||
|
|
mbarrier handoff and exact D=128 arithmetic. It must compare against the captured
|
||
|
|
prepared-byte fixtures before any canonical or integrated timing. Phase 2A does
|
||
|
|
not validate the heuristic `180-207 ms` mainloop screen.
|