4.7 KiB
Performance Model
Candidate attention latency values remain projections. Phase 2A capability kernels have measured resources, handoff cost, and concurrent issue, but no Vortex attention kernel has been implemented or timed.
Baseline
| Metric | Value |
|---|---|
| Mainloop p50 | 237.089 ms uninstrumented |
| Mainloop plus Sage quantization/preparation | approximately 257.7 ms |
| Full owned boundary including Q/K RMSNorm and RoPE | approximately 269.9 ms |
| Grid / CTA | 16,576 CTAs / 128 threads |
| KV iterations | 591 |
| Registers / shared memory | 255/thread / 32 KiB |
| Resident CTAs / warps | 2 / 8 |
| No-eligible cycles | 63.53% |
| Combined INT8 QK + FP8 PV tensor activity | 75.54% |
| L2 hit / memory throughput | 98.85% / approximately 31.74% |
The model targets shorter fixed-dependency chains and disjoint live-state ownership. It assigns no speedup to launch removal, DRAM, tails, or padding.
Resource Projection
| Candidate | Threads | Role-weighted registers/CTA | Shared memory | CTAs/SM projection | Active warps | QK owns RO? | PV owns m/d? |
|---|---|---|---|---|---|---|---|
| VEA-A | 256 | 24-32K | 51-52 KiB | 1 | 8 | No | No |
| VEA-B | 320 | 56,320 at uniform 176 registers | 50-52 KiB aliased | 1 | 10 | No | No |
| VEA-C pair | 384 total | 50-60K pair | 50+18 KiB | unverified cluster | 12 | No | No |
Logical ownership shortens per-warp dependency chains but does not prove role-weighted physical register allocation. VEA-B therefore uses the uniform largest-role projection; compiled allocation above 200 registers/thread or any material local spilling rejects it.
Latency Budget
| Component | VEA-A | VEA-B | VEA-C | Basis |
|---|---|---|---|---|
| Recovered dependency-stall fraction assumption | 22-35% | 30-42% | 38-50% | screening assumption, not measured |
| Handoff and synchronization | 8-18 ms | 6-15 ms | 12-30 ms | unmeasured named-barrier/DSM budget |
| Projected mainloop | 192-222 ms |
180-207 ms |
174-210 ms |
heuristic screen, not achieved |
| Non-overlapped exact preparation | 25-28 ms | 25-28 ms | 25-28 ms | P1 entry 12.33 + K mean 2.37 + K quant 3.83 + P2 V 6.39 = 24.92 ms, plus margin |
| Projected complete attention | 217-250 ms |
205-235 ms |
199-238 ms |
not achieved |
Phase 2A replaces the VEA-B synchronization assumption with capability evidence:
the selected 48-block mbarrier path measures 0.076800 ms p50 and 0.078880 ms
p95 for 591 barrier epochs. Its synthetic payload path measures 1.874944 ms
p50. These values clear the 11.85445 ms budget but must not be substituted into
the attention projection as if they were complete QK/softmax/PV work.
The screening range starts from 237.089 ms, assumes a candidate-specific
fraction of the measured 63.53% no-eligible region becomes useful overlap,
then adds the handoff budget. Scheduler percentages do not translate linearly
to time, so this is a rejection model rather than a performance prediction.
The remaining assumptions concern the complete exact attention implementation
and its latency, not VEA-B's basic SM121 capability.
Rejection Gates
- Compiled local spills materially above zero.
- Mainloop model lower bound cannot beat
220 ms. - Shared memory prevents projected resident work.
- QK and PV cannot issue concurrently across roles.
- Handoff requires a global N-by-N score/probability tensor.
- Handoff changes FP8 probability bytes, RO rescale order, tile order, or final reciprocal/conversion.
- p95 has no bounded synchronization path.
VEA-B is selected because it is the only intra-CTA design with a credible sub-190 range that does not depend on unverified DSM.
End-To-End Scenario Math
With baseline sampling 255.447 s and attention share approximately 159.3 s:
sampling(s) = 255.447 - 159.3 + 159.3 / attention_speedup
| Attention speedup scenario | Estimated sampling | Status |
|---|---|---|
| 1.15x minimum viable | 234.67 s |
target only |
| 1.30x strong | 218.70 s |
target only |
| 1.50x research-grade | 202.35 s |
target only |
| 2.00x breakthrough | 175.80 s |
target only |
| 2.70x stretch | 155.15 s |
target only; approximately half the original 310.11 s baseline |
These scenarios are not Vortex results and do not account for interactions, clock changes, or non-attention regressions.
Required Candidate Evidence
Executable candidates must use alternating AB/BA runs with one resident model and report p10/p50/p90/p95, registers, occupancy/resident CTAs, eligible and no-eligible cycles, INT8/FP8 utilization, instruction mix, shared traffic and conflicts, L2 requests/hit rate, off-chip proxy, spills, synchronization stalls, and kernel count. One-sample smoke timings cannot advance a candidate.