# Performance Model All candidate values are projections. No Vortex kernel has been implemented or timed. ## Baseline | Metric | Value | | --- | ---: | | Mainloop p50 | `237.089 ms` uninstrumented | | Mainloop plus Sage quantization/preparation | approximately `257.7 ms` | | Full owned boundary including Q/K RMSNorm and RoPE | approximately `269.9 ms` | | Grid / CTA | 16,576 CTAs / 128 threads | | KV iterations | 591 | | Registers / shared memory | 255/thread / 32 KiB | | Resident CTAs / warps | 2 / 8 | | No-eligible cycles | `63.53%` | | Combined INT8 QK + FP8 PV tensor activity | `75.54%` | | L2 hit / memory throughput | `98.85%` / approximately `31.74%` | The model targets shorter fixed-dependency chains and disjoint live-state ownership. It assigns no speedup to launch removal, DRAM, tails, or padding. ## Resource Projection | Candidate | Threads | Role-weighted registers/CTA | Shared memory | CTAs/SM projection | Active warps | QK owns RO? | PV owns m/d? | | --- | ---: | ---: | ---: | ---: | ---: | --- | --- | | VEA-A | 256 | 24-32K | 51-52 KiB | 1 | 8 | No | No | | VEA-B | 320 | 56,320 at uniform 176 registers | 50-52 KiB aliased | 1 | 10 | No | No | | VEA-C pair | 384 total | 50-60K pair | 50+18 KiB | unverified cluster | 12 | No | No | Logical ownership shortens per-warp dependency chains but does not prove role-weighted physical register allocation. VEA-B therefore uses the uniform largest-role projection; compiled allocation above 200 registers/thread or any material local spilling rejects it. ## Latency Budget | Component | VEA-A | VEA-B | VEA-C | Basis | | --- | ---: | ---: | ---: | --- | | Recovered dependency-stall fraction assumption | 22-35% | 30-42% | 38-50% | screening assumption, not measured | | Handoff and synchronization | 8-18 ms | 6-15 ms | 12-30 ms | unmeasured named-barrier/DSM budget | | Projected mainloop | `192-222 ms` | `180-207 ms` | `174-210 ms` | heuristic screen, not achieved | | Non-overlapped exact preparation | 25-28 ms | 25-28 ms | 25-28 ms | P1 entry 12.33 + K mean 2.37 + K quant 3.83 + P2 V 6.39 = 24.92 ms, plus margin | | Projected complete attention | `217-250 ms` | `205-235 ms` | `199-238 ms` | not achieved | The screening range starts from `237.089 ms`, assumes a candidate-specific fraction of the measured `63.53%` no-eligible region becomes useful overlap, then adds the handoff budget. Scheduler percentages do not translate linearly to time, so this is a rejection model rather than a performance prediction. Capability microbenchmarks must replace every assumption before implementation. ## Rejection Gates - Compiled local spills materially above zero. - Mainloop model lower bound cannot beat `220 ms`. - Shared memory prevents projected resident work. - QK and PV cannot issue concurrently across roles. - Handoff requires a global N-by-N score/probability tensor. - Handoff changes FP8 probability bytes, RO rescale order, tile order, or final reciprocal/conversion. - p95 has no bounded synchronization path. VEA-B is selected because it is the only intra-CTA design with a credible sub-190 range that does not depend on unverified DSM. ## End-To-End Scenario Math With baseline sampling `255.447 s` and attention share approximately `159.3 s`: ```text sampling(s) = 255.447 - 159.3 + 159.3 / attention_speedup ``` | Attention speedup scenario | Estimated sampling | Status | | ---: | ---: | --- | | 1.15x minimum viable | `234.67 s` | target only | | 1.30x strong | `218.70 s` | target only | | 1.50x research-grade | `202.35 s` | target only | | 2.00x breakthrough | `175.80 s` | target only | | 2.70x stretch | `155.15 s` | target only; approximately half the original `310.11 s` baseline | These scenarios are not Vortex results and do not account for interactions, clock changes, or non-attention regressions. ## Required Candidate Evidence Executable candidates must use alternating AB/BA runs with one resident model and report p10/p50/p90/p95, registers, occupancy/resident CTAs, eligible and no-eligible cycles, INT8/FP8 utilization, instruction mix, shared traffic and conflicts, L2 requests/hit rate, off-chip proxy, spills, synchronization stalls, and kernel count. One-sample smoke timings cannot advance a candidate.