Complete Vortex attention interface design
This commit is contained in:
parent
7d0608e857
commit
ff8f7672b3
1 changed files with 16 additions and 0 deletions
|
|
@ -166,3 +166,19 @@ boundary rather than a temporal schedule rearrangement.
|
|||
Stop immediately on changed KV order, any material score tensor, nonzero output
|
||||
difference, forced local spilling, inability to beat the `220 ms` model gate, or
|
||||
fallback to Sage2 for main computation.
|
||||
|
||||
## Interface And Future Ulysses Layout
|
||||
|
||||
The isolated API consumes projection-strided BF16 NHD views directly. Its work
|
||||
descriptor must carry `q_start`, `q_count`, `kv_count`, `head_start`,
|
||||
`head_count`, source token/head strides, destination token/head strides, and a
|
||||
future rank destination. No HND transpose or contiguous input copy is allowed.
|
||||
|
||||
Output stores land directly in token-major `[token, local_head*128+d]` rows for
|
||||
the local output projection. Ragged local token counts and nonzero destination
|
||||
offsets are legal, but each rank must receive the globally ordered KV sequence
|
||||
before attention. Independently normalized KV partials must never be merged.
|
||||
|
||||
VEA-B is optimized first for the single-GPU Q128 shape. A future SM120/SM100 or
|
||||
Ulysses schedule may choose Q64 or a different producer geometry, but it must
|
||||
retain this numerical and descriptor contract and pass separately tuned gates.
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue