186 lines
6.3 KiB
Markdown
186 lines
6.3 KiB
Markdown
|
|
# Vortex Research Agenda
|
||
|
|
|
||
|
|
## Flagship Research Project
|
||
|
|
|
||
|
|
**Vortex: Reference-Exact, Communication-Avoiding Video Diffusion on Commodity PCIe GPUs**
|
||
|
|
|
||
|
|
The core research question:
|
||
|
|
|
||
|
|
> Can a large multimodal video diffusion model achieve near-NVLink scaling on
|
||
|
|
> ordinary PCIe hardware while preserving reference-exact output?
|
||
|
|
|
||
|
|
The existing `86.1%` two-GPU efficiency is the starting evidence, not the final
|
||
|
|
contribution.
|
||
|
|
|
||
|
|
## Novel Architecture
|
||
|
|
|
||
|
|
Build an end-to-end H3 block where:
|
||
|
|
|
||
|
|
- NVFP4 QKV projections write directly into rank-destined communication buffers.
|
||
|
|
- Q, K, and V are never separately materialized or repacked.
|
||
|
|
- Ulysses communication begins at tile granularity while later projection tiles
|
||
|
|
are still computing.
|
||
|
|
- Video, audio, text, and conditioning tokens use different sharding policies.
|
||
|
|
- Small conditioning segments are replicated when replication is cheaper than
|
||
|
|
communication.
|
||
|
|
- Received attention output writes directly into the output-projection layout.
|
||
|
|
- Reference BF16 rounding boundaries are preserved exactly.
|
||
|
|
- The scheduler automatically accounts for PIX, PXB, PHB, NODE, and SYS topology.
|
||
|
|
|
||
|
|
The target pipeline is:
|
||
|
|
|
||
|
|
```text
|
||
|
|
NVFP4 projection
|
||
|
|
|
|
||
|
|
v
|
||
|
|
rank-destined QKV tiles
|
||
|
|
|
|
||
|
|
v overlapping
|
||
|
|
PCIe all-to-all
|
||
|
|
|
|
||
|
|
v overlapping
|
||
|
|
attention
|
||
|
|
|
|
||
|
|
v
|
||
|
|
output projection layout
|
||
|
|
```
|
||
|
|
|
||
|
|
There is no intermediate QKV pack, full synchronization point, or redundant
|
||
|
|
memory round trip.
|
||
|
|
|
||
|
|
Communication and computation overlap has already been explored by systems such
|
||
|
|
as Ulysses Unbound and topology-aware systems such as SwiftFusion. The
|
||
|
|
contribution therefore cannot simply be "we overlapped an all-to-all."
|
||
|
|
|
||
|
|
The novel combination is:
|
||
|
|
|
||
|
|
- Direct NVFP4 GEMM-to-collective dataflow.
|
||
|
|
- Ragged multimodal sequence parallelism.
|
||
|
|
- Segment-aware replication and sharding.
|
||
|
|
- Reference-exact numerical behavior.
|
||
|
|
- Automatic PCIe topology adaptation.
|
||
|
|
- Scaling across inexpensive non-NVLink hardware.
|
||
|
|
|
||
|
|
That combination appears meaningfully ahead of public H3 implementations.
|
||
|
|
|
||
|
|
## Rounding-Boundary-Aware Exact Fusion
|
||
|
|
|
||
|
|
Create a compiler or kernel-generation framework that understands where the
|
||
|
|
original model performs BF16 rounding.
|
||
|
|
|
||
|
|
Ordinary fusion can silently change output because fused multiply-add performs
|
||
|
|
one rounding where separate operations perform two. NVIDIA documents this
|
||
|
|
numerical distinction in its CUDA floating-point guidance.
|
||
|
|
|
||
|
|
The system would:
|
||
|
|
|
||
|
|
- Trace the reference H3 graph.
|
||
|
|
- Mark mandatory numerical materialization boundaries.
|
||
|
|
- Fuse everything between safe boundaries.
|
||
|
|
- Generate Triton or CUDA kernels.
|
||
|
|
- Prove equivalence with adversarial and randomized testing.
|
||
|
|
- Preserve bitwise behavior across SM100, SM120, and SM121 where possible.
|
||
|
|
|
||
|
|
Apply it to:
|
||
|
|
|
||
|
|
- AdaLN modulation.
|
||
|
|
- Residual gates.
|
||
|
|
- SwiGLU.
|
||
|
|
- QK normalization and RoPE.
|
||
|
|
- NVFP4 scaling and packing.
|
||
|
|
- GEMM epilogues.
|
||
|
|
|
||
|
|
This could be a broader contribution than Vortex itself. It would answer:
|
||
|
|
|
||
|
|
> How much of a mixed-precision transformer can be fused without altering its
|
||
|
|
> numerical program?
|
||
|
|
|
||
|
|
That is a serious systems and numerical-computing paper.
|
||
|
|
|
||
|
|
## Trajectory-Bounded Sparse Attention
|
||
|
|
|
||
|
|
This is the highest-risk project.
|
||
|
|
|
||
|
|
The Sol results show why simple threshold sparsity is not enough:
|
||
|
|
|
||
|
|
- It becomes fast at higher thresholds.
|
||
|
|
- Attention-level relative error rises to roughly `0.5-0.6`.
|
||
|
|
- Local attention error does not reveal what happens to the final video.
|
||
|
|
|
||
|
|
Instead of using one global threshold, allocate an error budget across:
|
||
|
|
|
||
|
|
- Denoising timestep.
|
||
|
|
- Transformer block.
|
||
|
|
- Attention head.
|
||
|
|
- Video, audio, text, and conditioning segments.
|
||
|
|
- Spatial and temporal regions.
|
||
|
|
|
||
|
|
The runtime would choose dense or sparse attention dynamically, with exact
|
||
|
|
fallback when its predicted error exceeds the remaining trajectory budget.
|
||
|
|
|
||
|
|
Existing research already covers distributed sparse attention, including DSA,
|
||
|
|
and feature reuse such as FasterCache. The potential Vortex contribution is
|
||
|
|
different:
|
||
|
|
|
||
|
|
> Optimize sparsity against final multimodal trajectory error, not isolated
|
||
|
|
> attention-tensor error.
|
||
|
|
|
||
|
|
That requires proper evaluation of image quality, temporal consistency, audio
|
||
|
|
quality, lip synchronization, and prompt adherence.
|
||
|
|
|
||
|
|
## Practical Ranking
|
||
|
|
|
||
|
|
| Research direction | Novelty | Technical risk | Commercial value | Recommendation |
|
||
|
|
| --- | --- | --- | --- | --- |
|
||
|
|
| PCIe-native ragged Ulysses | Very high | Medium | Exceptional | First |
|
||
|
|
| Rounding-aware exact fusion | Very high | Medium | Exceptional | Build alongside |
|
||
|
|
| Trajectory-bounded sparsity | Very high | Very high | Potentially exceptional | Later |
|
||
|
|
| NVFP4 block superkernel | Medium-high | Medium | Very high | Core component |
|
||
|
|
| Topology-autonomous planner | Medium | Low-medium | High | Supporting work |
|
||
|
|
| Conventional caching | Medium-low | Medium | High | Only with a novel error guarantee |
|
||
|
|
|
||
|
|
## Landmark Result
|
||
|
|
|
||
|
|
A compelling target is:
|
||
|
|
|
||
|
|
| Scale | Target efficiency | Interconnect |
|
||
|
|
| --- | ---: | --- |
|
||
|
|
| 2 GPUs | Above `90%` | PCIe, no NVLink |
|
||
|
|
| 4 GPUs | Above `80%` | PCIe |
|
||
|
|
| 8 GPUs | Above `70%` | PCIe or cross-NUMA |
|
||
|
|
| Output | Reference-exact or formally bounded | All scales |
|
||
|
|
| Portability | SM100, SM120, SM121 | Same runtime |
|
||
|
|
|
||
|
|
Also measure:
|
||
|
|
|
||
|
|
- End-to-end latency, not sampling alone.
|
||
|
|
- Exposed communication time.
|
||
|
|
- Bytes transferred per block.
|
||
|
|
- HBM traffic and achieved bandwidth.
|
||
|
|
- Tensor-core utilization.
|
||
|
|
- Cost per generated second.
|
||
|
|
- Energy per generated second.
|
||
|
|
- PIX/NODE/PHB/SYS topology sensitivity.
|
||
|
|
- Comparison against SGLang, xDiT, and LightX2V.
|
||
|
|
|
||
|
|
## Recommendation
|
||
|
|
|
||
|
|
Do not lead with sparse attention yet. Lead with:
|
||
|
|
|
||
|
|
> Reference-exact, communication-avoiding ragged Ulysses with direct NVFP4
|
||
|
|
> projection-to-collective dataflow.
|
||
|
|
|
||
|
|
The first fused modulation kernel establishes the exact-fusion infrastructure.
|
||
|
|
Use that infrastructure next to build the direct QKV-to-Ulysses path.
|
||
|
|
|
||
|
|
The measured single-GPU implementation order is direct QKV-to-attention layout,
|
||
|
|
direct attention-to-output-projection layout, streamed NVFP4 scale/pack/GEMM,
|
||
|
|
and exact residual-gate epilogues. This order attacks the measured `105.83 ms`
|
||
|
|
layout boundary first while preserving the larger projection-to-collective
|
||
|
|
architecture as the distributed destination.
|
||
|
|
|
||
|
|
If strong four-GPU and eight-GPU scaling over PCIe can be demonstrated while
|
||
|
|
retaining exact output, this stops being merely an excellent H3 runtime. It
|
||
|
|
becomes credible new research into how large generative models should be
|
||
|
|
executed without proprietary high-bandwidth interconnects.
|