Commit graph

17 commits

Author SHA1 Message Date
Daniel Maddern
807bd64a82 Add direct first/last-frame (fl2va) keyframe conditioning
Wire full fl2va into the direct H3 runtime so first/last keyframes flow
through VAE encode -> Qwen vision tokens -> DiT cond segments:

- vae_encoder.py: direct encoder-only H3 video VAE (causal 3D convs,
  reflect spatial padding, causal temporal padding, single-frame tap
  truncation, tiling, FP32 moments + mean/std normalization).
- qwen3vl_vision.py: Qwen3-VL-32B visual tower (27 blocks, 2D rotary,
  deepstack mergers) ported to match the Comfy reference exactly
  (head_dim=72, no-bias proj, LayerNorm blocks, split-half apply_rope),
  plus Qwen image preprocess, mrope ids/freqs, DiT token tags, keyframe
  resize (first=stretch / last=center cover-crop matching Comfy
  common_upscale), and build_fl2va_presentation.
- qwen3vl_text.py: split-half apply_rope, _embed_rows/_run_layers,
  optional mrope position_ids + DeepStack injection at the first three
  decoder layers at visual positions.
- packing.py: H3PromptPacker builds [text | cond | audio | video] with
  tag-run text spans, cond rows (first/last cond_t anchors,
  VISUAL_COND_TIMESTEP=0.999 noise augmentation via CPU-seeded RNG),
  three-timestep row table (t_row*3 + modality_tag), and rope positions.
- runtime.py: load VAE encoder + vision tower; generate() accepts
  first_frame/last_frame, builds the fl2va presentation, encodes keyframes,
  and passes text_token_tags/cond_latents/frame_count/seed to the sampler.
- sampler.py: thread pack kwargs + seed.
- serve_hot_runtime.py / direct_t2v_preview.py: /generate and
  --first-frame/--last-frame accept image paths or base64.
2026-08-19 20:17:41 +07:00
Daniel Maddern
8730920634 Default Spark runs to Sol attention 2026-08-15 03:37:50 +07:00
Daniel Maddern
c3cc04e98d Add approximate H3 cache modes 2026-08-14 20:38:48 +07:00
Daniel Maddern
fb257ef982 Add exact memory backend options 2026-08-14 20:34:41 +07:00
Daniel Maddern
eaf9324145 Add KJ Sage attention backends 2026-08-14 20:28:56 +07:00
Daniel Maddern
a388aa64ff Default video VAE decode to FP16 2026-08-14 19:50:36 +07:00
Daniel Maddern
53cd8bd4f4 Make runtime diagnostics opt in 2026-08-14 14:11:36 +07:00
Daniel Maddern
59b8302d4e Add direct H3 audio VAE decode 2026-08-14 13:39:46 +07:00
Daniel Maddern
6b0050adfa Add standalone H3 latent decode tooling 2026-08-14 00:11:17 +07:00
Daniel Maddern
c3d72d9f1e Timestamp direct memory profile output 2026-08-13 23:49:41 +07:00
Daniel Maddern
8ed9eecf17 Add fastsafetensors loader option 2026-08-13 23:35:01 +07:00
Daniel Maddern
ce022f06e9 Add direct mmap and memory profiling controls 2026-08-13 23:31:30 +07:00
Daniel Maddern
5f68f1dce9 Decode preview VAE in inference mode 2026-08-13 23:29:29 +07:00
Daniel Maddern
d289b3c497 Allow captured H3 timesteps in preview 2026-08-13 22:23:07 +07:00
Daniel Maddern
09f5d2ef65 Match Comfy joint AV noise initialization 2026-08-13 16:12:10 +07:00
Daniel Maddern
eed3a3d951 Add H3 parity diagnostics 2026-08-12 21:11:02 +07:00
Daniel Maddern
8851873fb0 Initial direct H3 runtime 2026-08-12 14:12:42 +07:00