Commit graph

15 commits

Author SHA1 Message Date
Daniel Maddern
9bb96a26e8 Fix FL2VA conditioning parity end to end 2026-08-20 16:43:22 +07:00
Daniel Maddern
0b7217485c nin_shortcut is a plain 1x1x1 conv (no causal padding) 2026-08-19 21:49:40 +07:00
Daniel Maddern
ea6ab87a34 Decouple spatial_padding from temporal_causal in causal conv 2026-08-19 21:47:00 +07:00
Daniel Maddern
b557cba173 Fix Downsample3D spatial reflect pad to (W,H) dims only 2026-08-19 21:41:16 +07:00
Daniel Maddern
136297f4ae Fix F.pad dim order (spatial reflect + no T pad) for 1-frame path 2026-08-19 21:36:40 +07:00
Daniel Maddern
c2d507691b Restore front-zero temporal pad for multi-frame causal conv; two-step F.pad (reflect spatial + constant T) 2026-08-19 21:35:07 +07:00
Daniel Maddern
f7c48b067a Remove permute bug: rely on F.conv3d spatial padding 2026-08-19 21:33:08 +07:00
Daniel Maddern
6fc7cf0ad6 Fix causal conv: correct F.pad spatial dim order + 1-frame kernel truncation 2026-08-19 21:31:05 +07:00
Daniel Maddern
3386975326 Fix VAE: use same multi-frame causal path for keyframes, add frame_pre_padding so 1 keyframe -> 1 latent 2026-08-19 21:20:32 +07:00
Daniel Maddern
a0274a9868 Fix tiled_encode latent x-overlap off-by-one (use [j-1] for left neighbor) 2026-08-19 21:13:40 +07:00
Daniel Maddern
65e80be1cf Thread single_frame flag so keyframe truncation does not apply to 1-frame tiles 2026-08-19 21:11:43 +07:00
Daniel Maddern
06fd79a8fd Fix causal temporal padding to match reference (2k-1 front zeros when spatial_padding>0) 2026-08-19 21:09:03 +07:00
Daniel Maddern
390135fca6 Fix VAE pixel normalization in-place bug 2026-08-19 21:07:34 +07:00
Daniel Maddern
9bec53ac36 Fix VAE encoder to load canonical checkpoint key names 2026-08-19 21:04:30 +07:00
Daniel Maddern
807bd64a82 Add direct first/last-frame (fl2va) keyframe conditioning
Wire full fl2va into the direct H3 runtime so first/last keyframes flow
through VAE encode -> Qwen vision tokens -> DiT cond segments:

- vae_encoder.py: direct encoder-only H3 video VAE (causal 3D convs,
  reflect spatial padding, causal temporal padding, single-frame tap
  truncation, tiling, FP32 moments + mean/std normalization).
- qwen3vl_vision.py: Qwen3-VL-32B visual tower (27 blocks, 2D rotary,
  deepstack mergers) ported to match the Comfy reference exactly
  (head_dim=72, no-bias proj, LayerNorm blocks, split-half apply_rope),
  plus Qwen image preprocess, mrope ids/freqs, DiT token tags, keyframe
  resize (first=stretch / last=center cover-crop matching Comfy
  common_upscale), and build_fl2va_presentation.
- qwen3vl_text.py: split-half apply_rope, _embed_rows/_run_layers,
  optional mrope position_ids + DeepStack injection at the first three
  decoder layers at visual positions.
- packing.py: H3PromptPacker builds [text | cond | audio | video] with
  tag-run text spans, cond rows (first/last cond_t anchors,
  VISUAL_COND_TIMESTEP=0.999 noise augmentation via CPU-seeded RNG),
  three-timestep row table (t_row*3 + modality_tag), and rope positions.
- runtime.py: load VAE encoder + vision tower; generate() accepts
  first_frame/last_frame, builds the fl2va presentation, encodes keyframes,
  and passes text_token_tags/cond_latents/frame_count/seed to the sampler.
- sampler.py: thread pack kwargs + seed.
- serve_hot_runtime.py / direct_t2v_preview.py: /generate and
  --first-frame/--last-frame accept image paths or base64.
2026-08-19 20:17:41 +07:00