vllm_omni.diffusion.models.seedvr2.pipeline_seedvr2 ¶
SeedVR2 restoration pipeline: input admission, one Euler step and colour transfer.
SeedVR2Input dataclass ¶
SeedVR2Pipeline ¶
Bases: Module
transformer instance-attribute ¶
transformer = SeedVR2NaDiT(
**SEEDVR2_3B_CONFIG, use_varlen_kernel=False
)
weights_sources instance-attribute ¶
weights_sources = [
DiffusersPipelineLoader.ComponentSource(
model_or_path=od_config.model,
subfolder=None,
revision=od_config.revision,
prefix=component + ".",
fall_back_to_pt=False,
allow_patterns_overrides=[filename],
)
for component, filename in (
("transformer", "seedvr2_ema_3b_fp16.safetensors"),
("vae", "ema_vae_fp16.safetensors"),
)
]
SourceVideo dataclass ¶
correct_video_color ¶
correct_video_color(
video: Tensor,
reference: Tensor,
method: str = DEFAULT_COLOR_CORRECTION_METHOD,
luminance_weight: float = DEFAULT_LUMINANCE_WEIGHT,
) -> Tensor
Transfer reference colour onto video; both are [B,C,T,H,W] in [0,1].
Frames are processed one at a time because histogram matching sorts every pixel of a channel. Inputs are never modified.
finish_video ¶
Colour-correct and quantize decoded [B,3,T,H,W] frames to uint8 [B,T,H,W,3].
One frame at a time: whole-clip float copies would dominate device memory on long clips, and a per-frame slice of one would need 64-bit indexing once the clip passes 2**31 pixels.
get_seedvr2_post_process_func ¶
get_seedvr2_post_process_func(
_od_config: OmniDiffusionConfig,
) -> Callable[[dict[str, object]], dict[str, object]]
get_seedvr2_pre_process_func ¶
get_seedvr2_pre_process_func(
od_config: OmniDiffusionConfig,
) -> Callable[[OmniDiffusionRequest], OmniDiffusionRequest]
prepare_request ¶
prepare_request(
request: OmniDiffusionRequest,
*,
frame_pixels: int = MAX_FRAME_PIXELS,
clip_pixels: int = MAX_CLIP_PIXELS,
) -> OmniDiffusionRequest
prepare_video ¶
Resize RGB frames to the model grid on device and repeat the final frame.
frames is uint8 [T,H,W,3] or float [T,3,H,W] in [0,1]. Frames are converted one at a time, so the device never holds a float copy of the whole clip, and returns fp16 [1,3,T',H,W] in [-1,1] with T' = 4n+1.
read_video ¶
read_video(
path: str | Path,
frame_pixels: int = MAX_FRAME_PIXELS,
clip_pixels: int = MAX_CLIP_PIXELS,
) -> tuple[Tensor, SourceVideo]
scaled_clip_pixels ¶
Largest padded clip the memory model admits on device_memory.
The result never falls below the calibrated budget or rises above the validated one. The worst case for a clip budget is a clip of full-size frames, so the model is evaluated at frame_pixels; smaller frames only need less.
sharded_budget ¶
Per-frame and padded-clip budgets for the VAE-sharded serving profile.
An unset clip budget scales with the smallest visible device's memory.