Skip to content

vllm_omni.diffusion.models.seedvr2.pipeline_seedvr2

SeedVR2 restoration pipeline: input admission, one Euler step and colour transfer.

CALIBRATED_SHARDED_CLIP_PIXELS module-attribute

CALIBRATED_SHARDED_CLIP_PIXELS = 5 * 2560 * 1472

CLIP_BYTES_PER_PIXEL module-attribute

CLIP_BYTES_PER_PIXEL = 8.5

DECODED_BYTES_PER_PIXEL module-attribute

DECODED_BYTES_PER_PIXEL = 28

DEFAULT_LUMINANCE_WEIGHT module-attribute

DEFAULT_LUMINANCE_WEIGHT = 0.8

ENCODER_BYTES_PER_PIXEL module-attribute

ENCODER_BYTES_PER_PIXEL = 1140

ENCODER_CHUNK_FRAMES module-attribute

ENCODER_CHUNK_FRAMES = 9

FRAGMENTATION module-attribute

FRAGMENTATION = 1.37

MAX_CLIP_PIXELS module-attribute

MAX_CLIP_PIXELS = 5 * MAX_FRAME_PIXELS

MAX_FRAME_PIXELS module-attribute

MAX_FRAME_PIXELS = 848 * 480

USABLE_MEMORY_FRACTION module-attribute

USABLE_MEMORY_FRACTION = 0.85

VALIDATED_SHARDED_CLIP_PIXELS module-attribute

VALIDATED_SHARDED_CLIP_PIXELS = 513 * 1536 * 2688

WAVELET_LEVELS module-attribute

WAVELET_LEVELS = 5

WEIGHT_BYTES module-attribute

WEIGHT_BYTES = 7 * 1024 ** 3

SeedVR2Input dataclass

frame_count instance-attribute

frame_count: int

frames instance-attribute

frames: Tensor

source class-attribute instance-attribute

source: SourceVideo | None = None

SeedVR2Pipeline

Bases: Module

device instance-attribute

device = get_local_device()

od_config instance-attribute

od_config = od_config

supports_request_batch class-attribute instance-attribute

supports_request_batch = True

transformer instance-attribute

transformer = SeedVR2NaDiT(
    **SEEDVR2_3B_CONFIG, use_varlen_kernel=False
)

vae instance-attribute

vae = SeedVR2VAE()

weights_sources instance-attribute

weights_sources = [
    DiffusersPipelineLoader.ComponentSource(
        model_or_path=od_config.model,
        subfolder=None,
        revision=od_config.revision,
        prefix=component + ".",
        fall_back_to_pt=False,
        allow_patterns_overrides=[filename],
    )
    for component, filename in (
        ("transformer", "seedvr2_ema_3b_fp16.safetensors"),
        ("vae", "ema_vae_fp16.safetensors"),
    )
]

forward

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

SourceVideo dataclass

audio instance-attribute

audio: Tensor | None

audio_sample_rate instance-attribute

audio_sample_rate: int | None

fps instance-attribute

fps: float

pts instance-attribute

pts: tuple[int, ...]

time_base instance-attribute

time_base: Fraction

correct_video_color

correct_video_color(
    video: Tensor,
    reference: Tensor,
    method: str = DEFAULT_COLOR_CORRECTION_METHOD,
    luminance_weight: float = DEFAULT_LUMINANCE_WEIGHT,
) -> Tensor

Transfer reference colour onto video; both are [B,C,T,H,W] in [0,1].

Frames are processed one at a time because histogram matching sorts every pixel of a channel. Inputs are never modified.

finish_video

finish_video(
    decoded: Tensor,
    sample: Tensor,
    frame_count: int,
    method: str,
) -> Tensor

Colour-correct and quantize decoded [B,3,T,H,W] frames to uint8 [B,T,H,W,3].

One frame at a time: whole-clip float copies would dominate device memory on long clips, and a per-frame slice of one would need 64-bit indexing once the clip passes 2**31 pixels.

get_seedvr2_post_process_func

get_seedvr2_post_process_func(
    _od_config: OmniDiffusionConfig,
) -> Callable[[dict[str, object]], dict[str, object]]

get_seedvr2_pre_process_func

get_seedvr2_pre_process_func(
    od_config: OmniDiffusionConfig,
) -> Callable[[OmniDiffusionRequest], OmniDiffusionRequest]

max_frames

max_frames() -> int

Decoder-work frame cap, shared by every serving profile.

prepare_request

prepare_request(
    request: OmniDiffusionRequest,
    *,
    frame_pixels: int = MAX_FRAME_PIXELS,
    clip_pixels: int = MAX_CLIP_PIXELS,
) -> OmniDiffusionRequest

prepare_video

prepare_video(
    frames: Tensor, height: int, width: int, device: device
) -> Tensor

Resize RGB frames to the model grid on device and repeat the final frame.

frames is uint8 [T,H,W,3] or float [T,3,H,W] in [0,1]. Frames are converted one at a time, so the device never holds a float copy of the whole clip, and returns fp16 [1,3,T',H,W] in [-1,1] with T' = 4n+1.

read_video

read_video(
    path: str | Path,
    frame_pixels: int = MAX_FRAME_PIXELS,
    clip_pixels: int = MAX_CLIP_PIXELS,
) -> tuple[Tensor, SourceVideo]

sample_noise

sample_noise(
    condition: Tensor, generator: Generator
) -> Tensor

scaled_clip_pixels

scaled_clip_pixels(
    frame_pixels: int, device_memory: int | None
) -> int

Largest padded clip the memory model admits on device_memory.

The result never falls below the calibrated budget or rises above the validated one. The worst case for a clip budget is a clip of full-size frames, so the model is evaluated at frame_pixels; smaller frames only need less.

sharded_budget

sharded_budget() -> tuple[int, int]

Per-frame and padded-clip budgets for the VAE-sharded serving profile.

An unset clip budget scales with the smallest visible device's memory.

validate_clip_size

validate_clip_size(
    frame_count: int,
    height: int,
    width: int,
    frame_pixels: int,
    clip_pixels: int,
) -> None

validate_seedvr2_config

validate_seedvr2_config(
    config: OmniDiffusionConfig,
) -> None