vllm_omni.diffusion.distributed.autoencoders.autoencoder_kl_ltx2 ¶
DistributedAutoencoderKLLTX2Video ¶
Bases: AutoencoderKLLTX2Video, DistributedVaeMixin
encode_tile_exec ¶
Encode one sample-space tile into latent moments.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
task | TileTask | Tile task whose tensor is a sample-space slice | required |
causal | bool | None | Optional causal-padding override forwarded to the encoder. | None |
Returns:
| Type | Description |
|---|---|
Tensor | Encoder output of shape ``(B, 2 * latent_channels, |
Tensor | (F - 1) // temporal_compression_ratio + 1, tile_h // scr, tile_w // scr)``. |
encode_tile_merge ¶
encode_tile_merge(
coord_tensor_map: dict[tuple[int, ...], Tensor],
grid_spec: GridSpec,
) -> Tensor
Blend and stitch encoded tiles into the full latent tensor.
Mirrors the sequential blending of diffusers' tiled_encode: each tile is blended against its upper/left neighbors over the latent-space blend_* extents, cropped to tile_latent_stride_*, concatenated, and finally cropped to (latent_height, latent_width). Returns latent moments of shape (B, 2 * latent_channels, F_latent, latent_height, latent_width).
encode_tile_split ¶
Split a sample-space video into overlapping spatial tiles for parallel encoding.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x | Tensor | Sample-space input of shape | required |
Returns:
| Name | Type | Description |
|---|---|---|
list[TileTask] | A | |
GridSpec | sample-space tile. | |
metadata | tuple[list[TileTask], GridSpec] |
|
tuple[list[TileTask], GridSpec] |
| |
tuple[list[TileTask], GridSpec] | crop). When every tile's spatial dims divide | |
tuple[list[TileTask], GridSpec] | it also carries | |
tuple[list[TileTask], GridSpec] | predicted as ``(B, 2 * latent_channels, | |
tuple[list[TileTask], GridSpec] | (F - 1) // temporal_compression_ratio + 1, tile_h // scr, tile_w // scr)`` — | |
tuple[list[TileTask], GridSpec] | which lets :class: | |
tuple[list[TileTask], GridSpec] | all-reduce. For tiles with non-divisible spatial dims the encoder output | |
tuple[list[TileTask], GridSpec] | shape is not predictable, so both keys are omitted and the executor falls | |
tuple[list[TileTask], GridSpec] | back to the dynamic metadata-gather path. |
tile_exec ¶
tile_merge ¶
tiled_decode ¶
tiled_decode(
z: Tensor,
temb: Tensor | None = None,
causal: bool | None = None,
return_dict: bool = True,
)
tiled_encode ¶
tiled_encode(
x: Tensor, causal: bool | None = None
) -> Tensor
Distributed tiled encode: split across the DiT group, encode, gather, merge.
Drop-in replacement for diffusers' sequential tiled_encode — same signature, same return (latent moments (B, 2 * latent_channels, F_latent, latent_height, latent_width)), numerically identical output. Falls back to the sequential parent implementation when distribution is disabled (vae_patch_parallel_size <= 1, tiling off, or no process group). The merged result is broadcast to all ranks because every rank's denoiser consumes the latents.
LTX2VaeExecutor ¶
Bases: DistributedVaeExecutor