vllm_omni.diffusion.models.seedvr2.vae ¶
SeedVR2 causal VAE with temporal tiling and spatial height sharding.
CausalConv3d ¶
Bases: Conv3d
Replicate the first frame to pad only the past temporal context.
Decoder3d ¶
Bases: Module
up_blocks instance-attribute ¶
up_blocks = nn.ModuleList(
[
DecoderBlock3d(
widths[max(0, i - 1)], width, groups, i
)
for i, width in enumerate(widths)
]
)
DecoderBlock3d ¶
Bases: Module
resnets instance-attribute ¶
resnets = nn.ModuleList(
[ResnetBlock3d(in_channels, out_channels, groups)]
+ [
ResnetBlock3d(out_channels, out_channels, groups)
for _ in range(2)
]
)
upsamplers instance-attribute ¶
upsamplers = nn.ModuleList(
[Upsample3d(out_channels, temporal=index < 2)]
if index < 3
else []
)
Downsample3d ¶
Bases: Module
conv instance-attribute ¶
conv = CausalConv3d(
channels,
channels,
(3 if temporal else 1, 3, 3),
(2 if temporal else 1, 2, 2),
(1 if temporal else 0, 0, 0),
)
Encoder3d ¶
Bases: Module
down_blocks instance-attribute ¶
down_blocks = nn.ModuleList(
[
EncoderBlock3d(
channels[max(0, i - 1)], width, groups, i
)
for i, width in enumerate(channels)
]
)
EncoderBlock3d ¶
Bases: Module
downsamplers instance-attribute ¶
downsamplers = nn.ModuleList(
[Downsample3d(out_channels, temporal=index >= 1)]
if index < 3
else []
)
resnets instance-attribute ¶
resnets = nn.ModuleList(
[
ResnetBlock3d(in_channels, out_channels, groups),
ResnetBlock3d(out_channels, out_channels, groups),
]
)
MidBlock3d ¶
Bases: Module
attentions instance-attribute ¶
attentions = nn.ModuleList(
[
Attention(
channels,
heads=1,
dim_head=channels,
eps=1e-06,
norm_num_groups=groups,
residual_connection=True,
bias=True,
upcast_softmax=True,
_from_deprecated_attn_block=True,
)
]
)
resnets instance-attribute ¶
resnets = nn.ModuleList(
[
ResnetBlock3d(channels, channels, groups)
for _ in range(2)
]
)
ResnetBlock3d ¶
Bases: Module
conv_shortcut instance-attribute ¶
conv_shortcut = (
CausalConv3d(
in_channels,
out_channels,
(1, 1, 1),
padding=(0, 0, 0),
)
if in_channels != out_channels
else None
)
SeedVR2VAE ¶
Bases: Module, DistributedVaeMixin
Whole-clip VAE; sampling uses the caller's request-local generator.
SpatialContext dataclass ¶
TemporalContext dataclass ¶
Past convolution inputs owned by one encode or decode call.