vllm_omni.diffusion.models.auk.pipeline_auk ¶
AuK audio-editing pipeline for vLLM-Omni.
Turns a layer-fused text condition (produced by the encoder stage) plus an optional source clip into a 24 kHz waveform: the source clip is encoded to VAE latents, a rectified-flow DiT integrates the target latents conditioned on both, and the BigVGAN-flow decoder renders the waveform.
AuKPipeline ¶
Bases: Module, SupportAudioInput, SupportAudioOutput, SupportsComponentDiscovery
Instruction-driven audio generation and editing with AuK.
One request per forward: the rectified-flow ODE runs over the whole target span and the batch dimension carries the CFG branches, so there is nothing to share between requests yet.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
od_config | OmniDiffusionConfig | OmniDiffusion configuration. | required |
prefix | str | Unused; kept for the pipeline construction contract. | '' |
cudagraph_wrapper instance-attribute ¶
cudagraph_wrapper = AuKCUDAGraphWrapper(
self.dit,
enabled=not od_config.enforce_eager,
max_graphs=max_dit_graphs,
)
flash_t_grid instance-attribute ¶
vae_decode instance-attribute ¶
vae_decode = AuKVAEDecodeGraph(
self.vae,
enabled=not od_config.enforce_eager,
**vae_decode_kwargs,
)
forward ¶
forward(
req: DiffusionRequestBatch,
) -> list[DiffusionOutput]
Generate one waveform.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
req | DiffusionRequestBatch | Request batch holding exactly one request. The prompt carries | required |
Returns:
| Type | Description |
|---|---|
list[DiffusionOutput] | One |
list[DiffusionOutput] |
|
list[DiffusionOutput] |
|
setup_compile ¶
Compile the DiT as configured, then capture the codec decode buckets.
Defining this hook replaces the runner's generic transformer compile, so the DiT's diffusion_compile_granularity is honoured here: full compiles the whole denoise step; regional compiles the double- and single-stream blocks. Compilation is lazy and happens in the DiT CUDA graph's warm-up, before capture, so each graph replays the compiled kernels. The startup cost worth paying is the VAE decode, whose buckets are compiled and captured before the first request.
get_auk_post_process_func ¶
get_auk_post_process_func(od_config: OmniDiffusionConfig)
Create the post-processing function for AuK audio output.
Tensor output types pass through; anything else becomes a numpy waveform. The sample rate is not attached here: the output formatter reads AuKPipeline.audio_sample_rate for audio-output pipelines.