Skip to content

vllm_omni.model_executor.models.cosyvoice3.runtime

Runtime toggles shared by the CosyVoice3 stages.

cosyvoice3_batch_flow_debug

cosyvoice3_batch_flow_debug() -> bool

Return whether Stage-1 batching diagnostics are enabled.

cosyvoice3_batch_flow_enabled

cosyvoice3_batch_flow_enabled() -> bool

Return whether cross-request Stage-1 flow batching is enabled.

cosyvoice3_batch_flow_profile

cosyvoice3_batch_flow_profile(name: str)

Create a profiler scope only when batching diagnostics are enabled.

cosyvoice3_full_response_enabled

cosyvoice3_full_response_enabled() -> bool

Opt-in Hopper full-response path; other devices retain the standard path.

cosyvoice3_packed_inference_enabled

cosyvoice3_packed_inference_enabled() -> bool

Either optimized profile needs live conditioning outside CUDA graphs.

cosyvoice3_packed_streaming_enabled

cosyvoice3_packed_streaming_enabled() -> bool

Opt-in Hopper packed, chunk-causal Flow with request-owned GPU HiFT state.

cosyvoice3_standard_sampling

cosyvoice3_standard_sampling(config: object) -> bool

Select ordinary sampling; RAS remains the default checkpoint behavior.

The optional HF override is shared by API stop configuration and the model sampler. Reject misspellings instead of silently selecting another policy.