vllm_omni.model_executor.models.cosyvoice3.runtime ¶
Runtime toggles shared by the CosyVoice3 stages.
cosyvoice3_batch_flow_debug ¶
cosyvoice3_batch_flow_debug() -> bool
Return whether Stage-1 batching diagnostics are enabled.
cosyvoice3_batch_flow_enabled ¶
cosyvoice3_batch_flow_enabled() -> bool
Return whether cross-request Stage-1 flow batching is enabled.
cosyvoice3_batch_flow_profile ¶
cosyvoice3_batch_flow_profile(name: str)
Create a profiler scope only when batching diagnostics are enabled.
cosyvoice3_full_response_enabled ¶
cosyvoice3_full_response_enabled() -> bool
Opt-in Hopper full-response path; other devices retain the standard path.
cosyvoice3_packed_inference_enabled ¶
cosyvoice3_packed_inference_enabled() -> bool
Either optimized profile needs live conditioning outside CUDA graphs.
cosyvoice3_packed_streaming_enabled ¶
cosyvoice3_packed_streaming_enabled() -> bool
Opt-in Hopper packed, chunk-causal Flow with request-owned GPU HiFT state.
cosyvoice3_standard_sampling ¶
Select ordinary sampling; RAS remains the default checkpoint behavior.
The optional HF override is shared by API stop configuration and the model sampler. Reject misspellings instead of silently selecting another policy.