vllm_omni.model_executor.models.common ¶
Modules:
| Name | Description |
|---|---|
alias_free_activation | Alias-free activation for BigVGAN-style speech decoders. |
cfg_pairing | Model-neutral classifier-free-guidance request pairing for the vLLM v1 scheduler. |
fused_sampling | Fused Gumbel selection after native nucleus probability reductions. |
ming | Components shared by the Ming-family model packages. |
nucleus_ras_sampling | Shared TTS sampling primitives: nucleus (top-p/top-k) and RAS. |
ops | Common operators shared across models and hardware backends. |
qwen3_code_predictor | Qwen3 Code Predictor -- optimized re-prefill, no KV cache. |
short_kv_attention | Short causal attention for the scratch KV cache inside a code predictor. |
snake_activation | Shared Snake/SnakeBeta activations for speech decoders. |
talker_first_audio | Shared eligibility; each model separately controls its first-audio opt-in. |
whisper_vq | WhisperVQEncoder: HF WhisperEncoder + VQ codebook + inter-layer pooling. |