vllm_omni.model_executor.models.moss_tts ¶
Modules:
| Name | Description |
|---|---|
audio_tokenizer | MOSS Audio Tokenizer — inference-only codec (encode waveform ↔ RVQ codes). |
audio_tokenizer_v2 | PyTorch MossAudioTokenizer model. |
configuration_moss_audio_tokenizer_v2 | MossAudioTokenizer model configuration |
configuration_moss_tts | MOSS-TTS model configuration. |
cuda_graph_streaming_decoder_wrapper | CUDA graphs for compact MOSS-TTS streaming codec batches. |
first_chunk_fast_path | Decode a stream's first codec chunk as soon as Stage 1 receives it. |
local_model_state | MOSS Local request-slot state, independent of each step's batch order. |
local_short_attention | Single-query attention for the frame-local MOSS depth transformer. |
modeling_moss_tts_codec | MOSS-TTS Stage-1 codec decoder: RVQ codes → 24 kHz waveform. |
modeling_moss_tts_local | Local depth transformer for MossTTSRealtime. |
modeling_moss_tts_local_depth | Per-frame depth transformer for MossTTSLocalModel (MOSS-TTS-Local-Transformer-v1.5). |
modeling_moss_tts_talker | MOSS-TTS Stage-0 talker: Qwen3 backbone + (n_vq+1) parallel AR heads. |
moss_codec_cudagraph | CUDA Graph acceleration for the MOSS Audio Tokenizer codec decoder. |
pipeline | Pipeline topology for all MOSS-TTS variants (2-stage: talker → codec). |
realtime_prompt | Prompt construction for MOSS-TTS-Realtime. |
reference_encoder | Reference-audio encoding + speaker cache for the MOSS-TTS-family talker. |
reference_encoder_graphs | CUDA graphs for batched MOSS reference-audio encoding. |
shared_reference_encoder | One reference-audio encoder shared by the API processes of a server. |
slot_attention | Direct-slot ring attention for MOSS streaming codec. |
streaming_attention | Masked streaming attention preserving ring-cache position semantics. |