vllm_omni.model_executor.models.cosyvoice3.ras_sampler ¶
Fused CosyVoice3 repetition-aware sampling for Model Runner V2.
One program per row replaces the ~70 small kernels (softmax, full-vocabulary sort, cumulative sum, masks, two Gumbel draws, history gather, rejection) of the tensor implementation. The nucleus is the top-p prefix of the full distribution capped at top-k, found by repeated argmax instead of a sort; a draw that repeats within the last win_size generated tokens is replaced by a draw from the full remaining distribution, from a disjoint noise stream.
fused_ras_sample ¶
fused_ras_sample(
logits: Tensor,
idx_mapping: Tensor,
temperature: Tensor,
top_k: Tensor | None,
top_p: Tensor | None,
seeds: Tensor,
pos: Tensor,
all_token_ids: Tensor,
total_len: Tensor,
prompt_len: Tensor,
*,
default_top_k: int,
default_top_p: float,
win_size: int,
tau_r: float,
eps: float,
) -> Tensor
Sample one token per logits row; per-request tensors are indexed by idx_mapping.