Skip to content

vllm_omni.model_executor.models.cosyvoice3.ras_sampler

Fused CosyVoice3 repetition-aware sampling for Model Runner V2.

One program per row replaces the ~70 small kernels (softmax, full-vocabulary sort, cumulative sum, masks, two Gumbel draws, history gather, rejection) of the tensor implementation. The nucleus is the top-p prefix of the full distribution capped at top-k, found by repeated argmax instead of a sort; a draw that repeats within the last win_size generated tokens is replaced by a draw from the full remaining distribution, from a disjoint noise stream.

MAX_FUSED_TOP_K module-attribute

MAX_FUSED_TOP_K = 64

fused_ras_sample

fused_ras_sample(
    logits: Tensor,
    idx_mapping: Tensor,
    temperature: Tensor,
    top_k: Tensor | None,
    top_p: Tensor | None,
    seeds: Tensor,
    pos: Tensor,
    all_token_ids: Tensor,
    total_len: Tensor,
    prompt_len: Tensor,
    *,
    default_top_k: int,
    default_top_p: float,
    win_size: int,
    tau_r: float,
    eps: float,
) -> Tensor

Sample one token per logits row; per-request tensors are indexed by idx_mapping.