Skip to content

vllm_omni.utils.seeded_exponential

Batched per-request exponential_(generator=...) that is bit-identical to torch.

Seeded sampling draws exponential noise one request at a time, which costs a kernel launch and host overhead per request and step. This kernel reproduces torch's CUDA distribution kernel exactly for float32 (the grid-stride Philox4x32-10 curand_uniform4 stream, the fast-math __logf transform and the per-call Philox offset increment), so a batch draws in one launch while every value and every generator state afterwards matches the per-request loop.

batched_seeded_exponential_supported

batched_seeded_exponential_supported(
    q: Tensor, generators: dict
) -> bool

Whether the NVIDIA CUDA kernel reproduces the per-row loop for q.

ROCm tensors also report is_cuda, but HIP uses a different distribution implementation and cannot compile the CUDA libdevice transform above. Keep those draws on Torch's original per-row path, including RNG state.

fill_exponential_rows

fill_exponential_rows(
    out: Tensor,
    generators: list,
    rows: list[int] | None = None,
) -> Tensor

Draw out[rows[i]].exponential_(generator=generators[i]) for all i in one launch.

out is a contiguous float32 matrix. Without rows, each generator owns an equal contiguous slice of the flattened output (which may span several codebook rows). None generators use the default generator in order. NVIDIA uses the batched CUDA kernel; other platforms retain ordered Torch draws, including direct Higgs Audio callers.

torch_exponential_policy

torch_exponential_policy(
    numel: int, device: device
) -> tuple[int, int]

(grid threads, Philox offset increment) of torch's CUDA exponential_.