vllm_omni.utils.seeded_exponential ¶
Batched per-request exponential_(generator=...) that is bit-identical to torch.
Seeded sampling draws exponential noise one request at a time, which costs a kernel launch and host overhead per request and step. This kernel reproduces torch's CUDA distribution kernel exactly for float32 (the grid-stride Philox4x32-10 curand_uniform4 stream, the fast-math __logf transform and the per-call Philox offset increment), so a batch draws in one launch while every value and every generator state afterwards matches the per-request loop.
batched_seeded_exponential_supported ¶
Whether the NVIDIA CUDA kernel reproduces the per-row loop for q.
ROCm tensors also report is_cuda, but HIP uses a different distribution implementation and cannot compile the CUDA libdevice transform above. Keep those draws on Torch's original per-row path, including RNG state.
fill_exponential_rows ¶
Draw out[rows[i]].exponential_(generator=generators[i]) for all i in one launch.
out is a contiguous float32 matrix. Without rows, each generator owns an equal contiguous slice of the flattened output (which may span several codebook rows). None generators use the default generator in order. NVIDIA uses the batched CUDA kernel; other platforms retain ordered Torch draws, including direct Higgs Audio callers.