Skip to content

vllm_omni.diffusion.layers.swiglu7

Interleaved clamped SwiGLU7 with a native compile path.

The Triton expression is adapted from SGLang revision 25536524af6701518d0b8ec0efca8109c23b3706 (Apache-2.0).

SwiGLU7

Bases: CustomOp

SwiGLU7 for packed [..., gate_0, up_0, gate_1, up_1, ...] inputs.

Arithmetic is evaluated in FP32 and cast once to the requested output dtype. Eager inference with contiguous inputs and the released constants uses Triton; compilation, autograd and other layouts retain the native expression so this pointwise operation does not become a fusion barrier.

forward_npu class-attribute instance-attribute

forward_npu = forward_native

forward_cuda

forward_cuda(
    x: Tensor,
    alpha: float = 1.702,
    limit: float = 7.0,
    out_dtype: dtype | None = None,
) -> Tensor

forward_native staticmethod

forward_native(
    x: Tensor,
    alpha: float = 1.702,
    limit: float = 7.0,
    out_dtype: dtype | None = None,
) -> Tensor