vllm.model_executor.layers.fused_moe.moe_output ¶
Output contract between a MoE layer and a consumer that fuses its tail.
Classes:
-
MoEOutput–A MoE layer's output with its final reduction still open.
-
UnfinalizedMoEOutput–Unfinalized output of a monolithic MoE kernel.
MoEOutput dataclass ¶
A MoE layer's output with its final reduction still open.
Returned by layers whose MoE runs un-reduced (reduce_results=False) so that the consumer -- typically the next layer's RMSNorm -- can fuse the tensor-parallel all-reduce into itself instead of paying for a standalone one. Keeping the shared-expert output and the routed scale separate leaves that reduction the consumer's to schedule; when the routed output is still unfinalized, the top-k reduction is open too and can fold into the same kernel.
Producers only leave the routed output unfinalized when a fused consumer can actually take that form -- the token ceiling and topology support are theirs to check -- so an UnfinalizedMoEOutput here means the fused path applies, and a consumer need not re-derive that.
Source code in vllm/model_executor/layers/fused_moe/moe_output.py
UnfinalizedMoEOutput dataclass ¶
Unfinalized output of a monolithic MoE kernel.
Kernels that can stop after GEMM2 (the TRTLLM-Gen do_finalize=False path) hand back their permuted, unweighted output plus the routing weights and the permute map, so that the top-k reduction can be fused with whatever follows -- the shared-expert add and the tensor-parallel all-reduce -- instead of running as its own kernel.
The buffers are consumed as-is by the fused kernels, which index gemm2_permuted by row: it must be densely packed at hidden_dim, and its row count (an autotuner-dependent padded value) is never referenced.