Breaking the Uniformity Trap Scaling Video Diffusion Model via SplitMoE

NeurIPS 2026 ✦Spotlight

✉ Corresponding author

Standard MoE · uniform load balancing
Generated by SplitMoE

One sparse model. Coherent motion.

Hover a clip to pause the reel and read its prompt · all clips come from the 27B-total / 14B-activated SplitMoE model

01 · Overview

TL;DR

Language-style MoE routes every token independently and pushes expert usage toward uniformity — a poor fit for video, whose tokens are spatiotemporally redundant and semantically long-tailed. SplitMoE splits the expert pool into semantic experts, steered by prototypes anchored in clean VAE features, and generic experts that keep flexible capacity for residual detail. Same 14B activated parameters as the dense model — better videos.

27B
total parametersupcycled from Wan2.2
14B
activated per tokensame as the dense model
20+80
semantic + generic expertsplus 1 shared expert per MoE layer
8/8
benchmark dimensions improvedover standard MoE at equal budget
~70%
of the training stepsto reach the dense model's loss

Abstract

Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and push-pull regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.

Standard MoE scatters patches of one object across experts; SplitMoE routes them to semantic expert groups and residual patches to generic experts
The idea in one picture. Standard MoE scatters patches of the same object across unrelated experts — spatiotemporal fragmentation. SplitMoE sends them to semantic expert groups under prototype guidance, while generic experts absorb residual appearance — semantic-aligned routing.
02 · Motivation

The Uniformity Trap

Load balancing, inherited from language MoEs, asks every expert to receive an equal share of tokens. That suits discrete, high-entropy text. Video tokens are different — densely correlated, spatially continuous and dominated by redundant background — so uniform routing splits one object across many experts while background patches crowd several others.

Self-similarity of text and visual tokens, semantic distinctiveness and routing dominance per semantic class
Evidence from routing logs. (a) Visual tokens are far more self-similar than text tokens. (b) Semantic distinctiveness — are different semantic classes sent to different experts? (c) Intra-sample routing dominance — are tokens of one region sent to the same expert? Statistics from 250 generated videos with SegFormer labels: SplitMoE's semantic experts beat the standard MoE (Wan-MoE) on routing dominance for every class and on distinctiveness for most.

What uniform routing looks like in practice

Top-1 expert maps across frames for the prompt “A horse is standing on grass in front of a barn.”

Baseline MoE routing maps: fragmented, striped and jittering assignments SplitMoE routing maps: coherent semantic regions and textured generic assignments
i

Temporal jitter

Similar regions switch experts from frame to frame.

ii

Spatial striping

Homogeneous areas are artificially split to satisfy load constraints.

iii

Semantic fragmentation

Adjacent tokens of one object scatter across unrelated experts.

Baseline MoE scatters correlated patches across experts under strict capacity constraints, causing fragmented routing and poor spatiotemporal consistency.

03 · Method

Split-role experts, guided by prototypes

SplitMoE upcycles a dense video DiT into a sparse one whose MoE layers have two branches with two jobs: a semantic branch that groups tokens by what they depict, and a generic branch that keeps flexible capacity for everything else.

SplitMoE pipeline: semantic MoE and generic MoE branches, semantic prototype guidance from clean VAE features, prototype learning with pull and push
Pipeline. Each video MoE layer is split into a semantic and a generic MoE branch. Soft targets induced by VAE-space prototypes guide semantic routing, while the generic branch preserves flexible residual modelling. Learnable prototypes are optimized in the VAE feature space with pull–push regularization and a global token bank.

One MoE layer, one token

Each token activates 2 semantic + 6 generic experts plus the always-on shared expert — exactly the Top-8 budget of the standard MoE baseline, so gains come from role-aware allocation, not extra compute.

Shared ×1
Semantic ×20 Top-2
Generic ×80 Top-6
8 × 1536 + 3072 = 15,360 active FFN width per token vs. 13,824 in the dense Wan2.2 FFN
01

Decoupled expert partitioning

The expert pool is split into two disjoint roles: semantic experts model high-level structure such as objects and layout; generic experts capture local appearance and generative residuals. A cosine router yields independent sigmoid affinities, so one token can match a semantic and a generic expert at once. Top-\(K\) is taken per group, with \(K = K_s + K_g\).

\[\mathbf{y}=\sum_{e\in\mathcal{I}^{\rm sem}} g^{\rm sem}_{e}\,E_e(\mathbf{x})+\sum_{e\in\mathcal{I}^{\rm gen}} g^{\rm gen}_{e}\,E_e(\mathbf{x})\]
02

Prototype-guided semantic routing

\(M_s\) learnable prototypes live in the clean-VAE feature space — a cost-free guidance signal with no external encoder. Comparing each token's VAE feature with the prototypes gives a fixed soft target \(\mathbf{q}\); the router, reading the DiT token, is aligned to it. A stop-gradient keeps the router from distorting the prototype topology.

\[\mathcal{L}_{\rm align}=\frac{1}{N}\sum_{i=1}^{N}\operatorname{KL}\!\left(\mathbf{q}_{i}\,\|\,\mathbf{r}_{i}\right)\]
03

Push–pull prototype learning

Prototypes behave like particles on a hypersphere. Pull attracts them to current features and to a circular bank of recent features, so no prototype drifts into a dead mode. Push repels prototypes whose similarity exceeds a margin \(m\), so they do not collapse onto redundant background regions.

\[\begin{aligned}\mathcal{L}_{\rm anchor}=\;&\underbrace{-\tfrac{1}{N}\sum_{i}\tau_c\log\sum_{m}e^{\cos(\mathbf{f}_i,\mathbf{p}_m)/\tau_c}+\tfrac{1}{M_s}\sum_{m}\big(1-\max_{\bar{\mathbf{f}}\in\mathcal{B}}\cos(\mathbf{p}_m,\bar{\mathbf{f}})\big)}_{\text{pull}}\\[2pt]&+\underbrace{\tfrac{1}{M_s(M_s-1)}\sum_{j\neq k}\max\!\big(0,\cos(\mathbf{p}_j,\mathbf{p}_k)-m\big)}_{\text{push}}\end{aligned}\]
04

Group-aware load balancing

Following DeepSeek-V3, only the generic experts use loss-free balancing: non-gradient biases steer the discrete Top-\(K\) choice while mixture weights come from clean scores. Semantic experts get no explicit balancing — they balance adaptively through semantic alignment. No auxiliary loss touches the diffusion gradients.

\[\mathcal{L}=\mathcal{L}_{\rm flow}+\lambda_{\rm align}\,\mathcal{L}_{\rm align}+\lambda_{\rm anchor}\,\mathcal{L}_{\rm anchor}\]
04 · Video results

See the difference

SplitMoE vs. standard MoE

Same Wan2.2 backbone, same data, same 80k steps, same 14B activated parameters — the only change is how tokens are routed.

Wan2.2-MoE · standard MoE
SplitMoE (Ours)

Comparison with state-of-the-art T2V models

VBench-2 prompts; every model is sampled with its default inference settings.

Frame-level comparison of Wan2.2, LongCat-Video, LTX-2, OmniWeaving, Dense Wan-FT, Wan-MoE and SplitMoE
Frame by frame. Watermelon cutting stresses object interaction: baselines show fused hand and fruit geometry (dashed boxes) and OmniWeaving blurs between frames. The stand-to-run horse stresses large motion: baselines grow extra legs. SplitMoE keeps boundaries stable and anatomy intact.

More text-to-video results

Long, cinematic prompts. Click any clip to enlarge it.

05 · Quantitative results

Same budget, better scores

VBench-2 measures intrinsic faithfulness over five capability dimensions; T2V-CompBench targets compositional generation.

Bold = best, underline = second best over all rows. A14B = MoE with 14B activated parameters. VBench-2 scores of LongCat-Video and HunyuanVideo follow the LongCat-Video paper; T2V-CompBench scores of CogVideoX-1.5 and Mochi follow the T2V-CompBench benchmark.

Full SplitMoE vs. same-source variants

Same checkpoint, data, steps and 14B activated budget.

Wan2.2-MoEFull SplitMoE

Efficiency

Parameters and per-iteration speed at the same 14B activated budget.

Both MoE models add a small per-iteration overhead over the dense model (routing and memory bandwidth); SplitMoE's cost over a standard MoE of the same size is negligible.

06 · Analysis

What the experts learn

Emergent behaviour

Coarse-to-fine, without being told

Tracking the routing weight of semantic experts over a 50-step denoising trajectory reveals a stage-dependent role. The weight peaks around steps 10–15, when the model establishes global layout, object placement and coarse semantics, then decays as denoising shifts to local appearance and detail. No timestep-conditioned routing constraint is used.

  • Early, high noise — semantic experts shape structure
  • Late, low noise — generic experts refine detail
Diffusion loss of SplitMoE versus dense baseline
Faster convergence. Against dense fine-tuning at the same activated size, SplitMoE reaches comparable validation loss with roughly 70% of the steps and stays lower throughout.
Diffusion loss of full model and ablations
Every term matters. Removing push, pull, or the whole semantic branch (standard MoE) raises the validation diffusion loss relative to the full model.
Prototype distribution in 2D spherical latent space
Prototype layout (Fig. 6c). Pull only: prototypes gather in dominant regions and leave the manifold under-covered. Push only: they are driven off the data manifold. Pull + push: balanced coverage.
Qualitative ablation: without prototype guidance, without push, without pull, and full model
Qualitative ablation. Without prototype guidance the horse never starts running; without push, the watermelon scene drifts between frames; without pull, the subject changes abruptly. Only the full model renders the action transition while keeping object semantics consistent.
07 · Citation

BibTeX

@inproceedings{xu2026splitmoe,
  title     = {Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE},
  author    = {Xu, Yu and Zhang, Yuxin and Yang, Xiao and Yang, Haotian and Wang, Yizhi and
               Huang, Xinwei and Lin, Minxuan and Wang, Angtian and Ma, Chongyang and Tang, Fan},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}