Research

San Francisco, California

Expert Coupling in MoE Pretraining

Radha Gulhane, Quentin Anthony, Beren Millidge

Mixture-of-Experts (MoE) models scale model capacity efficiently by routing each token to only a small subset of experts. But at large scale, the cost of moving tokens between GPUs can become a major bottleneck, accounting  for as much as 60.4% of total training step time in our experiments. In this work, we show that MoE routing contains substantial structure that can be exploited to reduce this communication. Experts selected by the same token are strongly correlated within a layer, while a token’s expert choices in one layer predict where it will be routed in later layers. We use these two properties to increase token–expert locality in two complementary ways: correlated expert placement co-locates experts that are frequently selected together, while token shuffling proactively moves tokens to the GPUs likely to contain their next experts. Neither technique changes the model architecture or its routing decisions. Moreover, we find that these routing patterns are cheap to estimate and emerge early enough in training to remain useful over long stretches of pretraining. In Megatron-LM experiments on MI300X GPUs, our methods reduce all-to-all communication time by up to 2.63× and improve end-to-end training step time by up to 1.41×.

Introduction

Mixture-of-Experts models have become an increasingly important way to scale language models. Instead of running every token through a single large feed-forward network, an MoE layer contains many independent experts and a router that sends each token to only a small subset of them. This allows model capacity to grow with the number of experts while keeping the amount of computation performed for each token relatively small.

However, at large scale, MoEs introduce a different bottleneck: all-to-all communication.

This arises because although increasing the number of experts does not increase the compute required in the forward pass, it increases the memory requirement needed to keep the experts in memory. As models become larger, this memory requirement expands beyond what can be supported in the HBM of a single GPU, necessitating parallelism schemes such as expert parallelism (EP). 

In expert parallelism, experts are distributed across GPUs. Whenever a token selects an expert on another GPU, its hidden state has to be sent across the network to that expert and then sent back. This happens in both the forward and backward pass, at every MoE layer.

As expert parallelism grows across nodes—and especially as models route each token to more experts—this communication can become a large fraction of training time. In our experiments, all-to-all communication accounts for as much as 60.4% of total step time. The standard approach to expert parallelism treats expert placement as essentially arbitrary: experts are distributed evenly across GPUs, without regard for which experts tend to process the same tokens.

In this work, we show that this leaves substantial performance on the table.

Crucially, expert routing is far from random. We find strong correlations both between experts within the same layer and between routing decisions across consecutive layers. By exploiting these correlations, we can substantially increase the probability that a token and the experts it needs are already on the same GPU.

The result is less data crossing the network, significantly faster all-to-all communication, and meaningful end-to-end training speedups—without changing the model architecture, the router's decisions, or the training objective. This means our technique can be applied to all MoE models without impacting output quality or requiring retraining. 

The Hidden Cost of Mixture-of-Experts

Under EP parallelism, experts are sharded across a group of GPUs. A token starts on one GPU, but the experts selected by its router may live on several others. The system therefore performs an all-to-all dispatch, sending each token representation to the GPUs containing its selected experts. After expert computation, another all-to-all returns the resulting activations. And both operations happen again, but transposed, in the backward pass.

As we scale EP, two things make the problem worse. First, a larger fraction of experts are remote. Second, once the EP group spans multiple machines, traffic crosses the significantly slower inter-node network. Increasing top-k compounds the problem because every token is routed to more destinations.

In our experiments on a cluster of 8 MI300x nodes, moving from EP8 to EP32 increases the all-to-all share of training time from 13% to 45% for top-2 routing, and from 24% to 60% for top-6 routing. At EP32 with top-6 routing, communication takes more time than expert computation and all of the model's non-MoE computation combined.


Figure 1: Training-step breakdown. Share of training step time by component at EP8 and EP32 with top-2 and top-6 routing, using the Megatron-LM all-to-all dispatcher and contiguous placement. Dashed boxes mark the all-to-all share.


So if we want larger MoE models to train efficiently, reducing computation is only part of the problem. We also need to reduce how much information moves between GPUs.

Expert Routing Is Not Random

Conventional expert placement largely ignores how the router actually behaves. Experts are typically assigned to GPUs by ID: expert 0 goes on the first GPU, followed by expert 1, expert 2, and so on. This keeps parameter counts balanced, but it implicitly assumes that any two experts are equally likely to be used together.

After training begins, we discover that this assumption is very wrong. Experts specialize, and routing develops considerable structure. We observe two kinds of correlation:

Experts Are Correlated Within a Layer

For top-k routing, some experts are selected together much more often than chance would predict. 

In one of our top-2 models, 45% of tokens select one of just 0.8% of all possible expert pairs. Within an individual layer, we find particular pairs that are selected together for a very large fraction of their traffic.

For example, in layer 8 of our top-2 model, just 64 of the 8,128 possible expert pairs account for 42% of tokens. If routing were independent while maintaining the same expert loads, those pairs would account for only 1.6%.


Figure 2: Expert correlation heatmap. Expert correlation in the top-2 model, measured on 524k tokens from iterations 1983–1995. Axes are expert IDs, and the diagonal in (a) is masked. Orange marks conditional probabilities above 0.3. Experts are correlated substantially more than would be expected by chance.


This immediately suggests an optimization.

If two experts are frequently selected by the same token, put them on the same GPU.

Then the token only needs to travel to one GPU rather than two.

Routing Is Also Correlated Across Layers

The routing decision at one layer also contains information about where a token will be routed next.

If a token visits a particular expert at layer l, its expert choices at layer l+1 are far from uniformly distributed. For the same model, 112 out of 128 experts at layer 8 send at least 30% of their tokens to one particular expert at layer 9, with a median of 48%.

This gives us a second opportunity.

Rather than only moving experts closer together, we can sometimes predict where a token will need to go next—and move the token there ahead of time.

These two forms of correlation lead to the two main techniques in our work.

Moving Correlated Experts Together

Our first technique is correlated expert placement.

For every MoE layer, we measure how frequently each pair of experts is selected by the same tokens. This produces a graph where experts are nodes and heavily correlated expert pairs have strong connections. We then partition this graph across GPUs while keeping the number of experts on every GPU fixed, attempting to place frequently co-selected experts together. This requires no additional expert replicas and does not change the router or the experts themselves. We are simply permuting where experts live.

Correlated placement becomes particularly useful when combined with a deduplicating dispatcher.

Normally, if a token selects three experts that happen to live on the same remote GPU, a dispatcher can send three copies of the token's hidden state—one for each expert. A deduplicating dispatcher instead sends the token to that GPU once, along with metadata indicating which local experts should process it.

Correlated placement deliberately creates more opportunities for this deduplication.


Figure 3: Correlated placement and deduplication. Baseline placement sends the same token to multiple destinations—or multiple times to the same destination. Correlated placement puts commonly co-selected experts together, allowing a single transfer to service multiple experts.


At EP8, deduplication alone removes only 6% of transmitted rows for top-2 routing because experts happen to share GPUs only by chance. With correlated placement, this rises to 26%. For top-6 routing, the reduction rises from 25% to 58%.

The intuition is simple: rather than accepting the communication pattern induced by routing, arrange the hardware layout around it.

Moving Tokens Toward Their Next Experts

However, expert placement solves only half of the locality problem. Large models and long-context training typically combine expert parallelism with tensor parallelism (TP), which shards each layer's computation across GPUs, since expert parallelism alone does not fit the model in memory. Under TP-EP, even if all of a token's likely experts live on the same GPU, the token itself may reside on a different one. 

Our second technique, token shuffling, exploits cross-layer routing correlations to address this.

The experts selected by a token in previous MoE layers are surprisingly predictive of the experts it will select next. We use these previous routing decisions to predict which GPU is most likely to contain the token's next experts. We then move that token to the predicted GPU before the next MoE router runs.

Naively, you would think that doing this would simply replace one communication operation with another and provide little benefit. But in common tensor-expert parallel training configurations, the model already performs a reduce-scatter after attention to distribute tokens across GPUs. We simply modify the indexing of this existing collective so that, instead of sending each token to its conventional sequence shard, it sends the token to the GPU predicted to hold most of its upcoming experts.

Because the collective was already moving the token, this reshuffling does not add network traffic. Before the next attention layer, a corresponding shuffle-aware all-gather restores canonical token order.


Figure 4: Token shuffling. We can use routing decisions from previous layers to predict which GPU will contain a token's next experts, then fold the move into the reduce-scatter that already occurs after attention, making the next MoE dispatch substantially more local.


This changes where computation happens without changing what computation happens.The router still selects exactly the same experts. The expert weights are unchanged. The training loss is unchanged.

We are simply placing tokens where their future computation is likely to be.

Turning Routing Correlation Into Locality

Together, expert placement and token shuffling attack communication from opposite directions.

Correlated placement moves experts toward one another while token shuffling moves tokens toward those experts.

The quantity we ultimately care about is token–expert locality: how often a token is already on the GPU containing one of its selected experts. Without token shuffling, the chance that an arbitrary selected expert is local is roughly (1/\mathrm{EP}). At EP8 this is only 12.5%, and at EP16 it is 6.3%. With token shuffling, locality increases dramatically. We measure 59% and 53% locality for top-2 routing at EP8 and EP16, and 57% and 46% for top-6.

Those local token–expert assignments never need to cross the network. As a result, relative to the baseline, the number of bytes sent over the network falls by as much as 3.1× in the configurations we evaluate.

The Correlations Are Cheap to Learn

A system optimization based on routing statistics would be much less compelling if it required continuously profiling enormous numbers of tokens and the correlations drifted constantly.

Fortunately, these correlations are remarkably easy to estimate and appear stable over time.

Using only 1,024 tokens, we can already recover expert placements that capture most of the benefit. With 4,096 tokens, the estimated co-selection matrices have Pearson correlations of 0.93 for top-2 and 0.98 for top-6 with statistics computed from a much larger sample. The resulting placement and routing predictions are within roughly one percentage point of those obtained using 524,000 tokens. In practice, a single microbatch from one rank is enough to estimate the correlation tables.

That means gathering the information required to optimize expert placement is extremely inexpensive relative to training.

They Also Emerge Early and Stay Stable

Routing correlations are not only easy to estimate; they stabilize relatively early in training. Our models train for 2.1 billion tokens. Statistics measured early in training already contain meaningful information about the routing structure present at the end.

Correlation tables estimated at step 199 provide roughly half of the final placement benefit. By step 1000, they are within a few percentage points of the final result, and from around step 1400 onward the correlation matrices closely match those measured at the end of training. Moreover, we find that cross-layer routing prediction improves on a similar trajectory.

This is an important systems property. Our approach does not require continually changing where every expert lives or profiling massive datasets. Routing structure can be estimated from a small number of tokens, reused for a substantial portion of training, and refreshed occasionally at low cost.

Reducing Communication Produces Real Training Speedups

Reducing bytes is useful only if it translates into wall-clock performance. We therefore implement our techniques directly in Megatron-LM and benchmark both communication time and full training-step time on MI300X GPU clusters.

With correlated expert placement and deduplication alone, we reduce all-to-all time by 1.16×–1.38× under top-2 routing and 1.35×–1.95× under top-6, depending on expert-parallel degree.

When token shuffling can also be applied, the reduction becomes as large as 1.74× for top-2 and 2.63× for top-6.


Figure 5: Correlated Placement Speedups. All-to-All Cost Under Expert Parallelism. A2A time (top) and communication volume per rank (bottom) at TP1 with EP8–EP64, for top-2 and top-6 routing. B is the baseline, D adds deduplication, and P adds correlated placement to deduplication.


These communication savings translate into substantial end-to-end improvements when all-to-all is a significant fraction of the training step. Across our pure expert-parallel configurations, correlated placement produces up to a 1.41× end-to-end training-step speedup, with the largest gains appearing for top-6 routing when communication crosses node boundaries. For the tensor-parallel plus expert-parallel layouts where token shuffling applies, we see up to a 1.34× end-to-end speedup, while all-to-all itself becomes as much as 2.63× faster.


Figure 6: Token Shuffling Speedups. All-to-All Cost on the TP = EP Layouts. Token shuffling is the fourth configuration (S). Token ownership can move inside the TP group. Token shuffling leaves the row count unchanged and keeps more rows on the device, so fewer rows cross the network. We observe that token-shuffling can substantially reduce end-to-end A2A cost.


These are not changes to model quality or sparsity. The same tokens visit the same experts and perform the same expert computation. The speedup comes from recognizing that where computation happens matters almost as much as which computation happens.

Why This Matters

MoE architectures are usually motivated by computational sparsity. A router chooses a small subset of experts, so a model can contain far more parameters than it evaluates for any individual token.

But on modern distributed systems, FLOPs are only part of the cost.

When experts are spread across many GPUs, routing decisions also define a communication graph. Conventional expert parallelism effectively ignores the structure of that graph: experts are distributed uniformly and tokens follow whatever communication pattern results.

Our results show that this routing graph contains substantial exploitable structure.

Experts specialize into correlated groups. Tokens follow predictable trajectories through those groups across layers. Those patterns emerge relatively early in training, remain stable, and can be estimated with only a few thousand tokens. Once we use that information to improve locality, a substantial fraction of MoE communication overhead simply disappears.

Importantly, these techniques are complementary to existing MoE design choices. They require no expert replication, do not constrain the router's choice of experts, and do not modify the model architecture or training objective.

More broadly, this suggests a different way of thinking about distributed MoE systems. Routing is usually treated as a model-level decision that the distributed system must react to. But routing is not arbitrary—it has structure. And once that structure is predictable, the system can organize itself around the model's computation rather than blindly communicating around a static hardware layout.

As MoE models continue to scale to more experts, higher top-k, and larger expert-parallel groups spanning many machines, this distinction becomes increasingly important. Sparse models should not only be sparse in computation; their communication should be sparse too.

© 2026 Zyphra Technologies Inc. All rights reserved.

© 2026 Zyphra Technologies Inc. All rights reserved.