SphereLab · Technical Report

What Makes Recurrence Effective
in Looped Language Models?

Understanding when to loop, where to loop,
and how to make every iteration count.

Xinlin Zhuang1,2Siyuan Wang1Imran Razzak2Weiyang Liu1,*

1The Chinese University of Hong Kong2MBZUAI

* Corresponding author

The question behind the loopsOriginal PDF ↗
Four panels compare knowledge and reasoning changes, ProofWriter and CLUTRR difficulty, and reasoning scores across physical-depth and loop-count configurations.
More loops, different outcomes. Extra recurrent computation can improve reasoning beyond the training horizon while knowledge performance declines. Benefits depend on the task, its difficulty, and the allocation of physical layers and loops.
2backbone families
3inference regimes
Up to 3×the training depth
Shared weightsmore compute, same backbone

Deeper computation.
But when is it useful?

TL;DR

Recurrence can unlock additional reasoning without adding backbone parameters, but effective depth alone does not predict success. History-state injection + timestep conditioning offers a lightweight way to improve robustness across inference budgets.

Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness.

Through controlled experiments, we systematically examine when recurrence helps, where it should be applied, and how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks.

We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benefit more. Performance also depends on how distinct layers and recurrent iterations are allocated. Non-recurrent output layers improve robustness to under-unrolling, while the preferred placement of input and output layers varies with inference budget.

Finally, conventional initial-state injection offers limited robustness to varying recurrence depth. We propose history-state injection as an alternative, and show that channel-wise history-state injection combined with timestep conditioning better preserves knowledge under extended unrolling while improving robustness across inference budgets.

Three questions for effective recurrence.

01 / WHEN

Reasoning can keep improving.

The training horizon is not a hard ceiling. Yet knowledge and reasoning respond differently, and harder reasoning problems do not consistently gain more.

Explore recurrent scaling ↗
02 / WHERE

Layer allocation matters.

Equal effective depth does not imply equal behavior. Output-side blocks help at smaller budgets; input-heavy allocations can support reasoning extrapolation.

Explore the architecture ↗
03 / HOW

Condition on the trajectory.

Recent state differences and timestep signals provide complementary information, making their combination effective under deep extrapolation.

Explore conditioning ↗

One shared core. A variable compute budget.

A looped decoder applies an input-side Prelude, repeatedly executes a parameter-shared recurrent core, and passes the result through an output-side Coda. Changing the number of loops changes inference compute.

Explore the compute budget
Tokens
Preludep layers
↻ 7 loops
Shared cores = 4 layers
Codac layers
Output
Training horizon
Physical depth 4Effective depth 28Inference / training 1.00×

Illustrative BaseLoop setting: p = c = 0, s = 4, training loops K = 7. Prelude and Coda are omitted in this setting. The slider illustrates compute, not predicted accuracy.

Controlled pretraining

Llama3.1-1B and Qwen3-0.6B backbone configurations, randomly initialized and trained on FineWeb-Edu with 2,048-token sequences. Models train with a fixed loop count and full backpropagation through all loops.

Knowledge & reasoning

Knowledge: SciQ, ARC-Easy, and PIQA. Reasoning: ARC-Challenge, WinoGrande, OpenBookQA, HellaSwag, CommonsenseQA, ProofWriter, CLUTRR, and BBH.

Physical depth vs. effective depth

Lphys = p + s + c · L(r) = p + r · s + c

The training horizon is not the finish line.

For Llama BaseLoop 4 × 5, increasing effective depth from 20 to 40 raises the reasoning score from 28.52 to 31.49, while knowledge falls from 62.80 to 52.11. Additional computation helps selectively.

The allocation of physical depth and loop count also changes the scaling curve. Under the same effective training depth of 20, BaseLoop 5 × 4 reaches a reasoning score of 31.88 at inference depth 30, exceeding the non-recurrent reference of 30.25 using additional test-time computation.

A practical takeaway. Evaluate knowledge and reasoning separately, and test across inference budgets. A single aggregate score at the training horizon can hide very different extrapolation behavior.

Give the boundaries their own role.

BaseLoop shares the entire Transformer stack across iterations. CoreLoop reserves independently parameterized Prelude and Coda blocks around the recurrent core. Their preferred allocation changes with the available compute.

BaseLoop vs. CoreLoop · LlamaOriginal PDF ↗
Overall, knowledge, and reasoning performance for BaseLoop and CoreLoop across four physical depths and different Prelude and Coda allocations.
Dashed lines mark the effective training depth of 20; shaded regions indicate extrapolation. Coda blocks improve knowledge robustness under under-unrolling, while Prelude-heavy allocations can better sustain reasoning beyond training.
r < K

Under-unrolling

Output-side layers help align partially executed recurrent states with the final readout.

r = K

Training horizon

Performance is generally strong, with smaller differences between boundary allocations.

r > K

Extrapolation

Input-heavy allocations can support reasoning, but the benefit depends on the configuration.

From a static anchor to an evolving history.

History-state injection

Reinjecting the initial state supplies the same reference at every loop. Instead, history-state injection conditions the current update on relative differences from recent completed recurrent states, exposing how the trajectory is evolving.

History-state injection

hℓ+1 = Fθ( hℓ + ∑j=1…mℓ Bj(hℓ−j − hℓ) )

mℓ = min{w, max(ℓ − 1, 0)}. History excludes h0; each lag-specific operator Bj is shared across iterations.
History-state injection · QwenOriginal PDF ↗
History-state injection results for scalar, channel-wise and dense parameterizations with different history window sizes, up to three times the training depth.
On Qwen BaseLoop 4 × 7, a short history window rescues deep extrapolation for dense injection. Low-capacity history conditioning alone gives limited gains, motivating complementary timestep signals.
Compare with conventional initial-state injection
Initial-state injection across Qwen configurations: dense injection shows pronounced knowledge degradation during extrapolation.
High-capacity dense initial-state injection can severely degrade knowledge beyond the training horizon. Original PDF ↗

Timestep conditioning

Continuous timestep signals let the same core adjust its computation across iterations. Loop Gating (LG) modulates the shared stack’s update; Branch Gating (BG) separately modulates attention and MLP residual branches. Both offer lightweight alternatives to channel-wise AdaLN modulation.

At inference, the time grid is rescaled to the requested loop count: more loops divide the same normalized interval [0, 1] into more, smaller steps. Gains depend on the architecture and inference budget.

Timestep conditioning · QwenOriginal PDF ↗
Comparison of Loop Gating, Branch Gating, AdaLN and BaseLoop across four Qwen recurrent configurations.
Time-dependent modulation interacts with physical depth and loop count. Improvements at the training horizon do not guarantee gains at larger inference budgets.

History and time work better together.

Channel-wise history-state injection with a window of two, combined with Loop Gating, gives the strongest pairwise result under deep extrapolation in the Qwen BaseLoop 4 × 7 study.

37.30%

Overall accuracy at depth 84

+1.16 pp

Over its stronger individual component

3×

Effective training depth

Combining conditioning mechanismsOriginal PDF ↗
Four panels compare initial plus history, initial plus timestep, history plus timestep, and all three conditioning mechanisms.
Complementary signals, stronger extrapolation. H = history-state injection; I = initial-state injection; LG = Loop Gating; BG = Branch Gating. All state conditioning here is channel-wise, with history window w = 2.
Pairwise conditioning at effective depth 84 · Qwen BaseLoop 4 × 7
ConditioningOverall accuracy (%) ↑
Initial-state + Loop Gating I + LG36.33
Initial-state + history-state I + H36.91
History-state + Loop Gating H + LG37.30

Training effective depth: 28. Adding initial-state conditioning to H + LG provides no further benefit in this comparison.

Useful loops build on earlier computation.

Small updates do not necessarily mean useful convergence. We probe computational interaction: how much a downstream block’s update changes when an earlier block occurrence is skipped.

At depth 84, H + LG has 18.6% and 10.4% higher mean cross-loop interaction over lags 1–6 than I + H and I + LG, respectively. This pattern is consistent with its stronger accuracy and suggests that history and timestep signals help later loops build on earlier computation.

Computational interaction across loopsOriginal PDF ↗
Skip-intervention analysis compares BaseLoop and CoreLoop responses, initial-state injection, pairwise conditioning, and three-way conditioning.
Interaction measures dependence between computations, rather than just the size of representation changes. The cross-loop pattern provides a possible explanation for the complementarity of history-state and timestep conditioning.

BibTeX

arXiv preprint · arXiv:2609.36636

BIBTEX
@misc{zhuang2026looplm,
  title  = {What Makes Recurrence Effective in Looped
            Language Models?},
  author = {Zhuang, Xinlin and Wang, Siyuan and
            Razzak, Imran and Liu, Weiyang},
  year   = {2026},
  eprint = {2609.36636},
  archivePrefix = {arXiv},
  url    = {https://arxiv.org/abs/2609.36636}
}

Figure preview

Scroll to inspect the full-resolution figure. Press Esc to close.