Video Generation · Diffusion Models · Efficient AI

Causal Forcing++ vs AnyFlow: Fixed Steps or Free?

Both distill Wan2.1 video diffusion for cheap sampling. Causal Forcing++ pins 1-2 steps per frame for 14.1 FPS real-time streaming; AnyFlow keeps steps free so quality scales from 4 to 32 NFE.

Causal Forcing++ vs AnyFlow: Fixed Steps or Free?

The one question they disagree on

Few-step video diffusion has one design fork: do you freeze the sampling budget at train time, or leave it free at inference time? Causal Forcing++ and AnyFlow, published two days apart in May 2026 and both built on the open Wan2.1 video backbone, take opposite sides of that fork.

Causal Forcing++ (Tsinghua) says: the user of an interactive video system cannot afford to choose a step count. A game-like world needs frames at a fixed rate no matter what, so the model should be distilled to a hard budget of 1-2 steps per frame, and everything else, training included, is engineered around making that fixed point as good and as cheap as possible.

AnyFlow (NVIDIA) says: practitioners expect diffusion to scale with compute, and consistency-distilled models broke that property by degrading at 16 or 32 steps. Fix the scaling law, not the budget, and one checkpoint serves every use case from 4-step previews to 32-step finals.

Same backbone family, same month, opposite answers. The comparison below uses only numbers the two papers actually report.

Key numbers

MeasurementCausal Forcing++AnyFlowSame harness?
Base modelWan2.1-1.3B (student), Wan2.1-14B as Stage 3 scorerWan2.1, 1.3B to 14BYes, Wan2.1 family
Sampling budgetFixed: 1-2 steps per frameFree: 4 / 16 / 32 NFE shownNo, different design goals
Throughput14.1 FPS at 2 steps frame-wise; 8.69 FPS at 4 steps frame-wise; 10.4 FPS at 4 steps chunk-wiseNot reported in available materialn/a
First-frame latency0.27s frame-wise vs 0.60s chunk-wise baseline (-50%)Not reportedn/a
VBench total84.14 at 2 steps vs 84.04 for 4-step Causal Forcing baselineNo single VBench figure publishedCF++ rows: same Wan2.1-1.3B harness
VBench quality / semantic84.89 / 81.13 at 2 steps vs 84.59 / 81.84 baselineNot reportedSame harness
VisionReward6.661 at 2 steps, 6.798 at 4 steps, vs 6.326 chunk-wise baselineNot reportedSame harness
Few-step training costStage 2 drops 11,600 to 2,900 A800 GPU hours (~4x), zero trajectory storageFlow Map Backward Simulation adds rollout decomposition; cost not quantified in abstractNo
Step-count scalingQuality flat between 2 and 4 steps (VBench moves ~0.1)Quality rises from 4 to 32 NFE; consistency baselines degrade insteadNo direct comparison exists
Causal / autoregressive setupFrame-wise AR is the main result; action-conditioned world-model variant stays 4-step chunk-wiseBoth bidirectional and causal AR covered; exposure bias targeted via on-policy supervisionNo

Read the table for what it does not contain: there is no experiment that runs both methods on the same benchmark. Causal Forcing++ reports hard latency, throughput, VBench and VisionReward numbers on a small 1.3B backbone; AnyFlow reports the shape of results across 1.3B-14B and 4-32 NFE but, in the material available, no headline benchmark score at all. Any table that gives you a single winner with a margin is inventing data.

What Causal Forcing++ is actually for

Causal Forcing++ is an efficiency paper, and it is unusually honest about it. Its problem statement is interactivity: for a world model you steer with actions, frames must stream out faster than human reaction time, and a bidirectional model that decides the whole clip before emitting anything cannot do that. The fix is architectural: go frame-wise autoregressive, so the first frame ships at 0.27 seconds while later frames are still being planned, then get each frame down to 2 denoising steps so the steady state runs at 14.1 FPS.

Its real contribution is not the AR part, which the earlier Causal Forcing already had, but making the few-step stage cheap. The predecessor initialized few-step training with causal ODE distillation, which precomputes and stores full denoising trajectories for every real training video. Causal Forcing++ replaces that with causal consistency distillation: a single online teacher ODE step between adjacent timesteps, which the paper argues learns the same autoregressive-conditional flow map without any precompute or storage. That is where the roughly 4x Stage 2 saving (11,600 to 2,900 A800 GPU hours) comes from.

The quality story is deliberately modest. VBench total moves from 84.04 to 84.14 when you cut from 4 steps to 2, which is inside noise; the win is holding quality flat while latency halves. The paper’s own caveats matter for buyers: the 1-step setting has good motion but weak semantic understanding (so 2 steps is the practical point), exposure bias still degrades late frames, and the action-conditioned world-model variant, the one true interactive simulators need, remains 4-step chunk-wise. “Fully real-time” is currently a claim about prompt-driven generation.

What AnyFlow is actually for

AnyFlow starts from a failure of consistency distillation that surprises people the first time they hit it: give a 4-step distilled video model 16 or 32 steps and it gets worse. The paper’s diagnosis is that consistency training replaces the probability-flow ODE trajectory with a separate consistency-sampling trajectory. The model learned the endpoint, not the path, so intermediate steps wander off-trajectory and accumulate error.

AnyFlow’s fix is to distill the flow map itself, the transition from z_t to z_r over arbitrary time intervals along the real ODE, instead of an endpoint shortcut. Every extra sampling step then stays on a path the model actually knows, which restores test-time scaling: more NFEs, more quality, from 4 up to the 32 shown on the project page. Because the model is trained on-trajectory via Flow Map Backward Simulation, which decomposes a full many-step Euler rollout into shortcut transitions, the supervision also lands on the states the student visits at inference. That directly targets the two named failure modes: few-step discretization error, and exposure bias in causal generation where per-frame errors compound across a rollout.

The catch is verification. The abstract and project page establish the qualitative claims, matches-or-surpasses consistency methods in the few-step regime, scales at 4/16/32 NFE, works from 1.3B to 14B in both bidirectional and causal setups, but no single VBench or FVD figure is published in the material available, and the extra training machinery is not costed. “Any-step” is also doing some work in the title: the demonstrated window is 4 to 32 steps, not an unbounded budget.

The disagreement, precisely

Strip away the engineering and the two papers disagree about what a distilled video model is.

Causal Forcing++ treats it as a real-time component with a service-level objective: 14.1 FPS, 0.27s first frame, quality within noise of the baseline. Under that contract, a free step count is a liability. If the model sometimes wants 16 steps to look good, you cannot promise 14.1 FPS, and an interactive product cannot make that trade at runtime. So it fixes the budget at 2, accepts the small quality cost of 1 step being unusable, and optimizes everything else, especially training cost, around the fixed point.

AnyFlow treats the distilled model as a general-purpose generator that should obey the same contract as the original diffusion model: spend compute, get quality. Under that contract, a hard-coded 2-step budget is the liability, because it forecloses the 16-step final render a film or asset pipeline would happily pay for. The failure it fixes is real and measurable in spirit (consistency models visibly degrade past their tuned regime), even if the paper does not publish the numbers to size the effect.

There is a deeper technical point where they rhyme: both papers converge on on-trajectory supervision as the cure for autoregressive error. Causal consistency distillation and flow map distillation are different formalisms, but both replace “train on precomputed or shortcut states” with “train on states tied to the true ODE path the student will traverse.” Causal Forcing++ uses that idea to kill a precompute pass; AnyFlow uses it to restore scaling. Neither invents it from nothing; both are refinements of the observation that off-trajectory supervision is where few-step and autoregressive video models quietly break.

When to use which

  • Interactive or streaming product with a frame-rate guarantee: Causal Forcing++. It is the only one of the two with published latency and throughput numbers, and 0.27s first frame / 14.1 FPS at 2 steps on Wan2.1-1.3B is a spec you can build against. Budget for the honest caveat that the action-conditioned variant is still 4-step chunk-wise.
  • Offline generation where the user picks quality vs speed per job: AnyFlow’s premise. One checkpoint serving 4-step drafts and 32-step finals is exactly the workflow image and video tools want. Verify the published numbers before committing; as of now the scaling claim is qualitative.
  • Training budget is the binding constraint: Causal Forcing++ has the measured answer, ~4x cheaper few-step stage with zero trajectory storage. AnyFlow’s on-policy rollout decomposition plausibly costs more than plain consistency distillation, but the paper does not quantify it.
  • Long autoregressive rollouts where error compounds: both papers explicitly target exposure bias, AnyFlow by name with on-policy flow-map supervision, Causal Forcing++ indirectly through AR training plus distillation. There is no head-to-head; treat this as an open empirical question.

Limits and open questions

The comparison everyone wants, same model, same benchmark, both methods, does not exist in the published material. Causal Forcing++‘s numbers are all on a 1.3B student, so it is unknown whether the causal CD equivalence and the training savings survive on 14B-scale backbones. AnyFlow’s range claim (1.3B to 14B) has no published per-scale scores, so the scaling behavior cannot be checked against VBench. Neither paper releases a number for the other’s operating point: Causal Forcing++ does not show what happens past 4 steps, and AnyFlow does not show a 2-step latency figure. And the evaluation metrics differ in spirit, Causal Forcing++ measured on VBench and VisionReward at fixed steps, AnyFlow demonstrating step scaling without a headline metric, so even a future head-to-head will need matched harnesses to mean anything.

FAQ

Which is faster at inference, Causal Forcing++ or AnyFlow?

Causal Forcing++ is the only one with a published speed number: 14.1 FPS at 2 steps per frame and 0.27s first-frame latency on Wan2.1-1.3B. AnyFlow does not publish latency or throughput figures, so no same-harness speed comparison exists.

Does AnyFlow get better with more sampling steps?

That is its central claim: quality rises across the shown 4 to 32 NFE range, while consistency-distilled baselines degrade past their tuned step count. The published material shows the scaling curve but no single headline VBench or FVD score, so the effect is demonstrated in shape more than in quantified magnitude.

Is Causal Forcing++ higher quality than AnyFlow?

There is no data that answers this. Causal Forcing++ reports VBench 84.14 at 2 steps on Wan2.1-1.3B, essentially tied with its own 4-step baseline. AnyFlow reports no VBench number in the available material. Anyone claiming a quality winner is comparing across missing data.

Do both methods work for streaming and real-time video?

Both target the causal (frame-by-frame autoregressive) regime where streaming lives. Causal Forcing++ ships it as the main result but leaves the action-conditioned world-model variant at 4 steps chunk-wise; AnyFlow covers causal generation and claims reduced exposure bias via on-policy supervision, without publishing a real-time configuration.

One line: fix the step budget and engineer everything around real time, or free the step budget and restore pay-more-get-more. They are answers to different products. Read Causal Forcing++ and AnyFlow for the full breakdowns.