Text-to-Image · Diffusion Models · Multimodal Models

DALL·E 2 vs Stable Diffusion vs Imagen, by the Numbers

DALL·E 2 diffuses in CLIP-embedding space with a learned prior, Stable Diffusion in a cheap latent, Imagen behind a frozen T5, and SD3 swaps the U-Net for a rectified-flow transformer.

DALL·E 2 vs Stable Diffusion vs Imagen, by the Numbers

The four papers and the one decision that splits them

Every text-to-image model answers the same question: where does the diffusion model actually do its work? The Latent Diffusion Model (LDM, 2021) answered “inside a compressed autoencoder latent, because pixel-space diffusion costs hundreds of GPU-days per training run.” DALL·E 2 (2022) answered “in CLIP’s joint image-text space, because that space already encodes semantics and style.” Imagen (2022) answered “in pixel space at 64×64, because the language encoder, not the denoiser, is the bottleneck.” And Stable Diffusion 3 (2024) kept LDM’s latent space but replaced almost everything else: rectified flow instead of the standard noise schedule, and a transformer with separate text and image weights instead of a U-Net.

Everything else (release strategy, resolution strategy, what they can spell) follows from those choices.

Key numbers

MeasurementLDM (SD base)DALL·E 2 (unCLIP)ImagenStable Diffusion 3Same protocol?
COCO 256×256 zero-shot FID ↓23.31 → 12.63 with classifier-free guidance at scale 1.510.39 (diffusion prior)7.27Not headline; reports CLIP-feature FID plus human preferenceMostly; same COCO validation split tradition, different guidance settings
Prior ablation, test-set FID ↓n/a9.16 full unCLIP stack, 7.99 text-embedding decoder, 16.55 text embedding fed to decoder directlyn/an/aWithin DALL·E 2 only
Prior typen/aDiffusion prior beats autoregressive: 10.39 vs 10.63n/an/aYes
Backbone1.45B-param U-Net with cross-attention conditioning3.5B-param GLIDE-style U-Net decoder2B-param Efficient U-Net baseMM-DiT transformer, 800M to 8B paramsNo; different parameterizations
Text encoderSmall transformer on the BERT tokenizer, trained with the modelFrozen CLIP text embedding + learned prior that maps text to a CLIP image embeddingFrozen T5-XXL, never sees an imageTwo CLIP encoders plus T5-XXL; T5 droppable at inferenceNo
Resolution path256×256 in one modelDecoder emits 64×64, upsamplers 700M + 300M go 64→256→1024Cascade: 2B base at 64×64, 600M to 256, 400M to 1024Native 1024×1024No
Human preference evidencen/aPreferred over GLIDE variants on diversity at matched photorealismPreferred over DALL·E 2, LDM, GLIDE, VQ-GAN+CLIP on DrawBench; alignment rated on par with COCO reference images8B model preferred over DALL·E 3, Midjourney v6, Ideogram on PartiPrompts (quality, prompt following, typography)Each paper’s own raters

Read the table the way it deserves: the FID column is the only row where all three 2021–2022 systems share a benchmark, and even there the guidance settings differ (LDM’s 12.63 uses classifier-free guidance; the DALL·E 2 and Imagen numbers are at their own tuned scales). Imagen’s 7.27 is the strongest number on the board, but its authors also ran the DrawBench human evaluations precisely because they did not trust FID to settle the question. SD3 went further and effectively stopped reporting headline COCO FID at all, judging itself on human preference instead.

Bet one: where the diffusion happens

LDM’s argument is economic. Pixel-space denoising evaluates a large network on a full-resolution tensor hundreds of times per sample; moving the process into a pretrained autoencoder’s latent (downsampling factor 8 for the text-to-image model) cuts both training and inference cost while a perceptual loss keeps reconstructions faithful. The cost shows up directly in the FID column: unguided LDM scores 23.31 zero-shot on COCO, but add classifier-free guidance at scale 1.5 and the same 1.45B model drops to 12.63, competitive with GLIDE’s 12.24 and within reach of Make-A-Scene’s 11.84, at a fraction of the training compute.

DALL·E 2 took the opposite bet: do not compress images, compress semantics. CLIP already maps an image and its caption to nearby points in a shared embedding space, so unCLIP inverts it: a prior generates a CLIP image embedding from the caption, and a diffusion decoder renders pixels from that embedding. The paper’s own ablation is the strongest evidence this is the load-bearing piece: the full unCLIP stack scores 9.16 on the test set versus 7.99 for a text-embedding decoder (slightly better FID, but humans prefer the full stack 57.0% of the time), while feeding a text embedding to the unCLIP decoder zero-shot collapses to 16.55. Skipping the prior saves nothing that matters.

And the prior itself is diffusion, not autoregression: 10.39 versus 10.63 COCO FID for the autoregressive variant, cheaper to run and better quality. That is two decisions (“use a prior” and “make the prior a diffusion model”), both backed by measured numbers, not taste.

Bet two: what reads the prompt

The second axis is the text encoder. LDM used a modest transformer over the BERT tokenizer, trained jointly with the image model, enough for LAION-scale captions, and the weakest link in hindsight. DALL·E 2 leaned on CLIP’s frozen text encoder but had to bridge the text-to-image-embedding gap with its prior. Imagen’s central finding is that this whole axis matters more than the denoiser: growing a frozen T5-XXL text encoder improved both fidelity and image-text alignment more than growing the 64×64 diffusion U-Net did, and T5-XXL never saw a single image during Imagen’s training. Prompt comprehension (clauses, attributes, relations) was bottlenecked by language modeling capacity, not by denoising capacity.

The practical consequence is that Imagen’s alignment held up where FID-focused systems slipped: human raters judged Imagen samples on par with the COCO reference images themselves on image-text alignment, and on the 200-prompt DrawBench suite they preferred Imagen over DALL·E 2, LDM, GLIDE, and VQ-GAN+CLIP on both quality and alignment.

SD3’s answer is to stop choosing. It stacks two CLIP encoders and T5-XXL, and its ablation quantifies exactly what T5 is worth: drop it and the model’s win rate against the full version falls to 50% on visual aesthetics (no effect), 46% on prompt adherence (small effect), and 38% on rendering written text (large effect). Typography, the capability that made SD3 famous, turns out to be mostly the T5’s doing.

Bet three: how you reach high resolution

None of these models generates 1024×1024 in one shot. LDM’s text-to-image model is a single 256×256 system (higher resolutions came later via concatenated conditioning and fine-tuning). DALL·E 2 and Imagen both cascade: DALL·E 2’s decoder emits 64×64 and two diffusion upsamplers (700M and 300M parameters) take it to 256 and then 1024, corrupting the conditioning image during upsampler training for robustness. Imagen uses the same coarse-to-fine philosophy but with text conditioning kept at every stage and noise-conditioning augmentation, so artifacts from the 64×64 base do not propagate cleanly into the final image.

SD3 is the break with the pattern: a native 1024×1024 latent model, no cascade, because the transformer backbone and the rectified-flow objective scale well enough that one model carries the full resolution.

What SD3 changed and why it held

The 2024 paper keeps LDM’s latent space and jettisons two things that had defined Stable Diffusion: the standard diffusion noise schedule and the U-Net. Rectified flow connects data and noise with a straight line instead of a curved path, which is cheaper to integrate, but the paper is honest that the straight line alone was not the win. The win was biasing training timesteps toward the perceptually hard middle of the path with a logit-normal sampler, which is what let rectified flow beat the established diffusion formulations in their sweep of 61 formulations and schedule combinations.

The MM-DiT architecture gives image tokens and text tokens separate weights, joined by bidirectional attention, on the argument that the two modalities are different distributions that a shared-weight transformer compromises. Then the paper runs the scaling study the earlier systems never did: validation loss falls smoothly from 800M to 8B parameters at 5×10²² training FLOPs, and lower validation loss tracks both automatic metrics and human preference, so quality is bought predictably rather than by lottery. The 8B model is rated above DALL·E 3, Midjourney v6, and Ideogram by human evaluators on PartiPrompts across visual quality, prompt following, and typography.

Notably, this is the same paper in which the earlier systems’ weaknesses converge: SD3 is latent (LDM’s lesson), stacks a strong text encoder and shows it carries typography (Imagen’s lesson), and evaluates itself on human preference rather than FID leadership (a lesson every earlier paper learned the hard way).

When to use which

  • You need to self-host, fine-tune, or build a product on open weights: the LDM line is the only option in this comparison, and SD3 is its current end state. LDM gave the ecosystem its efficiency, SD3 gave it typography, prompt following, and a scaling curve that had not flattened at 8B.
  • You want maximum prompt fidelity inside a closed product: the evidence chain runs Imagen → SD3. The frozen-T5 finding generalizes: prompt comprehension is bought with language-model capacity, so judge systems on hard compositional prompts (DrawBench, PartiPrompts), not on COCO FID.
  • You care about diversity at fixed realism: DALL·E 2’s specific claim. Its zero-shot COCO FID is worse than Imagen’s (10.39 vs 7.27), but its guidance behavior is better: classifier-free guidance degrades unCLIP’s FID far less than GLIDE’s, because the CLIP-embedding bottleneck preserves variety even when the decoder sharpens.
  • You are choosing a training recipe from scratch in 2024 or later: rectified flow with a logit-normal timestep bias, separate text and image weights with joint attention, and a strong frozen text encoder. SD3’s sweep of 61 alternatives is the most expensive ablation in this literature; its conclusion is the modern default.

Limits and open questions

No controlled, same-harness comparison exists across these four systems. Every number in the table comes from its own paper, with its own training data, guidance scale, and raters; even the COCO FID column mixes guidance settings. The human-preference results (DrawBench for Imagen, PartiPrompts for SD3) are pairwise studies sensitive to prompt selection and annotator pools, and they cannot be stacked into a single ranking. Two of the four systems (Imagen, and DALL·E 2 as originally released) never shipped open weights, so their numbers are not independently reproducible. And FID itself measures distribution match, not per-image correctness, which is exactly why both later papers moved their headline claims to human evaluators.

FAQ

Is Stable Diffusion better than DALL·E 2?

On the zero-shot COCO FID reported in the original papers, the 2021 LDM behind Stable Diffusion scores 12.63 with guidance versus DALL·E 2’s 10.39. DALL·E 2’s number is better, but the models were trained and sampled differently, so the gap is weak evidence. The decisive difference is practical: the LDM line ships open weights and runs on consumer GPUs, which is why an entire fine-tuning ecosystem exists around it and none exists around unCLIP.

Why does Imagen use T5 instead of CLIP as its text encoder?

Because Imagen’s ablation showed scaling a frozen, text-only T5 encoder improved both image fidelity and image-text alignment more than scaling the image diffusion model. A language model pretrained on far more text understands prompts (clauses, attributes, relations) better than an encoder limited to image-caption pairs. SD3 later kept the lesson: dropping T5-XXL from its three-encoder stack drops typography win rate to 38% while leaving aesthetics untouched.

What does the DALL·E 2 prior actually do?

It maps a text caption to a CLIP image embedding, which a separate diffusion decoder then renders into pixels. Without the prior (feeding the text embedding straight to the decoder) FID collapses from 9.16 to 16.55 in the paper’s ablation. The prior also explains unCLIP’s diversity: generation happens in CLIP embedding space, so guidance sharpens the render without collapsing the variety of embeddings the prior can produce.

How does Stable Diffusion 3 differ from the original Stable Diffusion architecture?

Same latent space, different everything else: rectified flow with a logit-normal timestep bias instead of the standard diffusion schedule, a multimodal diffusion transformer with separate image and text weights instead of a U-Net, three text encoders instead of one, and a scaling study from 800M to 8B parameters showing validation loss tracks human preference. It generates 1024×1024 natively instead of upsampling.

Can these models spell text inside images?

Mostly no, with one exception in this set. Spelled-out text was a known failure mode of LDM, unCLIP, and Imagen; SD3’s whole pitch is reliable typography. Its own ablation attributes that capability mostly to the T5-XXL encoder, not to the rectified-flow or transformer changes.