Self-Supervised Learning · Vision Foundation Models
SimCLR vs BYOL vs MAE vs DINOv2: ImageNet Numbers, Three Eras
SimCLR hit 69.3 ImageNet linear with negatives and 4096 batches; BYOL dropped negatives for 74.3 and small-batch robustness; MAE owns fine-tuning at 87.8 but freezes poorly; DINOv2's frozen 86.3 beats OpenCLIP 86.2.
The one-line answer
These four papers are not four entries on one leaderboard. They mark three different eras of self-supervised vision, and each one wins a different game. SimCLR proved that instance discrimination, pulling two augmentations of the same image together and pushing everything else apart, scales far better than anyone had shown. BYOL then removed the “push everything else apart” part, the negative pairs, and still improved. MAE abandoned instance discrimination entirely and learned to reconstruct masked pixels, a task that makes it the strongest fine-tuning pretext on ImageNet while producing weak frozen features. DINOv2 closed the loop by showing that self-supervision, given curated data at scale, produces frozen features that beat OpenCLIP, the strongest text-supervised alternative at the time. Pick the era that matches your constraint, not the biggest number in the table.
What each method actually changed
SimCLR’s contribution was a clean recipe, not a new objective. Contrastive instance discrimination existed before; SimCLR showed that with a strong augmentation composition, a nonlinear projection head, a batch size of 4096, and long training, a ResNet-50 learns features that reach 69.3 top-1 under linear evaluation on ImageNet. The batch size matters because negative examples come from the same batch, so the model’s ability to discriminate scales with how many distractors it sees at once.
BYOL kept the two augmented views but deleted the distractors. Each view passes through an online network and a target network; a small predictor on the online side is trained to regress the target’s output, and the target is an exponential moving average of the online weights. Nothing in the loss prevents collapse to a constant output except the predictor plus the moving-average target, a fact the paper demonstrates with ablations and a theoretical argument. The practical payoff is that BYOL no longer needs thousands of in-batch negatives, so it degrades gracefully when the batch gets small.
MAE changed the task itself. It masks roughly 75% of an image’s patches and trains a lightweight decoder to rebuild the missing pixels from the visible quarter, on a Vision Transformer backbone. The encoder never sees mask tokens, so pretraining compute drops by roughly 3x and the method scales to ViT-Huge on ImageNet-1k alone. Because the pretext task is reconstruction rather than discrimination, the features it learns are excellent after full fine-tuning and mediocre when frozen, the single most important fact in this comparison.
DINOv2’s claim is about the usage mode. It trains ViTs with a self-distillation objective on LVD-142M, a 142-million-image dataset that Meta curated automatically for quality and deduplication, then distills the 1B-parameter teacher into smaller models. The selling point is that you freeze the backbone, attach a linear classifier or even a kNN, and get features that match or beat text-supervised OpenCLIP on most image-level and pixel-level benchmarks.
Key numbers
| Measurement | SimCLR | BYOL | MAE | DINOv2 | Setting and source | Same harness? |
|---|---|---|---|---|---|---|
| ImageNet linear, ResNet-50 (1x) | 69.3 | 74.3 | n/a | n/a | BYOL paper Table 1, 1000-epoch training | Same arch and protocol family; both rows as reported by the BYOL authors |
| ImageNet linear, ResNet-50 (4x) | 76.5 | 78.6 | n/a | n/a | BYOL paper Table 1, 375M params | Same table, aligned protocol |
| ImageNet linear, ResNet-200 (2x) | n/a | 79.6 | n/a | n/a | BYOL paper Table 1 | n/a |
| ImageNet linear, ResNet-50, 300 epochs, batch 4096 | 67.9 (BYOL’s reproduction) | 72.5 | n/a | n/a | BYOL paper Table 16 | Yes, same data, codebase, and schedule |
| Same run at batch 256 | 64.3 | 71.8 | n/a | n/a | BYOL paper Table 16 | Yes |
| Semi-supervised, 1% of ImageNet labels | 48.3 | 53.2 | n/a | n/a | BYOL paper Table 2, ResNet-50 | Same protocol as stated by BYOL |
| Semi-supervised, 10% of labels | 65.6 | 68.8 | n/a | n/a | BYOL paper Table 2 | Same protocol |
| ImageNet-1k fine-tune, ViT | n/a | n/a | 83.6 (B) / 85.9 (L) / 86.9 (H) / 87.8 (H, 448px) | n/a | MAE paper Table 3 | MAE-internal scaling curve, no contrastive arm |
| ImageNet-1k linear probe, ViT | n/a | n/a | 68.0 (B) / 75.8 (L) / 76.6 (H) | 86.3 (ViT-L/14) / 86.5 (ViT-g/14) | MAE paper Table 12; DINOv2 paper Table 4, which also lists MAE ViT-H at 76.6 | No: MAE trained on IN1K only, DINOv2 on LVD-142M |
| ImageNet-1k kNN on frozen features, ViT | n/a | n/a | 49.4 (ViT-H) | 82.1 (ViT-B/14) / 83.5 (ViT-L/14) | DINOv2 paper Table 4, single eval harness | Yes for the DINOv2 family; MAE row measured in that same table |
| ADE20k segmentation, frozen linear head (mIoU) | n/a | n/a | 33.3 (ViT-H) | 47.7 (ViT-L/14) | DINOv2 paper Table 10; OpenCLIP ViT-G scores 39.3 in the same table | Yes, one eval table across methods |
Read the “Same harness?” column before ranking anything. The cleanest rows are the middle ones: the BYOL authors retrained SimCLR themselves under an identical 300-epoch schedule and report both numbers, and the gap there is 4.6 points at batch 4096 and 7.5 points at batch 256. The ViT rows are claims from different papers with different pretraining data, and the DINOv2 table’s value is precisely that it put MAE, DINO, iBOT, and OpenCLIP into one evaluation harness, which is the closest thing to a controlled comparison this literature has.
Why dropping negatives was a big deal
SimCLR’s implicit bet was that instance discrimination needs a large supply of negatives, and its ablations backed that: shrinking the batch hurts because each image has fewer things to be distinguished from. BYOL’s reproduction quantifies the stakes. At 300 epochs with batch 4096, their SimCLR reimplementation reaches 67.9, and BYOL reaches 72.5. Shrink the batch to 256 and SimCLR falls to 64.3 while BYOL only slips to 71.8. In other words, most of SimCLR’s dependence on giant batches comes from the negatives themselves, not from contrastive learning in general.
The mechanism the paper offers is the predictor. Without it, or without the moving-average target, representations collapse. With both, the online network is encouraged to produce features from which a one-layer network can predict the target’s features, which implicitly structures the representation space without any repulsion term. Later analyses added caveats: BatchNorm statistics in the target encoder leak information, and the loss landscape has flat directions that make training finicky. But the empirical claim stood up well enough that “no negatives” became a whole subfield, and DINOv2’s self-distillation objective is a descendant of the same family.
MAE wins the fine-tuning game and loses the frozen one
MAE’s headline numbers sit at opposite ends of the table. Fine-tuned, ViT-Huge/448 reaches 87.8 top-1 on ImageNet-1k using only ImageNet-1k data, above supervised pretraining of the same backbone at the time of publication. Frozen, the same model family gets 76.6 linear and 49.4 kNN, and even ViT-Base only reaches 68.0 linear. In DINOv2’s unified evaluation, MAE ViT-H’s 49.4 kNN is not a typo-level gap but a 33-point chasm against DINOv2 ViT-B’s 82.1.
The reason is the pretext task. Pixel reconstruction forces the network to keep low-level information alive, which is exactly what full fine-tuning can exploit and exactly what a frozen linear classifier does not need. Contrastive and distillation objectives throw away pixel detail and keep what discriminates images, which is what frozen probes reward. So the 87.8 and the 49.4 are the same model saying the same thing: MAE’s features are a great initialization, not a great endpoint. Its other practical virtue is cost, since the encoder only processes the visible 25% of patches, pretraining runs roughly 3x faster than feeding every patch through.
DINOv2 closes the frozen-feature gap
DINOv2’s central table is the one worth bookmarking. On ImageNet-1k linear evaluation, DINOv2 ViT-L/14 reaches 86.3 and ViT-g/14 reaches 86.5, against 86.2 for text-supervised OpenCLIP ViT-G/14 trained on LAION-2B, a dataset roughly 14 times larger than LVD-142M. The kNN column tells the same story more sharply: DINOv2 ViT-B/14 gets 82.1 with no classifier training at all, while MAE ViT-H gets 49.4. On pixel-level tasks the gap widens further: frozen linear segmentation on ADE20k gives DINOv2 ViT-L 47.7 mIoU versus MAE ViT-H 33.3 and OpenCLIP ViT-G 39.3 in the same table.
The lesson the paper draws, and the one worth copying, is about data curation. Earlier self-supervised work assumed scale washes out noise; DINOv2 argues quality is the lever, and built an automatic filtering pipeline to assemble LVD-142M. The objective itself is a recombination of existing self-distillation techniques rather than a new one, which suggests the recipe, large batches, KoLeo regularizer, iBOT masked-token loss, long schedule, matters as much as the idea.
When to use which
- You will fine-tune on a labeled target set and want the best ImageNet-scaled initialization: MAE. 87.8 fine-tune with ViT-H/448 on IN1K-only data remains the reference point, and the 3x pretraining speedup makes it cheap to reproduce at moderate scale.
- You need features with zero downstream labels, for retrieval, segmentation, depth, or probing: DINOv2. Its frozen linear and kNN numbers beat every alternative in its own table, and the distilled small models are the practical tier.
- You are pretraining a ResNet-style model and cannot afford 4096-image batches: BYOL. The same-harness batch sweep is the evidence: 71.8 at batch 256 against SimCLR’s 64.3.
- You want the simplest possible baseline with public recipes everywhere: SimCLR. It is the method every later paper reimplemented, which is why its numbers appear in everyone else’s tables. Use it as the reference, not the ceiling.
Limits and open questions
The four papers never run each other’s models, so almost every cross-method cell above is a claim placed side by side, not a head-to-head. The exceptions are the rows from the BYOL paper’s own SimCLR reproduction and DINOv2’s unified evaluation table, and even the latter compares models trained on very different data regimes, IN1K only versus LVD-142M, so you cannot attribute the gap to the objective alone. All four papers also evaluate on ImageNet and friends; what happens on medical, satellite, or scientific imagery is extrapolation, and frozen-feature quality says nothing about fairness or safety. Finally, the field has moved on: newer work combines masked reconstruction with distillation, and CLIP-style text supervision has kept improving, so treat DINOv2’s edge over OpenCLIP as a 2023 snapshot rather than a permanent ordering.
FAQ
Which is better on ImageNet, SimCLR or BYOL?
Under the protocol the BYOL paper used, BYOL is better at every point of comparison: 74.3 versus 69.3 on ResNet-50 linear evaluation at 1000 epochs, and 72.5 versus 67.9 in BYOL’s own 300-epoch reproduction at batch 4096. The advantage grows as the batch shrinks, at batch 256 BYOL holds 71.8 while the SimCLR reproduction drops to 64.3, because SimCLR’s negatives live in the batch and BYOL has none.
Can BYOL really work without negative samples?
Yes, and the paper explains why it does not collapse. The online network has an extra predictor layer trained to match the target network’s output, and the target is an exponential moving average of the online weights rather than an identical copy. Ablations show that removing either ingredient collapses the representation. The predictor-and-EMA combination imposes enough structure that no repulsion term is needed.
Why is MAE’s linear probe so much lower than its fine-tune accuracy?
Because pixel reconstruction keeps low-level detail that a frozen classifier cannot use but full fine-tuning can. MAE ViT-H/448 reaches 87.8 top-1 when fine-tuned on ImageNet-1k, yet the same ViT-H backbone gets only 76.6 linear and 49.4 kNN in frozen evaluations. If your pipeline freezes the backbone, MAE is the wrong choice despite the headline number.
Is DINOv2 better than CLIP?
On DINOv2’s own ImageNet-1k table, yes at the frozen-feature game: DINOv2 ViT-L/14 reaches 86.3 linear and its ViT-g/14 reaches 86.5, against 86.2 for OpenCLIP ViT-G/14, despite OpenCLIP training on LAION-2B, a much larger dataset. The pixel-level margins are larger: 47.7 mIoU on ADE20k segmentation with a frozen linear head, versus 39.3 for OpenCLIP in the same table. Whether it is better for your task depends on whether you need text-aligned semantics, which self-supervised features do not provide for free.
Which should I use, frozen features or fine-tuning?
Use frozen features from DINOv2 when you have no labels and need general-purpose features fast. Use MAE pretraining as an initialization when you have a labeled target set and the compute to fine-tune, that is where its 87.8 comes from. Use BYOL over SimCLR if you are working in the ResNet regime with limited batch size.