Small Language Models · Hugging Face
DistilBERT: A Smaller, Faster BERT by Distillation
DistilBERT distills BERT during pretraining into a 6-layer model with 40% fewer parameters that runs about 60% faster and keeps roughly 97% of BERT-base's GLUE score.