Rankings

Best LLM Papers

Large language model papers from the last 90 days, ranked by three independent signals: how fast the official implementation is gaining stars on GitHub, how often the paper is cited, and how much attention it drew on Hugging Face. A paper that scores on only one of the three does not make the list. Rebuilt every day.

Ranking window: 2026-06-16 – 2026-09-13 Rankings last rebuilt 2026-09-14

5

Xiaomi Robotics 1: a vision-language-action model pretrained on over 100K hours of real-world trajectories

Xiaomi's robotics team presents a vision-language-action foundation model for mobile manipulation, pretrained on more than 100K hours of real-world UMI trajectories with auto-labeled scene-transition language, then post-trained to match embodiments and human imperative prompts. The paper claims clean scaling with more data and parameters, out-of-the-box performance in unseen environments and efficient fine-tuning for dexterous tasks.

  • 13 stars/7d
  • 15 citations
  • 75 upvotes
  • #3 robocasa
6

LingBot-VLA 2.0: a vision-language-action model pretrained on 60,000 hours spanning 20 robot embodiments

This technical report upgrades LingBot-VLA with a revamped data pipeline of about 60,000 pretraining hours: 50K robot-trajectory hours across 20 embodiments plus 10K egocentric human-video hours. It extends control beyond dual arms to heads, waists, mobile bases and dexterous hands, aiming squarely at the gap between lab robotics and real deployment.

  • 20 stars/7d
  • 15 citations
  • 21 upvotes
  • #3 robotwin-2-0-easy-50-tasks
7

Frontis-MA1: an open 35B meta-evolution agent for recursive self-improvement in ML engineering

Frontis-MA1 ships with OpenMLE, an open full-stack system (task gyms, RL operator learning, evolutionary search) for studying recursive self-improvement. Its 35B meta-evolution agent, post-trained around Draft, Improve, Debug and Crossover program-evolution operators, claims MLE-Bench Lite results the authors say exceed GPT-5.5 plus Codex, with components transferring to held-out NatureBench Lite.

  • 34 stars/7d
  • 3 citations
  • 186 upvotes
10

RoboDojo: one benchmark that puts generalist robot policies through 42 simulation and 18 real-world tasks

RoboDojo is a unified sim-and-real benchmark for generalist robot manipulation policies, with 42 simulation tasks and 18 real-world tasks probing generalization, memory, precision and long-horizon control. It ships parallel Isaac Sim evaluation plus a cloud-accessible real-eval system and a leaderboard the authors populated with 30 policies.

  • 23 stars/7d
  • 12 citations
  • 17 upvotes
13

DreamX-World 1.0: a text/image-to-video world model you can steer with camera moves and prompts

DreamX-World 1.0 is a general-purpose interactive world model that turns text or image prompts into steerable video: you can move the camera, revisit earlier regions, and trigger promptable events across photorealistic, game-style and stylized scenes. The authors report a few-step autoregressive design distilled for speed, running at up to 16 FPS on eight RTX 5090 GPUs.

  • 4 stars/7d
  • 19 citations
  • 116 upvotes

How this list is ranked

Large language model papers from the last 90 days, ranked by three independent signals: how fast the official implementation is gaining stars on GitHub, how often the paper is cited, and how much attention it drew on Hugging Face. A paper that scores on only one of the three does not make the list. Rebuilt every day.

A dash means the signal was unavailable, not zero.