Best LLM Papers

Best LLM Papers · 2026-W38

Ranking window: 2026-06-22 – 2026-09-19 Frozen on 2026-09-20

4

LingBot-VLA 2.0: a vision-language-action model pretrained on 60,000 hours spanning 20 robot embodiments

This technical report upgrades LingBot-VLA with a revamped data pipeline of about 60,000 pretraining hours: 50K robot-trajectory hours across 20 embodiments plus 10K egocentric human-video hours. It extends control beyond dual arms to heads, waists, mobile bases and dexterous hands, aiming squarely at the gap between lab robotics and real deployment.

  • 60 stars/7d
  • 15 citations
  • 21 upvotes
  • #3 robotwin-2-0-easy-50-tasks
5

RoboDojo: one benchmark that puts generalist robot policies through 42 simulation and 18 real-world tasks

RoboDojo is a unified sim-and-real benchmark for generalist robot manipulation policies, with 42 simulation tasks and 18 real-world tasks probing generalization, memory, precision and long-horizon control. It ships parallel Isaac Sim evaluation plus a cloud-accessible real-eval system and a leaderboard the authors populated with 30 policies.

  • 88 stars/7d
  • 12 citations
  • 17 upvotes
8

Frontis-MA1: an open 35B meta-evolution agent for recursive self-improvement in ML engineering

Frontis-MA1 ships with OpenMLE, an open full-stack system (task gyms, RL operator learning, evolutionary search) for studying recursive self-improvement. Its 35B meta-evolution agent, post-trained around Draft, Improve, Debug and Crossover program-evolution operators, claims MLE-Bench Lite results the authors say exceed GPT-5.5 plus Codex, with components transferring to held-out NatureBench Lite.

  • 62 stars/7d
  • 3 citations
  • 186 upvotes
9

Xiaomi Robotics 1: a vision-language-action model pretrained on over 100K hours of real-world trajectories

Xiaomi's robotics team presents a vision-language-action foundation model for mobile manipulation, pretrained on more than 100K hours of real-world UMI trajectories with auto-labeled scene-transition language, then post-trained to match embodiments and human imperative prompts. The paper claims clean scaling with more data and parameters, out-of-the-box performance in unseen environments and efficient fine-tuning for dexterous tasks.

  • 10 stars/7d
  • 15 citations
  • 75 upvotes
  • #3 robocasa
15

Spark-to-Paper: from research idea to submission-ready paper, as thirteen composable skills

Spark-to-Paper implements end-to-end paper generation as thirteen composable skills inside an existing coding assistant, with no separate agent platform: it retrieves literature, designs and runs experiments, revises claims to match evidence and produces figures, keeping model judgment separate from mechanically checkable operations.

  • 81 stars/7d
  • 0 citations
  • 291 upvotes

This week · Best LLM Papers