Best LLM Papers

Best LLM Papers · 2026-W39

Ranking window: 2026-06-29 – 2026-09-26 Frozen on 2026-09-27

3

LingBot-VLA 2.0: a vision-language-action model pretrained on 60,000 hours spanning 20 robot embodiments

This technical report upgrades LingBot-VLA with a revamped data pipeline of about 60,000 pretraining hours: 50K robot-trajectory hours across 20 embodiments plus 10K egocentric human-video hours. It extends control beyond dual arms to heads, waists, mobile bases and dexterous hands, aiming squarely at the gap between lab robotics and real deployment.

  • 42 stars/7d
  • 15 citations
  • 21 upvotes
  • #3 robotwin-2-0-easy-50-tasks
5

RoboDojo: one benchmark that puts generalist robot policies through 42 simulation and 18 real-world tasks

RoboDojo is a unified sim-and-real benchmark for generalist robot manipulation policies, with 42 simulation tasks and 18 real-world tasks probing generalization, memory, precision and long-horizon control. It ships parallel Isaac Sim evaluation plus a cloud-accessible real-eval system and a leaderboard the authors populated with 30 policies.

  • 46 stars/7d
  • 12 citations
  • 17 upvotes
6

Xiaomi Robotics 1: a vision-language-action model pretrained on over 100K hours of real-world trajectories

Xiaomi's robotics team presents a vision-language-action foundation model for mobile manipulation, pretrained on more than 100K hours of real-world UMI trajectories with auto-labeled scene-transition language, then post-trained to match embodiments and human imperative prompts. The paper claims clean scaling with more data and parameters, out-of-the-box performance in unseen environments and efficient fine-tuning for dexterous tasks.

  • 7 stars/7d
  • 15 citations
  • 75 upvotes
  • #3 robocasa
7

Frontis-MA1: an open 35B meta-evolution agent for recursive self-improvement in ML engineering

Frontis-MA1 ships with OpenMLE, an open full-stack system (task gyms, RL operator learning, evolutionary search) for studying recursive self-improvement. Its 35B meta-evolution agent, post-trained around Draft, Improve, Debug and Crossover program-evolution operators, claims MLE-Bench Lite results the authors say exceed GPT-5.5 plus Codex, with components transferring to held-out NatureBench Lite.

  • 39 stars/7d
  • 3 citations
  • 186 upvotes
11

WROP: a cognitive-science exam that tests whether video world models understand object permanence — and a 16B model trained to pass it

When an object leaves the frame, does a video generation model remember it exists? The authors built WROP, 150 cognitive-science tasks across six categories rendered in Blender with randomized lighting, speed, and camera angle — a 1.5M-sample training corpus plus a 300-question exam used to test 14 video models. They then trained PWM-WROP, a 16B world model on their native PyTorch stack for AWS Trainium2, releasing data, exam, scores, and weights. In a blind pairwise Elo study, it ranked first among continuation models and effectively tied the best reference-to-video models.

  • 109 stars/7d
  • 0 citations
  • 193 upvotes
12

Spark-to-Paper: from research idea to submission-ready paper, as thirteen composable skills

Spark-to-Paper implements end-to-end paper generation as thirteen composable skills inside an existing coding assistant, with no separate agent platform: it retrieves literature, designs and runs experiments, revises claims to match evidence and produces figures, keeping model judgment separate from mechanically checkable operations.

  • 74 stars/7d
  • 0 citations
  • 291 upvotes
13

Realtime-Venus: two open 9B full-duplex models that keep talking while delegating tool work to a background harness

Realtime-Venus is a proactive full-duplex interaction stack built from two separately trained 9B models: Realtime-Venus-Omni handles audio-visual conversation and Realtime-Venus-Audio handles spoken dialogue, each integrating continuous perception, conversational control and native speech generation on one shared causal timeline. A dual-loop runtime lets the foreground conversation continue while Realtime-Venus-Harness executes delegated tasks asynchronously and folds results back into the live dialogue; the authors report Realtime-Venus-Omni leads evaluated online models on six of eight video benchmarks and Realtime-Venus-Audio beats Gemini 3.1 Live and GPT-4o on all three Full-Duplex-Bench continuation metrics.

  • 91 stars/7d
  • 0 citations
  • 195 upvotes
14

LimiX-2: a tabular foundation model that learns data-generating mechanisms instead of just predicting targets

LimiX-2 is a tabular foundation model built on Contextual Mechanism Networks: instead of the usual in-context objective of predicting y given x and context tables, it learns the joint structure p(x, y | context) of how the data was generated. It is pretrained with context-conditional masked modeling on synthetic datasets built from structural causal models spanning varied graph structures and mechanisms. The authors report state-of-the-art results against both dataset-specific models and other tabular foundation models on TabArena, TALENT, and BCCO, and show the mechanism-oriented training has a side effect: the model's feature attention encodes direct causal relationships, allowing it to recover causal skeletons from data.

  • 80 stars/7d
  • 0 citations
  • 188 upvotes
15

VideoChat3: a fully open 4B video MLLM that adapts frame resolution on the fly for streaming video

Most open video-language models are either narrowly specialized or only partly reproducible. VideoChat3 goes fully open — code, training strategy, and three curated corpora (Academic2M for general video, LV116K for long-form, OL617K for streaming). Its I3D-ViT encoder inflates a 2D vision transformer into 3D for cheaper spatiotemporal processing, and an adaptive frame-resolution scheme spends compute only where the streaming input needs it. The authors report outperforming earlier open-source models of equal or larger size with only 4B parameters.

  • 6 stars/7d
  • 4 citations
  • 172 upvotes
  • #2 tempcompass-mcq

This week · Best LLM Papers