Top AI Research Papers

Top AI Research Papers · 2026-W37

Ranking window: 2026-06-16 – 2026-09-13 Frozen on 2026-09-14

2

LingBot-World 2.0: a game world model with an unbounded interaction horizon and a 60 fps real-time variant

LingBot-World 2.0 upgrades the game world model to an unbounded interaction horizon with consistent output quality via a causal pretraining paradigm, distills a real-time variant fast enough to drive 720p video at 60 fps, and greatly expands the action set (attacking, archery, spell-casting, shooting) plus text-driven interactive elements.

  • 149 stars/7d
  • 17 citations
  • 47 upvotes
6

LLM-as-a-Verifier: continuous verification scores read from scoring-token logits, no judge prompts needed

LLM-as-a-Verifier turns a language model into a verifier by computing expectations over its scoring-token logits instead of prompting for discrete judgments, yielding continuous feedback with no extra training. The authors report state-of-the-art results on benchmarks including Terminal-Bench V2 and SWE-Bench Verified, and ship a Claude Code extension that watches agents and supplies dense rewards for RL.

  • 134 stars/7d
  • 11 citations
  • 19 upvotes
8

Xiaomi Robotics 1: a vision-language-action model pretrained on over 100K hours of real-world trajectories

Xiaomi's robotics team presents a vision-language-action foundation model for mobile manipulation, pretrained on more than 100K hours of real-world UMI trajectories with auto-labeled scene-transition language, then post-trained to match embodiments and human imperative prompts. The paper claims clean scaling with more data and parameters, out-of-the-box performance in unseen environments and efficient fine-tuning for dexterous tasks.

  • 13 stars/7d
  • 15 citations
  • 75 upvotes
  • #3 robocasa
9

Alaya-EVOKE: linear-scaling supervision with an explicit world-state bank for endlessly evolving worlds

Alaya-EVOKE externalizes persistent scene geometry into a camera-indexed world state bank so the denoiser context stays bounded, and trains with a sparse-attention teacher that supervises long horizons at linear cost, exposing content drift and supporting prompt changes mid-sequence. The teacher distills into a three-step student that the authors report producing each 1.5-second chunk in 2.11s on a single H200 GPU.

  • 93 stars/7d
  • 2 citations
  • 166 upvotes
  • #2 wbench-navigation-split
11

LingBot-VLA 2.0: a vision-language-action model pretrained on 60,000 hours spanning 20 robot embodiments

This technical report upgrades LingBot-VLA with a revamped data pipeline of about 60,000 pretraining hours: 50K robot-trajectory hours across 20 embodiments plus 10K egocentric human-video hours. It extends control beyond dual arms to heads, waists, mobile bases and dexterous hands, aiming squarely at the gap between lab robotics and real deployment.

  • 20 stars/7d
  • 15 citations
  • 21 upvotes
  • #3 robotwin-2-0-easy-50-tasks
14

ABot-World-0: an action-conditioned world model streaming interactive 720p worlds on a single desktop GPU

ABot-World-0 is an action-conditioned video world model trained on multi-source data from AAA games, simulation engines and internet videos, controlled through raw keyboard input with reference-character memory for consistent third-person rollouts. The authors report streaming 720p at up to 16 FPS on one RTX 5090 with 1.2s action-to-first-frame latency, targeting real-time long-horizon closed-loop interaction.

  • 11 stars/7d
  • 6 citations
  • 313 upvotes
15

OSWorld 2.0: 108 long-horizon computer-use workflows where even the best agent finishes only a fifth

OSWorld 2.0 from the xlang-ai team replaces short scripted tasks with 108 long-horizon computer-use workflows built on authentic artifacts and stateful user profiles; humans take a median of about 1.6 hours per task. The headline finding is how far agents still are: the best model, Claude Opus 4.8, completes only 20.6% of tasks, often losing track of constraints or skipping verification.

  • 19 stars/7d
  • 15 citations
  • 25 upvotes
18

A survey of self-improving agents: how a foundation model plus scaffold turns experience into capability

This survey decomposes a self-improving agent into a foundation model and an operational scaffold of prompts, memory, tools and control logic, then defines improvement as a self-induced update operator that converts experience into accumulated capability with minimal human input. It organizes the growing literature by update target and driving signal, from prompt and memory edits to weight-level self-modification.

  • 24 stars/7d
  • 9 citations
  • 35 upvotes
19

DreamX-World 1.0: a text/image-to-video world model you can steer with camera moves and prompts

DreamX-World 1.0 is a general-purpose interactive world model that turns text or image prompts into steerable video: you can move the camera, revisit earlier regions, and trigger promptable events across photorealistic, game-style and stylized scenes. The authors report a few-step autoregressive design distilled for speed, running at up to 16 FPS on eight RTX 5090 GPUs.

  • 4 stars/7d
  • 19 citations
  • 116 upvotes
20

Frontis-MA1: an open 35B meta-evolution agent for recursive self-improvement in ML engineering

Frontis-MA1 ships with OpenMLE, an open full-stack system (task gyms, RL operator learning, evolutionary search) for studying recursive self-improvement. Its 35B meta-evolution agent, post-trained around Draft, Improve, Debug and Crossover program-evolution operators, claims MLE-Bench Lite results the authors say exceed GPT-5.5 plus Codex, with components transferring to held-out NatureBench Lite.

  • 34 stars/7d
  • 3 citations
  • 186 upvotes

This week · Top AI Research Papers