Best AI Agent Papers

Best AI Agent Papers · 2026-W38

Ranking window: 2026-06-22 – 2026-09-19 Frozen on 2026-09-20

1

LingBot-World 2.0: a game world model with an unbounded interaction horizon and a 60 fps real-time variant

LingBot-World 2.0 upgrades the game world model to an unbounded interaction horizon with consistent output quality via a causal pretraining paradigm, distills a real-time variant fast enough to drive 720p video at 60 fps, and greatly expands the action set (attacking, archery, spell-casting, shooting) plus text-driven interactive elements.

  • 87 stars/7d
  • 17 citations
  • 47 upvotes
5

Dream-RSI: coding agents that get better at exploring by replaying their own past discoveries offline

Dream-RSI is a framework for letting coding agents improve their own exploration strategy instead of keeping a fixed one. It builds a replay simulator from the agent's past discovery trees, lets the exploration policy practice inside that simulator for cheap off-policy feedback, then redeploys the refined strategy online so new discoveries feed the next round. The authors report competitive or better discovery quality on algorithm engineering, math optimization, and GPU kernel tasks, with substantially lower discovery cost in several settings.

  • 900 stars/7d
  • 0 citations
  • 270 upvotes
6

ABot-World-0: an action-conditioned world model streaming interactive 720p worlds on a single desktop GPU

ABot-World-0 is an action-conditioned video world model trained on multi-source data from AAA games, simulation engines and internet videos, controlled through raw keyboard input with reference-character memory for consistent third-person rollouts. The authors report streaming 720p at up to 16 FPS on one RTX 5090 with 1.2s action-to-first-frame latency, targeting real-time long-horizon closed-loop interaction.

  • 16 stars/7d
  • 6 citations
  • 313 upvotes
8

Atria Dawn: a foundation agentic model for scientific research, trained on verifiable experience

Atria Dawn introduces a foundation agentic language model built for scientific research and engineering workflows. Its training pipeline, which the authors call the Verifiable Experience Pipeline, ties tool-mediated interactions to executable environments and externally checked outcomes; the paper also studies human-AI collaboration using 769 task records from 56 researchers working with the agent.

  • 509 stars/7d
  • 0 citations
  • 421 upvotes
9

OSWorld 2.0: 108 long-horizon computer-use workflows where even the best agent finishes only a fifth

OSWorld 2.0 from the xlang-ai team replaces short scripted tasks with 108 long-horizon computer-use workflows built on authentic artifacts and stateful user profiles; humans take a median of about 1.6 hours per task. The headline finding is how far agents still are: the best model, Claude Opus 4.8, completes only 20.6% of tasks, often losing track of constraints or skipping verification.

  • 20 stars/7d
  • 15 citations
  • 25 upvotes
10

LLM-as-a-Verifier: continuous verification scores read from scoring-token logits, no judge prompts needed

LLM-as-a-Verifier turns a language model into a verifier by computing expectations over its scoring-token logits instead of prompting for discrete judgments, yielding continuous feedback with no extra training. The authors report state-of-the-art results on benchmarks including Terminal-Bench V2 and SWE-Bench Verified, and ship a Claude Code extension that watches agents and supplies dense rewards for RL.

  • 33 stars/7d
  • 11 citations
  • 19 upvotes
11

ZGCM-1: a fully open 7B model pairing long thinking with tool use for math and agentic search

ZGCM-1 is a fully open 7B dense foundation model aimed at math reasoning and agentic search. The idea is that a smaller model cannot memorize the whole web, so it couples long internal thinking with active tool use, using interleaved gated sliding-window and full attention, an FP8 Muon optimizer, and curriculum training that scales context to 256K. The authors say they release stage-wise weights, training code, per-stage data recipes, and W&B logs.

  • 475 stars/7d
  • 0 citations
  • 300 upvotes
12

A survey of self-improving agents: how a foundation model plus scaffold turns experience into capability

This survey decomposes a self-improving agent into a foundation model and an operational scaffold of prompts, memory, tools and control logic, then defines improvement as a self-induced update operator that converts experience into accumulated capability with minimal human input. It organizes the growing literature by update target and driving signal, from prompt and memory edits to weight-level self-modification.

  • 21 stars/7d
  • 9 citations
  • 35 upvotes
13

SoL-Pi: letting an auto-research loop redesign the coding-agent harness itself

SoL-Pi applies recursive self-improvement to the agent harness rather than the model: an automated research loop scaled across many environments discovers reusable mechanisms for action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench benchmark, the authors report matching Pi-level performance with 44.7-49.0% less recorded token traffic and roughly one-third lower API cost.

  • 975 stars/7d
  • 0 citations
  • 66 upvotes

This week · Best AI Agent Papers