AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
Original reporting by arXiv (cs.AI)

AREX-2 refers to a new research initiative aimed at significantly advancing the self-improving capabilities of large language model (LLM) agents, specifically their ability to iteratively refine solutions during real-time use. This ambitious goal hinges on two critical, complementary skills: reflection, which enables an agent to generate solutions superior to its current attempts, and long-horizon execution, ensuring that these iterative improvements remain effective over numerous rounds. Researchers hypothesized that both these capabilities are domain-agnostic, meaning they can be learned and generalized from targeted training scenarios.
Training for Iteration
To test this hypothesis, the team synthesized extensive 'long-horizon improvement trajectories' from machine learning and algorithmic programming tasks. These domains were chosen for their clear, verifiable feedback mechanisms and the inherent reward for sustained, iterative problem-solving. By training an agent built on Qwen3.8-27B with this specialized data, AREX-2 demonstrated remarkable performance. It achieved strong scores on standard benchmarks like MLE-bench Lite (81.8) and Frontier-CS (70.7), but more importantly, exhibited impressive transferability to complex deep research tasks, including BrowseComp (84.0), GAIA (92.2), and DeepSearchQA (93.8). Crucially, the system proved its core self-improvement mettle by consistently enhancing its solutions as its allocated budget of refinement rounds increased, underscoring the effectiveness of long-horizon reflective data in forging truly self-improving AI agents.
The AREX-2 initiative marks a significant stride in the development of truly self-improving LLM agents. By meticulously integrating advanced reflection with long-horizon execution capabilities, AREX-2, powered by Qwen3.8-27B, demonstrates a robust capacity for iterative solution refinement at test time. Its impressive performance across diverse, challenging benchmarks—from machine learning engineering to deep research tasks like BrowseComp and GAIA—validates the efficacy of training with synthesized long-horizon reflective data. Crucially, the agent's ability to maintain and enhance its performance over numerous improvement rounds underscores a pivotal breakthrough: sustainable, autonomous improvement is not just theoretical but demonstrably achievable, setting a new precedent for agentic AI that learns and adapts.
Beyond Benchmarks
This development extends far beyond mere benchmark scores, pointing towards a future where AI systems can independently tackle complex, multi-step problems with increasing sophistication. The self-improving nature of AREX-2 suggests a paradigm shift from static, pre-trained models to dynamic, adaptive agents capable of continuous learning and optimization without constant human oversight. Such agents could accelerate discovery in scientific research, streamline intricate engineering processes, and deliver increasingly personalized and effective solutions across diverse industries. Ultimately, the AREX-2 project offers a compelling blueprint for the next generation of AI, moving closer to systems that can autonomously enhance their own capabilities, pushing the boundaries of what intelligent machines can achieve.
Frequently asked questions
- What is AREX-2 and what is its primary contribution to large language model agents?
- AREX-2 is an advanced large language model (LLM) agent designed to enhance self-improvement at test time. Its primary contribution is demonstrating an effective method for LLM agents to iteratively refine solutions. This involves learning domain-agnostic capabilities like reflection and sustained long-horizon execution, leading to continuously improving performance across diverse complex tasks. AREX-2 showcases that specialized training data can effectively cultivate these crucial self-improving capabilities in AI.
- How do LLM agents like AREX-2 achieve self-improvement during task execution?
- LLM agents like AREX-2 achieve self-improvement through two core capabilities: reflection and long-horizon execution. Reflection enables the agent to evaluate its current solution and generate a better one. Long-horizon execution ensures that these iterative refinements remain effective over many rounds, preventing degradation or getting stuck. By learning these capabilities from specialized training scenarios, agents can autonomously enhance their performance and accuracy on complex problems during the problem-solving process.
- Which domains or tasks were used to train AREX-2 for its self-improving abilities?
- AREX-2 was specifically trained using data synthesized from machine learning (ML) and algorithmic programming tasks. These domains were chosen because they provide verifiable feedback and reward sustained iteration, making them ideal for supervising the learning of self-improvement capabilities. After training, AREX-2 demonstrated strong transferability, successfully applying its learned reflective and long-horizon execution skills to complex deep research tasks like BrowseComp, HLE, GAIA, and DeepSearchQA.