BAAI Releases AREX-2: A 27B Long-Horizon Agent Model for Iterative Self-Improvement
The Beijing Academy of Artificial Intelligence (BAAI) has released AREX-2, a 27B-parameter long-horizon agent model built on Qwen3.8 architecture. The model is capable of iteratively refining solutions at inference time through a continuous loop of proposing, measuring, reflecting, and revising. AREX-2 demonstrates that training agents on verifiable coding and machine learning tasks allows test-time self-improvement to effectively transfer to deep research tasks. This highlights a scalable path for open-source AI agents to sustain productive reasoning as test-time computational budgets scale. The model features a dense Qwen3.8-compatible multimodal architecture with a 262,144-token context window. It evaluates scores, logs, errors, and timing data during test time to decide subsequent revisions without requiring new search trajectories.
## BACKGROUND
Long-horizon AI agents are designed to execute complex, multi-step tasks autonomously over extended periods. Test-time refinement techniques allow large language models to continuously improve their outputs by inspecting feedback and execution results before delivering a final solution.