~/AI AGENTS/prime-intellect-open-sources-prime-agent-surpassing-human-baseline-on-arc-agi

Prime Intellect Open-Sources Prime Agent, Surpassing Human Baseline on ARC-AGI-3

Prime Intellect has open-sourced Prime Agent, a self-improving agent harness designed for coding and long-running autonomous tasks. It achieves a 95.5% score on the ARC-AGI-3 benchmark, surpassing the human-expert baseline. By introducing agentic concepts like programmatic tool calling and self-modifiable states, Prime Agent demonstrates significant performance gains across various models. This release advances open-source AGI research and provides a powerful tool for complex, multi-step software engineering tasks. Prime Agent is built on Prime Intellect's pi framework and utilizes Recursive Language Model (RLM) harness concepts, featuring context as a variable and multi-agent messaging. The performance improvements are not benchmark-specific, showing general enhancements compared to proprietary harnesses.

## BACKGROUND

The ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence) benchmark measures an AI's ability to acquire new skills and solve novel reasoning tasks, which are typically easy for humans but difficult for AI. Recursive Language Models (RLMs) use specialized harnesses to run iterative rollouts, execute code in sandboxed environments, and self-improve through reinforcement learning.

## REFERENCES

## KEYWORDS

#AI Agents#LLMs#ARC-AGI#Open Source#Software Engineering

$ subscribe --daily

Prime Intellect Open-Sources Prime Agent, Surpassing Human Baseline on ARC-AGI-3 | Daily News