Latent Space Evaluates Autonomous AI Engineer GPT-6 Astra Across 20B Tokens
Latent Space published an empirical deep dive sharing learnings from consuming over 20 billion tokens while testing an autonomous AI software engineer named GPT-6 Astra. The evaluation explores the practical capabilities of running an automated software developer at an operational cost of less than $6 per hour. This extensive testing offers rare empirical insight into how autonomous coding agents perform at extreme scale and operational efficiency. It highlights the growing feasibility of deploying low-cost AI agents for continuous, high-volume software engineering tasks. The study stress-tested the AI agent across long-context workloads to observe performance degradation, reasoning limits, and cost structure over 20 billion tokens. The agent is designed to handle complex engineering routines independently at standard hourly rates under $6.
## BACKGROUND
Autonomous AI software engineers represent an evolution beyond inline code autocomplete tools like GitHub Copilot. Systems like Devin and advanced LLM agents can independently navigate repositories, execute code, run tests, and resolve bugs with minimal human intervention. Evaluating these models requires large-scale token benchmarking to measure real-world reliability and economic viability.