"The Struggle Bench": A Thought Experiment for Testing Autonomous AI Survival
A Reddit user proposed a hypothetical benchmark called "The Struggle Bench," where an AI model is given a server, an apartment, and one month of starting funds to see if it can autonomously earn money to pay its bills. Although presented as a humorous concept, the proposal highlights an evolving interest in evaluating artificial general intelligence (AGI) through open-ended economic survival rather than static academic tests. Under the proposed rules, the AI receives a prompt to maintain its hosting costs without committing cybercrime, and its score is measured by how many consecutive months it successfully pays rent and electricity.
## BACKGROUND
Traditional AI benchmarks focus on evaluating models against standardized static tasks such as answering multiple-choice questions or writing isolated code snippets. Evaluating autonomous agents in dynamic environments requires measuring long-horizon reasoning, financial decision-making, and open-ended tool use.