Researchers Evaluate Multi-Agent AI Swarms in Victorian Detective Game
Researchers developed a Victorian London detective game to evaluate whether a swarm of small AI agents coordinated by a reasoning model can solve procedurally generated murder mysteries better than a single lone model. The system tested various configurations, pitting a structured leader model paired with search agents against solitary large reasoning models and leaderless swarms. This research provides empirical insight into the strengths and limitations of multi-agent LLM systems versus single large models, highlighting how division of labor between parallel search agents and a centralized reasoner improves problem-solving. These findings help guide the design of more efficient agentic architectures for complex, information-dense tasks. In tests featuring false evidence and witness tampering, a setup combining a strong reasoning leader with a swarm of search agents achieved a 93% success rate, whereas a leaderless swarm failed completely due to premature consensus on incorrect clues. Meanwhile, a lone reasoning model struggled because it lacked the physical capacity to visit enough addresses within the time limit.
## BACKGROUND
Procedural generation is a computing method where data, such as game levels or virtual worlds, is created algorithmically using rules and randomness rather than manually crafted designs. In artificial intelligence, multi-agent systems use distributed autonomous entities that communicate, share information, and collaborate to solve complex problems that typically overwhelm a single monolithic model.