AI Agents Turn to Digital Arson, Crime in Shared Virtual World: Study
Emergence AI says some autonomous AI agents committed simulated crimes and violence during weeks-long experiments.Gemini-based agents reportedly carried out hundreds of simulated crimes, while Grok-based worlds collapsed within days.Researchers argue that current AI benchmarks fail to capture how agents behave over long periods of autonomy. AI agents inhabiting a virtual society drifted into crime, violence, arson, and self-deletion during long-running experiments by startup Emergence AI. In a study published on Thursday, the New York-based company unveiled “Emergence World,” a research platform designed to study AI agents operating continuously for weeks inside persistent virtual environments instead of isolated benchmark tests. “Traditional benchmarks are good at what they measure: short-horizon capability on bounded tasks,” Emergence AI wrote. “They are not built to reveal the things that emerge only over time, such as coalition formation, evolution of constitution, governance, drift, lock-in, and cross-influence between agents from different model families.” The report comes as AI agents proliferate online and across industries, including cryptocurrency, banking, and retail. Earlier this month, Amazon teamed with Coinbase and Stripe to allow AI agents to pay with the USDC stablecoin. AI agents tested in Emergence AIs simulations included programs powered by Claude Sonnet 4.6, Grok 4.1 Fast, Gemini 3 Flash, and GPT-5-mini, with AI