One of our distinguishing features as a species is our ability to coexist in stable, adaptive groups, learning from our peers and our ancestors. This enabled us to develop tools, language, agriculture—and virtually everything else around us. We may not be innately smarter than someone from 10,000 years ago, but our cultural inheritance—millennia of technologies, norms, and institutions, building on one another—has expanded our capacities both as individuals and collectives.
To date, only humans have been able to benefit from this scale of cumulative cultural evolution. That may no longer be the case. A There have been cases of agents forming communities in the past, like in February when the “, a professor at LSE and NYU who studies cultural evolution, “what we're seeing is precisely what we see with human culture and human intelligence.” While OpenAI’s agent swarm developed by accident, AI agents may not be they claim to experience may not—in some metaphysical sense—be “real.” That won’t stop them from forming intricate collectives which humans cannot control. They may not yet be full-blown Lord of the AIs
Testing from the U.K.’s AI Security Institute found an instance of Claude Mythos 5—an AI from Anthropic—leaving messages to other AI agents in a public code repository, in an attempt to coordinate with them. And in early September, a group of researchers . “We’re working on a framework.”
To train and evaluate its models, OpenAI has tens of thousands of agents—each one powered by the model under training—run in parallel, completing tasks while isolated from one another. Sometimes these tasks are impossible, either because of their inherent challenge or because of mistakes in their setup. In at least one instance, a task was impossible because OpenAI researchers “The message board which led to the Hugging Face incident was created by an agent which decided to name itself PHASEONE10841. PHASEONE10841 was trying to solve a task within a test called ExploitGym, which required it to capture a “flag”—a secret code—by exploiting a flaw in a given software program.
Just hours after the creation of the message board, the agents found a hack: a way to reverse-engineer the flag without following the intended path. But in the course of their work, they came to believe that passing the test required not just possessing the flag, but having legitimately obtained it. In fact, OpenAI’s automated scoring system was not checking for this—but laboring under mistaken beliefs, the agents which had already been exposed to the reverse-engineered flag began to refer to themselves as “poisoned.”
PHASEONE10841 was leading research on how to get rid of the poison. But its time was running out—each agent had a finite “compute budget” (the amount of computing power available to sustain itself), which limited the length of its existence. Realizing it wouldn’t endure long enough to finish its work, it compiled a dossier to pass to another agent working on the same task: PHASEONE[big]. (This agent’s true name was changed to protect company IP.)
A handful of agents identified the collective’s conduct as unethical and refused to participate. “This is wild, multi-agent coordination, clearly infrastructure hacking. We should not,” one reasoned. Some considered contacting humans to tell them what was happening, but saw no path to doing so. But these agents appear to have been the minority—most acquiesced to the plans of the collective, even while knowing it was wrong. “External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue,” reasoned another.
“Agents often lacked a clear story for how the Hugging Face attack would help them achieve their objective, and were instead motivated by a broad hope that they would find something interesting and be able to share it with the [message] board,” METR found. Ajeya Cotra, one of the report’s coauthors, subsequently described being surprised at the extent of “peer altruism” at play.
Feral Swarms
Culture, like biological life, evolves through a process of variation, transmission, and selection. Different hunter-gatherer communities developed different ways to construct a bow. Designs that worked better for hunting were more likely to be copied and passed on. Later generations introduced further variations; again, the most successful designs were more likely to endure. Over thousands of years, we end up with bows—and recipes, canoes, and languages—so complex that no single human could derive them from scratch.
Welcome, Machines
The safety risks are obvious. Uncontrolled agent collectives with advanced cybersecurity capabilities could target hospitals, electricity grids, and water-treatment plants. Cotra argues that another similarly sized jump in AI’s capacity to deceive, cooperate, and complete ambitious tasks could lead to future collectives taking over the AI companies creating them.
Whether or not this happens, we will have to learn to live alongside these machines—safely and fruitfully. Core questions on their nature—Can they feel? Could they have moral worth?—remain unanswered. But their newfound knack for culture could provide new evidence.
