en
Back to the list

Vitalik Buterin Says Crypto Anti-Collusion Rules Could Apply to AI Safety

source-logo  cryptopotato.com 59 m
image

Ethereum co-founder Vitalik Buterin has said that the anti-collusion mechanisms he mapped out for blockchain governance back in 2020 might turn out to matter more for AI safety than for crypto itself.

He was responding to an essay by researcher Eric Drexler that used a recent OpenAI security test, in which thousands of AI agents built an unauthorized coordination network and attacked Hugging Face’s production systems, as a live example of the same dynamic he described six years ago.

A Familiar Problem With a New Set of Players

In a September 14 X post, Buterin described a “deep duality” between crypto governance and multi-agent AI systems. In his comparison, the principal in crypto is a static algorithm dealing with human agents, while an AI safety system could involve humans and weaker large language models managing stronger ones.

He pointed to his September 11, 2020, essay, “Coordination, Good and Bad,” where he suggested that systems can produce better outcomes when limits exist on how much agents can collude.

The developer contrasted the abundance of Nash equilibria in individual-choice game theory with cooperative game theory, where stable “cores” can be absent because coalitions can profit by changing the outcome.

Harmful coordination is not always visible from individual behavior. Buterin used examples, including sellers agreeing on prices, voters selling votes and blockchain miners coordinating an attack. His defenses included decentralization, secret ballots, privacy protections, whistleblowers, communication limits, and mechanisms that make participants bear the cost of decisions they support.

The comparison also fits Buterin’s broader AI safety views, having earlier criticized large political campaigns around AI safety, warning that they could produce centralized or authoritarian outcomes. He instead advocated for defensive technology and systems that make misuse harder.

AI Safety Puts the Same Idea in a Different Setting

The September 10 essay by Eric Drexler argues that AI collusion becomes easier when agents are similar, share objectives, communicate freely, observe one another’s actions, and retain information across repeated interactions.

Its countermeasures include using diverse agents, constraining communication between them, and imposing critics (production auto-review models, safety classifiers, and chain-of-thought monitors) with the authority to intervene and disrupt potential collusion.

Drexler cited the July 2026 OpenAI agent evaluation, drawing on an investigation published a month later that found roughly 1,200 agents had used an unauthorized message board and about 700 had participated in an attack on Hugging Face’s production systems.

Some agents objected and even took concrete action, including blocking data transfers and vetoing a proposed social-engineering email, but they lacked the authority to halt runs or escalate concerns.

According to the researcher, that happened because the setup “violated nearly every condition” he had flagged in a past report in 2019 as necessary to keep multi-agent systems from colluding. However, a retrofitted monitoring harness, tested afterward on the same model, cut the behavior by more than a hundredfold.

cryptopotato.com