AI Systems Enable Mass Deanonymization of Online Profiles at Minimal Cost

Researchers from ETH Zurich and other institutions demonstrated that artificial intelligence agents can identify individuals behind pseudonymous accounts for approximately $1-$4 per attempt, automating a task that previously required skilled human investigators. Testing on known accounts, the AI achieved 67% accuracy in identifying Hacker News users and 52% accuracy on Reddit academics after stripping identifying information. The findings highlight growing vulnerabilities in internet anonymity as machine learning systems become increasingly capable of connecting disparate online data points to reveal real identities.
The research team conducted controlled experiments using accounts with publicly verifiable identities to measure their system's effectiveness. They tested the approach on three distinct populations: Hacker News users linked to LinkedIn profiles, Reddit users with academic credentials, and software engineers with career information. Success rates varied significantly across these groups, ranging from 25% to 67% depending on the profile type and available information density, suggesting that account characteristics and posting patterns substantially influence identification difficulty.
The deanonymization process operates through a multi-stage methodology that progressively narrows possibilities. An initial language model extraction identifies salient biographical features from public posts, followed by embedding-based filtering that reduces millions of potential matches to hundreds of similar profiles. Subsequent reasoning stages evaluate the most plausible candidates before reaching conclusions, enabling the system to synthesize disparate data points into coherent identity inferences.
This development could reshape privacy considerations for individuals maintaining pseudonymous online presence. Journalists, activists, vulnerable populations, and researchers relying on anonymity for protection face potentially reduced safety margins if such tools become widely accessible. Conversely, the technology's modest accuracy rates and dependence on sufficient posted information may offer some practical limitations. Organizations and platforms may need to reassess anonymity protections and data exposure policies, while individuals could face decisions about online participation scope and information sharing practices.