Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification

August 23, 2026 Β· Grace Period Β· πŸ› The 2026 Conference on Empirical Methods in Natural Language Processing

⏳ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Qian Ma, Anna Squicciarini, Sarah Rajtmajer arXiv ID 2608.22161 Category cs.AI: Artificial Intelligence Cross-listed cs.CL Citations 0 Venue The 2026 Conference on Empirical Methods in Natural Language Processing
Abstract
Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation methods optimize privacy independently for each document, leaving them blind to cross-document correlations that make aggregation dangerous. We propose Aggregation-Aware Synthetic Text Generation (AAST), a framework that addresses this gap by jointly selecting synthetic texts at the bundle level rather than optimizing each text in isolation. AAST targets attribution and verification attacks, including cross-genre settings where attacker references come from a genre not observed during generation or selection. Experiments across same-genre, cross-genre, neural, and independent non-neural stylometric attacks show that AAST lowers account-level linkability as bundle size grows, while preserving semantic quality, linguistic acceptability, and sentiment alignment.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Artificial Intelligence