AI rewriting tools homogenize writing style and obscure author traits
An analysis of over 880,000 texts found that large language models reduce stylistic variation in writing by 21% to 50% while preserving meaning. This homogenization makes it harder for computer models to infer authors' personal characteristics such as age, extraversion, and neuroticism. The findings suggest that reliance on AI rewriting could obscure important linguistic clues about identity and mental health.
The study examined over 880,000 texts spanning Reddit posts, news articles, academic papers, essays, social media content, and political speeches. Three LLMs—GPT-3.5, Llama 3 70B, and Gemini Pro—were tasked with rewriting thousands of human-written samples. The analysis revealed that stylistic variation in writing complexity dropped by 21% to 50%, while meaning preservation remained high, with 87% of rewrites scoring above 0.95 on similarity measures.
The homogenization effect weakened computer models' ability to detect author traits by an average of six percentage points. Some linguistic associations, such as pronoun use with extraversion and future-focused words with age, faded after rewriting. However, other markers persisted, including negative-emotion words linked to neuroticism and social words tied to gender, suggesting LLMs selectively preserve certain identity signals.
The growing reliance on AI-assisted writing could subtly erode the linguistic fingerprints that researchers, clinicians, and hiring professionals use to understand individuals. If LLM rewriting becomes widespread, language-based assessments in psychology, mental health screening, and recruitment may become less reliable, potentially affecting how people are evaluated and served. While some identity markers survive rewriting, the overall trend toward homogenization could reduce the richness of personal expression, with implications for both automated systems and human readers who depend on stylistic cues.