Authors leave distinctive linguistic traces that reflect their identity through consistent writing styles, particularly in morphosyntactic patterns and lexical choices. However, the increasing use of AI writing tools challenges authorship attribution because machine-generated text can imitate or obscure individual writing characteristics. This study investigates linguistic features that effectively identify authors and differentiate human-written from AI-generated texts. A corpus comprising 2,074,125 tokens and 63,414 word types was compiled from collaborative digital platforms, including instant messaging and social media. Lexical and stylistic features were extracted to develop hybrid authorship-classification models, while N-gram tracing was used to identify salient patterns. The findings demonstrate that lexical choice is the most reliable indicator for distinguishing human and AI-generated texts. Character-level N-gram analysis further demonstrates that authorship can be identified through delicate patterns involving letters, capitalization, punctuation, and other non-alphabetic characters. Diction appeared as the strongest factor in differentiating individual authors. These results enhance the reliability of authorship attribution methods and provide valuable insights for forensic investigations of digitally mediated communication involving human–AI interaction.
Copyrights © 2026