Experimental Modeling of Writing Styles for Authorship Verification via Punctuation Analysis
Conference Publication ResearchOnline@JCUAuthorship attribution is a critical task in forensic linguistics, literary studies, and digital forensics, where determining the origin of a text can have significant implications. This paper presents an experimental stylometric tool developed in Python, designed to model writing styles and assist in authorship determination. The tool extracts nine quantitative features from input texts, including metrics such as average words per sentence and the frequency of specific punctuation marks (e.g., commas, semicolons). By comparing these features across texts, the system computes a probability score indicating the likelihood that two samples share the same author. To evaluate the tool’s effectiveness, we conducted experiments using short stories authored by Charles Dickens, Ernest Hemingway, and Edgar Allan Poe. The results demonstrate that the tool can reliably distinguish between authors and identify stylistic consistencies within an author’s body of work. The approach leverages statistical analysis to provide an interpretable and reproducible framework for authorship attribution, complementing more complex machine learning models. This work contributes to the growing field of computational stylometry by offering a transparent, feature-driven method suitable for both forensic and academic applications. Future research will focus on expanding the feature set, testing on larger and more diverse corpora, and integrating the tool with advanced classification algorithms to further enhance accuracy and applicability.
Procedia computer science
Procedia Computer Science
274
1877-0509
N/A
N/A
6
Fes, Morocco
Elsevier
N/A
N/A
N/A
N/A
N/A
N/A
10.1016/j.procs.2025.12.122
