Standard orthography suppresses most of the phonetic detail present in speech, yet non-standard writing — including dialect literature, eye dialect, transcribed oral narratives, computer-mediated communication, and historical spelling — often preserves systematic traces of allophonic variation. This paper proposes a replicable procedure for identifying and categorizing such allophonic markers in non-standard orthographic material. The identification procedure moves through five stages, from corpus assembly and deviation extraction to frequency screening, phonological environment coding, and hypothesis cross-checking, and includes explicit tests for excluding eye dialect, orthographic borrowing, and typographical error. Identified markers are then classified along four independent dimensions: phonological process type, structural domain, distributional status, and sociolinguistic motivation. Four illustrative case studies demonstrate how the procedure distinguishes genuine, systematically conditioned markers from superficially similar but non-phonetic respellings. The paper also discusses methodological challenges, including limited corpus size, hypercorrection, genre-internal heterogeneity, and the risk of confirmation bias during cross-checking, and outlines applications of the framework in historical linguistics, sociolinguistics, language documentation, and computational sociolinguistics.
Copyrights © 2026