The goal of Automatic Swabian Recognition (ASR) is to make the Arno Ruoff archive fully usable for both the scientific community and the wider public by creating a curated, open dataset with parallel tiers of (i) sound recordings, (ii) existing transcriptions in German orthography, and (iii) phonetic transcriptions in IPA produced by a dialect-adapted automatic speech recognition pipeline.
Developing robust automatic speech recognition models on archival dialect data enables transfer to contemporary applications facing similar challenges: regional variation, spontaneous speech, and limited training data. The project addresses a critical gap: current automatic speech recognition systems fail on regional dialects, limiting cognitive assessment for neurodegenerative diseases. By validating dialect-aware models on the TREND cohort (1,200 participants, Tübingen region, longitudinal recordings), we establish speech as a non-invasive digital biomarker while ensuring robustness across patient populations.
Building on this dual foundation of historical dialect archives and contemporary clinical speech data, the project integrates machine-learning methods, quantitative models of linguistic change, and qualitative methods of cultural-anthropological analyses of archival meaning and access.