Multimodal text-audio sentiment in clinical aphasia speech using NLP. Front Digit Health 2026;8:1740270
Date
08/25/2026Pubmed ID
42638756Pubmed Central ID
PMC13500762DOI
10.3389/fdgth.2026.1740270Abstract
INTRODUCTION: Aphasia affects expressive and receptive communication and may influence the affective tone expressed during clinical speech tasks. This study presents an exploratory weakly supervised NLP analysis of positive/negative affective-tone proxies in AphasiaBank transcripts with paired audio.
METHODS: We extracted sentence embeddings from DistilBERT (
RESULTS: Under the recording-disjoint split, both the text-only and fused text-audio models achieved 97.9% accuracy and 0.791 macro-F1, while the audio-only model achieved 55.4% accuracy and 0.388 macro-F1. These results indicate that classification performance was primarily driven by textual embeddings, while the recording-level acoustic summaries did not improve performance over text-only features. Across the pseudo-labeled corpus, aphasic utterances were more often labeled negative than control utterances. Age-stratified summaries showed subgroup variation in pseudo-label distributions, but these patterns were treated descriptively because labels were model-derived. Human-rater agreement was low for aphasic utterances, indicating that affective-tone interpretation in fragmented clinical speech is ambiguous.
DISCUSSION: These findings should be interpreted as exploratory evidence about weakly supervised affective-tone proxies, not as validated clinical sentiment recognition. The results highlight both the promise of clinical NLP for aphasia discourse analysis and the need for independent human-labeled validation, utterance-aligned acoustic features, and careful control of domain, task, age, and topic bias.









