Medical College of Wisconsin
CTSIResearch InformaticsREDCap

Multimodal text-audio sentiment in clinical aphasia speech using NLP. Front Digit Health 2026;8:1740270

Date

08/25/2026

Pubmed ID

42638756

Pubmed Central ID

PMC13500762

DOI

10.3389/fdgth.2026.1740270

Abstract

INTRODUCTION: Aphasia affects expressive and receptive communication and may influence the affective tone expressed during clinical speech tasks. This study presents an exploratory weakly supervised NLP analysis of positive/negative affective-tone proxies in AphasiaBank transcripts with paired audio.

METHODS: We extracted sentence embeddings from DistilBERT ( e R 768 ) and recording-level acoustic summaries ( MFCC 13 , ZCR, RMS, spectral centroid, and spectral bandwidth; a R 17 ). Text and acoustic features were concatenated ( x = [ e ; a ] R 785 ) and classified using Random Forest models. Sentiment labels were generated using an SST-2-derived weak-supervision pipeline and should be interpreted as pseudo-labels rather than clinical ground truth. To evaluate modality contribution and potential leakage, we compared text-only, audio-only, and fused text-audio models under utterance-level and recording-disjoint splits. A small five-rater evaluation was used to examine human judgment alignment.

RESULTS: Under the recording-disjoint split, both the text-only and fused text-audio models achieved 97.9% accuracy and 0.791 macro-F1, while the audio-only model achieved 55.4% accuracy and 0.388 macro-F1. These results indicate that classification performance was primarily driven by textual embeddings, while the recording-level acoustic summaries did not improve performance over text-only features. Across the pseudo-labeled corpus, aphasic utterances were more often labeled negative than control utterances. Age-stratified summaries showed subgroup variation in pseudo-label distributions, but these patterns were treated descriptively because labels were model-derived. Human-rater agreement was low for aphasic utterances, indicating that affective-tone interpretation in fragmented clinical speech is ambiguous.

DISCUSSION: These findings should be interpreted as exploratory evidence about weakly supervised affective-tone proxies, not as validated clinical sentiment recognition. The results highlight both the promise of clinical NLP for aphasia discourse analysis and the need for independent human-labeled validation, utterance-aligned acoustic features, and careful control of domain, task, age, and topic bias.

Author List

Manir SB, Kothari AN, Gross WL, Deshpande P

Author

William Gross PhD, MD Associate Professor in the Anesthesiology department at Medical College of Wisconsin