Skip to main navigation Skip to search Skip to main content

Multimodal Emotion Recognition and Contextual Analysis in Therapy Sessions Using Video Large Language Models

  • Rabia Jafri*
  • , Sushant Patil
  • , Pranav Krishnakumar
  • , Syed Omar Ali
  • , Syed Abid Ali
  • , Syed Fawad Hussain
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Tracking patients’ emotions during therapy is crucial for accurate diagnosis and treatment planning, yet current automated emotion recognition methods are limited in that they either rely on a single modality or output only coarse emotion labels which are insufficient for therapeutic contexts. Recently, some multimodal techniques have employed video large language models (VLLMs) to produce emotional descriptions rather than mere labels; however, the potential of such approaches for therapeutic use remains underexplored. To address these gaps, we propose a VLLM-based, multimodal emotion detection system tailored for therapy dialogue. Our method isolates patient-only segments from session recordings, extracting both textual transcripts and audio features for each clip, which are processed using pretrained models to generate preliminary emotion labels. These, along with the segment transcript and preceding conversational context, are embedded in a structured prompt, which is passed to a VLLM with the corresponding video segment. The VLLM then produces an overall emotion classification, an explanation highlighting salient cues, a list of key observations, and a confidence estimate. Finally, the outputs across all patient segments are consolidated into a timeline reflecting the patient’s emotional progression throughout the session. 

Our key contribution is a narrative-based, interpretable emotion tracking system that offers therapists deeper insight into patient affect. By leveraging prompting for multimodal fusion, our approach avoids the need for domain-specific training and large datasets, which are particularly difficult to obtain in therapy contexts due to confidentiality. Preliminary results are promising and demonstrate the feasibility of applying VLLMs for enriched emotion analysis in therapeutic settings.

Original languageEnglish
Title of host publicationHCI International 2025 – Late Breaking Posters
Subtitle of host publication27th International Conference on Human-Computer Interaction, HCII 2025, Gothenburg, Sweden, June 22–27, 2025, Proceedings, Part II
EditorsConstantine Stephanidis, Margherita Antona, Stavroula Ntoa, George Margetis, Gavriel Salvendy
PublisherSpringer
Pages110-120
Number of pages11
ISBN (Electronic)9783032127679
ISBN (Print)9783032127662
DOIs
Publication statusPublished - 3 Jan 2026
Event27th International Conference on Human-Computer Interaction - Gothenburg, Sweden
Duration: 22 Jun 202527 Jun 2025
https://2025.hci.international/

Publication series

NameCommunications in Computer and Information Science
Volume2772 CCIS
ISSN (Print)1865-0929
ISSN (Electronic)1865-0937

Conference

Conference27th International Conference on Human-Computer Interaction
Abbreviated titleHCI International 2025
Country/TerritorySweden
CityGothenburg
Period22/06/2527/06/25
Internet address

Bibliographical note

Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • Affective Computing
  • Mental Health Technology
  • Multimodal Emotion Recognition
  • Prompt Engineering
  • Therapy Session Analysis
  • Video Large Language Models

ASJC Scopus subject areas

  • General Computer Science
  • General Mathematics

Fingerprint

Dive into the research topics of 'Multimodal Emotion Recognition and Contextual Analysis in Therapy Sessions Using Video Large Language Models'. Together they form a unique fingerprint.

Cite this