Understanding users’ emotional states in immersive virtual reality (VR) remains a key challenge in human–computer interaction (HCI), particularly when relying on non-intrusive signals that naturally arise during interaction. Although extended reality (XR) systems generate rich embodied behavioral data, such as gaze, body movement, and interaction context, these signals are often modeled as static features, limiting their ability to capture temporal and contextual emotional dynamics. This paper proposes a conceptual and methodological framework that models embodied VR interaction as a language-like, sequential process, in which emotional meaning emerges from temporally structured patterns of behavior rather than isolated cues. Inspired by principles of natural language processing, multimodal interaction logs are represented as sequences of token-like interaction units and analyzed using a transformer-based architecture as a sequence-to-sequence model. The aim is not to introduce a new emotion classifier, but to examine the feasibility and interpretability of this representation paradigm. An empirical VR study demonstrates that embodied behavioral signals alone contain structured emotional information that can be learned by sequential models, particularly for stable affective states. This work contributes a reusable design pattern for modeling embodied interaction in emotion-aware XR systems.

VrT: A Framework to Model Embodied VR Interaction as a Language-Like Sequence for Emotion Inference

Foresti G. L.
2026-01-01

Abstract

Understanding users’ emotional states in immersive virtual reality (VR) remains a key challenge in human–computer interaction (HCI), particularly when relying on non-intrusive signals that naturally arise during interaction. Although extended reality (XR) systems generate rich embodied behavioral data, such as gaze, body movement, and interaction context, these signals are often modeled as static features, limiting their ability to capture temporal and contextual emotional dynamics. This paper proposes a conceptual and methodological framework that models embodied VR interaction as a language-like, sequential process, in which emotional meaning emerges from temporally structured patterns of behavior rather than isolated cues. Inspired by principles of natural language processing, multimodal interaction logs are represented as sequences of token-like interaction units and analyzed using a transformer-based architecture as a sequence-to-sequence model. The aim is not to introduce a new emotion classifier, but to examine the feasibility and interpretability of this representation paradigm. An empirical VR study demonstrates that embodied behavioral signals alone contain structured emotional information that can be learned by sequential models, particularly for stable affective states. This work contributes a reusable design pattern for modeling embodied interaction in emotion-aware XR systems.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11390/1337368
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
social impact