SoundSphere: Multimodal Captioning for Equitable and Collaborative VR Offices

Citations

SCOPUS

0

초록

We present SoundSphere, a multimodal captioning system designed to support equitable and collaborative work within Virtual Reality (VR) offices. While VR enables immersive professional collaboration, existing captioning solutions often lack non-verbal context and conversation history, resulting in structural communication barriers for users with hearing impairments during fast-paced meetings. SoundSphere addresses these challenges through a hybrid architecture that integrates on-device processing with cloud-based intelligence. The system employs YAMNet for real-time environmental sound classification (e.g., alarms) and orchestrates Whisper and Large Language Models (LLMs) to provide accurate transcription, real-time translation, and context-aware meeting summarization. By integrating verbal and non-verbal cues, SoundSphere aims to promote more equitable participation in collaborative VR workspaces. A preliminary system validation with XR domain experts indicates that SoundSphere's hybrid architecture is technically viable and effective in reducing cognitive load during complex discussions. Rather than framing captioning as a supplementary assistive feature, our work positions it as an infrastructural component of inclusive and ethically grounded VR collaboration. Through a heuristic evaluation of verbal and non-verbal cue integration, this study establishes a validated technical foundation as a preparatory step toward future deployment in inclusive virtual offices. © 2026 IEEE.

키워드

AccessibilityEquityInclusive DesignMultimodal CaptioningXR Collaboration
제목
SoundSphere: Multimodal Captioning for Equitable and Collaborative VR Offices
저자
Kim, HaneolKim, HyojungPark, Sanghun
DOI
10.1109/VRW70859.2026.00028
발행일
2026
유형
Conference paper
저널명
Proceedings - 2026 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops, VRW 2026
페이지
128 ~ 133