상세 보기
SoundSphere: Multimodal Captioning for Equitable and Collaborative VR Offices
- Kim, Haneol;
- Kim, Hyojung;
- Park, Sanghun
SCOPUS
0초록
We present SoundSphere, a multimodal captioning system designed to support equitable and collaborative work within Virtual Reality (VR) offices. While VR enables immersive professional collaboration, existing captioning solutions often lack non-verbal context and conversation history, resulting in structural communication barriers for users with hearing impairments during fast-paced meetings. SoundSphere addresses these challenges through a hybrid architecture that integrates on-device processing with cloud-based intelligence. The system employs YAMNet for real-time environmental sound classification (e.g., alarms) and orchestrates Whisper and Large Language Models (LLMs) to provide accurate transcription, real-time translation, and context-aware meeting summarization. By integrating verbal and non-verbal cues, SoundSphere aims to promote more equitable participation in collaborative VR workspaces. A preliminary system validation with XR domain experts indicates that SoundSphere's hybrid architecture is technically viable and effective in reducing cognitive load during complex discussions. Rather than framing captioning as a supplementary assistive feature, our work positions it as an infrastructural component of inclusive and ethically grounded VR collaboration. Through a heuristic evaluation of verbal and non-verbal cue integration, this study establishes a validated technical foundation as a preparatory step toward future deployment in inclusive virtual offices. © 2026 IEEE.
키워드
- 제목
- SoundSphere: Multimodal Captioning for Equitable and Collaborative VR Offices
- 저자
- Kim, Haneol; Kim, Hyojung; Park, Sanghun
- 발행일
- 2026
- 유형
- Conference paper
- 저널명
- Proceedings - 2026 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops, VRW 2026
- 페이지
- 128 ~ 133