Multimodal Contrastive Learning for Dialogue Embeddings with Global and Local Views

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Recent dialogues increasingly include not only text but also images. However, research on multimodal dialogue embeddings remains limited, and existing methods face three primary challenges: i) they often reflect only the context of specific moments when images are shared, failing to capture the overall meaning of the dialogue, ii) they require large batch sizes to achieve high performance when using contrastive learning, and iii) they generally lack evaluation of intrinsic tasks that directly measure the structural quality and semantic consistency of embeddings, as they rely heavily on extrinsic tasks. To address these issues, we propose MMCDE, multimodal contrastive learning for dialogue embeddings with global and local views. Our method constructs contrastive pairs by leveraging a global view that considers the context of the entire dialogue and a local view that captures interactions between images and text within the dialogue. We are the first to include both extrinsic and intrinsic tasks in the evaluation of performance in multimodal dialogue embedding research. Furthermore, we demonstrate that our approach achieves superior performance across three tasks, even with limited memory and batch sizes. Our code is available at https://github.com/subeenc/MMCDE our GitHub repository.

키워드

Dialogue EmbeddingMultimodalContrastive Learning
제목
Multimodal Contrastive Learning for Dialogue Embeddings with Global and Local Views
저자
Choe, SubeenOh, Ji hyeonYang, Ji hoon
DOI
10.1007/978-981-96-8180-8_13
발행일
2025-06
유형
Proceedings Paper
저널명
Lecture Notes in Computer Science
15872 LNAI
페이지
155 ~ 166