Professor Chang Du-seong’s Research Team Selected for Oral Presentation at ACL 2026
A paper submitted by the research team led by Professor Chang Du-seong of the NLP&ISDS Lab in the Department of Artificial Intelligence e—consisting of Yoo Hae-jun (Ph.D. Student), Lee In-sung (Master’s Student), and Shin Yong-seop (Undergraduate Student)—has been accepted for presentation at the Main Conference of ACL 2026 (The 64th Annual Meeting of the Association for Computational Linguistics), the world’s most prestigious academic conference in the field of natural language processing.
The published paper, titled “Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval,” presents an audio-text retrieval model that leverages multimodal large language models (LLMs) to effectively capture and respond to real-world user search intentions.
Existing audio-text retrieval models have been evaluated based on queries in the form of detailed descriptions (captions) of speech and sound; however, this approach differs significantly from how people actually search, leading to the limitation that it fails to properly measure practical robustness. In particular, with the widespread adoption of large language models, user queries are becoming increasingly complex, incorporating interrogative, imperative, and exclusionary conditions. To reflect these real-world search behaviors, the research team proposed a new audio search evaluation benchmark called User-Intent Queries (UIQ), consisting of five query types: “questions,” “commands,” “keywords,” “paraphrases,” and “negative exclusion queries.”
The Omni-Embed-Audio (OEA) model proposed by the research team encodes text and audio together into a single multimodal language model and aligns them in a shared embedding space. As a result, while matching the performance of the state-of-the-art model (M2D-CLAP) in traditional text-audio retrieval, it achieved a 22% relative performance improvement in text-text retrieval and demonstrated a clear advantage in distinguishing “hard negatives”—audio clips that are acoustically similar but semantically different (HNSR@10 +4.3%p, TFR@10 relative +34.7%). This demonstrates that large-scale language models can be used to develop encoders with superior semantic understanding capabilities for complex queries.
The paper received a “strong accept” rating—placing it in the top 15% of the 12,148 submissions from around the world—and was selected for an oral presentation. It is scheduled to be presented on July 6 at ACL 2026, to be held in San Diego, USA, in July 2026.
[SEO Keyword]
ACL 2026 oral presentation, audio-text retrieval model, user intent query benchmark
The published paper, titled “Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval,” presents an audio-text retrieval model that leverages multimodal large language models (LLMs) to effectively capture and respond to real-world user search intentions.
Existing audio-text retrieval models have been evaluated based on queries in the form of detailed descriptions (captions) of speech and sound; however, this approach differs significantly from how people actually search, leading to the limitation that it fails to properly measure practical robustness. In particular, with the widespread adoption of large language models, user queries are becoming increasingly complex, incorporating interrogative, imperative, and exclusionary conditions. To reflect these real-world search behaviors, the research team proposed a new audio search evaluation benchmark called User-Intent Queries (UIQ), consisting of five query types: “questions,” “commands,” “keywords,” “paraphrases,” and “negative exclusion queries.”
The Omni-Embed-Audio (OEA) model proposed by the research team encodes text and audio together into a single multimodal language model and aligns them in a shared embedding space. As a result, while matching the performance of the state-of-the-art model (M2D-CLAP) in traditional text-audio retrieval, it achieved a 22% relative performance improvement in text-text retrieval and demonstrated a clear advantage in distinguishing “hard negatives”—audio clips that are acoustically similar but semantically different (HNSR@10 +4.3%p, TFR@10 relative +34.7%). This demonstrates that large-scale language models can be used to develop encoders with superior semantic understanding capabilities for complex queries.
The paper received a “strong accept” rating—placing it in the top 15% of the 12,148 submissions from around the world—and was selected for an oral presentation. It is scheduled to be presented on July 6 at ACL 2026, to be held in San Diego, USA, in July 2026.
[SEO Keyword]
ACL 2026 oral presentation, audio-text retrieval model, user intent query benchmark