Hybrid multimodal GenAI for solving math problems containing various figures

초록

Mathematical problem solving with visual components remains a significant challenge for AI systems, as conventional OCR pipelines often fail to extract critical structural cues such as axes, tick marks, and spatial relationships in diagrams or graphs. While recent advances in generative models have improved multimodal reasoning, the integration of visual perception and symbolic inference is still underexplored. In this paper, we propose a lightweight hybrid pipeline that combines ColPali for vision-language understanding with open-source LLMs like LLaMA for symbolic reasoning. We evaluate our system on the MathVision dataset, leveraging its category-level statistics to measure performance across diverse visual math tasks. Our method improves accuracy by up to 29.3\% compared to standalone models, demonstrating the effectiveness of integrating visual structure into mathematical reasoning.

키워드

Generative AImultimodal reasoningvision language modelsmathematical problem solving
제목
Hybrid multimodal GenAI for solving math problems containing various figures
저자
Sangsoo Lee백재현김종락
발행일
2026-03
유형
Y
저널명
International Journal of Fuzzy Logic and Intelligent systems
26
1
페이지
1 ~ 9