Hybrid Multimodal Generative Artificial Intelligence Approach for Mathematical Problem-Solving with Figures

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Mathematical problem-solving with visual components remains a significant challenge for artificial intelligence (AI) systems. This is because conventional OCR pipelines often fail to extract critical structural cues such as axes, tick marks, and spatial relationships in diagrams or graphs. Despite recent advances in generative models, a cohesive framework that integrates visual perception with symbolic inference is still lacking. Therefore, this study proposes a lightweight hybrid pipeline that combines ColPali for vision-language understanding with open-source large language models (LLMs) such as LLaMA for symbolic reasoning. We evaluated our system on the MathVision dataset, leveraging its category-level statistics to measure the performance across diverse visual math tasks. Our method improves accuracy by up to 29.3% compared to standalone models, demonstrating the effectiveness of integrating visual structures into mathematical reasoning. © The Korean Institute of Intelligent Systems. This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/3.0/) which permits unrestricted non commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

키워드

Generative AIMathematical problem solvingMultimodal reasoningVision language models
제목
Hybrid Multimodal Generative Artificial Intelligence Approach for Mathematical Problem-Solving with Figures
저자
Lee, SangsooBaek, Jae-HyunKim, Jon-Lark
DOI
10.5391/IJFIS.2026.26.1.1
발행일
2026-03
유형
Article
저널명
International Journal of Fuzzy Logic and Intelligent systems
26
1
페이지
1 ~ 9