상세 보기
Hybrid multimodal GenAI for solving math problems containing various figures
- Sangsoo Lee;
- 백재현;
- 김종락
초록
Mathematical problem solving with visual components remains a significant challenge for AI systems, as conventional OCR pipelines often fail to extract critical structural cues such as axes, tick marks, and spatial relationships in diagrams or graphs. While recent advances in generative models have improved multimodal reasoning, the integration of visual perception and symbolic inference is still underexplored. In this paper, we propose a lightweight hybrid pipeline that combines ColPali for vision-language understanding with open-source LLMs like LLaMA for symbolic reasoning. We evaluate our system on the MathVision dataset, leveraging its category-level statistics to measure performance across diverse visual math tasks. Our method improves accuracy by up to 29.3\% compared to standalone models, demonstrating the effectiveness of integrating visual structure into mathematical reasoning.
키워드
- 제목
- Hybrid multimodal GenAI for solving math problems containing various figures
- 저자
- Sangsoo Lee; 백재현; 김종락
- 발행일
- 2026-03
- 유형
- Y
- 권
- 26
- 호
- 1
- 페이지
- 1 ~ 9