lupantech/MathVista

MathVista: data, code, and evaluation for Mathematical Reasoning in Visual Contexts

What it solves

MathVista is a benchmark designed to evaluate the mathematical reasoning capabilities of Large Multimodal Models (LMMs) and Large Language Models (LLMs) when presented with visual contexts. It addresses the gap in systematic study of how foundation models handle tasks that require both deep visual understanding and compositional mathematical reasoning.

How it works

The benchmark consists of 6,141 examples curated from 28 existing multimodal datasets and 3 newly created datasets (IQTest, FunctionQA, and PaperQA). It provides a standardized way to test models on a variety of mathematical tasks, including geometry, algebra, and statistics, using a testmini subset of 1,000 examples for rapid evaluation and a full test set for comprehensive assessment.

Who it’s for

This project is for AI researchers and developers building Large Multimodal Models who need a rigorous way to measure their models' ability to solve visually-rich mathematical problems.

Highlights

  • Comprehensive Dataset: Combines tasks from 31 different sources to create a diverse evaluation suite.
  • Human Baseline: Provides a human performance baseline (60.3% accuracy) to measure model progress.
  • Detailed Analysis: Includes tools for dataset exploration, visualization, and a leaderboard to track state-of-the-art (SOTA) performance.
  • Broad Model Support: Evaluates a wide range of models including OpenAI o1, GPT-4o, Claude 3.5, and Gemini.

Related

  • Project
  • Project
  • Project
  • Project
  • Project