A benchmark for figure-to-code generation

From Pixel to Coding:
Evaluating the Figure Reproduction Capabilities of MLLMs

Zijian Chen1,2, Zhengyu Chen1, Bohan Liang2, Lirong Deng3, Yushuo Zheng1,2, Yanwei Jiang1,2, Qi Jia2, Kaiwei Zhang2, Wenjun Zhang1, Guangtao Zhai1,2,*

1 Shanghai Jiao Tong University2 Shanghai AI Laboratory3 Macao Polytechnic University

* Corresponding author

Evaluating MLLMs on figure reproduction,
integrating multimodal comprehension and generation.

FigCodeBench overview of scientific figure reproduction and the evaluation pipeline
From a scientific figure to reproducing code: the FigCodeBench task and evaluation overview.

Perception meets generation

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages (Python, Matlab, R, and Latex). We further categorize figure reproduction into three tiers (Easy, Medium, Hard) with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity (e.g., PSNR, SSIM, LPIPS) and syntactic isomorphism (e.g., abstract syntax tree similarity), that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages. Our benchmark offers (1) a comprehensive dataset and multi-dimensional evaluation pipeline; (2) an up-to-date leaderboard on MLLM figure reproduction proficiency; and (3) a nuanced understanding of the perceiving and reasoning bottlenecks hindering unified comprehension and generation in current MLLMs.

Beyond describing an image

Why figure reproduction?

Researchers often have a scientific figure and need working code that reproduces it. For a model, this means turning visual structure into a rendered program, not merely describing the image or answering questions about it. We argue that text-only coding tests and image-understanding tests leave this perception-to-generation step unmeasured.

FigCodeBench assesses the ability to observe complex charts—such as line graphs, bar charts, scatter plots, and architecture diagrams—and generate corresponding reproducing code. This tests visual perception and fine-grained understanding, while challenging logical reasoning and alignment in code generation.

Core contributions

6,194

Image-code pairs

A comprehensive collection covering four languages: Python, MATLAB, R, and LaTeX.

Multi-dimensional

Automatic evaluation

Image-oriented and code-oriented metrics, Mean Machine Opinion Score, and Figure-Code Fidelity.

24

Evaluated MLLMs

Experiments examine capability boundaries, adaptability across languages, and common error patterns.

The figure-code collection

Dataset

6,194 image-code pairs across four programming languages.
Explore the collection on Hugging Face.

Difficulty distribution

Dataset difficulty distribution across the language collections
Difficulty distribution of the FigCodeBench collections.

Language collections

DirectoryLanguageSource code extension
python/Python.py
Matlab/MATLAB.m
R/R.R
latex/LaTeX.tex

Figure images are stored in PNG format. Auxiliary information is provided in text files whose names end with _auxinfo.txt.

Subsets and versions

Each language directory contains exemplary/ and user_generated/ subsets. Both follow the same six-directory structure:

<language>/
├── exemplary/
│   ├── code_base/
│   ├── code_variant1/
│   ├── code_variant2/
│   ├── image_base/
│   ├── image_variant1/
│   └── image_variant2/
└── user_generated/
    ├── code_base/
    ├── code_variant1/
    ├── code_variant2/
    ├── image_base/
    ├── image_variant1/
    └── image_variant2/

Base and variant collections

code_base/ contains source code and auxiliary information for the base collection; image_base/ contains its figure images.

code_variant1/ and code_variant2/ contain source code and auxiliary information for the first and second variants. Their figure images are in image_variant1/ and image_variant2/.

Variant collections contain modified figure-code examples. Their sizes may differ from the base collection: do not assume every base example has both variants.

Visualization categories

The Python, MATLAB, and R collections organize files into six categories within each code and image directory:

Composition
Composition-based visualizations
Geospatial
Geographic plots
Mathematical
Mathematical graphics
Relational
Relationship-based charts
Statistical
Statistical graphics
Temporal
Time-oriented visualizations

The LaTeX collection (Conceptual) uses a flatter structure: source code, auxiliary information, and images are stored directly in their respective version directories, without the six category subdirectories.

Multi-dimensional evaluation

Metrics

FigCodeBench combines image-oriented and code-oriented evaluation with machine opinion scoring and figure-code fidelity.

03

Mean Machine Opinion Score

MMOS results, detailed score distributions, and qualitative examples.

View MMOS results

The evaluated model landscape

Evaluated models

24 mainstream MLLMs: 11 proprietary and 13 open-source models.

Overview of the 24 evaluated MLLMs, including proprietary and open-source models
The models selected for evaluation in FigCodeBench.

Figure reproduction proficiency

FigCodeBench leaderboard

Performance, evaluation cost, and results across image categories.

Leaderboard showing MMOS versus average cost per problem and six representative models across image categories
Left: Mean Machine Opinion Score (MMOS) versus average cost per problem for various models. Right: Performance comparison of six representative models on different image categories.

Expand a result below. Select any research figure to view it at full size.

Evaluation cost
Evaluation cost comparison for models in FigCodeBench
Evaluation cost.
Results on image-oriented metrics
Model results on image-oriented metrics
Image-oriented metric results.
Results on code-oriented metrics
Model results on code-oriented metrics
Code-oriented metric results.
Results on Mean Machine Opinion Score (MMOS)
Mean Machine Opinion Score results using MLLM judges
MMOS evaluation with MLLM judges.

Detailed MMOS distribution

Fine-grained distribution of Mean Machine Opinion Scores
Detailed MMOS distribution.

Qualitative results

Qualitative figure reproduction examples for MMOS evaluation
Qualitative MMOS results.
Results on Figure-Code Fidelity (FCF)
Results on the Figure-Code Fidelity metric
Figure-Code Fidelity (FCF) results.

Reproducing the figures

Environments

Reference runtimes and rendering settings for the four language collections.

CollectionRuntime / toolchainMain libraries or settings
PythonPythonPrimarily matplotlib and seaborn; the execution workflow automatically detects and installs missing packages.
MATLABMATLAB R2023aAdditional toolboxes may be required by individual scripts.
RR 4.6.0The execution workflow automatically detects and installs missing R packages.
LaTeXA TeX distribution providing XeLaTeXxelatex is the reference rendering engine; 300 DPI for rasterized output.

LaTeX examples require the packages and fonts they use. pdflatex may work for compatible examples, but is not the reference engine; XeLaTeX-specific font or Unicode features may not compile with it.

Consult the repository environment documentation for setup context and check the files available in your checkout before running a generation workflow.

Contact

Please contact the first author of this paper for queries.

Zijian Chen
zijian.chen@sjtu.edu.cn

Citation

Please feel free to cite our paper.

BibTeX
@misc{chen2026pixelcodingevaluatingfigure,
      title={From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs}, 
      author={Zijian Chen and Zhengyu Chen and Bohan Liang and Lirong Deng and Yushuo Zheng and Yanwei Jiang and Qi Jia and Kaiwei Zhang and Wenjun Zhang and Guangtao Zhai},
      year={2026},
      eprint={2610.10066},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.10066}, 
}

Research figure

Scroll to inspect the full-size image. Press Escape or select Close to return.

Open original image in a new tab