Zijian Chen 「陈子健」

I am currently a third-year PhD Student of Multimedia Lab at Shanghai Jiao Tong University (SJTU), advised by Prof. Guangtao Zhai and Prof. Wenjun Zhang. Previously, I received my M.E. degree from East China University of Science and Technology in June 2023, and B.S. degree in EE from Wenzhou University in June 2020.

I'm generally interested in Visual Quality Assessment, Agentic AI, especially All-round Evaluation, and Digital Humanities (Oracle bone character processing). My ultimate goal is to explore the limits of AI's capabilities.

I'm always eager to communicate and cooperate, so feel free to contact me!!!

Email: zijian.chen@sjtu.edu.cn          

Email  /  Google Scholar  /  Github  /  Zhihu  /  Team Web

profile photo
News

Selected Publications

For the full publication list, please refer to my Google Scholar.

* denotes equal contribution and † denotes corresponding author (noted per paper).

Showing 20 of 20 publications.

FigCodeBench

From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

Zijian Chen, Zhengyu Chen, Bohan Liang, Lirong Deng, Yushuo Zheng, Yanwei Jiang, Qi Jia, Kaiwei Zhang, Wenjun Zhang, Guangtao Zhai
arxiv, 2026. (NEW)

We propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation.

Squid Game

Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models

Zijian Chen, Wenjun Zhang, Guangtao Zhai†
arXiv, 2025. (NEW)

In this paper, we introduce SQUID GAME, a dynamic and adversarial evaluation environment with resource-constrained and asymmetric information settings elaborated to evaluate LLMs through interactive gameplay against other LLM opponents.

MACEval

MACEval: A Multi-Agent Continual Evaluation Network for Large Models

Zijian Chen, Yuze Sun, Yuan Tian, Wenjun Zhang, Guangtao Zhai†
arXiv, 2025. (NEW)

In this paper, we introduce MACEval, a Multi-Agent Continual Evaluation network for dynamic evaluation of large models, and define a AUC-inspired metric to quantify performance longitudinally and sustainably.

CMSBench

Benchmarking Cross-Scale Perception Ability of Large Multimodal Models in Material Science

Yuting Zheng, Zijian Chen†, Qi Jia
ICME, 2026. (Oral presentation)

In this paper, we introduce CSMBench, a dataset comprising 1,041 highquality figures curated from premier journals up to September 2025. CSMBench categorizes data into four scientifically distinct regimes: atomic, micro, meso, and macro scales, strictly aligning with the focus and definitions in materials study.

BioMotion Arena

Can Large Models Fool the Eye? A New Turing Test for Biological Animation

arXiv, 2025. (NEW)

In this paper, we introduce BioMotion Arena, the first biological motion-based visual preference evaluation framework for large models. We focus on ten typical human motions and introduce fine-grained control over gender, weight, mood, and direction. More than 45k votes for 53 mainstream LLMs and MLLMs on 90 biological motion variants are collected.

LMM-JND

Just Noticeable Difference for Large Multimodal Models

arXiv, 2025. (NEW)

In this paper, we propose a novel concept, LMM-JND, to quantify the perceptual redundancy characteristic for LMMs and a well-designed pipeline for its determination. We also construct a large-scale dataset, named VPA-JND, which contains 21.5k reference images with over 489k stimuli across 12 distortion types, to facilitate LMM-JND studies.

Puzzlebench

PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving

arXiv, 2025.

In this paper, we construct PuzzleBench, a dynamic and scalable benchmark comprising 11,840 VQA samples, which features six carefully designed puzzle tasks targeting three core LMM competencies, visual recognition, logical reasoning, and context understanding.

debanding

Joint Luminance-Chrominance Learning for Image Debanding

Zijian Chen, Wei Sun, Jun Jia, Ru Huang, Fangfang Lu, Ying Chen, Xiongkuo Min, Guangtao Zhai†, Wenjun Zhang
IEEE Transactions on Circuits and Systems for Video Technology, 2025.

In this paper, we propose a unified deep neural network that explicitly disentangles the luminance and chrominance channels, and simultaneously recovers intensity gradients and color discontinuity from detection-free measurement in an end-to-end manner.

GAIA

GAIA: Rethinking Action Quality Assessment for AI-Generated Videos

NeurIPS, 2024. (Spotlight Presentation)

In this work, we construct GAIA, a Generic AI-generated Action dataset, by conducting a large-scale subjective evaluation from a novel causal reasoning-based perspective, resulting in 971,244 ratings among 9,180 video-action pairs, and evaluate a suite of popular text-to-video models on their ability to generate visually rational actions.

AGIN

Study of Subjective and Objective Naturalness Assessment of AI-Generated Images

IEEE Transactions on Circuits and Systems for Video Technology, 2025.

In this work, we construct the AI-Generated Image Naturalness (AGIN) dataset and propose the Joint Objective Image Naturalness evaluaTor (JOINT) to automatically assess the naturalness of AIGIs that align with human opinions.

Band2k

BAND-2k: Banding Artifact Noticeable Database for Banding Detection and Quality Assessment

Zijian Chen, Wei Sun, Jun Jia, Fangfang Lu, Zicheng Zhang, Jing Liu, Ru Huang, Xiongkuo Min†, Guangtao Zhai†
IEEE Transactions on Circuits and Systems for Video Technology, 2024.

In this work, we build the Banding Artifact Noticeable Database (BAND-2k), which consists of 2,000 banding images generated by 15 compression and quantization schemes.

fsband

FS-BAND: A frequency-sensitive banding detector

Zijian Chen, Wei Sun, Zicheng Zhang, Ru Huang, Fangfang Lu, Xiongkuo Min, Guangtao Zhai†, Wenjun Zhang
IEEE International Symposium on Circuits and Systems (ISCAS), 2024.

In this paper, we develop a no-reference banding evaluator for banding detection and quality assessment by leveraging its frequency characteristics.

ROOTS

ROOTS: Recognizing Oracle Bone Inscriptions via an Organized Tree Structure

Ziqi Li*, Zijian Chen*, Ziyi Yang, Zhiji Liu, Guangtao Zhai, Tingzhu Chen†
npj heritage science, 2026. (NEW)

We propose ROOTS (Recognizing Oracle Bone Inscriptions via an Organized Tree Structure), a hierarchical framework that restructures the original fine-grained classes into size-balanced superclasses and proceeds coarse-to-fine recognition, where a shared root branches into semantic superclasses before resolving individual characters, with a winner-take-all mask confining predictions to the relevant branch.

S-OBI

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

Ziqi Li*, Zijian Chen*†, Tingzhu Chen†, Guangtao Zhai
ICIG, 2026. (NEW)

We introduce S-OBI, a novel benchmark for evaluating MLLMs in Sentence-level OBI understanding. S-OBI synthesizes standardized sentence-level OBI instances through glyph substitution and composition, consisting of semantic matching, semantic slot extraction, and contextual reasoning tasks.

OBI Survey

Oracle Bone Inscriptions Information Processing: A Comprehensive Survey

Zijian Chen, Wenjie Hua, Jinhao Li, Yucheng Zhu, Xiaona Zhi, Zhiji Liu, Tingzhu Chen†,Wenjun Zhang, Guangtao Zhai†
npj Heritage Science, 2026. (NEW)

We conduct a comprehensive survey into Oracle bone inscriptions information processing works over the past 20 years, reviewing more than 150 related articles.

PictOBI-20k

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters

Zijian Chen*, Wenjie Hua*, Jinhao Li, Lirong Deng, Fan Du, Tingzhu Chen†, Guangtao Zhai†
ICASSP, 2026.

In this paper, we introduce PictOBI-20k, a dataset designed to evaluate LMMs on the visual decipherment tasks of pictographic OBCs. It includes 20k meticulously collected OBC and real object images, forming over 15k multi-choice questions. We also conduct subjective annotations to investigate the consistency of the reference point between humans and LMMs in visual reasoning.

obi-bench

OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?

ICLR, 2025. (NEW)

In this work, we introduce OBI-Bench, a holistic benchmark crafted to systematically evaluate large multi-modal models (LMMs) on whole-process oracle bone inscriptions (OBI) processing tasks demanding expert-level domain knowledge and deliberate cognition.

Oracle-P15k

Mitigating Long-tail Distribution in Oracle Bone Inscriptions: Dataset, Model, and Benchmark

ACM MM, 2025. (Oral Recommendation)

In this paper, we present the Oracle-P15K, a structure-aligned OBI dataset for OBI generation and denoising, consisting of 14,542 images infused with domain knowledge from OBI experts. Based on this, we propose a diffusion model-based pseudo OBI generator, called OBIDiff, to achieve realistic and controllable OBI generation.

OBIFormer

OBIFormer: A fast attentive denoising framework for oracle bone inscriptions

Displays, 2025.

In this work, we propose OBIFormer, a fast attentive framework for high-precision Oracle bone inscriptions denoising.

LiveProteinBench

LiveProteinBench: A Contamination-Free Benchmark for Assessing Models' Specialized Capabilities in Protein Science

Dingyi Rong*, Zijian Chen*, Qi Jia*, Kaiwei Zhang, Haotian Liu, Guangtao Zhai†, Ning Liu†
arXiv, 2025.

In this paper, we introduce LiveProteinBench, a contamination-free, multimodal benchmark of 12 tasks for evaluating LLM performance on protein property and function prediction.

Reviewer Service
  • International Conference on Machine Learning (ICML 2026)
  • IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2026)
  • Annual Conference on Neural Information Processing Systems (NeurIPS 2025, 2026)
  • International Conference on Learning Representations (ICLR 25-27)
  • ACM Multimedia (ACM MM 2025, 2026)
  • Annual AAAI Conference on Artificial Intelligence (AAAI 26-27)
  • European Conference on Computer Vision (ECCV 2026)
  • IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI 2026)
  • IEEE Transactions on Multimedia (TMM 2025)
  • IEEE Transactions on Circuits and Systems for Video Technology (TCSVT 2026)
  • Pattern Recognition (2026)
  • ACM Transactions on Multimedia Computing Communications and Applications (ACM TOMM 2025)
  • IEEE Intelligent Transportation Systems Magazine (ITSM 2024)
  • IEEE Signal Processing Letters (2026)
  • Signal, Image and Video Processing (2026)
  • Journal of Supercomputing 2025
  • Scientific Reports 2026
  • Displays 2025
Talks
  • [2025.12] Large Multimodal Model-Driven Oracle Bone Inscriptions Information Processing (The 1st CCF AI for Humanities Conference)
  • [2024.11] AI+Virtual Simulation: Empowering Display Device Development (The 16th China Display Academic Conference)
Blogs
Awards
  • [2025] Doctoral Student Program of the Young S&T Talents Cultivation Project, CAST (中国科协青年科技人才培育工程博士生专项计划)
  • [2025] National Scholarship (for PhD students)
  • [2023] Excellent Graduates in Shanghai (for postgraduates)
  • [2021] National Second Prize (National Graduate Electronics Design Contest)
  • [2019] First Prize in Zhejiang Province (National Undergraduate Electronics Design Contest)

Updated in Apr. 2026

Thanks Jon Barron for this amazing website template.