Prior to that, I received my M.Eng. from Shanghai Jiao Tong University
(SJTU) in 2023, supervised by
Prof. Xinyi Le and
Prof. Yu Zheng,
and my B.Eng. from the same university in 2020.
My current research interest mainly lies in
generation models, multi-modal systems, vision-centric agent, and low-level vision.
Harnessing Image Diffusion Prior for Photo-Realistic Video Restoration
Fanghua Yu, Hongyu An, Jinfan Hu, Xinqi Lin, Zhiyuan You, Jian Wang, Hongyang Li, Chao Dong†, Jinjin Gu†
Neural Information Processing Systems (
NeurIPS), 2026
[Paper]
We propose HYPVR, which transfers powerful image generation priors to video restoration while synthesizing fine-grained details with temporal consistency.
LongSCOP: Semantically Consistent Long Video Outpainting
Yu Tang*, Zhiyuan You*, Zhixiao Wang, Fanghua Yu, Qingyu Zhang, Xiang Yin, Chao Dong†, Jinjin Gu†
Neural Information Processing Systems (
NeurIPS), 2026
[Paper]
We introduce LongSCOP, a framework for stable and semantically consistent long-video outpainting with long-horizon generation and VLM-based reference retrieval.
Evaluation of Visual Processing Should Be Human-Centered, Not Metric-Centered
Fanghua Yu*, Jinfan Hu*, Zhiyuan You, Xiang Yin, Hongyu An, Xinqi Lin, Chao Dong†, Jinjin Gu†
Neural Information Processing Systems (
NeurIPS), 2026, Position Paper Track
[Paper]
We argue that visual processing evaluation should move beyond metric-centered benchmarks toward human-centered, context-aware, and fine-grained assessment.
Generalizable Blind Image Quality Assessment with Retrieval Augmentation and Contrastive Vision-Text Local Meta-Training
Yan Zhong, Zhiyuan You, Xinping Zhao, Li Zhang, Xinyuan Song, Ruoyu Zhao, Lei Shi, Tingting Jiang
ACM Multimedia (
ACM MM), 2026
[Paper]
We propose Relocat-IQA, a retrieval-enhanced blind image quality assessment framework with contrastive vision-text learning and local meta-training for improved generalization.
PhotoAgent: Agentic Photo Editing with Exploratory Visual Aesthetic Planning
Mingde Yao, Zhiyuan You, King-Man Tam, Menglu Wang, and Tianfan Xue†
We propose PhotoAgent, an autonomous photo editing agent that performs multi-step aesthetic planning and refinement to transform images with minimal user input.
RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-Agents
Jize Wang, Han Wu, Zhiyuan You, Yiming Song, Yijun Wang, Zifei Shan, Yining Li, Songyang Zhang, Xinyi Le†, Cailian Chen†, Xinping Guan, Dacheng Tao
Association for Computational Linguistics (
ACL), 2026
[Paper]
[Code]
We propose RouteMoA, an efficient mixture-of-agents framework with dynamic routing. RouteMoA uses a lightweight scorer to predict coarse-grained performance, finding high-potential candidates without inference.
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
Xin Cai, Zhiyuan You, Zhoutong Zhang, Tianfan Xue†
We introduce Detail-Aligned VAE (DA-VAE), a method that increases the compression ratio of a pretrained VAE while requiring only lightweight adaptation for the pretrained diffusion backbone.
PhotoFramer: Multi-modal Image Composition Instruction
Zhiyuan You, Ke Wang, He Zhang, Xin Cai, Jinjin Gu, Tianfan Xue†, Chao Dong†, Zhoutong Zhang
We introduce PhotoFramer, a multi-modal composition instruction framework, to provide composition guidance. Given a poorly composed image, PhotoFramer first describes how to improve the composition in natural language and then generates a well-composed example image.
How Far Have We Gone in Generative Image Restoration? A Study on Its Capability, Limitations and Evaluation Practices
Xiang Yin, Jinfan Hu, Zhiyuan You, Kainan Yan, Yu Tang, Chao Dong, Jinjin Gu
We present a large-scale and multi-dimensional study on generative image restoration, covering diverse architectures, revealing critical performance disparities.
Revisiting the Generalization Problem of Low-level Vision Models Through the Lens of Image Deraining
Jinfan Hu*, Zhiyuan You*, Jinjin Gu, Kaiwen Zhu, Tianfan Xue, Chao Dong†
Pattern Recognition (
PR), 2026
[Paper]
We investigate the underlying mechanism of the generalization failure to unseen degradations in low-level vision, and show that the key to better generalization lies in guiding the network to learn the image content rather than the degradation.
RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
Xuming He*, Zhiyuan You*, Junchao Gong, Couhua Liu, Xiaoyu Yue, Peiqin Zhuang, Wenlong Zhang†, Lei Bai†
Neural Information Processing Systems (
NeurIPS), 2025
[Paper]
[Code]
[Data]
We introduce an MLLM-based weather forecast quality analysis method, RadarQA, integrating key physical attributes with detailed assessment reports.
ReinAD: Towards Real-world Industrial Anomaly Detection with a Comprehensive Contrastive Dataset
Xu Wang, Jingyuan Zhuo, Zhiyuan You, Zhiyu Tan, Yikuan Yu, Siyu Wang, Xinyi Le†
We introduce ReinAD dataset, a comprehensive, contrast-based, fine-grained, unaligned, and large-scale dataset towards real-world industrial anomaly detection.
Harnessing Diffusion-Yielded Score Priors for Image Restoration
Xinqi Lin, Fanghua Yu, Jinfan Hu, Zhiyuan You, Wu Shi, Jimmy S. Ren, Jinjin Gu†, Chao Dong†
We introduce HYPIR, a simple yet powerful paradigm: fine-tuning a pre-trained diffusion model with adversarial (GAN) loss — no diffusion sampling, no extra adapters, achieving an unprecedented balance of speed, fidelity, and quality.
DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
Junjie Gao, Runze Liu, Yingzhe Peng, Shujian Yang, Jin Zhang, Kai Yang†, Zhiyuan You†
International Conference on Computer Vision Workshop (
ICCVW), 2025
[Paper]
[Code]
Our DeQA-Doc wins the Championship in the VQualA 2025 DIQA (Document Image Quality Assessment) Challenge, ICCV 2025 Workshop.
Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution
Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue†, Chao Dong†
We introduce DeQA-Score, a distribution-based depicted image quality assessment model for score regression.
UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion
Zixuan Chen*, Yujin Wang*, Xin Cai, Zhiyuan You, Zheming Lu, Fan Zhang, Shi Guo, Tianfan Xue†
We propose UltraFusion, the first exposure fusion technique that can merge input with 9 stops differences.
Interpreting Low-level Vision Models with Causal Effect Maps
Jinfan Hu, Jinjin Gu, Shiyao Yu, Fanghua Yu, Zheyuan Li, Zhiyuan You, Chaochao Lu, Chao Dong†
IEEE Transactions on Pattern Analysis and Machine Intelligence (
TPAMI), 2025
[Paper]
[Code]
We propose Causal Effect Map (CEM), a model-agnostic and task-agnostic interpreting method for low-level vision models.
An Intelligent Agentic System for Complex Image Restoration Problems
Kaiwen Zhu*, Jinjin Gu*, Zhiyuan You, Yu Qiao, Chao Dong†
We introduce AgenticIR, an LLM-based agentic system that utilize various tools for complex image restoration problems.
SAIL: Sample-Centric In-Context Learning for Document Information Extraction
Jinyu Zhang*, Zhiyuan You*, Jize Wang, Xinyi Le†
Association for the Advancement of Artificial Intelligence (
AAAI), 2025,
Oral
[Paper]
[Code]
We introduce SAIL, a sample-centric approach selecting tailored in-context examples for training-free document information extraction.
Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset
Zhiyuan You, Jinjin Gu, Xin Cai, Zheyuan Li, Kaiwen Zhu, Chao Dong†, Tianfan Xue†
We introduce DepictQA-Wild, also named Enhanced DepictQA (EDQA), a multi-functional in-the-wild descriptive image quality assessment model.
PhoCoLens: Photorealistic and Consistent Reconstruction in Lensless Imaging
Xin Cai, Zhiyuan You, Hailong Zhang, Wentao Liu, Jinwei Gu, Tianfan Xue†
We introduce PhoCoLens, a novel two-stage approach for consistent and photorealistic lensless image reconstruction.
Depicting Beyond Scores: Advancing Image Quality Assessment through Multi-modal Language Models
Zhiyuan You*, Zheyuan Li*, Jinjin Gu*, Zhenfei Yin, Tianfan Xue†, Chao Dong†
We introduce DepictQA, leveraging multi-modal large language models, allowing for detailed, language-based, and human-like evaluation of image quality.
MaskMA: Towards Zero-Shot Multi-Agent Decision Making with Mask-Based Collaborative Learning
Jie Liu*, Yinmin Zhang*, Chuming Li, Zhiyuan You, Zhanhui Zhou, Chao Yang, Yaodong Yang†, Yu Liu, Wanli Ouyang
Transactions on Machine Learning Research (
TMLR), 2024
[Paper]
We release MaskMA, a masked pretraining framework for multi-agent decision-making.
Few-shot Object Counting with Similarity-Aware Feature Enhancement
Zhiyuan You, Kai Yang, Wenhan Luo, Xin Lu, Lei Cui, Xinyi Le†
Winter Conference on Applications of Computer Vision (
WACV), 2023
[Paper]
[Code]
[Video]
We propose a novel SAFECount block, equipped with a similarity comparison module and a feature enhancement module for few-shot object counting.
A Unified Model for Multi-class Anomaly Detection
Zhiyuan You, Lei Cui, Yujun Shen, Kai Yang, Xin Lu, Yu Zheng, Xinyi Le†
Neural Information Processing Systems (
NeurIPS), 2022,
Spotlight
[Paper]
[Code]
We present UniAD that accomplishes anomaly detection for multiple classes with a unified framework.
ADTR: Anomaly Detection Transformer with Feature Reconstruction
Zhiyuan You, Kai Yang, Wenhan Luo, Lei Cui, Yu Zheng, Xinyi Le†
International Conference on Neural Information Processing (
ICONIP), 2022,
Oral
[Paper]
We propose ADTR, a transformer that reconstructs pre-trained features for anomaly detection and extends to image- and pixel-level labels.