← Search

Zhanyu Ma

40 accepted papers

2026

Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation

CVPR 2026

Text-to-Image (T2I) generation has achieved remarkable progress in recent years. Meanwhile, reinforcement learning methods, particularly those based on Group Relative Policy Optimization (GRPO), have attracted widespread attention and been successfully applied to T2I tasks. However, the uniform samp

Cited by 0SourcecodeScholar
2026

EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior Understanding

CVPR 2026

Vision-language models (VLMs) have achieved remarkable success across numerous domains, yet they lag significantly in animal behavior understanding due to severe data scarcity. Annotated animal behavior videos are prohibitively expensive and time-consuming to collect, requiring domain expertise and

Cited by 0SourcecodeScholar
2026

Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers

ICLR 2026poster

Recent advances in diffusion models have significantly improved image editing. However, challenges persist in handling geometric transformations, such as translation, rotation, and scaling, particularly in complex scenes. Existing approaches suffer from two main limitations: (1) difficulty in achiev…

Cited by 0SourceScholar
2026

IncreFA: Breaking the Static Wall of Generative Model Attribution

CVPR 2026

As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive generators appear almost monthly, making existing watermark, classifier and inversion methods obsolete upon release. The core problem lies not in model r

Cited by 0SourcecodeScholar
2026

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

ICML 2026poster

Evaluation benchmarks play a central role in assessing vision–language models (VLMs). However, most existing multimodal benchmarks are static, making them increasingly vulnerable to data contamination, temporal staleness, and high construction costs. In this work, we introduce MMBench-Live, a multi-…

Cited by 0SourceScholar
2026

MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision

AAAI 2026technical

Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding pipelines still rely on supervised fine-tuning with explicit

Cited by 0SourcePDFScholar
2026

Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion Guidance

AAAI 2026technical

Clean images are crucial for visual tasks such as small object detection, especially at high resolutions. However, real-world images are often degraded by adverse weather, and weather restoration methods may sacrifice high-frequency details critical for analyzing small objects. A natural solution is

Cited by 0SourcePDFScholar
2026

Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding

CVPR 2026

Fine-grained visual understanding is shifting from static classification to knowledge-augmented reasoning, where models must justify as well as recognise. Existing approaches remain limited by closed-set taxonomies and single-label prediction, leading to significant degradation under open-set or con

Cited by 0SourcecodeScholar
2026

Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding

CVPR 2026

The task of visual grounding (i.e., VG) aims to locate or segment objects in images based on referring expressions. Existing research on VG primarily focuses on large objects. However, these images often contain objects at various scales. Although large objects are usually the visual focus, small ob

Cited by 0SourcecodeScholar
2025

CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation

NeurIPS 2025poster

Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of…

Cited by 0SourcecodeScholar
2025

ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer

CVPR 2025poster

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to…

2025

FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models

ICCV 2025poster

Image generation has achieved remarkable progress with the development of large-scale text-to-image models, especially diffusion-based models. However, generating human images with plausible details, such as faces or hands, remains challenging due to insufficient supervision of local regions during…

2025

PIDSR: Complementary Polarized Image Demosaicing and Super-Resolution

CVPR 2025poster

Polarization cameras can capture multiple polarized images with different polarizer angles in a single shot, bringing convenience to polarization-based downstream tasks. However, their direct outputs are color-polarization filter array (CPFA) raw images, requiring demosaicing to reconstruct full-res…

2025

PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface Reconstruction

CVPR 2025poster

Reflective and textureless surfaces remain a challenge in multi-view 3D reconstruction. Both camera pose calibration and shape reconstruction often fail due to insufficient or unreliable cross-view visual features. To address these issues, we present PMNI (Pose-free Multi-view Normal Integration), a…

2025

PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface Reconstruction

ICCV 2025poster

Efficient shape reconstruction for surfaces with complex reflectance properties is crucial for real-time virtual reality. While 3D Gaussian Splatting (3DGS)-based methods offer fast novel view rendering by leveraging their explicit surface representation, their reconstruction quality lags behind tha…

Cited by 0SourcePDFScholar
2025

PolarAnything: Diffusion-based Polarimetric Image Synthesis

ICCV 2025poster

Polarization images facilitate image enhancement and 3D reconstruction tasks, but the limited accessibility of polarization cameras hinders their broader application. This gap drives the need for synthesizing photorealistic polarization images. The existing polarization simulator Mitsuba relies on a…

Cited by 0SourcePDFScholar
2025

Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face Restoration

NeurIPS 2025poster

Old-photo face restoration poses significant challenges due to compounded degradations such as breakage, fading, and severe blur. Existing pre-trained diffusion-guided methods either rely on explicit degradation priors or global statistical guidance, which struggle with localized artifacts or face c…

Cited by 0SourcecodeScholar
2025

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

ICCV 2025poster

Recent advances in diffusion models have significantly advanced image generation; however, existing models remain task-specific, limiting their efficiency and generalizability. While universal models attempt to address these limitations, they face critical challenges, including generalizable instruc…

2024

Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding

NeurIPS 2024poster

With the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural wo…

2024

Benchmarking Segmentation Models with Mask-Preserved Attribute Editing

CVPR 2024poster

When deploying segmentation models in practice it is critical to evaluate their behaviors in varied and complex scenes. Different from the previous evaluation paradigms only in consideration of global attribute variations (e.g. adverse weather) we investigate both local and global attribute variatio…

2024

DemoFusion: Democratising High-Resolution Image Generation With No $$$

CVPR 2024poster

High-resolution image generation with Generative Artificial Intelligence (GenAI) has immense potential but due to the enormous capital investment required for training it is increasingly centralised to a few large corporations and hidden behind paywalls. This paper aims to democratise high-resolutio…

2024

Dual-Prior Augmented Decoding Network for Long Tail Distribution in HOI Detection

AAAI 2024technical

Human object interaction detection aims at localizing human-object pairs and recognizing their interactions. Trapped by the long-tailed distribution of the data, existing HOI detection methods often have difficulty recognizing the tail categories. Many approaches try to improve the recognition of HO…

2024

Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT

NeurIPS 2024poster

Lumina-T2X is a nascent family of Flow-based Large Diffusion Transformers (Flag-DiT) that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters chall…

2024

NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized Images

CVPR 2024poster

We present NeRSP a Neural 3D reconstruction technique for Reflective surfaces with Sparse Polarized images. Reflective surface reconstruction is extremely challenging as specular reflections are view-dependent and thus violate the multiview consistency for multiview stereo. On the other hand sparse…

Cited by 7SourcePDFScholar
2023

An Erudite Fine-Grained Visual Classification Model

CVPR 2023poster

Current fine-grained visual classification (FGVC) models are isolated. In practice, we first need to identify the coarse-grained label of an object, then select the corresponding FGVC model for recognition. This hinders the application of the FGVC algorithm in real-life scenarios. In this paper, we…

2023

Bi-directional Feature Reconstruction Network for Fine-Grained Few-Shot Image Classification

AAAI 2023technical

The main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting --…

2023

Multi-View Active Fine-Grained Visual Recognition

ICCV 2023poster

Despite the remarkable progress of Fine-grained visual classification (FGVC) with years of history, it is still limited to recognizing 2 images. Recognizing objects in the physical world (i.e., 3D environment) poses a unique challenge -- discriminative information is not only present in visible loca…

Cited by 9PDFcodeScholar
2023

On-the-Fly Category Discovery

CVPR 2023poster

Although machines have surpassed humans on visual recognition problems, they are still limited to providing closed-set answers. Unlike machines, humans can cognize novel categories at the first observation. Novel category discovery (NCD) techniques, transferring knowledge from seen categories to dis…

2023

Semantic Memory Guided Image Representation for Polyp Segmentation

ICASSP 2023accepted

Polyp segmentation is important in the early diagnosis and treatment of colorectal cancer. Since polyps vary in shape, size, color, and texture, accurate polyp segmentation is very challenging. One promising solution is to model the contextual relation for each pixel. However, previous methods only…

Cited by 0SourceScholar
2023

Super-Resolution Information Enhancement for Crowd Counting

ICASSP 2023accepted

Crowd counting is a challenging task due to the heavy occlusions, scales, and density variations. Existing methods handle these challenges effectively while ignoring low-resolution (LR) circumstances. The LR circumstances weaken the counting performance deeply for two crucial reasons: 1) limited det…

Cited by 0SourceScholar
2023

Task-aware Adaptive Learning for Cross-domain Few-shot Learning

ICCV 2023poster

Although existing few-shot learning works yield promising results for in-domain queries, they still suffer from weak cross-domain generalization. Limited support data requires effective knowledge transfer, but domain-shift makes this harder. Towards this emerging challenge, researchers improved adap…

Cited by 15PDFcodeScholar
2022

HCLD: A Hierarchical Framework for Zero-shot Cross-lingual Dialogue System

COLING 2022main

Recently, many task-oriented dialogue systems need to serve users in different languages. However, it is time-consuming to collect enough data of each language for training. Thus, zero-shot adaptation of cross-lingual task-oriented dialog systems has been studied. Most of existing methods consider t…

2022

Learning Invariant Visual Representations for Compositional Zero-Shot Learning

ECCV 2022poster

"Compositional Zero-Shot Learning (CZSL) aims to recognize novel compositions using knowledge learned from seen attribute-object compositions in the training set. Previous works mainly project an image and a composition into a common embedding space to measure their compatibility score. However, bot…

2021

Dual Graph Convolutional Networks for Aspect-based Sentiment Analysis

ACL 2021long

Aspect-based sentiment analysis is a fine-grained sentiment classification task. Recently, graph neural networks over dependency trees have been explored to explicitly model connections between aspects and opinion words. However, the improvement is limited due to the inaccuracy of the dependency par…

2021

Your "Flamingo" is My "Bird": Fine-Grained, or Not

CVPR 2021poster

Whether what you see in Figure 1 is a "flamingo" or a "bird", is the question we ask in this paper. While fine-grained visual classification (FGVC) strives to arrive at the former, for the majority of us non-experts just "bird" would probably suffice. The real question is therefore -- how can we tai…

Cited by 145PDFcodeScholar
2020

Fine-Grained Visual Classification via Progressive Multi-Granularity Training of Jigsaw Patches

ECCV 2020poster

Fine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works mainly tackle this problem by focusing on how to locate the most discriminative parts, more complementary parts, and parts o…

2020

GINet: Graph Interaction Network for Scene Parsing

ECCV 2020poster

Recently, context reasoning using image regions beyond local convolution has shown great potential for scene parsing. In this work, we explore how to incorperate the linguistic knowledge to promote context reasoning over image regions by proposing a Graph Interaction unit (GI unit) and a Semantic Co…

2018

SketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval

CVPR 2018poster

We propose a deep hashing framework for sketch retrieval that, for the first time, works on a multi-million scale human sketch dataset.Leveraging on this large dataset, we explore a few sketch-specific traits that were otherwise under-studied in prior literature. Instead of following the conventiona…

Cited by 149SourcePDFScholar