← Search

Zeyu Wang

63 accepted papers

2026

Breaking Task Boundaries: A Unified Model for 3D Medical Image Fusion and Segmentation Guided by Manifold Perspective

AAAI 2026technical

3D medical image fusion (MIF) and segmentation (MIS) are critical and inherently synergistic tasks in medical image analysis. However, fundamentally integrating them remains highly challenging, since effective collaborative paradigms are still scarce and their optimization objectives fundamentally d

Cited by 0SourcePDFScholar
2026

Breaking the Passive Learning Trap: An Active Perception Strategy for Human Motion Prediction

AAAI 2026technical

Forecasting 3D human motion is an important embodiment of fine-grained understanding and cognition of human behavior by artificial agents. Current approaches excessively rely on implicit network modeling of spatiotemporal relationships and motion characteristics, falling into the passive learning tr

Cited by 0SourcePDFScholar
2026

Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions

CVPR 2026

Large Vision-Language Models (VLMs) often answer classic visual illusions "correctly" on original images, yet persist with the same responses when illusion factors are inverted, even though the visual change is obvious to humans. This raises a fundamental question: do VLMs perceive visual changes or

Cited by 0SourceScholar
2026

EasyCreator: Empowering 4D Creation through Video Inpainting

ICLR 2026poster

We introduce EasyCreator, a novel 4D video creation framework capable of both generating and editing 4D content from a single monocular video input. By leveraging a powerful video inpainting foundation model as a generative prior, we reformulate 4D video creation as a video inpainting task, enabling…

Cited by 0SourceScholar
2026

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

AAAI 2026technical

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to brid

Cited by 0SourcePDFScholar
2026

EvDiff3D: Event-Aware Diffusion Repair for High-Fidelity Event-Based 3D Reconstruction

AAAI 2026technical

Event cameras are bio-inspired sensors that capture visual information through asynchronous brightness changes, offering distinct advantages including high temporal resolution and wide dynamic range. While prior research has investigated event-based 3D reconstruction for extreme scenarios, existing

Cited by 0SourcePDFScholar
2026

GardenDesigner: Encoding Aesthetic Principles into Jiangnan Garden Construction via a Chain of Agents

CVPR 2026

Jiangnan gardens, a prominent style of Chinese classical gardens, hold great potential as digital assets for film and game production and digital tourism. However, manual modeling of Jiangnan gardens heavily relies on expert experience for layout design and asset creation, making the process time-co

Cited by 0SourcecodeScholar
2026

ReCoG: Relational and Compact Context Graph Learning for Few-shot Molecular Property Prediction

ICML 2026poster

Few-shot molecular property prediction (FSMPP) is essential in drug discovery and materials design, where high-quality labeled data are often scarce and expensive to obtain. Despite the promising performance of existing methods, especially in the context-aware methods, they still face two-fold sever…

Cited by 0SourceScholar
2026

Reassessing Layer Pruning in LLMs: New Insights and Methods

ICLR 2026poster

Although large language models (LLMs) have achieved remarkable success across various domains, their considerable scale necessitates substantial computational resources, posing significant challenges for deployment in resource-constrained environments. Layer pruning, as a simple yet effective compre…

Cited by 0SourcecodeScholar
2026

SigFusion: Unified Signal-Level Self-Supervised Learning Paradigm for Image Fusion

AAAI 2026technical

Image Fusion (IF) aims to integrate complementary features from multiple source images into a single image. However, a key challenge in this field is the lack of large-scale real-world training datasets. Existing models typically rely on either small datasets or synthetic, less realistic datasets. T

Cited by 0SourcePDFScholar
2026

Tea-Adapter: Teacher Adapter for Efficient Conditional Generation

CVPR 2026

We propose Tea-Adapter, a plug-and-play adapter designed to efficiently integrate conditional knowledge from a smaller teacher model into a larger student video diffusion model. Existing controllable video DiT methods face critical challenges: full fine-tuning of billion-parameter models is extremel

Cited by 0SourceScholar
2026

UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenes

CVPR 2026

Illumination inconsistency is a fundamental challenge in multi-view 3D reconstruction. Variations in sunlight direction, cloud cover, and shadows break the constant-lighting assumption underlying both classical multi-view stereo (MVS) and structure from motion (SfM) pipelines and recent neural rende

Cited by 0SourceScholar
2026

UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections

ICLR 2026poster

We present UP2You, the first tuning-free solution for reconstructing high-fidelity 3D clothed portraits from extremely unconstrained in-the-wild 2D photos. Unlike previous approaches that require "clean" inputs (e.g., full-body images with minimal occlusions, or well calibrated cross-view captures),…

Cited by 0SourcecodeScholar
2026

UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based Deformation

AAAI 2026technical

Joint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately,

Cited by 0SourcePDFScholar
2026

VQ-VA World: Towards High-Quality Visual Question-Visual Answering

CVPR 2026

This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question---an ability that has recently emerged in proprietary systems such as NanoBanana and GPT-Image. To also bring this capability to open-source models, we introduce VQ-VA

Cited by 0SourcecodeScholar
2026

When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought

CVPR 2026

We propose MIRA (Multimodal Imagination for Reasoning Assessment), a new benchmark designed to evaluate models in scenarios where generating intermediate visual images is essential for successful reasoning. Unlike traditional Chain-of-thought (CoT) methods that rely solely on text, tasks in MIRA req

Cited by 0SourcecodeScholar
2025

Conditional Semantic Textual Similarity via Conditional Contrastive Learning

COLING 2025main

Conditional semantic textual similarity (C-STS) assesses the similarity between pairs of sentence representations under different conditions. The current method encounters the over-estimation issue of positive and negative samples. Specifically, the similarity within positive samples is excessively…

2025

Data-Free Model Extraction for Black-box Recommender Systems via Graph Convolutions

NeurIPS 2025poster

Privacy and security concerns are becoming increasingly critical for recommender systems, as model extraction attack provides an effective way to probe system robustness by replicating the model’s recommendation logic — potentially exposing sensitive user preferences and proprietary algorithmic know…

Cited by 0SourcecodeScholar
2025

DiT4Edit: Diffusion Transformer for Image Editing

AAAI 2025technical

Despite recent advances in UNet-based image editing, methods for shape-aware object editing in high-resolution images are still lacking. Compared to UNet, Diffusion Transformers (DiT) demonstrate superior capabilities to effectively capture the long-range dependencies among patches, leading to highe…

2025

Event-Guided Consistent Video Enhancement with Modality-Adaptive Diffusion Pipeline

NeurIPS 2025poster

Recent advancements in low-light video enhancement (LLVE) have increasingly leveraged both RGB and event cameras to improve video quality under challenging conditions. However, existing approaches share two key drawbacks. First, they are tuned for steady low-light scenes, so their performance drops…

Cited by 0SourceScholar
2025

GS-ID: Illumination Decomposition on Gaussian Splatting via Adaptive Light Aggregation and Diffusion-Guided Material Priors

ICCV 2025poster

Gaussian Splatting (GS) has emerged as an effective representation for photorealistic rendering, but the underlying geometry, material, and lighting remain entangled, hindering scene editing. Existing GS-based methods struggle to disentangle these components under non-Lambertian conditions, especial…

Cited by 0SourcePDFScholar
2025

Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image Fusion

ICCV 2025poster

Infrared and visible image fusion (VIS-IR) aims to integrate complementary information from both source images to produce a fused image with enriched details. However, most existing fusion models lack controllability, making it difficult to customize the fused output according to user preferences. T…

2025

Learning Fused State Representations for Control from Multi-View Observations

ICML 2025poster

Multi-View Reinforcement Learning (MVRL) seeks to provide agents with multi-view observations, enabling them to perceive environment with greater effectiveness and precision. Recent advancements in MVRL focus on extracting latent representations from multiview observations and leveraging them in con…

2025

MFANet: Multi-Feature Aggregation Network for Multi-focus Image Fusion

ICASSP 2025accepted

Existing deep learning-based Multi-focus Image Fusion (MFIF) methods often rely on loss functions derived from linear combinations of image quality metrics, leading to complexities in training and only marginal improvements in image quality. Recognizing this, our study identifies input space and sca…

Cited by 0SourceScholar
2025

Mamba YOLO: A Simple Baseline for Object Detection with State Space Model

AAAI 2025technical

Driven by the rapid development of deep learning technology, the YOLO series has set a new benchmark for real-time object detectors. Additionally, transformer-based structures have emerged as the most powerful solution in the field, greatly extending the model's receptive field and achieving signifi…

2025

MambaInst: Lightweight State Space Model for Real-Time Instance Segmentation

ICASSP 2025accepted

In this paper, we propose a lightweight and efficient state-space model-based instance segmentation network named MambaInst, which extracts deep semantic features through a LightSSM Block consisting of gating mechanisms and residual connectivity to model long-distance spatial dependencies with linea…

Cited by 0SourceScholar
2025

Multi-Modal Medical Image Fusion via 3D Manifold Fitting and Dual-Domain Cross-Attention

ICASSP 2025accepted

Medical image fusion (MIF) aims to extract complementary features from multi-modal source images and fuse them into a single image to assist in clinical diagnostics. Despite its importance, MIF faces two primary challenges: the lack of tailored paradigms for CMSF extraction and insufficient dual exp…

Cited by 0SourceScholar
2025

Multi-Stage Multimodal Distillation for Audio-Visual Speaker Tracking

ICASSP 2025accepted

Speaker tracking plays a crucial role in various human-robot interaction applications. Recently, leveraging multimodal information, such as audio and visual signals, has become an important strategy for enhancing the robustness of the tracking system. However, current methods face challenges in effe…

Cited by 0SourceScholar
2025

RemDet: Rethinking Efficient Model Design for UAV Object Detection

AAAI 2025technical

Object detection in Unmanned Aerial Vehicle (UAV) images has emerged as a focal area of research, which presents two significant challenges: i) objects are typically small and dense within vast images; ii) computational resource constraints render most models unsuitable for real-time deployment. Cur…

2025

RestorMamba: An Enhanced Synergistic State Space Model for Image Restoration

ICASSP 2025accepted

In this paper, we introduce an image inpainting method based on the State Space Model (SSM), named Restoration Mamba (RestorMamba). This approach incorporates effi-cient long-range dependency modeling within the network, which is particularly suited for the complexities of high-texture and high-reso…

Cited by 0SourceScholar
2025

Spiking Point Transformer for Point Cloud Classification

AAAI 2025technical

Spiking Neural Networks (SNNs) offer an attractive and energy-efficient alternative to conventional Artificial Neural Networks (ANNs) due to their sparse binary activation. When SNN meets Transformer, it shows great potential in 2D image processing. However, their application for 3D point cloud rem…

2025

Text2VDM: Text to Vector Displacement Maps for Expressive and Interactive 3D Sculpting

ICCV 2025poster

Professional 3D asset creation often requires diverse sculpting brushes to add surface details and geometric structures.Despite recent progress in 3D generation, producing reusable sculpting brushes compatible with artists' workflows remains an open and challenging problem.These sculpting brushes ar…

Cited by 0SourcePDFScholar
2025

Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts

NeurIPS 2025poster

Single-source Domain Generalized Object Detection (SDGOD), as a cutting-edge research topic in computer vision, aims to enhance model generalization capability in unseen target domains through single-source domain training. Current mainstream approaches attempt to mitigate domain discrepancies via d…

Cited by 0SourceScholar
2025

VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion Transformers

AAAI 2025technical

The Diffusion Transformers Models (DiTs) have transitioned the network architecture from traditional UNets to transformers, demonstrating exceptional capabilities in image generation. Although DiTs have been widely applied to high-definition video generation tasks, their large parameter size hinders…

Cited by 8SourcePDFScholar
2025

ViM-VQ: Efficient Post-Training Vector Quantization for Visual Mamba

ICCV 2025poster

Visual Mamba networks (ViMs) extend the selective state space model (Mamba) to various vision tasks and demonstrate significant potential. As a promising compression technique, vector quantization (VQ) decomposes network weights into codebooks and assignments, significantly reducing memory usage and…

Cited by 0SourcePDFScholar
2025

What If We Recaption Billions of Web Images with LLaMA-3?

ICML 2025poster

Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investi…

Cited by 38SourcePDFScholar
2024

A Robust Quantile Huber Loss with Interpretable Parameter Adjustment in Distributional Reinforcement Learning

ICASSP 2024accepted

Distributional Reinforcement Learning (RL) estimates return distribution mainly by learning quantile values via minimizing the quantile Huber loss function, entailing a threshold parameter often selected heuristically or via hyperparameter search, which may not generalize well and can be suboptimal.…

Cited by 0SourceScholar
2024

An Intelligent Robotic Endoscope Control System Based on Fusing Natural Language Processing and Vision Models

ICRA 2024poster

In recent years, the area of Robot-Assisted Minimally Invasive Surgery (RAMIS) is standing on the the verge of a new wave of innovations. However, autonomy in RAMIS is still in a primitive stage. Therefore, most surgeries still require manual control of the endoscope and the robotic instruments, res…

Cited by 3SourceScholar
2024

DreamCatcher: A Wearer-aware Multi-modal Sleep Event Dataset Based on Earables in Non-restrictive Environments

NeurIPS 2024spotlight

Poor quality sleep can be characterized by the occurrence of events ranging from body movement to breathing impairment. Widely available earbuds equipped with sensors (also known as earables) can be combined with a sleep event detection algorithm to offer a convenient alternative to laborious clinic…

2024

LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video Reconstruction

NeurIPS 2024poster

Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video (E2V) reconstruction to bridge event-based and standard comput…

Cited by 2SourcePDFScholar
2024

MemoNav: Working Memory Model for Visual Navigation

CVPR 2024highlight

Image-goal navigation is a challenging task that requires an agent to navigate to a goal indicated by an image in unfamiliar environments. Existing methods utilizing diverse scene memories suffer from inefficient exploration since they use all historical observations for decision-making without cons…

2024

Rejuvenating image-GPT as Strong Visual Representation Learners

ICML 2024oral

This paper enhances image-GPT (iGPT), one of the pioneering works that introduce autoregressive pretraining to predict the next pixels for visual representation learning. Two simple yet essential changes are made. First, we shift the prediction target from raw pixels to semantic tokens, enabling a h…

2024

Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training

CVPR 2024poster

Contrastive learning has emerged as a promising paradigm for 3D open-world understanding i.e. aligning point cloud representation to image and text embedding space individually. In this paper we introduce MixCon3D a simple yet effective method aiming to sculpt holistic 3D representation in contrasti…

2024

SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting

CVPR 2024poster

We present SplattingAvatar a hybrid 3D representation of photorealistic human avatars with Gaussian Splatting embedded on a triangle mesh which renders over 300 FPS on a modern GPU and 30 FPS on a mobile device. We disentangle the motion and appearance of a virtual human with explicit mesh geometry…

2024

Voila-A: Aligning Vision-Language Models with User's Gaze Attention

NeurIPS 2024spotlight

In recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs). However, existing VLMs face challenges in handling real-world applications with complex scenes and multiple object…

Cited by 10SourcePDFScholar
2023

A Step Towards Conditional Autonomy - Robotic Appendectomy

RA-L 2023

In recent years, Robot-Assisted Minimally Invasive Surgery (RAMIS) has been widely adopted worldwide due to its high precision, improved ergonomics and intuitive control. With the advances in artificial intelligence and surgical robot technologies, it is anticipated that the cognitive load on the su

Cited by 13SourceScholar
2023

DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge Distillation

ICCV 2023poster

3D perception based on the representations learned from multi-camera bird's-eye-view (BEV) is trending as cameras are cost-effective for mass production in autonomous driving industry. However, there exists a distinct performance gap between multi-camera BEV and LiDAR based 3D object detection. One…

Cited by 36PDFcodeScholar
2023

Masked Autoencoders Enable Efficient Knowledge Distillers

CVPR 2023poster

This paper studies the potential of distilling knowledge from pre-trained models, especially Masked Autoencoders. Our approach is simple: in addition to optimizing the pixel reconstruction loss on masked inputs, we minimize the distance between the intermediate feature map of the teacher model and t…

2023

Scenario Diffusion: Controllable Driving Scenario Generation With Diffusion

NeurIPS 2023poster

Automated creation of synthetic traffic scenarios is a key part of scaling the safety validation of autonomous vehicles (AVs). In this paper, we propose Scenario Diffusion, a novel diffusion-based architecture for generating traffic scenarios that enables controllable scenario generation. We combine…

Cited by 35SourcePDFScholar
2023

Thermal Infrared Image Inpainting Via Edge-Aware Guidance

ICASSP 2023accepted

Image inpainting has achieved fundamental advances with deep learning. However, almost all existing inpainting methods aim to process natural images, while few target Thermal Infrared (TIR) images, which have widespread applications. When applied to TIR images, conventional inpainting methods usuall…

Cited by 0SourceScholar
2022

Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

EMNLP 2022main

Language models increasingly rely on massive web crawls for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and news often serve as anchors for automatically selecting web text most suitable for language modeling, a process typic…

Cited by 31SourcePDFScholar
2020

Towards Fairness in Visual Recognition: Effective Strategies for Bias Mitigation

CVPR 2020poster

Computer vision models learn to perform a task by capturing relevant statistics from training data. It has been shown that models learn spurious age, gender, and race correlations when trained for seemingly unrelated tasks like activity recognition or image captioning. Various mitigation techniques…

Cited by 441PDFcodeScholar
2020

Towards Unique and Informative Captioning of Images

ECCV 2020poster

Despite considerable progress, state of the art image captioning models produce generic captions, leaving out important image details. Furthermore, these systems may even misrepresent the image in order to produce a simpler caption consisting of common concepts. In this paper, we first analyze both…