← Search

Yueyi Zhang

41 accepted papers

2026

Seeing the Unseen: Zooming in the Dark with Event Cameras

AAAI 2026technical

This paper addresses low-light video super-resolution (LVSR), aiming to restore high-resolution videos from low-light, low-resolution (LR) inputs. Existing LVSR methods often struggle to recover fine details due to limited contrast and insufficient high-frequency information. To overcome these chall

Cited by 0SourcePDFScholar
2026

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

AAAI 2026technical

Universal multimodal embedding models are essential in various tasks. Existing approaches typically use in-batch mining to identify hard negatives by measuring the similarity of query-candidate pairs. However, these methods often struggle to capture subtle semantic differences among candidates and l

Cited by 0SourcePDFScholar
2025

D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution Refinement

ICLR 2025spotlight

We introduce D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR models. D-FINE comprises two key components: Fine-grained Distribution Refinement (FDR) and Global Optimal Localization Self-Distillation…

2025

Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach

AAAI 2025technical

Event cameras have recently been introduced into image semantic segmentation, owing to their high temporal resolution and other advantageous properties. However, existing event-based semantic segmentation methods often fail to fully exploit the complementary information provided by frames and events…

2025

Event-Enhanced Blurry Video Super-Resolution

AAAI 2025technical

In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insu…

2025

Event-boosted Deformable 3D Gaussians for Dynamic Scene Reconstruction

ICCV 2025poster

Deformable 3D Gaussian Splatting (3D-GS) is limited by missing intermediate motion information due to the low temporal resolution of RGB cameras. To address this, we introduce the first approach combining event cameras, which capture high-temporal-resolution, continuous motion data, with deformable…

2025

Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving

ICCV 2025poster

Existing benchmarks for Vision-Language Model (VLM) in autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grained tasks, which remain insufficient to assess capabilities in complex driving scenarios. To this end, we introduce VLAD…

2025

GenFlow3D: Generative Scene Flow Estimation and Prediction on Point Cloud Sequences

ICCV 2025poster

Scene flow provides the fundamental information of the scene dynamics. Existing scene flow estimation methods typically rely on the correlation between only a consecutive point cloud pair, which makes them limited to the instantaneous state of the scene and face challenges in real-world scenarios wi…

2025

Generalizable Non-Line-of-Sight Imaging with Learnable Physical Priors

ICCV 2025poster

Non-line-of-sight (NLOS) imaging, recovering the hidden volume from indirect reflections, has attracted increasing attention due to its potential applications. Despite promising results, existing NLOS reconstruction approaches are constrained by the reliance on empirical physical priors, e.g., singl…

2025

Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space Model

CVPR 2025poster

Brain tumor segmentation plays a crucial role in clinical diagnosis, yet the frequent unavailability of certain MRI modalities poses a significant challenge. In this paper, we introduce the Learnable Sorting State Space Model (LS3M), a novel framework designed to maximize the utilization of availabl…

Cited by 0SourcePDFScholar
2025

Learning Neural Scene Representation from iToF Imaging

ICCV 2025poster

Indirect Time-of-Flight (iToF) cameras are popular for 3D perception because they are cost-effective and easy to deploy. They emit modulated infrared signals to illuminate the scene and process the received signals to generate amplitude and phase images. The depth is calculated from the phase using…

Cited by 0SourcePDFScholar
2025

S2D-LFE: Sparse-to-Dense Light Field Event Generation

CVPR 2025poster

In this paper, we present S2D-LFE, an innovative approach for sparse-to-dense light field event generation. For the first time to our knowledge, S2D-LFE enables controllable novel view synthesis only from sparse-view light field event (LFE) data, and addresses three critical challenges for the LFE g…

2025

Spiking Point Transformer for Point Cloud Classification

AAAI 2025technical

Spiking Neural Networks (SNNs) offer an attractive and energy-efficient alternative to conventional Artificial Neural Networks (ANNs) due to their sparse binary activation. When SNN meets Transformer, it shows great potential in 2D image processing. However, their application for 3D point cloud rem…

2024

Cross-Dimension Affinity Distillation for 3D EM Neuron Segmentation

CVPR 2024poster

Accurate 3D neuron segmentation from electron microscopy (EM) volumes is crucial for neuroscience research. However the complex neuron morphology often leads to over-merge and over-segmentation results. Recent advancements utilize 3D CNNs to predict a 3D affinity map with improved accuracy but suffe…

2024

EvTexture: Event-driven Texture Enhancement for Video Super-Resolution

ICML 2024poster

Event-based vision has drawn increasing attention due to its unique characteristics, such as high temporal resolution and high dynamic range. It has been used in video super-resolution (VSR) recently to enhance the flow estimation and temporal alignment. Rather than for motion learning, we propose i…

2024

Event-Adapted Video Super-Resolution

ECCV 2024poster

"Introducing event cameras into video super-resolution (VSR) shows great promise. In practice, however, integrating event data as a new modality necessitates a laborious model architecture design. This not only consumes substantial time and effort but also disregards valuable insights from successfu…

Cited by 6SourcePDFScholar
2024

Event-assisted Low-Light Video Object Segmentation

CVPR 2024poster

In the realm of video object segmentation (VOS) the challenge of operating under low-light conditions persists resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras characterized by their high dynamic ran…

2024

High-Resolution and Few-shot View Synthesis from Asymmetric Dual-lens Inputs

ECCV 2024poster

"Novel view synthesis has achieved remarkable quality and efficiency by the paradigm of 3D Gaussian Splatting (3D-GS), but still faces two challenges: 1) significant performance degradation when trained with only few-shot samples due to a lack of geometry constraint, and 2) incapability of rendering…

2024

Image Captioning with Multi-Context Synthetic Data

AAAI 2024technical

Image captioning requires numerous annotated image-text pairs, resulting in substantial annotation costs. Recently, large models (e.g. diffusion models and large language models) have excelled in producing high-quality images and text. This potential can be harnessed to create synthetic image-text p…

Cited by 13SourcePDFScholar
2024

Scene Adaptive Sparse Transformer for Event-based Object Detection

CVPR 2024poster

While recent Transformer-based approaches have shown impressive performances on event-based object detection tasks their high computational costs still diminish the low power consumption advantage of event cameras. Image-based works attempt to reduce these costs by introducing sparse Transformers. H…

2024

TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing Modalities

AAAI 2024technical

Numerous techniques excel in brain tumor segmentation using multi-modal magnetic resonance imaging (MRI) sequences, delivering exceptional results. However, the prevalent absence of modalities in clinical scenarios hampers performance. Current approaches frequently resort to zero maps as substitutes…

Cited by 5SourcePDFScholar
2024

Toward Dynamic Non-Line-of-Sight Imaging with Mamba Enforced Temporal Consistency

NeurIPS 2024poster

Dynamic reconstruction in confocal non-line-of-sight imaging encounters great challenges since the dense raster-scanning manner limits the practical frame rate. A fewer pioneer works reconstruct high-resolution volumes from the under-scanning transient measurements but overlook temporal consistency…

2024

Visual Perception by Large Language Model’s Weights

NeurIPS 2024poster

Existing Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (LLMs) and concatenating visual tokens with text tokens to form a unified sequence input for LLMs. These methods demonstra…

2023

A Soma Segmentation Benchmark in Full Adult Fly Brain

CVPR 2023poster

Neuron reconstruction in a full adult fly brain from high-resolution electron microscopy (EM) data is regarded as a cornerstone for neuroscientists to explore how neurons inspire intelligence. As the central part of neurons, somas in the full brain indicate the origin of neurogenesis and neural func…

2023

Better and Faster: Adaptive Event Conversion for Event-Based Object Detection

AAAI 2023technical

Event cameras are a kind of bio-inspired imaging sensor, which asynchronously collect sparse event streams with many advantages. In this paper, we focus on building better and faster event-based object detectors. To this end, we first propose a computationally efficient event representation Hyper Hi…

Cited by 18SourcePDFScholar
2023

Deep Non-line-of-sight Imaging from Under-scanning Measurements

NeurIPS 2023poster

Active confocal non-line-of-sight (NLOS) imaging has successfully enabled seeing around corners relying on high-quality transient measurements. However, acquiring spatial-dense transient measurement is time-consuming, raising the question of how to reconstruct satisfactory results from under-scannin…

2023

Depth Estimation From Indoor Panoramas With Neural Scene Representation

CVPR 2023poster

Depth estimation from indoor panoramas is challenging due to the equirectangular distortions of panoramas and inaccurate matching. In this paper, we propose a practical framework to improve the accuracy and efficiency of depth estimation from multi-view indoor panoramic images with the Neural Radian…

2023

GET: Group Event Transformer for Event-Based Vision

ICCV 2023poster

Event cameras are a type of novel neuromorphic sen-sor that has been gaining increasing attention. Existing event-based backbones mainly rely on image-based designs to extract spatial information within the image transformed from events, overlooking important event properties like time and polarity.…

Cited by 107PDFcodeScholar
2023

Learning Cross-Representation Affinity Consistency for Sparsely Supervised Biomedical Instance Segmentation

ICCV 2023poster

Sparse instance-level supervision has recently been explored to address insufficient annotation in biomedical instance segmentation, which is easier to annotate crowded instances and better preserves instance completeness for 3D volumetric datasets compared to common semi-supervision.In this paper,…

Cited by 8PDFcodeScholar
2023

NLOST: Non-Line-of-Sight Imaging With Transformer

CVPR 2023poster

Time-resolved non-line-of-sight (NLOS) imaging is based on the multi-bounce indirect reflections from the hidden objects for 3D sensing. Reconstruction from NLOS measurements remains challenging especially for complicated scenes. To boost the performance, we present NLOST, the first transformer-base…

Cited by 29SourcePDFScholar
2023

Progressive Spatio-Temporal Alignment for Efficient Event-Based Motion Estimation

CVPR 2023poster

In this paper, we propose an efficient event-based motion estimation framework for various motion models. Different from previous works, we design a progressive event-to-map alignment scheme and utilize the spatio-temporal correlations to align events. In detail, we progressively align sampled event…

2022

Biological Instance Segmentation with a Superpixel-Guided Graph

IJCAI 2022poster

Recent advanced proposal-free instance segmentation methods have made significant progress in biological images. However, existing methods are vulnerable to local imaging artifacts and similar object appearances, resulting in over-merge and over-segmentation. To reduce these two kinds of errors, we…

2022

Degradation-Agnostic Correspondence From Resolution-Asymmetric Stereo

CVPR 2022poster

In this paper, we study the problem of stereo matching from a pair of images with different resolutions, e.g., those acquired with a tele-wide camera system. Due to the difficulty of obtaining ground-truth disparity labels in diverse real-world systems, we start from an unsupervised learning perspec…

Cited by 10PDFScholar
2022

Exploiting Rigidity Constraints for LiDAR Scene Flow Estimation

CVPR 2022poster

Previous LiDAR scene flow estimation methods, especially recurrent neural networks, usually suffer from structure distortion in challenging cases, such as sparse reflection and motion occlusions. In this paper, we propose a novel optimization method based on a recurrent neural network to predict LiD…

Cited by 40PDFScholar
2021

Training Spiking Neural Networks with Accumulated Spiking Flow

AAAI 2021technical

The fast development of neuromorphic hardwares promotes Spiking Neural Networks (SNNs) to a thrilling research avenue. Current SNNs, though much efficient, are less effective compared with leading Artificial Neural Networks (ANNs) especially in supervised learning tasks. Recent efforts further demon…

2020

Spatial Hierarchy Aware Residual Pyramid Network for Time-of-Flight Depth Denoising

ECCV 2020poster

Time-of-Flight (ToF) sensors have been increasingly used on mobile devices for depth sensing. However, the existence of noise, such as Multi-Path Interference (MPI) and shot noise, degrades the ToF imaging quality. Previous CNN-based methods remove ToF depth noise without considering the spatial hie…