← Search

Yu Zhu

40 accepted papers

2026

CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation

AAAI 2026technical

Recent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveal

Cited by 0SourcePDFScholar
2026

NeuroFlow: Toward Unified Visual Encoding and Decoding from Neural Activity

CVPR 2026

Visual encoding and decoding models act as gateways to understanding the neural mechanisms underlying human visual perception. Typically, visual encoding models that predict brain activity from stimuli and decoding models that reproduce stimuli from brain activity are treated as distinct tasks, requ

Cited by 0SourcecodeScholar
2026

Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework

CVPR 2026

To continuously enhance model adaptability in surgical video scene parsing, recent studies incrementally update it to progressively learn to segment an increasing number of surgical instruments over time. However, prior works constantly overlooked the potential of positive forward knowledge transfer

Cited by 0SourceScholar
2025

CSV-Occ: Fusing Multi-frame Alignment for Occupancy Prediction with Temporal Cross State Space Model and Central Voting Mechanism

ICML 2025poster

Recently, image-based 3D semantic occupancy prediction has become a hot topic in 3D scene understanding for autonomous driving. Compared with the bounding box form of 3D object detection, the ability to describe the fine-grained contours of any obstacles in the scene is the key insight of voxel occ…

Cited by 0SourcePDFScholar
2025

MPBR: Multimodal Progressive Bidirectional Reasoning for Open-Set Fine-Grained Recognition

ICCV 2025poster

Open-set fine-grained recognition (OSFGR) is the core exploration of building open-world intelligent systems. The challenge lies in the gradual semantic drift during the transition from coarse-grained to fine-grained categories. However, although existing methods leverage hierarchical representation…

Cited by 0SourcePDFScholar
2025

Multi-Modal Latent Variables for Cross-Individual Primary Visual Cortex Modeling and Analysis

AAAI 2025technical

Elucidating the functional mechanisms of the primary visual cortex (V1) remains a fundamental challenge in systems neuroscience. Current computational models face two critical limitations, namely the challenge of cross-modal integration between partial neural recordings and complex visual stimuli, a…

Cited by 0SourcePDFScholar
2025

Neural Representational Consistency Emerges from Probabilistic Neural-Behavioral Representation Alignment

ICML 2025poster

Individual brains exhibit striking structural and physiological heterogeneity, yet neural circuits can generate remarkably consistent functional properties across individuals, an apparent paradox in neuroscience. While recent studies have observed preserved neural representations in motor cortex thr…

2025

PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis

NeurIPS 2025poster

Pairwise camera pose estimation from sparsely overlapping image pairs remains a critical and unsolved challenge in 3D vision. Most existing methods struggle with image pairs that have small or no overlap. Recent approaches attempt to address this by synthesizing intermediate frames using video inte…

Cited by 0SourceScholar
2025

Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse Views

CVPR 2025poster

We present a Gaussian Splatting method for surface reconstruction using sparse input views. Previous methods relying on dense views struggle with extremely sparse Structure-from-Motion points for initialization. While learning-based Multi-view Stereo (MVS) provides dense 3D points, directly combinin…

2025

SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning

NeurIPS 2025poster

Deciphering how visual stimuli are transformed into cortical responses is a fundamental challenge in computational neuroscience. This visual-to-neural mapping is inherently a one-to-many relationship, as identical visual inputs reliably evoke variable hemodynamic responses across trials, contexts, a…

Cited by 0SourcecodeScholar
2025

Token-Level Accept or Reject: A Micro Alignment Approach for Large Language Models

IJCAI 2025

With the rapid development of Large Language Models (LLMs), aligning these models with human preferences and values is critical to ensuring ethical and safe applications. However, existing alignment techniques such as RLHF or DPO often require direct fine-tuning on LLMs with billions of parameters,

2024

Diffevent: Event Residual Diffusion for Image Deblurring

ICASSP 2024accepted

Traditional frame-based cameras inevitably suffer from non-uniform blur in real-world scenarios. Event cameras that record the intensity changes with high temporal resolution provide an effective solution for image deblurring. In this paper, we formulate the event-based image deblurring as an image…

Cited by 0SourceScholar
2024

GSDD: Generative Space Dataset Distillation for Image Super-resolution

AAAI 2024technical

Single image super-resolution (SISR), especially in the real world, usually builds a large amount of LR-HR image pairs to learn representations that contain rich textural and structural information. However, relying on massive data for model training not only reduces training efficiency, but also ca…

Cited by 3SourcePDFScholar
2024

Memory-Efficient Prompt Tuning for Incremental Histopathology Classification

AAAI 2024technical

Recent studies have made remarkable progress in histopathology classification. Based on current successes, contemporary works proposed to further upgrade the model towards a more generalizable and robust direction through incrementally learning from the sequentially delivered domains. Unlike previou…

Cited by 1SourcePDFScholar
2024

Multiple Object Tracking Based on Occlusion-Aware Embedding Consistency Learning

ICASSP 2024accepted

The Joint Detection and Embedding (JDE) framework has achieved remarkable progress for multiple object tracking. Existing methods often employ extracted embeddings to re-establish associations between new detections and previously disrupted tracks. However, the reliability of embeddings diminishes w…

Cited by 0SourceScholar
2024

Retrospective for the Dynamic Sensorium Competition for predicting large-scale mouse primary visual cortex activity from videos

NeurIPS 2024poster

Understanding how biological visual systems process information is challenging because of the nonlinear relationship between visual input and neuronal responses. Artificial neural networks allow computational neuroscientists to create predictive models that connect biological and machine vision. Ma…

Cited by 3SourcePDFScholar
2023

A Unified HDR Imaging Method With Pixel and Patch Level

CVPR 2023poster

Mapping Low Dynamic Range (LDR) images with different exposures to High Dynamic Range (HDR) remains nontrivial and challenging on dynamic scenes due to ghosting caused by object motion or camera jitting. With the success of Deep Neural Networks (DNNs), several DNNs-based methods have been proposed t…

Cited by 39SourcePDFScholar
2023

Boosting No-Reference Super-Resolution Image Quality Assessment with Knowledge Distillation and Extension

ICASSP 2023accepted

Deep learning (DL) based image super-resolution (SR) tech-niques have been well investigated for recent years. However, studies dedicated to SR image quality assessment (SR-IQA) have not been fully developed, which is even more difficult if pristine high-resolution (HR) images are lacking as a refer…

Cited by 0SourceScholar
2023

Hierarchical Semi-Implicit Variational Inference with Application to Diffusion Model Acceleration

NeurIPS 2023poster

Semi-implicit variational inference (SIVI) has been introduced to expand the analytical variational families by defining expressive semi-implicit distributions in a hierarchical manner. However, the single-layer architecture commonly used in current SIVI methods can be insufficient when the target p…

2023

Learning To Fuse Monocular and Multi-View Cues for Multi-Frame Depth Estimation in Dynamic Scenes

CVPR 2023poster

Multi-frame depth estimation generally achieves high accuracy relying on the multi-view geometric consistency. When applied in dynamic scenes, e.g., autonomous driving, this consistency is usually violated in the dynamic areas, leading to corrupted estimations. Many multi-frame methods handle dynami…

2023

SMAE: Few-Shot Learning for HDR Deghosting With Saturation-Aware Masked Autoencoders

CVPR 2023poster

Generating a high-quality High Dynamic Range (HDR) image from dynamic scenes has recently been extensively studied by exploiting Deep Neural Networks (DNNs). Most DNNs-based methods require a large amount of training data with ground truth, requiring tedious and time-consuming work. Few-shot HDR ima…

Cited by 19SourcePDFScholar
2022

An Online Time-Optimal Trajectory Planning Method for Constrained Multi-Axis Trajectory With Guaranteed Feasibility

RA-L 2022

Online time-optimal trajectory planning exists in a wide range of applications such as computer numerical control (CNC) manufacturing, robotics and autonomous vehicles. Generally, the methods to generate such time-parameterized trajectory can be categorized as offline methods and online methods. Off

Cited by 30SourceScholar
2022

Exploring and Evaluating Image Restoration Potential in Dynamic Scenes

CVPR 2022poster

In dynamic scenes, images often suffer from dynamic blur due to superposition of motions or low signal-noise ratio resulted from quick shutter speed when avoiding motions. Recovering sharp and clean result from the captured images heavily depends on the ability of restoration methods and the quality…

Cited by 13PDFcodeScholar
2022

Hypergraphs with Edge-Dependent Vertex Weights: Spectral Clustering Based on the 1-Laplacian

ICASSP 2022accepted

We propose a flexible framework for defining the 1-Laplacian of a hypergraph that incorporates edge-dependent vertex weights. These weights are able to reflect varying importance of vertices within a hyperedge, thus conferring the hypergraph model higher expressivity than homogeneous hypergraphs. We…

Cited by 0SourceScholar
2022

SKFlow: Learning Optical Flow with Super Kernels

NeurIPS 2022accept

Optical flow estimation is a classical yet challenging task in computer vision. One of the essential factors in accurately predicting optical flow is to alleviate occlusions between frames. However, it is still a thorny problem for current top-performing optical flow estimation methods due to insuff…

2022

Task Space Contouring Error Estimation and Precision Iterative Control of Robotic Manipulators

RA-L 2022

The task space contouring performance is significant for the machining accuracy of industrial robotic manipulators, but the contouring control of end-effector which is important in the industry has received scant attention. In this letter, a novel task space contouring error estimation and control s

Cited by 19SourceScholar
2021

CPG-Based Hierarchical Locomotion Control for Modular Quadrupedal Robots Using Deep Reinforcement Learning

RA-L 2021

Modular robots have the potential for an unmatched ability to perform versatile and robust locomotion. However, designing effective and adaptive locomotion controllers for modular robots is challenging, resulting in a number of model-based methods that typically require various forms of prior knowle

Cited by 32SourceScholar
2021

Personalized Adaptive Meta Learning for Cold-start User Preference Prediction

AAAI 2021technical

A common challenge in personalized user preference prediction is the cold-start problem. Due to the lack of user-item interactions, directly learning from the new users' log data causes serious over-fitting problem. Recently, many existing studies regard the cold-start personalized preference predic…

Cited by 76SourcePDFScholar
2020

Blindly Assess Image Quality in the Wild Guided by a Self-Adaptive Hyper Network

CVPR 2020poster

Blind image quality assessment (BIQA) for authentically distorted images has always been a challenging problem, since images captured in the wild include varies contents and diverse types of distortions. The vast majority of prior BIQA methods focus on how to predict synthetic image quality, but fai…

Cited by 788PDFcodeScholar
2020

GINet: Graph Interaction Network for Scene Parsing

ECCV 2020poster

Recently, context reasoning using image regions beyond local convolution has shown great potential for scene parsing. In this work, we explore how to incorperate the linguistic knowledge to promote context reasoning over image regions by proposing a Graph Interaction unit (GI unit) and a Semantic Co…

2019

Estimation of Network Processes via Blind Graph Multi-filter Identification

ICASSP 2019accepted

We study the problem of jointly estimating several network processes that are driven by the same input, recasting it as one of blind identification of a bank of graph filters. More precisely, we consider the observation of several graph signals - i.e., signals defined on the nodes of a graph - and w…

Cited by 0SourceScholar
2015

Modeling Deformable Gradient Compositions for Single-Image Super-Resolution

CVPR 2015poster

We propose a single-image super-resolution method based on the gradient reconstruction. To predict the gradient field, we collect a dictionary of gradient patterns from an external set of images. We observe that there are patches representing singular primitive structures (e.g. a single edge), and n…

Cited by 43SourcePDFScholar