← Search

Liyuan Pan

27 accepted papers

2026

DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera Conditions

CVPR 2026

Human Action Recognition (HAR) is a fundamental computer vision task with diverse real-world applications. Practical deployments often involve low-light environments and unconstrained 6-DoF camera motion, conditions that degrade visual quality, disrupt temporal coherence, and compromise reliability

Cited by 0SourceScholar
2026

EmoThinker: Advancing Visual-Acoustic Emotion Analysis via Structural Token Selection and Chain-of-Thought Reasoning

CVPR 2026

Multimodal Emotion Analysis (MEA) is crucial for human-centric AI, yet current methods struggle with two core challenges: the sparse nature of emotional cues across modalities and their inherent temporal asynchrony. Existing approaches, which often rely on implicit fusion, consequently suffer from d

Cited by 0SourceScholar
2026

GeoFree-CoSeg: Unsupervised Point Cloud-Image Cross-Modal Co-Segmentation Without Geometric Alignment

CVPR 2026

Co-segmentation aims to identify and segment common objects across a set of point clouds or images. Existing methods focus on single-modal co-segmentation. However, the limited semantics of a single modality restrict the discovery of common objects, leading to costly and labor-intensive segmentation

Cited by 0SourceScholar
2026

Language-Guided One-Step Diffusion Model for Nighttime Flare Removal

CVPR 2026

Nighttime photography is susceptible to flare caused by strong light sources, which degrades visual quality and disrupts structural information required by downstream vision tasks. Existing nighttime flare removal methods generally lack semantic priors for flare-occluded regions and thus tend to int

Cited by 0SourceScholar
2025

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding

CVPR 2025poster

The challenge in LLM-based video understanding lies in preserving visual and semantic information in long videos while maintaining a memory-affordable token count. However, redundancy and correspondence in videos have hindered the performance potential of existing methods. Through statistical learni…

Cited by 2SourcePDFScholar
2025

ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks

ACL 2025finding

Solving expert-level multimodal tasks is a key milestone in general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to evolve, evaluation of frontier multimodal intelligence becomes necessary yet challenging. In this work, we introduce ProBench, a benchmark of…

Cited by 0SourcePDFScholar
2025

Storyboard-guided Alignment for Fine-grained Video Action Recognition

NeurIPS 2025poster

Fine-grained video action recognition can be formulated as a video–text matching problem. Previous approaches primarily rely on global video semantics to consolidate video embeddings, often leading to misaligned video–text pairs due to inaccurate atomic-level action understanding. This inaccuracy ar…

Cited by 0SourceScholar
2025

Task-Specific Gradient Adaptation for Few-Shot One-Class Classification

CVPR 2025poster

Optimization-based meta-learning methods for few-shot one-class classification (FS-OCC) aim to fine-tune a meta-trained model to classify the positive and negative samples using only a few positive samples by adaptation. However, recent approaches primarily focus on adjusting existing meta-learning…

Cited by 0SourcePDFScholar
2024

Event-based Few-shot Fine-grained Human Action Recognition

IROS 2024poster

Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as…

Cited by 1SourceScholar
2024

LDP: Language-driven Dual-Pixel Image Defocus Deblurring Network

CVPR 2024poster

Recovering sharp images from dual-pixel (DP) pairs with disparity-dependent blur is a challenging task. Existing blur map-based deblurring methods have demonstrated promising results. In this paper we propose to the best of our knowledge the first framework to introduce the contrastive language-imag…

Cited by 12SourcePDFScholar
2023

K3DN: Disparity-Aware Kernel Estimation for Dual-Pixel Defocus Deblurring

CVPR 2023poster

The dual-pixel (DP) sensor captures a two-view image pair in a single snapshot by splitting each pixel in half. The disparity occurs in defocus blurred regions between the two views of the DP pair, while the in-focus sharp regions have zero disparity. This motivates us to propose a K3DN framework fo…

Cited by 12SourcePDFScholar
2023

L2T-DLN: Learning to Teach with Dynamic Loss Network

NeurIPS 2023poster

With the concept of teaching being introduced to the machine learning community, a teacher model start using dynamic loss functions to teach the training of a student model. The dynamic intends to set adaptive loss functions to different phases of student model learning. In existing works, the teach…

Cited by 3SourcePDFScholar
2021

Dual Pixel Exploration: Simultaneous Depth Estimation and Image Restoration

CVPR 2021poster

The dual-pixel (DP) hardware works by splitting each pixel in half and creating an image pair in a single snapshot. Several works estimate depth/inverse depth by treating the DP pair as a stereo pair. However, dual-pixel disparity only occurs in image regions with the defocus blur. The heavy defocus…

Cited by 43PDFScholar
2021

Stereo Hybrid Event-Frame (SHEF) Cameras for 3D Perception

IROS 2021poster

Stereo camera systems play an important role in robotics applications to perceive the 3D world. However, conventional cameras have drawbacks such as low dynamic range, motion blur and latency due to the underlying frame- based mechanism. Event cameras address these limitations as they report the bri…

Cited by 29SourcecodeScholar
2019

Bringing a Blurry Frame Alive at High Frame-Rate With an Event Camera

CVPR 2019oral

Event-based cameras can measure intensity changes (called 'events') with microsecond accuracy under high-speed motion and challenging lighting conditions. With the active pixel sensor (APS), the event camera allows simultaneous output of the intensity frames. However, the output images are captured…

Cited by 312PDFScholar
2019

Phase-Only Image Based Kernel Estimation for Single Image Blind Deblurring

CVPR 2019poster

The image motion blurring process is generally modelled as the convolution of a blur kernel with a latent image. Therefore, the estimation of the blur kernel is essentially important for blind image deblurring. Unlike existing approaches which focus on approaching the problem by enforcing various pr…

Cited by 80PDFScholar