← Search

Lai-Man Po

8 accepted papers

2026

From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) for Multimodal Large Language Models (MLLMs) is highly dependent on high-quality labeled data, which is often scarce and prone to substantial annotation noise in real-world scenarios. Existing unsupervised RLVR methods, including pure entropy min

Cited by 0SourcecodeScholar
2024

Joint Learning of Identity and Vein Features for Enhanced Representations in Vascular Biometrics

ICASSP 2024accepted

Vascular biometrics have shown great promise for secure authentication applications and have received increased attention in recent years. This paper proposes a novel framework for joint identity and segmentation feature learning to enrich representations and improve verification performance. The fr…

Cited by 0SourceScholar
2024

Motion Transfer-Driven Intra-Class Data Augmentation for Finger Vein Recognition

ICASSP 2024accepted

Finger vein recognition (FVR) has emerged as a secure biometric technique because of the confidentiality of vascular bio-information. Recently, deep learning-based FVR has gained increased popularity and achieved promising performance. However, the limited size of public vein datasets has caused ove…

Cited by 0SourceScholar
2023

Bidirectionally Deformable Motion Modulation For Video-based Human Pose Transfer

ICCV 2023poster

Video-based human pose transfer is a video-to-video generation task that animates a plain source human image based on a series of target human poses. Considering the difficulties in transferring highly structural patterns on the garments and discontinuous poses, existing methods often generate unsat…

Cited by 26PDFcodeScholar
2022

Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video Representation

AAAI 2022technical

Spatio-temporal representation learning is critical for video self-supervised representation. Recent approaches mainly use contrastive learning and pretext tasks. However, these approaches learn representation by discriminating sampled instances via feature similarity in the latent space while ignor…

2022

D2HNet: Joint Denoising and Deblurring with Hierarchical Network for Robust Night Image Restoration

ECCV 2022poster

"Night imaging with modern smartphone cameras is troublesome due to low photon count and unavoidable noise in the imaging system. Directly adjusting exposure time and ISO ratings cannot obtain sharp and noise-free images at the same time in low-light conditions. Though many methods have been propose…

2016

Face liveness detection and recognition using shearlet based feature descriptors

ICASSP 2016accepted

Face recognition is a widely used biometric technology due to its convenience but it is vulnerable to spoofing attacks made by non-real faces such as a photograph or video of valid user. Face liveness detection is a core technology to make sure that the input face is a live person. However, this is…

Cited by 0SourceScholar
2015

Dynamic ROI based on K-means for remote photoplethysmography

ICASSP 2015accepted

Remote imaging photoplethysmography (RIPPG) can achieve contactless human vital signs monitoring. Though the remote operation mode brings a great convenience for RIPPG applications, the RIPPG signal quality is limited by the remote nature. Improving the RIPPG signal quality becomes an essential task…

Cited by 0SourceScholar