← Search

Kin-Man Lam

13 accepted papers

2025

Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based Attacks

ICCV 2025poster

Vision-Language Models (VLMs) extend the capabilities of Large Language Models (LLMs) by incorporating visual information, yet they remain vulnerable to jailbreak attacks, especially when processing noisy or corrupted images. Although existing VLMs adopt security measures during training to mitigate…

2025

See In Detail: Enhancing Sparse-view 3D Gaussian Splatting with Local Depth and Semantic Regularization

ICASSP 2025accepted

3D Gaussian Splatting (3DGS) has shown remarkable performance in novel view synthesis. However, its rendering quality deteriorates with sparse inphut views, leading to distorted content and reduced details. This limitation hinders its practical application. To address this issue, we propose a sparse…

Cited by 0SourceScholar
2024

AMSP-UOD: When Vortex Convolution and Stochastic Perturbation Meet Underwater Object Detection

AAAI 2024technical

In this paper, we present a novel Amplitude-Modulated Stochastic Perturbation and Vortex Convolutional Network, AMSP-UOD, designed for underwater object detection. AMSP-UOD specifically addresses the impact of non-ideal imaging factors on detection accuracy in complex underwater environments. To mit…

2024

Learning Equilibrium Transformation for Gamut Expansion and Color Restoration

ECCV 2024poster

"Existing imaging systems support wide-gamut images like ProPhoto RGB, but most images are typically encoded in a narrower gamut space (e.g., sRGB). To this end, these images can be enhanced by learning to recover the original color values beyond the sRGB gamut, or out-of-gamut values. Current metho…

2024

Towards Progressive Multi-Frequency Representation for Image Warping

CVPR 2024poster

Image warping a classic task in computer vision aims to use geometric transformations to change the appearance of images. Recent methods learn the resampling kernels for warping through neural networks to estimate missing values in irregular grids which however fail to capture local variations in de…

2023

Efficient Feature Fusion for Learning-Based Photometric Stereo

ICASSP 2023accepted

How to handle an arbitrary number for input images is a fundamental problem of learning-based photometric stereo methods. Existing approaches adopt max-pooling or observation map to fuse an arbitrary number of extracted features. However, these methods discard a large amount of the features from the…

Cited by 0SourceScholar
2022

A Hybrid Egocentric Activity Anticipation Framework via Memory-Augmented Recurrent and One-Shot Representation Forecasting

CVPR 2022poster

Egocentric activity anticipation involves identifying the interacted objects and target action patterns in the near future. A standard activity anticipation paradigm is recurrently forecasting future representations to compensate the missing activity semantics of the unobserved sequence. However, th…

Cited by 28PDFScholar
2021

Structure-Enhanced Attentive Learning For Spine Segmentation From Ultrasound Volume Projection Images

ICASSP 2021accepted

Automatic spine segmentation, based on ultrasound volume projection imaging (VPI), is of great value in clinical applications to diagnose scoliosis in teenagers. In this paper, we propose a novel framework to improve the segmentation accuracy on spine images via structure-enhanced attentive learning…

Cited by 0SourceScholar
2020

Pay Attention to Devils: A Photometric Stereo Network for Better Details

IJCAI 2020poster

We present an attention-weighted loss in a photometric stereo neural network to improve 3D surface recovery accuracy in complex-structured areas, such as edges and crinkles, where existing learning-based methods often failed. Instead of using a uniform penalty for all pixels, our method employs the…

Cited by 0SourcePDFScholar
2018

Pyramid Dilated Deeper ConvLSTM for Video Salient Object Detection

ECCV 2018poster

This paper proposes a fast video salient object detection model, based on a novel recurrent network architecture, named Pyramid Dilated Bidirectional ConvLSTM (PDB-ConvLSTM). A Pyramid Dilated Convolution (PDC) module is first designed for simultaneously extracting spatial features at multiple scale…

Cited by 593SourcePDFScholar