← Search

Son Lam Phung

18 accepted papers

2025

Localizing Before Answering: A Benchmark for Grounded Medical Visual Question Answering

IJCAI 2025

Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in

Cited by 0SourcePDFScholar
2025

Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding

ICCV 2025poster

Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and…

Cited by 0SourcePDFScholar
2023

Bilateral Coarse-to-Fine Network for Point Cloud Completion

ICASSP 2023accepted

Point cloud completion aims to accurately estimate complete point clouds from partial observations. Existing methods of-ten directly infer the missing points from the partial shape, but they suffer from limited structural information. To address this, we propose the Bilateral Coarse-to-Fine Network…

Cited by 0SourceScholar
2022

Class Similarity Weighted Knowledge Distillation for Continual Semantic Segmentation

CVPR 2022poster

Deep learning models are known to suffer from the problem of catastrophic forgetting when they incrementally learn new classes. Continual learning for semantic segmentation (CSS) is an emerging field in computer vision. We identify a problem in CSS: A model tends to be confused between old and new c…

Cited by 65PDFScholar
2022

DirecFormer: A Directed Attention in Transformer Approach to Robust Action Recognition

CVPR 2022poster

Human action recognition has recently become one ofthe popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in thetask of video action recognition with competitive results.However, these methods…

Cited by 75PDFcodeScholar
2021

BiMaL: Bijective Maximum Likelihood Approach to Domain Adaptation in Semantic Scene Segmentation

ICCV 2021poster

Semantic segmentation aims to predict pixel-level labels. It has become a popular task in various computer vision applications. While fully supervised segmentation methods have achieved high accuracy on large-scale vision datasets, they are unable to generalize on a new test environment or a new dom…

Cited by 43PDFcodeScholar
2020

A Variational Bayesian Approach for Multichannel Through-Wall Radar Imaging with Low-Rank and Sparse Priors

ICASSP 2020accepted

This paper considers the problem of multichannel through-wall radar (TWR) imaging from a probabilistic Bayesian perspective. Given the observed radar signals, a joint distribution of the observed data and latent variables is formulated by incorporating two important beliefs: low-dimensional structur…

Cited by 0SourceScholar
2019

Radar Stationary and Moving Indoor Target Localization with Low-rank and Sparse Regularizations

ICASSP 2019accepted

This paper proposes a low-rank and sparse regularized optimization model to address the problem of wall clutter mitigation, stationary, and moving target indications using through-wall radar. The task of wall clutter suppression and target image reconstruction is formulated as a nuclear and ℓ <sub x…

Cited by 0SourceScholar
2018

A Matrix Completion Approach for Wall-Clutter Mitigation in Compressive Radar Imaging of Indoor Targets

ICASSP 2018accepted

This paper presents a low-rank matrix completion approach to tackle the problem of wall clutter mitigation for through-wall radar imaging in the compressive sensing context. In particular, the task of wall clutter removal is reformulated as a matrix completion problem in which a low-rank matrix cont…

Cited by 0SourceScholar
2018

Anatomy-Guided Inverse-Gradient Susceptibility Artifact Correction Method for High-Resolution FMRI

ICASSP 2018accepted

Functional Magnetic Resonance Imaging (fMRI) is a widely used and non-invasive technique for recording changes in brain activity. However, susceptibility artifacts are ubiquitous distortions in fMRI, especially strong in high-resolution images, causing the misrepresentation of brain function and str…

Cited by 0SourceScholar
2018

Human Motion Classification with Micro-Doppler Radar and Bayesian-Optimized Convolutional Neural Networks

ICASSP 2018accepted

In recent years, Doppler radar has emerged as an alternative sensing modality for human gait classification since it measures not only the target speed, but also the local dynamics of the moving body parts, thereby creating a unique spectral signature. This paper presents a learning-based method for…

Cited by 0SourceScholar
2018

Scalable Hierarchical Mixture of Gaussian Processes for Pattern Classification

ICASSP 2018accepted

This paper introduces a novel Gaussian process (GP) classification method that combines advantages of global and local GP approximators through a two-layer hierarchical model. The upper layer consists of a global sparse GP to coarsely model the entire dataset. The lower layer is a mixture of GP expe…

Cited by 0SourceScholar
2017

Enhanced pixel-wise voting for image vanishing point detection in road scenes

ICASSP 2017accepted

Vanishing point estimation is a crucial task in vision-based road detection. This paper presents a new texture-based voting scheme, which enhances both accuracy and speed of vanishing point estimation. In the proposed method, color tensors analysis is adopted to calculate local orientations and colo…

Cited by 0SourceScholar
2017

Human interaction recognition using low-rank matrix approximation and super descriptor tensor decomposition

ICASSP 2017accepted

Audio-visual recognition systems rely on efficient feature extraction. Many spatio-temporal interest point detectors for visual feature extraction are either too sparse, leading to loss of information, or too dense resulting in noisy and redundant information. Furthermore, interest point detectors d…

Cited by 0SourceScholar
2016

Radar imaging of stationary indoor targets using joint low-rank and sparsity constraints

ICASSP 2016accepted

This paper introduces a joint low-rank and sparsity-based model to address the problem of wall-clutter mitigation in compressed through-the-wall radar imaging. The proposed model is motivated by two observations that wall reflections reside in a low-rank subspace, and target signals tend to be spars…

Cited by 0SourceScholar
2016

Variational inference for infinite mixtures of sparse Gaussian processes through KL-correction

ICASSP 2016accepted

We propose a new approximation method for Gaussian process (GP) regression based on the mixture of experts structure and variational inference. Our model is essentially an infinite mixture model in which each component is composed of a Gaussian distribution over the input space, and a Gaussian proce…

Cited by 0SourceScholar
2015

Depth image super-resolution using internal and external information

ICASSP 2015accepted

The fast development of 3-D imaging techniques has increased demands for high-resolution depth images. Conventional depth super-resolution methods reconstruct the high-resolution image by accessing high frequency information, either internally from a high-resolution intensity image or externally fro…

Cited by 0SourceScholar
2015

Multi-view indoor scene reconstruction from compressed through-wall radar measurements using a joint bayesian sparse representation

ICASSP 2015accepted

This paper addresses the problem of scene reconstruction, incorporating wall-clutter mitigation, for compressed multi-view through-the-wall radar imaging. We consider the problem where the scene is sensed using different reduced sets of frequencies at different antennas. A joint Bayesian sparse reco…

Cited by 0SourceScholar