← Search

Minh N. Do

23 accepted papers

2025

Care-PD: A Multi-Site Anonymized Clinical Dataset for Parkinson’s Disease Gait Assessment

NeurIPS 2025poster

Objective gait assessment in Parkinson’s Disease (PD) is limited by the absence of large, diverse, and clinically annotated motion datasets. We introduce Care-PD, the largest publicly available archive of 3D mesh gait data for PD, and the first multi-site collection spanning 9 cohorts from 8 clinica…

Cited by 0SourceScholar
2025

Robult: Leveraging Redundancy and Modality-Specific Features for Robust Multimodal Learning

IJCAI 2025

Addressing missing modalities and limited labeled data is crucial for advancing robust multimodal learning. We propose Robult, a scalable framework designed to mitigate these challenges by preserving modality-specific information and leveraging redundancy through a novel information-theoretic approa

Cited by 0SourcePDFScholar
2024

Making Vision Transformers Truly Shift-Equivariant

CVPR 2024poster

In the field of computer vision Vision Transformers (ViTs) have emerged as a prominent deep learning architecture. Despite being inspired by Convolutional Neural Networks (CNNs) ViTs are susceptible to small spatial shifts in the input data - they lack shift-equivariance. To address this shortcoming…

Cited by 9SourcePDFScholar
2024

MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation

AAAI 2024technical

Few-shot instance segmentation extends the few-shot learning paradigm to the instance segmentation task, which tries to segment instance objects from a query image with a few annotated examples of novel categories. Conventional approaches have attempted to address the task via prototype learning, kn…

2024

Persistent Test-time Adaptation in Recurring Testing Scenarios

NeurIPS 2024poster

Current test-time adaptation (TTA) approaches aim to adapt a machine learning model to environments that change continuously. Yet, it is unclear whether TTA methods can maintain their adaptability over prolonged periods. To answer this question, we introduce a diagnostic setting - **recurring TTA**…

2022

Learnable Polyphase Sampling for Shift Invariant and Equivariant Convolutional Networks

NeurIPS 2022accept

We propose learnable polyphase sampling (LPS), a pair of learnable down/upsampling layers that enable truly shift-invariant and equivariant convolutional networks. LPS can be trained end-to-end from data and generalizes existing handcrafted downsampling layers. It is widely applicable as it can be i…

2019

Dense 3D Reconstruction for Visual Tunnel Inspection using Unmanned Aerial Vehicle

IROS 2019poster

Advances in Unmanned Aerial Vehicle (UAV) opens venues for application such as tunnel inspection. Owing to its versatility to fly inside the tunnels, it can quickly identify defects and potential problems related to safety. However, long tunnels, especially with repetitive or uniform structures pose…

Cited by 18SourceScholar
2019

Learning Motion in Feature Space: Locally-Consistent Deformable Convolution Networks for Fine-Grained Action Detection

ICCV 2019oral

Fine-grained action detection is an important task with numerous applications in robotics and human-computer interaction. Existing methods typically utilize a two-stage approach including extraction of local spatio-temporal features followed by temporal modeling to capture long-term dependencies. Wh…

Cited by 46PDFcodeScholar
2018

Image Restoration with Deep Generative Models

ICASSP 2018accepted

Many image restoration problems are ill-posed in nature, hence, beyond the input image, most existing methods rely on a carefully engineered image prior, which enforces some local image consistency in the recovered image. How tightly the prior assumptions are fulfilled has a big impact on the result…

Cited by 0SourceScholar
2018

Time-Frequency Networks for Audio Super-Resolution

ICASSP 2018accepted

Audio super-resolution (a.k.a. bandwidth extension) is the challenging task of increasing the temporal resolution of audio signals. Recent deep networks approaches achieved promising results by modeling the task as a regression problem in either time or frequency domain. In this paper, we introduced…

Cited by 0SourceScholar
2017

Semantic Image Inpainting With Deep Generative Models

CVPR 2017poster

Semantic image inpainting is a challenging task where large missing regions have to be filled based on the available visual data. Existing methods which extract information from only a single image generally produce unsatisfactory results due to the lack of high level context. In this paper, we pro…

Cited by 1485PDFScholar
2015

DASC: Dense Adaptive Self-Correlation Descriptor for Multi-Modal and Multi-Spectral Correspondence

CVPR 2015poster

Establishing dense visual correspondence between multiple images is a fundamental task in many applications of computer vision and computational photography. Classical approaches, which aim to estimate dense stereo and optical flow fields for images adjacent in viewpoint or in time, have been dramat…

Cited by 119SourcePDFScholar
2015

PatchMatch-Based Automatic Lattice Detection for Near-Regular Textures

ICCV 2015poster

In this work, we investigate the problem of automatically inferring the lattice structure of near-regular textures (NRT) in real-world images. Our technique leverages the PatchMatch algorithm for finding k-nearest-neighbor (kNN) correspondences in an image. We use these kNNs to recover an initial es…

Cited by 28PDFScholar
2015

SPM-BP: Sped-up PatchMatch Belief Propagation for Continuous MRFs

ICCV 2015oral

Markov random fields are widely used to model many computer vision problems that can be cast in an energy minimization framework composed of unary and pairwise potentials. While computationally tractable discrete optimizers such as Graph Cuts and belief propagation (BP) exist for multi-label discret…

Cited by 111PDFScholar