← Search

Thanh-Toan Do

33 accepted papers

2026

Align-SAM: Seeking Flatter Minima for Better Cross-Subset Alignment

ICLR 2026poster

Sharpness-Aware Minimization (SAM) has proven effective in enhancing deep neural network training by simultaneously minimizing the training loss and the sharpness of the loss landscape, thereby guiding models toward flatter minima that are empirically linked to improved generalization. From another…

Cited by 0SourceScholar
2026

Beyond Uniformity: Sample and Frequency Meta Weighting for Post-Training Quantization of Diffusion Models

ICLR 2026poster

Post-training quantization (PTQ) is an attractive approach for compressing diffusion models to speed up the sampling process and reduce the memory footprint. Most existing PTQ methods uniformly sample data from various time steps in the denoising process to construct a calibration set for quantizati…

Cited by 0SourceScholar
2026

Coverage-Constrained Human-AI Cooperation with Multiple Experts

AAAI 2026technical

Human-AI cooperative classification (HAI-CC) aims to develop hybrid intelligent systems that enhance decision-making in various high-stakes real-world scenarios by leveraging both human expertise and AI capabilities. Current HAI-CC methods primarily focus on learning-to-defer (L2D), where decisions

Cited by 0SourcePDFScholar
2026

Gradient-Aligned Calibration for Post-Training Quantization of Diffusion Models

ICLR 2026poster

Diffusion models have shown remarkable performance in image synthesis by progressively estimating a smooth transition from a Gaussian distribution of noise to a real image. Unfortunately, their practical deployment is limited by slow inference speed, high memory usage, and the computational demands…

Cited by 0SourceScholar
2026

Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization

ICLR 2026poster

Direct Preference Optimization (DPO) has emerged as a popular algorithm for aligning pretrained large language models with human preferences, owing to its simplicity and training stability. However, DPO suffers from the recently identified squeezing effect (also known as likelihood displacement), wh…

Cited by 0SourcecodeScholar
2025

Enhancing Dataset Distillation via Non-Critical Region Refinement

CVPR 2025poster

Dataset distillation has gained popularity as a technique for compressing large datasets into smaller, more efficient representations while retaining essential information for model training. Data features can be broadly divided into two types: instance-specific features, which capture unique, fine…

2025

Geometry-Aware Collaborative Multi-Solutions Optimizer for Model Fine-Tuning with Parameter Efficiency

NeurIPS 2025poster

We propose a framework grounded in gradient flow theory and informed by geometric structure that provides multiple diverse solutions for a given task, ensuring collaborative results that enhance performance and adaptability across different tasks. This framework enables flexibility, allowing for eff…

Cited by 0SourceScholar
2025

MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora

EMNLP 2025

Continually updating model-based indexes in generative retrieval with new documents remains challenging, as full retraining is computationally expensive and impractical under resource constraints. We propose MixLoRA-DSI, a novel framework that combines an expandable mixture of Low-Rank Adaptation ex

2025

Multi-Perspective Data Augmentation for Few-shot Object Detection

ICLR 2025poster

Recent few-shot object detection (FSOD) methods have focused on augmenting synthetic samples for novel classes, show promising results to the rise of diffusion models. However, the diversity of such datasets is often limited in representativeness because they lack awareness of typical and hard sam…

2025

Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation

CVPR 2025poster

Recent approaches leveraging multi-modal pre-trained models like CLIP for Unsupervised Domain Adaptation (UDA) have shown significant promise in bridging domain gaps and improving generalization by utilizing rich semantic knowledge and robust visual representations learned through extensive pre-trai…

Cited by 0SourcePDFScholar
2025

Probabilistic Learning to Defer: Handling Missing Expert Annotations and Controlling Workload Distribution

ICLR 2025oral

Recent progress in machine learning research is gradually shifting its focus towards *human-AI cooperation* due to the advantages of exploiting the reliability of human experts and the efficiency of AI models. One of the promising approaches in human-AI cooperation is *learning to defer* (L2D), wher…

Cited by 0SourcePDFScholar
2025

Self-supervised Learning for Acoustic Few-Shot Classification

ICASSP 2025accepted

Labelled data are limited and self-supervised learning is one of the most important approaches for reducing labelling requirements. While it has been extensively explored in the image domain, it has so far not received the same amount of attention in the acoustic domain. Yet, reducing labelling is a…

Cited by 0SourceScholar
2024

Instance-dependent Noisy-label Learning with Graphical Model Based Noise-rate Estimation

ECCV 2024poster

"Deep learning faces a formidable challenge when handling noisy labels, as models tend to overfit samples affected by label noise. This challenge is further compounded by the presence of instance-dependent noise (IDN), a realistic form of label noise arising from ambiguous sample information. To add…

2024

Learning to Complement and to Defer to Multiple Users

ECCV 2024poster

"With the development of Human-AI Collaboration in Classification (HAI-CC), integrating users and AI predictions becomes challenging due to the complex decision-making process. This process has three options: 1) AI autonomously classifies, 2) learning to complement, where AI collaborates with users,…

2024

MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation

AAAI 2024technical

Few-shot instance segmentation extends the few-shot learning paradigm to the instance segmentation task, which tries to segment instance objects from a query image with a few annotated examples of novel categories. Conventional approaches have attempted to address the task via prototype learning, kn…

2024

MetaAug: Meta-Data Augmentation for Post-Training Quantization

ECCV 2024poster

"Post-Training Quantization (PTQ) has received significant attention because it requires only a small set of calibration data to quantize a full-precision model, which is more practical in real-world applications in which full access to a large training set is not available. However, it often leads…

2024

Sharpness-Aware Data Generation for Zero-shot Quantization

ICML 2024poster

Zero-shot quantization aims to learn a quantized model from a pre-trained full-precision model with no access to original real training data. The common idea in zero-shot quantization approaches is to generate synthetic data for quantizing the full-precision model. While it is well-known that deep n…

Cited by 0SourcePDFScholar
2023

Flat Seeking Bayesian Neural Networks

NeurIPS 2023poster

Bayesian Neural Networks (BNNs) provide a probabilistic interpretation for deep learning models by imposing a prior distribution over model parameters and inferring a posterior distribution based on observed data. The model sampled from the posterior distribution can be used for providing ensemble p…

Cited by 9SourcePDFScholar
2023

Model and Feature Diversity for Bayesian Neural Networks in Mutual Learning

NeurIPS 2023poster

Bayesian Neural Networks (BNNs) offer probability distributions for model parameters, enabling uncertainty quantification in predictions. However, they often underperform compared to deterministic neural networks. Utilizing mutual learning can effectively enhance the performance of peer BNNs. In thi…

Cited by 4SourcePDFScholar
2023

Optimal Transport Model Distributional Robustness

NeurIPS 2023poster

Distributional robustness is a promising framework for training deep learning models that are less vulnerable to adversarial examples and data distribution shifts. Previous works have mainly focused on exploiting distributional robustness in the data space. In this work, we explore an optimal transp…

2022

A4T: Hierarchical Affordance Detection for Transparent Objects Depth Reconstruction and Manipulation

RA-L 2022

Transparent objects are widely used in our daily lives and therefore robots need to be able to handle them. However, transparent objects suffer from light reflection and refraction, which makes it challenging to obtain the accurate depth maps required to perform handling tasks. In this letter, we pr

Cited by 40SourceScholar
2022

Logic Rules Meet Deep Learning: A Novel Approach for Ship Type Classification (Extended Abstract)

IJCAI 2022poster

The shipping industry is an important component of the global trade and economy. In order to ensure law compliance and safety, it needs to be monitored. In this paper, we present a novel ship type classification model that combines vessel transmitted data from the Automatic Identification System, wi…

Cited by 0SourcePDFScholar
2020

Direct Quantization for Training Highly Accurate Low Bit-width Deep Neural Networks

IJCAI 2020poster

This paper proposes two novel techniques to train deep convolutional neural networks with low bit-width weights and activations. First, to obtain low bit-width weights, most existing methods obtain the quantized weights by performing quantization on the full-precision network weights. However, this…

2019

A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric Learning

CVPR 2019poster

We propose a method that substantially improves the efficiency of deep distance metric learning based on the optimization of the triplet loss function. One epoch of such training process based on a na"ive optimization of the triplet loss function has a run-time complexity O(N^3), where N is the numb…

Cited by 77PDFScholar
2019

Compact Trilinear Interaction for Visual Question Answering

ICCV 2019poster

In Visual Question Answering (VQA), answers have a great correlation with question meaning and visual contents. Thus, to selectively utilize image, question and answer information, we propose a novel trilinear interaction model which simultaneously learns high level associations between these three…

Cited by 91PDFcodeScholar
2019

SDRSAC: Semidefinite-Based Randomized Approach for Robust Point Cloud Registration Without Correspondences

CVPR 2019oral

This paper presents a novel randomized algorithm for robust point cloud registration without correspondences. Most existing registration approaches require a set of putative correspondences obtained by extracting invariant descriptors. However, such descriptors could become unreliable in noisy and c…

Cited by 114PDFcodeScholar
2019

Scalable Place Recognition Under Appearance Change for Autonomous Driving

ICCV 2019oral

A major challenge in place recognition for autonomous driving is to be robust against appearance changes due to short-term (e.g., weather, lighting) and long-term (seasons, vegetation growth, etc.) environmental variations. A promising solution is to continuously accumulate images to maintain an ade…

Cited by 82PDFScholar
2018

AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection

ICRA 2018poster

We propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an object detection branch to localize and classify the object, and an affordance detection branch to assign each pixel in the o…

Cited by 366SourcecodeScholar
2018

Bayesian Semantic Instance Segmentation in Open Set World

ECCV 2018poster

This paper addresses the semantic instance segmentation task in the open-set conditions, where input images can contain known and unknown object classes. The training process of existing semantic instance segmentation methods requires annotation masks for all object instances, which is expensive to…

2018

SceneCut: Joint Geometric and Object Segmentation for Indoor Scenes

ICRA 2018poster

This paper presents SceneCut, a novel approach to jointly discover previously unseen objects and non-object surfaces using a single RGB-D image. SceneCut's joint reasoning over scene semantics and geometry allows a robot to detect and segment object instances in complex scenes where modern deep lear…

Cited by 50SourceScholar
2017

Simultaneous Feature Aggregating and Hashing for Large-Scale Image Search

CVPR 2017poster

In most state-of-the-art hashing-based visual search systems, local image descriptors of an image are first aggregated as a single feature vector. This feature vector is then subjected to a hashing function that produces a binary hash code. In previous work, the aggregating and the hashing processes…

Cited by 40PDFScholar
2015

FAemb: A Function Approximation-Based Embedding Method for Image Retrieval

CVPR 2015poster

The objective of this paper is to design an embedding method mapping local features describing image (e.g. SIFT) to a higher dimensional representation used for image retrieval problem. By investigating the relationship between the linear approximation of a nonlinear function in high dimensional spa…

Cited by 37SourcePDFScholar