← Search

Cong Phuoc Huynh

8 accepted papers

2026

Composite-Attribute Person Re-Identification via Pose-Guided Disentanglement

CVPR 2026

Recent advancements in vision-language models have enabled multi-modal person re-identification (Re-ID), where the system takes both an image and a text query to identify matching individuals. While previous state-of-the-art methods perform well with detailed, sentence-level descriptions, we found t

Cited by 0SourceScholar
2026

Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

CVPR 2026

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settin

Cited by 0SourceScholar
2025

Beyond Speaker Identity: Text Guided Target Speech Extraction

ICASSP 2025accepted

Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker’s identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided TSE model StyleTSE that uses natural language descriptions of speaking style in…

Cited by 0SourceScholar
2025

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing

EMNLP 2025

We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X util

Cited by 0SourcePDFScholar
2023

LEMaRT: Label-Efficient Masked Region Transform for Image Harmonization

CVPR 2023poster

We present a simple yet effective self-supervised pretraining method for image harmonization which can leverage large-scale unannotated image datasets. To achieve this goal, we first generate pre-training data online with our Label-Efficient Masked Region Transform (LEMaRT) pipeline. Given an image,…

Cited by 23SourcePDFScholar
2023

PADCLIP: Pseudo-labeling with Adaptive Debiasing in CLIP for Unsupervised Domain Adaptation

ICCV 2023poster

Traditional Unsupervised Domain Adaptation (UDA) leverages the labeled source domain to tackle the learning tasks on the unlabeled target domain. It can be more challenging when a large domain gap exists between the source and the target domain. A more practical setting is to utilize a large-scale p…

Cited by 71PDFScholar
2019

Material Segmentation in Hyperspectral Images with a Spatio-spectral Texture Descriptor

ICASSP 2019accepted

In this paper, we address the problem of ground-based hyperspectral image segmentation by combining pixel-level and region-level classification with a region boundary refinement approach. To this end, we represent the spatio-spectral feature of image regions by a descriptor based on Vector of Locall…

Cited by 0SourceScholar