← Search

Mingzhen Huang

9 accepted papers

2025

HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models

CVPR 2025poster

We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (e.g. text, objects, partia…

Cited by 1SourcePDFScholar
2025

Your Text Encoder Can Be An Object-Level Watermarking Controller

ICCV 2025poster

Invisible watermarking of AI-generated images can help with copyright protection, enabling detection and identification of AI-generated media. In this work, we present a novel approach to watermark images of T2I Latent Diffusion Models (LDMs). By only fine-tuning text token embeddings \mathcal W _*,…

2024

Exposing Text-Image Inconsistency Using Diffusion Models

ICLR 2024poster

In the battle against widespread online misinformation, a growing problem is text-image inconsistency, where images are misleadingly paired with texts with different intent or meaning. Existing classification-based methods for text-image inconsistency can identify contextual inconsistencies but fail…

2024

ParallelEdits: Efficient Multi-Aspect Text-Driven Image Editing with Attention Grouping

NeurIPS 2024poster

Text-driven image synthesis has made significant advancements with the development of diffusion models, transforming how visual content is generated from text prompts. Despite these advances, text-driven image editing, a key area in computer graphics, faces unique challenges. A major challenge is ma…

Cited by 2SourcePDFScholar
2023

Tracking Multiple Deformable Objects in Egocentric Videos

CVPR 2023poster

Most existing multiple object tracking (MOT) methods that solely rely on appearance features struggle in tracking highly deformable objects. Other MOT methods that use motion clues to associate identities across frames have difficulty handling egocentric videos effectively or efficiently. In this wo…

Cited by 14SourcePDFScholar
2022

Forward Propagation, Backward Regression, and Pose Association for Hand Tracking in the Wild

CVPR 2022poster

We propose HandLer, a novel convolutional architecture that can jointly detect and track hands online in unconstrained videos. HandLer is based on Cascade-RCNNwith additional three novel stages. The first stage is Forward Propagation, where the features from frame t-1 are propagated to frame t based…

Cited by 13PDFcodeScholar
2022

Text-Image De-Contextualization Detection Using Vision-Language Models

ICASSP 2022accepted

Text-image de-contextualization, which uses inconsistent image-text pairs, is an emerging form of misinformation and drawing increasing attention due to the great threat to information authenticity. With real content but semantic mismatch in multiple modalities, the detection of de-contextualization…

Cited by 0SourceScholar
2022

Whose Hands Are These? Hand Detection and Hand-Body Association in the Wild

CVPR 2022poster

We study a new problem of detecting hands and finding the location of the corresponding person for each detected hand. This task is helpful for many downstream tasks such as hand tracking and hand contact estimation. Associating hands with people is challenging in unconstrained conditions since mult…

Cited by 26PDFcodeScholar
2021

Variational Feature Disentangling for Fine-Grained Few-Shot Classification

ICCV 2021poster

Data augmentation is an intuitive step towards solving the problem of few-shot classification. However, ensuring both discriminability and diversity in the augmented samples is challenging. To address this, we propose a feature disentanglement framework that allows us to augment features with random…

Cited by 76PDFcodeScholar