← Search

Yongkang Wong

18 accepted papers

2026

MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer

CVPR 2026

3D pose transfer aims to transfer the pose-style of a source mesh to a target character while preserving both the target's geometry and the source's pose characteristic. Existing methods are largely restricted to characters with similar structures and fail to generalize to category-free settings (e.

Cited by 0SourceScholar
2026

Mitigating Noise-Induced Layout Priors for Object Counting in Diffusion Models

ICML 2026poster

Despite remarkable progress in text-to-image diffusion models, accurately generating the specified number of objects remains a persistent challenge. We identify the initial noise as a primary determinant of spatial layout formation, with early-stage cross-attention serving as the key mechanism that …

Cited by 0SourceScholar
2026

Object-Centric Framework for Video Moment Retrieval

AAAI 2026technical

Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing mo

Cited by 0SourcePDFScholar
2024

Finetuning Text-to-Image Diffusion Models for Fairness

ICLR 2024oral

The rapid adoption of text-to-image diffusion models in society underscores an urgent need to address their biases. Without interventions, these biases could propagate a skewed worldview and restrict opportunities for minority groups. In this work, we frame fairness as a distributional alignment pro…

2024

Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition Learning

ICML 2024poster

Previous efforts using frozen Large Language Models (LLMs) for visual understanding, via image captioning or image-text retrieval tasks, face challenges when dealing with complex multimodal scenarios. In order to enhance the capabilities of Multimodal Large Language Models (MLLM) in comprehending th…

2024

MCM: Multi-condition Motion Synthesis Framework

IJCAI 2024poster

Conditional human motion synthesis (HMS) aims to generate human motion sequences that conform to specific conditions. Text and audio represent the two predominant modalities employed as HMS control conditions. While existing research has primarily focused on single conditions, the multi-condition hu…

Cited by 1SourcePDFScholar
2024

TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

NeurIPS 2024spotlight

Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability of substantial web video-text data. This difficulty primarily arises from the inherent complexity of videos and the inef…

2022

Chairs Can Be Stood On: Overcoming Object Bias in Human-Object Interaction Detection

ECCV 2022poster

"Detecting Human-Object Interaction (HOI) in images is an important step towards high-level visual comprehension. Existing work often shed light on improving either human and object detection, or interaction recognition. However, due to the limitation of datasets, these methods tend to fit well on f…

2022

Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action Segmentation

NeurIPS 2022accept

We propose Differentiable Temporal Logic (DTL), a model-agnostic framework that introduces temporal constraints to deep networks. DTL treats the outputs of a network as a truth assignment of a temporal logic formula, and computes a temporal logic loss reflecting the consistency between the output an…

Cited by 40SourcePDFScholar
2021

Learning Causal Representation for Training Cross-Domain Pose Estimator via Generative Interventions

ICCV 2021poster

3D pose estimation has attracted increasing attention with the availability of high-quality benchmark datasets. However, prior works show that deep learning models tend to learn spurious correlations, which fail to generalize beyond the specific dataset they are trained on. In this work, we take a s…

Cited by 39PDFScholar
2021

Learning to Predict Trustworthiness with Steep Slope Loss

NeurIPS 2021poster

Understanding the trustworthiness of a prediction yielded by a classifier is critical for the safe and effective use of AI models. Prior efforts have been proven to be reliable on small-scale datasets. In this work, we study the problem of predicting trustworthiness on real-world large-scale dataset…

2021

Unsupervised Motion Representation Learning with Capsule Autoencoders

NeurIPS 2021poster

We propose the Motion Capsule Autoencoder (MCAE), which addresses a key challenge in the unsupervised learning of motion representations: transformation invariance. MCAE models motion in a two-level hierarchy. In the lower level, a spatio-temporal motion signal is divided into short, local, and sema…

2020

n-Reference Transfer Learning for Saliency Prediction

ECCV 2020poster

Benefiting from deep learning research and large-scale datasets, saliency prediction has achieved significant success in the past decade. However, it still remains challenging to predict saliency maps on images in new domains that lack sufficient data for data-hungry models. To solve this problem, w…

2019

Learning to Detect Human-Object Interactions With Knowledge

CVPR 2019poster

The recent advances in instance-level detection tasks lay a strong foundation for automated visual scenes understanding. However, the ability to fully comprehend a social scene still eludes us. In this work, we focus on detecting human-object interactions (HOIs) in images, an essential step towards…

Cited by 191PDFScholar
2018

Unsupervised Learning of View-invariant Action Representations

NeurIPS 2018poster

The recent success in human action recognition with deep learning methods mostly adopt the supervised learning paradigm, which requires significant amount of manually labeled data to achieve good performance. However, label collection is an expensive and time-consuming process. In this work, we prop…

Cited by 136SourcePDFScholar