← Search

Kim-Hui Yap

22 accepted papers

2026

Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep

CVPR 2026

Diffusion-based video editing has emerged as an important paradigm for high-quality and flexible content generation. However, despite their generality and strong modeling capacity, Diffusion Transformers (DiT) remain computationally expensive due to the iterative denoising process, posing challenges

Cited by 0SourcecodeScholar
2026

KinemaDiff: Towards Diffusion for Coherent and Physically Plausible Human Motion Prediction

ICLR 2026poster

Stochastic Human Motion Prediction (HMP) has become an essential task for the realm of computer vision, for its capacity to anticipate accurate and diverse future human trajectories. Current diffusion-based techniques typically enforce skeletal consistency by encoding structural priors into network…

Cited by 0SourceScholar
2025

A Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with Mamba

ICCV 2025poster

Recent Mamba-based methods for the pose-lifting task tend to model joint dependencies by 2D-to-1D mapping with diverse scanning strategies. Though effective, they struggle to model intricate joint connections and uniformly process all joint motion trajectories while neglecting the intrinsic differen…

Cited by 0SourcePDFScholar
2025

Rectification-specific Supervision and Constrained Estimator for Online Stereo Rectification

CVPR 2025poster

Online stereo rectification is critical for autonomous vehicles and robots in dynamic environments, where factors such as vibration, temperature fluctuations, and mechanical stress can affect rectification accuracy and severely degrade downstream stereo depth estimation. Current dominant approaches…

Cited by 0SourcePDFScholar
2024

CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question Answering

CVPR 2024poster

Diagram Question Answering (DQA) is a challenging task requiring models to answer natural language questions based on visual diagram contexts. It serves as a crucial basis for academic tutoring technical support and more practical applications. DQA poses significant challenges such as the demand for…

Cited by 6SourcePDFScholar
2024

Contextual Human Object Interaction Understanding from Pre-Trained Large Language Model

ICASSP 2024accepted

Existing human object interaction (HOI) detection methods have introduced zero-shot learning techniques to recognize unseen interactions, but they still have limitations in understanding context information and comprehensive reasoning. To overcome these limitations, we propose a novel HOI learning f…

Cited by 0SourceScholar
2024

Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

EMNLP 2024main

In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challen…

2024

Multi-Modality Action Recognition Based on Dual Feature Shift in Vehicle Cabin Monitoring

ICASSP 2024accepted

Driver Action Recognition (DAR) is crucial in vehicle cabin monitoring systems. In real-world applications, it is common for vehicle cabins to be equipped with cameras featuring different modalities. However, multi-modality fusion strategies for the DAR task within car cabins have rarely been studie…

Cited by 0SourceScholar
2023

Bitstream-Corrupted JPEG Images Are Restorable: Two-Stage Compensation and Alignment Framework for Image Restoration

CVPR 2023poster

In this paper, we study a real-world JPEG image restoration problem with bit errors on the encrypted bitstream. The bit errors bring unpredictable color casts and block shifts on decoded image contents, which cannot be trivially resolved by existing image restoration methods mainly relying on pre-de…

2023

Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and Method

NeurIPS 2023poster

The past decade has witnessed great strides in video recovery by specialist technologies, like video inpainting, completion, and error concealment. However, they typically simulate the missing content by manual-designed error masks, thus failing to fill in the realistic video loss in video communica…

2023

TAPS3D: Text-Guided 3D Textured Shape Generation From Pseudo Supervision

CVPR 2023poster

In this paper, we investigate an open research task of generating controllable 3D textured shapes from the given textual descriptions. Previous works either require ground truth caption labeling or extensive optimization time. To resolve these issues, we present a novel framework, TAPS3D, to train a…

2022

Learning Transferable Human-Object Interaction Detector With Natural Language Supervision

CVPR 2022poster

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors…

Cited by 66PDFcodeScholar
2021

Discovering Human Interactions With Large-Vocabulary Objects via Query and Multi-Scale Detection

ICCV 2021poster

In this work, we study the problem of human-object interaction (HOI) detection with large vocabulary object categories. Previous HOI studies are mainly conducted in the regime of limit object categories (e.g., 80 categories). Their solutions may face new difficulties in both object detection and int…

Cited by 32PDFScholar
2020

Discovering Human Interactions With Novel Objects via Zero-Shot Learning

CVPR 2020poster

We aim to detect human interactions with novel objects through zero-shot learning. Different from previous works, we allow unseen object categories by using its semantic word embedding. To do so, we design a human-object region proposal network specifically for the human-object interaction detection…

Cited by 50PDFcodeScholar
2020

Multi-Path Region Mining for Weakly Supervised 3D Semantic Segmentation on Point Clouds

CVPR 2020poster

Point clouds provide intrinsic geometric information and surface context for scene understanding. Existing methods for point cloud segmentation require a large amount of fully labeled data. Using advanced depth sensors, collection of large scale 3D dataset is no longer a cumbersome process. However,…

Cited by 178PDFcodeScholar
2020

What Does Plate Glass Reveal About Camera Calibration?

CVPR 2020poster

This paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image conten…

Cited by 19PDFScholar
2019

The Unusual Effectiveness of Averaging in GAN Training

ICLR 2019poster

We examine two different techniques for parameter averaging in GAN training. Moving Average (MA) computes the time-average of parameters, whereas Exponential Moving Average (EMA) computes an exponentially discounted sum. Whilst MA is known to lead to convergence in bilinear settings, we provide the…

2017

Laplace gradient based Discriminative and Contrast Invertible descriptor

ICASSP 2017accepted

The performance of local descriptors such as SIFT drops under severe illumination changes. In this paper, we propose a Discriminative and Contrast Invertible (DCI) local feature descriptor. In order to increase the discriminative ability of the descriptor under illumination changes, a Laplace gradie…

Cited by 0SourceScholar