← Search

Wentao Liu

41 accepted papers

2025

AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks

AAAI 2025technical

Automated machine learning (AutoML) is a collection of techniques designed to automate the machine learning development process. While traditional AutoML approaches have been successfully applied in several critical steps of model development (e.g. hyperparameter optimization), there lacks a AutoML…

2025

F-LMM: Grounding Frozen Large Multimodal Models

CVPR 2025poster

Endowing Large Multimodal Models (LMMs) with visual grounding capability can significantly enhance AIs' understanding of the visual world and their interaction with humans. However, existing methods typically fine-tune the parameters of LMMs to learn additional segmentation tokens and overfit ground…

2025

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

ICCV 2025poster

Unifying visual understanding and generation within a single multimodal framework remains a significant challenge, as the two inherently heterogeneous tasks require representations at different levels of granularity. Current approaches that utilize vector quantization (VQ) or variational autoencoder…

2025

NADER: Neural Architecture Design via Multi-Agent Collaboration

CVPR 2025poster

Designing effective neural architectures poses a significant challenge in deep learning. While Neural Architecture Search (NAS) automates the search for optimal architectures, existing methods are often constrained by predetermined search spaces and may miss critical neural architectures. In this pa…

Cited by 0SourcePDFScholar
2025

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries

AAAI 2025technical

Existing research on human-centric video understanding typically focuses on analyzing specific moments or entire videos. However, many applications require higher precision at the frame level. In this work, we propose a novel task, BestShot, which aims to locate highlight frames within human-centric…

2025

Unsupervised Continual Domain Shift Learning with Multi-Prototype Modeling

CVPR 2025highlight

In real-world applications, deep neural networks may encounter constantly changing environments, where the test data originates from continually shifting unlabeled target domains. This problem, known as Unsupervised Continual Domain Shift Learning (UCDSL), poses practical difficulties. Existing meth…

Cited by 0SourcePDFScholar
2024

CLIM: Contrastive Language-Image Mosaic for Region Representation

AAAI 2024technical

Detecting objects accurately from a large or open vocabulary necessitates the vision-language alignment on region representations. However, learning such a region-text alignment by obtaining high-quality box annotations with text labels or descriptions is expensive and infeasible. In contrast, colle…

2024

CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

ICLR 2024spotlight

Open-vocabulary dense prediction tasks including object detection and image segmentation have been advanced by the success of Contrastive Language-Image Pre-training (CLIP). CLIP models, particularly those incorporating vision transformers (ViTs), have exhibited remarkable generalization ability in…

2024

GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition

ECCV 2024poster

"Multi-Label Image Recognition (MLIR) is a challenging task that aims to predict multiple object labels in a single image while modeling the complex relationships between labels and image regions. Although convolutional neural networks and vision transformers have succeeded in processing images as r…

2024

KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension

NeurIPS 2024poster

Recent advancements in Multimodal Large Language Models (MLLMs) have greatly improved their abilities in image understanding. However, these models often struggle with grasping pixel-level semantic details, e.g., the keypoints of an object. To bridge this gap, we introduce the novel challenge of Sem…

Cited by 1SourcePDFScholar
2024

Leveraging Frame Affinity for sRGB-to-RAW Video De-rendering

CVPR 2024poster

Unprocessed RAW video has shown distinct advantages over sRGB video in video editing and computer vision tasks. However capturing RAW video is challenging due to limitations in bandwidth and storage. Various methods have been proposed to address similar issues in single image RAW capture through de-…

Cited by 2SourcePDFScholar
2024

PROGRAM: PROtotype GRAph Model based Pseudo-Label Learning for Test-Time Adaptation

ICLR 2024poster

Test-time adaptation (TTA) aims to adapt a pre-trained model from a source domain to a target domain only using online unlabeled target data during testing, without accessing to the source data or modifying the original training process. Among the various TTA methods, pseudo-labeling has gained popu…

Cited by 30SourcePDFScholar
2024

PhoCoLens: Photorealistic and Consistent Reconstruction in Lensless Imaging

NeurIPS 2024spotlight

Lensless cameras offer significant advantages in size, weight, and cost compared to traditional lens-based systems. Without a focusing lens, lensless cameras rely on computational algorithms to recover the scenes from multiplexed measurements. However, current algorithms struggle with inaccurate for…

Cited by 7SourcePDFScholar
2024

When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset

ECCV 2024poster

"Recent years have witnessed increasing research attention towards pedestrian detection by taking the advantages of different sensor modalities (RGB, IR, Depth, LiDAR and Event). However, designing a unified generalist model that can effectively process diverse sensor modalities remains a challenge.…

2023

Aligning Bag of Regions for Open-Vocabulary Object Detection

CVPR 2023poster

Pre-trained vision-language models (VLMs) learn to align vision and language representations on large-scale datasets, where each image-text pair usually contains a bag of semantic concepts. However, existing open-vocabulary object detectors only align region embeddings individually with the correspo…

2023

Automatic Animation of Hair Blowing in Still Portrait Photos

ICCV 2023poster

We propose a novel approach to animate human hair in a still portrait photo. Existing work has largely studied the animation of fluid elements such as water and fire. However, hair animation for a real image remains underexplored, which is a challenging problem, due to the high complexity of hair st…

Cited by 10PDFcodeScholar
2022

3D Interacting Hand Pose Estimation by Hand De-Occlusion and Removal

ECCV 2022poster

"Estimating 3D interacting hand pose from a single RGB image is essential for understanding human actions. Unlike most previous works that directly predict the 3D poses of two interacting hands simultaneously, we propose to decompose the challenging interacting hand pose estimation task and estimate…

2022

A New Robotic Knee Impedance Control Parameter Optimization Method Facilitated by Inverse Reinforcement Learning

RA-L 2022

Recent efforts in the design of intelligent controllers for configuring robotic prostheses have demonstrated new possibilities in improving mobility and restoring locomotion for individuals with lower-limb disabilities. In these efforts, personalizing the controller of the robotic device is a crucia

Cited by 18SourceScholar
2022

Characterizing Prosthesis Control Fault During Human-Prosthesis Interactive Walking Using Intrinsic Sensors

RA-L 2022

The physical interactions between wearable lower limb robots and humans have been investigated to inform effective robot design for walking augmentation. However, human-robot interactions when internal faults occur within robots have not been systematically reported, but it is essential to improve t

Cited by 8SourceScholar
2022

Inferring Human-Robot Performance Objectives During Locomotion Using Inverse Reinforcement Learning and Inverse Optimal Control

RA-L 2022

Quantitatively characterizing a locomotion performance objective for a human-robot system is an important consideration in the assistive wearable robot design towards human-robot symbiosis. This problem, however, has only been addressed sparsely in the literature. In this study, we propose a new inv

Cited by 20SourceScholar
2022

Multiscale Attention Aggregation Network for 2D Vessel Segmentation

ICASSP 2022accepted

Vessel segmentation is essential for clinical diagnosis and surgical planning. However, it is quite challenging for automatic blood vessel segmentation due to low contrast, complex structure, and variable scale, especially when the annotated data is scarce. In this paper, we propose a novel multisca…

Cited by 0SourceScholar
2022

Not All Tokens Are Equal: Human-Centric Visual Analysis via Token Clustering Transformer

CVPR 2022oral

Vision transformers have achieved great successes in many computer vision tasks. Most methods generate vision tokens by splitting an image into a regular and fixed grid and treating each cell as a token. However, not all regions are equally important in human-centric vision tasks, e.g., the human bo…

Cited by 167PDFcodeScholar
2022

Pose for Everything: Towards Category-Agnostic Pose Estimation

ECCV 2022poster

"Existing works on 2D pose estimation mainly focus on a certain category, e.g. human, animal, and vehicle. However, there are lots of application scenarios that require detecting the poses/keypoints of the unseen class of objects. In this paper, we introduce the task of Category-Agnostic Pose Estima…

2022

PoseTrans: A Simple yet Effective Pose Transformation Augmentation for Human Pose Estimation

ECCV 2022poster

"Human pose estimation aims to accurately estimate a wide variety of human poses. However, existing datasets often follow a long-tailed distribution that unusual poses only occupy a small portion, which further leads to the lack of diversity of rare poses. These issues result in the inferior general…

2022

Pseudo-Labeled Auto-Curriculum Learning for Semi-Supervised Keypoint Localization

ICLR 2022poster

Localizing keypoints of an object is a basic visual problem. However, supervised learning of a keypoint localization network often requires a large amount of data, which is expensive and time-consuming to obtain. To remedy this, there is an ever-growing interest in semi-supervised learning (SSL), wh…

Cited by 20SourcePDFScholar
2022

Reinforcement Learning Impedance Control of a Robotic Prosthesis to Coordinate With Human Intact Knee Motion

RA-L 2022

This study aims to demonstrate reinforcement learning tracking control for automatically configuring the impedance parameters of a robotic knee prosthesis. While our previous studies involving human subjects have focused on tuning the impedance control parameters to meet a fixed, subjectively prescr

Cited by 28SourceScholar
2021

Graph-Based 3D Multi-Person Pose Estimation Using Multi-View Images

ICCV 2021poster

This paper studies the task of estimating the 3D human poses of multiple persons from multiple calibrated camera views. Following the top-down paradigm, we decompose the task into two stages, i.e. person localization and pose estimation. Both stages are processed in coarse-to-fine manners. And we pr…

Cited by 66PDFcodeScholar
2021

Human Pose Regression With Residual Log-Likelihood Estimation

ICCV 2021poster

Heatmap-based methods dominate in the field of human pose estimation by modelling the output distribution through likelihood heatmaps. In contrast, regression-based methods are more efficient but suffer from inferior performance. In this work, we explore maximum likelihood estimation (MLE) to develo…

Cited by 277PDFcodeScholar
2021

Joint Depth and Normal Estimation from Real-world Time-of-flight Raw Data

IROS 2021poster

We present a novel approach to joint depth and normal estimation for time-of-flight (ToF) sensors. Our model learns to predict the high-quality depth and normal maps jointly from ToF raw sensor data. To achieve this, we meticulously constructed the first large-scale dataset (named ToF-100) with pair…

Cited by 4SourceScholar
2021

ViPNAS: Efficient Video Pose Estimation via Neural Architecture Search

CVPR 2021poster

Human pose estimation has achieved significant progress in recent years. However, most of the recent methods focus on improving accuracy using complicated models and ignoring real-time efficiency. To achieve a better trade-off between accuracy and efficiency, we propose a novel neural architecture s…

Cited by 74PDFcodeScholar
2021

When Human Pose Estimation Meets Robustness: Adversarial Algorithms and Benchmarks

CVPR 2021poster

Human pose estimation is a fundamental yet challenging task in computer vision, which aims at localizing human anatomical keypoints. However, unlike human vision that is robust to various data corruptions such as blur and pixelation, current pose estimators are easily confused by these corruptions.…

Cited by 82PDFcodeScholar
2020

Differentiable Hierarchical Graph Grouping for Multi-Person Pose Estimation

ECCV 2020poster

Multi-person pose estimation is challenging because it localizes body keypoints for multiple persons simultaneously. Previous methods can be divided into two streams, \ie top-down and bottom-up methods. The top-down methods localize keypoints after human detection, while the bottom-up methods locali…

2020

HMOR: Hierarchical Multi-Person Ordinal Relations for Monocular Multi-Person 3D Pose Estimation

ECCV 2020poster

Remarkable progress has been made in 3D human pose estimation from a monocular RGB camera. However, only a few studies explored 3D multi-person cases. In this paper, we attempt to address the lack of a global perspective of the top-down approaches by introducing a novel form of supervision - Hierarc…

Cited by 75SourcePDFScholar
2020

Omni-sourced Webly-supervised Learning for Video Recognition

ECCV 2020poster

We introduce OmniSource, a novel framework for leveraging web data to train video recognition models. OmniSource overcomes the barriers between data formats, such as images, short videos, and long untrimmed videos for webly-supervised learning. First, data samples with multiple formats, curated by t…

2020

SMAP: Single-Shot Multi-Person Absolute 3D Pose Estimation

ECCV 2020poster

Recovering multi-person 3D poses with absolute scales from a single RGB image is a challenging problem due to the inherent depth and scale ambiguity from a single view. Addressing this ambiguity requires to aggregate various cues over the entire image, such as body sizes, scene layouts, and inter-pe…

Cited by 132SourcePDFScholar
2020

Whole-Body Human Pose Estimation in the Wild

ECCV 2020poster

This paper investigates the task of 2D human whole-body pose estimation, which aims to localize dense landmarks on the entire human body including face, hands, body, and feet. As existing datasets do not have whole-body annotations, previous methods have to assemble different deep models trained ind…

2019

TRB: A Novel Triplet Representation for Understanding 2D Human Body

ICCV 2019oral

Human pose and shape are two important components of 2D human body. However, how to efficiently represent both of them in images is still an open question. In this paper, we propose the Triplet Representation for Body (TRB) --- a compact 2D human body representation, with skeleton keypoints capturin…

Cited by 20PDFcodeScholar
2019

Weakly-Supervised Discovery of Geometry-Aware Representation for 3D Human Pose Estimation

CVPR 2019oral

Recent studies have shown remarkable advances in 3D human pose estimation from monocular images, with the help of large-scale in-door 3D datasets and sophisticated network architectures. However, the generalizability to different environments remains an elusive goal. In this work, we propose a geome…

Cited by 139PDFScholar
2018

Person Search in Videos with One Portrait Through Visual and Temporal Links

ECCV 2018poster

In real-world applications, e.g. law enforcement and video retrieval, one often needs to search a certain person in long videos with just one portrait. This is much more challenging than the conventional settings for person re-identification, as the search may need to be carried out in an environmen…