← Search

Qiong Cao

10 accepted papers

2025

Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution

AAAI 2025technical

Human preference alignment can significantly enhance the capabilities of Multimodal Large Language Models (MLLMs). However, collecting high-quality preference data remains costly. One promising solution is the self-evolution strategy, where models are iteratively trained on data they generate. Curre…

2024

MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models

IJCAI 2024poster

Foundation models have demonstrated significant emergent abilities, holding great promise for enhancing embodied agents' reasoning and planning capacities. However, the absence of a comprehensive benchmark for evaluating embodied agents with multimodal observations in complex environments remains a…

2024

Towards Variable and Coordinated Holistic Co-Speech Motion Generation

CVPR 2024poster

This paper addresses the problem of generating lifelike holistic co-speech motions for 3D avatars focusing on two key aspects: variability and coordination. Variability allows the avatar to exhibit a wide range of motions even with similar speech content while coordination ensures a harmonious align…

2023

Sharper Bounds for Uniformly Stable Algorithms with Stationary Mixing Process

ICLR 2023poster

Generalization analysis of learning algorithms often builds on a critical assumption that training examples are independently and identically distributed, which is often violated in practical problems such as time series prediction. In this paper, we use algorithmic stability to study the generaliza…

Cited by 5SourcePDFScholar
2023

TriDet: Temporal Action Detection With Relative Boundary Modeling

CVPR 2023poster

In this paper, we present a one-stage framework TriDet for temporal action detection. Existing methods often suffer from imprecise boundary predictions due to the ambiguous action boundaries in videos. To alleviate this problem, we propose a novel Trident-head to model the action boundary via an est…

2022

DearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers

CVPR 2022poster

Transformers have been successfully applied to computer vision due to its powerful modelling capacity with self-attention. However, the good performance of transformers heavily depends on enormous training images. Thus, a data-efficient transformer solution is urgently needed. In this work, we propo…

Cited by 99PDFScholar
2022

ReAct: Temporal Action Detection with Relational Queries

ECCV 2022poster

"This work aims at advancing temporal action detection (TAD) using an encoder-decoder framework with action queries, similar to DETR, which has shown great success in object detection. However, the framework suffers from several problems if directly applied to TAD: the insufficient exploration of in…

2022

View Vertically: A Hierarchical Network for Trajectory Prediction via Fourier Spectrums

ECCV 2022poster

"Understanding and forecasting future trajectories of agents are critical for behavior analysis, robot navigation, autonomous cars, and other related applications. Previous methods mostly treat trajectory prediction as time sequence generation. Different from them, this work studies agents’ trajecto…

2019

MMFace: A Multi-Metric Regression Network for Unconstrained Face Reconstruction

CVPR 2019poster

We propose to address the face reconstruction in the wild by using a multi-metric regression network, MMFace, to align a 3D face morphable model (3DMM) to an input image. The key idea is to utilize a volumetric sub-network to estimate an intermediate geometry representation, and a parametric sub-net…

Cited by 54PDFScholar
2019

Non-Local Recurrent Neural Memory for Supervised Sequence Modeling

ICCV 2019oral

Typical methods for supervised sequence modeling are built upon the recurrent neural networks to capture temporal dependencies. One potential limitation of these methods is that they only model explicitly information interactions between adjacent time steps in a sequence, hence the high-order intera…

Cited by 13PDFcodeScholar