← Search

Qingqiu Huang

19 accepted papers

2026

From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

ICML 2026poster

Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action. Existing generative policies typically adopt a "Generation-from-Noise" paradig…

Cited by 0SourceScholar
2026

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

CVPR 2026

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak at spatial grounding and understanding, and VLA systems bui

Cited by 0SourceScholar
2025

STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation

IROS 2025

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature misalignment due to inadequate decoupling of spatio-temporal dynami

Cited by 5SourcecodeScholar
2023

CLIP2: Contrastive Language-Image-Point Pretraining From Real-World Point Cloud Data

CVPR 2023poster

Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. However, due to the limited Text-3D data pairs, adapting the success of 2D Vision-Language Models (VLM) to the 3D space remain…

Cited by 107SourcePDFScholar
2023

PARTNER: Level up the Polar Representation for LiDAR 3D Object Detection

ICCV 2023poster

Recently, polar-based representation has shown promising properties in perceptual tasks. In addition to Cartesian-based approaches, which separate point clouds unevenly, representing point clouds as polar grids has been recognized as an alternative due to (1) its advantage in robust performance unde…

Cited by 10PDFcodeScholar
2022

TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection With Transformers

CVPR 2022poster

LiDAR and camera are two important sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, the robustness against inferior image conditions, e.g., bad illumination and sensor misalignment, is under-explored. Existing fusion methods are…

Cited by 806PDFcodeScholar
2020

A Local-to-Global Approach to Multi-Modal Movie Scene Segmentation

CVPR 2020poster

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of movies. This is very challenging - compared to the videos st…

Cited by 155PDFcodeScholar
2020

A Unified Framework for Shot Type Classification Based on Subject Centric Lens

ECCV 2020poster

In film making, shot has a profound influence on how the story is delivered and how the audiences are echoed. As different scale and movement types of shots can express different emotions and contents, recognizing shots and their attributes is important to the understanding of movies as well as gene…

Cited by 85SourcePDFScholar
2020

Caption-Supervised Face Recognition: Training a State-of-the-Art Face Model without Manual Annotation

ECCV 2020poster

The advances over the past several years have pushed the performance of face recognition to an amazing level. This great success, to a large extent, is built on top of millions of annotated samples. However, as we endeavor to take the performance to the next level, the reliance on annotated data bec…

Cited by 28SourcePDFScholar
2020

Distribution-Balanced Loss for Multi-Label Classification in Long-Tailed Datasets

ECCV 2020poster

We present a new loss function called Distribution-Balanced Loss for the multi-label recognition problems that exhibit long-tailed class distributions. Compared to conventional single-label classification problem, multi-label recognition problems are often more challenging due to two significant iss…

2020

Learn to Propagate Reliably on Noisy Affinity Graphs

ECCV 2020poster

Recent works have shown that exploiting unlabeled data through label propagation can substantially reduce the labeling cost, which has been a critical issue in developing visual recognition models. Yet, how to propagate labels reliably, especially on a dataset with unknown outliers, remains an open…

Cited by 16SourcePDFScholar
2020

MovieNet: A Holistic Dataset for Movie Understanding

ECCV 2020poster

Recent years have seen remarkable advances in visual understanding. However, how to understand a story-based long video with artistic styles, e.g. movie, remains challenging. In this paper, we introduce MovieNet -- a holistic dataset for movie understanding. MovieNet contains 1,100 movies with a lar…

2020

Online Multi-modal Person Search in Videos

ECCV 2020poster

The task of searching certain people in videos has seen increasing potential in real-world applications, such as video organization and editing. Most existing approaches are devised to work in an offline manner, where identifies can only be inferred after an entire video is examined. This working ma…

Cited by 35SourcePDFScholar
2020

Placepedia: Comprehensive Place Understanding with Multi-Faceted Annotations

ECCV 2020poster

Place is an important element in visual understanding. Given a photo of a building, people can often tell its functionality, e.g. a restaurant or a shop, its cultural style, e.g. Asian or European, as well as its economic type, e.g. industry oriented or tourism oriented. While place recognition has…

Cited by 8SourcePDFScholar
2019

A Graph-Based Framework to Bridge Movies and Synopses

ICCV 2019oral

Inspired by the remarkable advances in video analytics, research teams are stepping towards a greater ambition - movie understanding. However, compared to those activity videos in conventional datasets, movies are significantly different. Generally, movies are much longer and consist of much richer…

Cited by 78PDFcodeScholar
2018

Find and Focus: Retrieve and Localize Video Events with Natural Language Queries

ECCV 2018poster

The thriving of video sharing services brings new challenges to video retrieval, e.g. the rapid growth in video duration and content diversity. Meeting such challenges calls for new techniques that can effectively retrieve videos with natural language queries. Existing methods along this line, which…

Cited by 90SourcePDFScholar
2018

Person Search in Videos with One Portrait Through Visual and Temporal Links

ECCV 2018poster

In real-world applications, e.g. law enforcement and video retrieval, one often needs to search a certain person in long videos with just one portrait. This is much more challenging than the conventional settings for person re-identification, as the search may need to be carried out in an environmen…