← Search

Beihao Xia

11 accepted papers

2026

TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding

AAAI 2026technical

Spatio-Temporal Video Grounding (STVG) aims to localize a spatio-temporal tube that corresponds to a given language query in an untrimmed video. This is a challenging task since it involves complex vision-language understanding and spatiotemporal reasoning. Recent works have explored weakly-superv

Cited by 0SourcePDFScholar
2025

A Multi-modal Hand Imitation Dataset for Dexterous Hand

IROS 2025

Multimodal data is indispensable for advancing imitation learning, particularly in the context of dexterous hands. However, existing datasets predominantly rely on single-modality inputs, such as RGB images, which inherently lack the capacity to capture the spatial and temporal dynamics essential fo

Cited by 0SourcecodeScholar
2025

Resonance: Learning to Predict Social-Aware Pedestrian Trajectories as Co-Vibrations

ICCV 2025poster

Learning to forecast trajectories of intelligent agents has caught much more attention recently. However, it remains a challenge to accurately account for agents' intentions and social behaviors when forecasting, and in particular, to simulate the unique randomness within each of those components in…

2024

Efficient Backdoor Attacks for Deep Neural Networks in Real-world Scenarios

ICLR 2024poster

Recent deep neural networks (DNNs) have came to rely on vast amounts of training data, providing an opportunity for malicious attackers to exploit and contaminate the data to carry out backdoor attacks. However, existing backdoor attack methods make unrealistic assumptions, assuming that all trainin…

2024

Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels

CVPR 2024poster

This paper focuses on open-ended video question answering which aims to find the correct answers from a large answer set in response to a video-related question. This is essentially a multi-label classification task since a question may have multiple answers. However due to annotation costs the labe…

Cited by 2SourcePDFScholar
2024

SocialCircle: Learning the Angle-based Social Interaction Representation for Pedestrian Trajectory Prediction

CVPR 2024poster

Analyzing and forecasting trajectories of agents like pedestrians and cars in complex scenes has become more and more significant in many intelligent systems and applications. The diversity and uncertainty in socially interactive behaviors among a rich variety of agents make this task more challengi…

2023

TODE-Trans: Transparent Object Depth Estimation with Transformer

ICRA 2023poster

Transparent objects are widely used in industrial automation and daily life. However, robust visual recognition and perception of transparent objects have always been a major challenge. Currently, most commercial-grade depth cameras are still not good at sensing the surfaces of transparent objects d…

Cited by 24SourcecodeScholar
2022

Recent Advances in Concept Drift Adaptation Methods for Deep Learning

IJCAI 2022poster

In the ``Big Data'' age, the amount and distribution of data have increased wildly and changed over time in various time-series-based tasks, e.g weather prediction, network intrusion detection. However, deep learning models may become outdated facing variable input data distribution, which is called…

2022

View Vertically: A Hierarchical Network for Trajectory Prediction via Fourier Spectrums

ECCV 2022poster

"Understanding and forecasting future trajectories of agents are critical for behavior analysis, robot navigation, autonomous cars, and other related applications. Previous methods mostly treat trajectory prediction as time sequence generation. Different from them, this work studies agents’ trajecto…

2021

FREE: Feature Refinement for Generalized Zero-Shot Learning

ICCV 2021poster

Generalized zero-shot learning (GZSL) has achieved significant progress, with many efforts dedicated to overcoming the problems of visual-semantic domain gaps and seen-unseen bias. However, most existing methods directly use feature extraction models trained on ImageNet alone, ignoring the cross-dat…

Cited by 244PDFcodeScholar