← Search

Huijuan Xu

13 accepted papers

2024

Stratified Avatar Generation from Sparse Observations

CVPR 2024poster

Estimating 3D full-body avatars from AR/VR devices is essential for creating immersive experiences in AR/VR applications. This task is challenging due to the limited input from Head Mounted Devices which capture only sparse observations from the head and hands. Predicting the full-body avatars parti…

Cited by 4SourcePDFScholar
2023

MEID: Mixture-of-Experts with Internal Distillation for Long-Tailed Video Recognition

AAAI 2023technical

The long-tailed video recognition problem is especially challenging, as videos tend to be long and untrimmed, and each video may contain multiple classes, causing frame-level class imbalance. The previous method tackles the long-tailed video recognition only through frame-level sampling for class re…

2022

Disentangled Action Recognition with Knowledge Bases

NAACL 2022long

Action in video usually involves the interaction of human with objects. Action labels are typically composed of various combinations of verbs and nouns, but we may not have training data for all possible combinations. In this paper, we aim to improve the generalization ability of the compositional a…

2022

Syntax Controlled Knowledge Graph-to-Text Generation with Order and Semantic Consistency

NAACL 2022findings

The knowledge graph (KG) stores a large amount of structural knowledge, while it is not easy for direct human understanding. Knowledge graph-to-text (KG-to-text) generation aims to generate easy-to-understand sentences from the KG, and at the same time, maintains semantic consistency between generat…

2022

Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic Actions

ECCV 2022poster

"Action understanding has evolved into the era of fine granularity, as most human behaviors in real life have only minor differences. To detect these fine-grained actions accurately in a label-efficient way, we tackle the problem of weakly-supervised fine-grained temporal action detection in videos…

2021

Meta-Baseline: Exploring Simple Meta-Learning for Few-Shot Learning

ICCV 2021poster

Meta-learning has been the most common framework for few-shot learning in recent years. It learns the model from collections of few-shot classification tasks, which is believed to have a key advantage of making the training objective consistent with the testing objective. However, some recent works…

Cited by 520PDFScholar
2021

Temporal Action Detection With Multi-Level Supervision

ICCV 2021poster

Training temporal action detection in videos requires large amounts of labeled data, yet such annotation is expensive to collect. Incorporating unlabeled or weakly-labeled data to train action detection model could help reduce annotation cost. In this work, we first introduce the Semi-supervised Act…

Cited by 16PDFcodeScholar
2020

Auxiliary Task Reweighting for Minimum-data Learning

NeurIPS 2020poster

Supervised learning requires a large amount of training data, limiting its application where labeled data is scarce. To compensate for data scarcity, one possible method is to utilize auxiliary tasks to provide additional supervision for the main task. Assigning and optimizing the importance weights…

Cited by 39SourcePDFScholar
2020

Learning Canonical Representations for Scene Graph to Image Generation

ECCV 2020poster

Generating realistic images of complex visual scenes becomes challenging when one wishes to control the structure of the generated images. Previous approaches showed that scenes with few entities can be controlled using scene graphs, but this approach struggles as the complexity of the graph (the nu…

2020

Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction Networks

CVPR 2020poster

Human action is naturally compositional: humans can easily recognize and perform actions with objects that are different from those used in training demonstrations. In this paper, we study the compositionality of action by looking into the dynamics of subject-object interactions. We propose a novel…

Cited by 218PDFScholar
2020

Weakly-Supervised Action Localization with Expectation-Maximization Multi-Instance Learning

ECCV 2020poster

Weakly-supervised action localization requires training a model to localize the action segments in the video given only video level action label. It can be solved under the Multiple Instance Learning (MIL) framework, where a bag (video) contains multiple instances (action segments). Since only the b…

2019

Learning Instance Activation Maps for Weakly Supervised Instance Segmentation

CVPR 2019poster

Discriminative region responses residing inside an object instance can be extracted from networks trained with image-level label supervision. However, learning the full extent of pixel-level instance response in a weakly supervised manner remains unexplored. In this work, we tackle this challenging…

Cited by 95PDFScholar