← Search

Xiaojun Hou

6 accepted papers

2025

CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion

ICASSP 2025accepted

Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand in today’s video proliferation era. Multi-modal video summarization that accomodates user input has become a research hot…

Cited by 0SourceScholar
2025

LITE: A Learning-Integrated Topological Explorer for Multi-Floor Indoor Environments

IROS 2025

This work focuses on multi-floor indoor exploration, which remains an open area of research. Compared to traditional methods, recent learning-based explorers have demonstrated significant potential due to their robust environmental learning and modeling capabilities, but most are restricted to 2D en

Cited by 0SourceScholar
2024

A Robotic-centric Paradigm for 3D Human Tracking Under Complex Environments Using Multi-modal Adaptation

IROS 2024poster

The goal of this paper is to strike a feasible tracking paradigm that can make 3D human trackers applicable on robot platforms and enable more high-level tasks. Till now, two fundamental problems haven’t been adequately addressed. One is the computational cost lightweight enough for robotic deployme…

Cited by 0SourceScholar
2024

Multi-modal 3D Human Tracking for Robots in Complex Environment with Siamese Point-Video Transformer

ICRA 2024poster

Tracking a specific person in 3D scene is gaining momentum due to its numerous applications in robotics. Currently, most 3D trackers focus on driving scenarios with neglected jitter and uncomplicated surroundings, which results in their severe degeneration in complex environments, especially on jolt…

Cited by 3SourceScholar
2024

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

CVPR 2024poster

Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers which was inefficient and lacked generalized representation due to the scarcity of multimodal data. Therefore recent studies have ut…

2023

PANet: LiDAR Panoptic Segmentation with Sparse Instance Proposal and Aggregation

IROS 2023poster

Reliable LiDAR panoptic segmentation (LPS), including both semantic and instance segmentation, is vital for many robotic applications, such as autonomous driving. This work proposes a new LPS framework named PANet to eliminate the dependency on the offset branch and improve the performance on large…

Cited by 5SourcecodeScholar