← Search

Xinglong Xu

5 accepted papers

2026

GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models

CVPR 2026

Unified Multimodal Models (UMMs) are redefining the landscape of artificial intelligence by coupling perception and generation across language, vision, and structured reasoning. Yet, despite their growing sophistication, a critical gap persists in evaluation: existing benchmarks largely measure disc

Cited by 0SourceScholar
2024

Exploiting Multi-Modal Synergies for Enhancing 3D Multi-Object Tracking

RA-L 2024

3D Multi-Object Tracking (MOT) aims to establish and maintain consistent object trajectories in continuously dynamic environments. At present, the tracking-by-detection has emerged as a dominant paradigm for 3D MOT, due to its simplicity and efficiency. However, this paradigm depends heavily on the

Cited by 5SourceScholar
2024

GroupTrack: Multi-Object Tracking by Using Group Motion Patterns

IROS 2024poster

The main challenge of Multi-Object Tracking (MOT) lies in maintaining a distinctive identity for each target in dense crowds or occluded scenarios. Although the existing methods have achieved significantly progress by using robust object detectors or complex association strategies, they cannot effec…

Cited by 0SourceScholar
2024

Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation

ECCV 2024poster

"Skeleton-based Temporal Action Segmentation (STAS) aims to densely segment and classify human actions in long, untrimmed skeletal motion sequences. Existing STAS methods primarily model spatial dependencies among joints and temporal relationships among frames to generate frame-level one-hot classif…

2024

MLPER: Multi-Level Prompts for Adaptively Enhancing Vision-Language Emotion Recognition

IROS 2024poster

In the field of robotics, vision-based Emotion Recognition (ER) has achieved significant progress, but it still faces the challenge of poor generalization ability under unconstrained conditions (e.g., occlusions and pose variations). In this work, we propose MLPER model, which introduces Vision-Lang…

Cited by 1SourceScholar