← Search

Xin Meng

14 accepted papers

2026

Enhanced Probabilistic Collision Detection for Motion Planning under Sensing Uncertainty

ICRA 2026poster

Probabilistic collision detection (PCD) is essential in motion planning for robots operating in unstructured environments, where considering sensing uncertainty helps prevent damage. Existing PCD methods mainly used simplified geometric models and addressed only position estimation errors. This pape…

2026

Optimal Design of Integrated Aerial Platforms with Passive Joints

ICRA 2026poster

The Integrated Aerial Platform (IAP) uses multiple quadrotor sub-vehicles, acting as independent thrust generators, connected to a central platform via passive joints. This setup allows the sub-vehicles to collectively apply forces and torques to the central platform, achieving full six-degree-of-fr…

Cited by 0SourceScholar
2026

PRIMP: PRobabilistically-Informed Motion Primitives for Efficient Affordance Learning from Demonstration (Abstract Reprint)

AAAI 2026technical

This paper proposes a learning-from-demonstration (LfD) method using probability densities on the workspaces of robot manipulators. The method, named PRobabilistically-Informed Motion Primitives (PRIMP), learns the probability distribution of the end effector trajectories in the 6D workspace that in

Cited by 0SourcePDFScholar
2026

SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System

CVPR 2026

Recent advancements in multimodal large language models (MLLMs) and video agent systems have significantly improved general video understanding. However, when applied to scientific video understanding and educating--a domain that demands external professional knowledge integration and rigorous step-

Cited by 0SourceScholar
2025

EventGPT: Event Stream Understanding with Multimodal Large Language Models

CVPR 2025poster

Event cameras capture visual information as asynchronous pixel change streams, excelling in challenging lighting and high-dynamic scenarios. Existing multimodal large language models (MLLMs) concentrate on natural RGB images, failing in scenarios where event data fits better. In this paper, we intro…

Cited by 3SourcePDFScholar
2025

Real-Time Human-Drone Interaction via Active Multimodal Gesture Recognition Under Limited Field of View in Indoor Environments

RA-L 2025

Gesture recognition, an important method for Human-Drone Interaction (HDI), is often constrained by sensor limitations, such as sensitivity to lighting variations and field of view (FoV) restrictions. This paper proposes a real-time drone control system that integrates multimodal fusion gesture reco

Cited by 2SourcecodeScholar
2023

Data Level Lottery Ticket Hypothesis for Vision Transformers

IJCAI 2023poster

The conventional lottery ticket hypothesis (LTH) claims that there exists a sparse subnetwork within a dense neural network and a proper random initialization method, called the winning ticket, such that it can be trained from scratch to almost as good as the dense counterpart. Meanwhile, the resear…

2023

HotBEV: Hardware-oriented Transformer-based Multi-View 3D Detector for BEV Perception

NeurIPS 2023poster

The bird's-eye-view (BEV) perception plays a critical role in autonomous driving systems, involving the accurate and efficient detection and tracking of objects from a top-down perspective. To achieve real-time decision-making in self-driving scenarios, low-latency computation is essential. While re…

Cited by 5SourcePDFScholar
2023

Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training

AAAI 2023technical

Vision transformers (ViTs) have recently obtained success in many applications, but their intensive computation and heavy memory usage at both training and inference time limit their generalization. Previous compression algorithms usually start from the pre-trained dense models and only focus on eff…

2023

Prepare the Chair for the Bear! Robot Imagination of Sitting Affordance to Reorient Previously Unseen Chairs

RA-L 2023

In this letter, a paradigm for the classification and manipulation of novel objects is established and demonstrated with the example of chairs. Our approach leverages the robot's understanding of object stability, perceptibility, and affordance to prepare previously unseen and randomly oriented chai

Cited by 4SourceScholar
2023

SpeedDETR: Speed-aware Transformers for End-to-end Object Detection

ICML 2023poster

Vision Transformers (ViTs) have continuously achieved new milestones in object detection. However, the considerable computation and memory burden compromise their efficiency and generalization of deployment on resource-constraint devices. Besides, efficient transformer-based detectors designed by ex…

Cited by 3SourcePDFScholar
2022

Put the Bear on the Chair! Intelligent Robot Interaction with Previously Unseen Chairs via Robot Imagination

ICRA 2022poster

In this paper, we study the problem of autonomously seating a teddy bear on a previously unseen chair. To achieve this goal, we present a novel method for robots to imagine the sitting pose of the bear by physically simulating a virtual humanoid agent sitting on the chair. We also develop a robotic…

Cited by 8SourcecodeScholar
2022

SPViT: Enabling Faster Vision Transformers via Latency-Aware Soft Token Pruning

ECCV 2022poster

"Recently, Vision Transformer (ViT) has continuously established new milestones in the computer vision field, while the high computation and memory cost makes its propagation in industrial production difficult. Considering the computation complexity, the internal data pattern of ViTs, and the edge d…

2022

Transporters with Visual Foresight for Solving Unseen Rearrangement Tasks

IROS 2022poster

Rearrangement tasks have been identified as a crucial challenge for intelligent robotic manipulation, but few methods allow for precise construction of unseen structures. We propose a visual foresight model for pick-and-place rearrangement manipulation which is able to learn efficiently. In addition…

Cited by 16SourcecodeScholar