← Search

Jilong Wang

17 accepted papers

2026

CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution

AAAI 2026technical

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual cues often conflict, requiring models to perform structured

Cited by 0SourcePDFScholar
2026

Gait Transformer: End-to-End Transformer Backbone for Gait Recognition

AAAI 2026technical

Gait recognition has emerged as a promising biometric technique for long-distance and non-intrusive human identification. While Transformers have revolutionized vision tasks, their adaptation to gait recognition remains underexplored due to domain-specific challenges such as sparse silhouette modali

Cited by 0SourcePDFScholar
2026

Humanoid Generative Pre-Training for Zero-Shot Motion Tracking

CVPR 2026

We introduce Humanoid-GPT, a GPT-style Transformer with causal attention trained on a billion-scale motion corpus for whole-body control. Unlike prior shallow MLP trackers constrained by scarce data and an agility-generalization trade-off, Humanoid-GPT is pre-trained on a 2B-frame retargeted corpus

Cited by 0SourcecodeScholar
2026

Unleashing Humanoid Reaching Potential Via Real-World-Ready Skill Space

ICRA 2026poster

Humans possess a large reachable space in the 3D world, enabling interactions with objects at varying heights and distances. However, realizing such large-space reaching on humanoids is a complex whole-body control (WBC) problem. Learning from scratch often leads to optimization difficulty and poor …

2026

Unleashing Humanoid Reaching Potential via Real-World-Ready Skill Space

RA-L 2026

Humans possess a large reachable space in the 3D world, enabling interactions with objects at varying heights and distances. However, realizing such large-space reaching on humanoids is a complex whole-body control (WBC) problem. Learning from scratch often leads to optimization difficulty and poor

Cited by 23SourcecodeScholar
2025

Bridging Gait Recognition and Large Language Models Sequence Modeling

CVPR 2025poster

Gait sequences exhibit sequential structures and contextual relationships similar to those in natural language, where each element--whether a word or a gait step--is connected to its predecessors and successors. This similarity enables the transformation of gait sequences into "texts" containing ide…

Cited by 1SourcePDFScholar
2025

DUNE: Sim2Real Transfer for Depth-based Navigation in Unstructured Dynamic Indoor Environments

ICASSP 2025accepted

Collision-free navigation in dynamic environments, especially with moving pedestrians, is crucial for mobile robots. This paper introduces DUNE, a depth-based policy trained in simulation for collision-free navigation of Ackermann mobile robots in unstructured indoor environments. DUNE uses a CNN-LS…

Cited by 0SourceScholar
2025

FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real

CoRL 2025oral

Generalizable object fetching in cluttered scenes remains a fundamental and application-critical challenge in embodied AI. Closely packed objects cause inevitable occlusions, making safe action generation particularly difficult. Under such partial observability, effective policies must not only gene…

Cited by 0SourceScholar
2025

MobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic Data

CVPR 2025poster

This paper introduces MobileH2R, a framework for learning generalizable vision-based human-to-mobile-robot (H2MR) handover skills. Unlike traditional fixed-base handovers, this task requires a mobile robot to reliably receive objects in a large workspace enabled by its mobility. Our key insight is t…

Cited by 0SourcePDFScholar
2025

QuadWBG: Generalizable Quadrupedal Whole-Body Grasping

ICRA 2025

Legged robots with advanced manipulation capabilities have the potential to significantly improve household duties and urban maintenance. Despite considerable progress in developing robust locomotion and precise manipulation methods, seamlessly integrating these into cohesive whole-body control for

Cited by 6SourcecodeScholar
2025

Watch Less, Feel More: Sim-to-Real RL for Generalizable Articulated Object Manipulation via Motion Adaptation and Impedance Control

ICRA 2025

Articulated object manipulation poses a unique challenge compared to rigid object manipulation as the object itself represents a dynamic environment. In this work, we present a novel RL-based pipeline equipped with variable impedance control and motion adaptation leveraging observation history for g

Cited by 4SourcecodeScholar
2024

GAMMA: Graspability-Aware Mobile MAnipulation Policy Learning based on Online Grasping Pose Fusion

ICRA 2024poster

Mobile manipulation constitutes a fundamental task for robotic assistants and garners significant attention within the robotics community. A critical challenge inherent in mobile manipulation is the effective observation of the target while approaching it for grasping. In this work, we propose a gra…

Cited by 24SourcecodeScholar
2021

Hierarchical Terrain-Aware Control for Quadrupedal Locomotion by Combining Deep Reinforcement Learning and Optimal Control

IROS 2021poster

Quadruped robots possess advantages on different terrains over other types of mobile robots by virtue of their flexible choices of foothold points. It is crucial to integrate terrain perception with motion planning to exploit the potential of quadruped robots. We propose a novel hierarchical terrain…

Cited by 10SourceScholar
2021

Terrain-Aware Risk-Assessment-Network-Aided Deep Reinforcement Learning for Quadrupedal Locomotion in Tough Terrain

IROS 2021poster

When it comes to the control system of quadruped robots, deep reinforcement learning (DRL) is considered to be a promising solution. Despite years of development in this field, difficulties remain in guaranteeing the action stability of DRL-based quadruped robots’ locomotion, especially in tough ter…

Cited by 6SourceScholar
2020

Graph Stochastic Neural Networks for Semi-supervised Learning

NeurIPS 2020poster

Graph Neural Networks (GNNs) have achieved remarkable performance in the task of the semi-supervised node classification. However, most existing models learn a deterministic classification function, which lack sufficient flexibility to explore better choices in the presence of kinds of imperfect ob…