← Search

Lei Jin

17 accepted papers

2025

CCIN: Compositional Conflict Identification and Neutralization for Composed Image Retrieval

CVPR 2025highlight

Composed Image Retrieval (CIR) is a multi-modal task that seeks to retrieve target images by harmonizing a reference image with a modified instruction. A key challenge in CIR lies in compositional conflicts between the reference image (e.g., blue, long sleeve) and the modified instruction (e.g., gre…

2025

JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems

CVPR 2025poster

Unmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we int…

Cited by 0SourcePDFScholar
2025

RoboScape: Physics-informed Embodied World Model

NeurIPS 2025spotlight

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models exhibit limited physical awareness, particularly in modelin…

Cited by 0SourcecodeScholar
2025

StickMotion: Generating 3D Human Motions by Drawing a Stickman

CVPR 2025poster

Text-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient diffusion-based network designed for multi-condition scenarios, wh…

2024

E3V-K5: An Authentic Benchmark for Redefining Video-Based Energy Expenditure Estimation

ECCV 2024poster

"Accurately estimating energy expenditure (EE) is crucial for optimizing athletic training, monitoring daily activity levels, and preventing sports-related injuries. Estimating energy expenditure based on video (E3 V) is an appealing research direction. This paper introduces E3V-K5, an authentic dat…

2024

Error Identification and Accuracy Compensation Algorithm for Improved 2RPU/UPR+R+P Hybrid Robot

RA-L 2024

To improve the precision of the 2RPU/UPR+R+P hybrid robot and fulfill production requirements, error compensation was explored. The robot's fixed coordinate system was first used to analyze the workbench error mapping matrix. Next, the correlation between the joint geometric errors and the moving pl

Cited by 2SourceScholar
2024

Investigating Stability Outcomes Across Diverse Gait Patterns in Quadruped Robots: A Comparative Analysis

RA-L 2024

Quadruped robots have gained attention for their potential to navigate various terrains. However, the stability of these robots in different gait sequences remains an open question. This study investigates the relationship between different gait sequences and the motion stability of quadruped robots

Cited by 4SourceScholar
2024

SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement

CVPR 2024poster

Predicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily co…

2024

Unified Single-Stage Transformer Network for Efficient RGB-T Tracking

IJCAI 2024poster

Most existing RGB-T tracking networks extract modality features in a separate manner, which lacks interaction and mutual guidance between modalities. This limits the network's ability to adapt to the diverse dual-modality appearances of targets and the dynamic relationships between the modalities. A…

2023

GPLight: Grouped Multi-agent Reinforcement Learning for Large-scale Traffic Signal Control

IJCAI 2023poster

The use of multi-agent reinforcement learning (MARL) methods in coordinating traffic lights (CTL) has become increasingly popular, treating each intersection as an agent. However, existing MARL approaches either treat each agent absolutely homogeneous, i.e., same network and parameter for each agent…

Cited by 26SourcePDFScholar
2023

RefCLIP: A Universal Teacher for Weakly Supervised Referring Expression Comprehension

CVPR 2023poster

Referring Expression Comprehension (REC) is a task of grounding the referent based on an expression, and its development is greatly limited by expensive instance-level annotations. Most existing weakly supervised methods are built based on two-stage detection networks, which are computationally expe…

2022

Learning Quality-Aware Representation for Multi-Person Pose Regression

AAAI 2022technical

Off-the-shelf single-stage multi-person pose regression methods generally leverage the instance score (i.e., confidence of the instance localization) to indicate the pose quality for selecting the pose candidates. We consider that there are two gaps involved in existing paradigm: 1) The instance sco…

Cited by 17SourcePDFScholar
2022

QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query

NeurIPS 2022accept

We propose a sparse end-to-end multi-person pose regression framework, termed QueryPose, which can directly predict multi-person keypoint sequences from the input image. The existing end-to-end methods rely on dense representations to preserve the spatial detail and structure for precise keypoint lo…

2022

Single-Stage Is Enough: Multi-Person Absolute 3D Pose Estimation

CVPR 2022poster

The existing multi-person absolute 3D pose estimation methods are mainly based on two-stage paradigm, i.e., top-down or bottom-up, leading to redundant pipelines with high computation cost. We argue that it is more desirable to simplify such two-stage paradigm to a single-stage one to promote both e…

Cited by 56PDFScholar
2020

Geometric Structure Based and Regularized Depth Estimation From 360 Indoor Imagery

CVPR 2020poster

Motivated by the correlation between the depth and the geometric structure of a 360 indoor image, we propose a novel learning-based depth estimation framework that leverages the geometric structure of a scene to conduct depth estimation. Specifically, we represent the geometric structure of an indoo…

Cited by 86PDFScholar
2020

P²Net: Patch-match and Plane-regularization for Unsupervised Indoor Depth Estimation

ECCV 2020poster

This paper tackles the unsupervised depth estimation task in indoor environments. The task is extremely challenging because of the vast areas of non-texture regions in these scenes. These areas could overwhelm the optimization process in the commonly used unsupervised depth estimation framework prop…