← Search

Qijun Chen

40 accepted papers

2026

A&B-LO: Continuous-Time LiDAR Odometry with Adaptive Non-Uniform B-Spline Trajectory Representation

ICRA 2026poster

LiDAR odometry, fused by inertial measurement units (IMU), is an essential task in robotics navigation. Unlike the mainstream methods compensate the motion distortion of LiDAR data by high frequency inertial sensors, this paper deals with the distortion with continuous-time trajectory representation…

Cited by 0Scholar
2026

BEVDrive-E2E: Imitation With Bird's Eye View Perception for Interpretable End-to-End Autonomous Driving

RA-L 2026

Imitation learning (IL) for end-to-end autonomous driving (E2E-AD) has made great progress recently in the closed-loop evaluation of the CARLA simulator. However, the causal confusion remains an open problem. To address this issue, we propose the BEVDrive-E2E to explore the interpretability of the e

Cited by 0SourceScholar
2026

Bootstrapping MLLM for Weakly‑Supervised Class‑Agnostic Object Counting

ICLR 2026poster

Object counting is a fundamental task in computer vision, with broad applicability in many real-world scenarios. Fully-supervised counting methods require costly point-level annotations per object. Few weakly-supervised methods leverage only image-level object counts as supervision and achieve fairl…

Cited by 0SourcecodeScholar
2026

Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy Adaptation

ICML 2026poster

Real-world dynamics shifts pose a critical challenge for reinforcement learning, yet prior methods typically rely on encoding explicitly identified physical parameters into a latent context, a rigid parameterization that proves brittle to unmodeled or compound dynamics variations. We instead investi…

Cited by 0SourceScholar
2026

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

CVPR 2026

This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on variational autoencoders (VAEs) to generate discrete occupancy tokens, which inherently limit representational capacity, our approach predicts multi-fra

Cited by 0SourcecodeScholar
2026

TerrFlat: Physics-Driven Geometry Representation for Structure-Aware Freespace Detection

ICRA 2026poster

Freespace detection in autonomous driving is limited by the lack of explicit geometric modeling, hindering generalization across complex terrains. Existing approaches are predominantly data-driven and neglect the physical structure of drivable surfaces. We propose Terrain Flat (TerrFlat), a physics-…

Cited by 0SourceScholar
2025

CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation

ICCV 2025poster

In the effort to achieve robust and generalizable category-level object pose estimation, recent methods primarily focus on learning fundamental representations from data. However, the inherent biases within the data are often overlooked: the repeated training samples and similar environments may mis…

2025

LGPR: Local Feature Learning Brings More Generalizable Visual Place Recognition

IROS 2025

We propose a Visual Place Recognition (VPR) framework by sharing lightweight keypoint extraction modules for local features. Current research on the joint learning of local keypoint matching and VPR is relatively scarce, and the application deployment of real-time spatial computing on edge devices h

Cited by 0SourcecodeScholar
2025

LSW-Net: A Spatio-temporal Self-Supervised Framework for 2D LiDAR-Based Environment Perception

IROS 2025

In the deep learning era, 2D LiDAR perception is often overlooked as research prioritizes 3D point clouds. Yet, 2D LiDAR remains essential for low-cost robotic systems due to its affordability. Despite its simplicity, it faces major challenges-not only in perception but also in learning itself, as d

Cited by 0SourcecodeScholar
2025

Reinforcement Learning-based Optimization of Humanoid Joint Motion Control via Text-driven Human Motion Mapping

IROS 2025

Human motion retargeting for humanoid robots, transferring human motion data to robots for imitation, presents significant challenges but offers considerable potential for real-world applications. Traditionally, this process relies on human demonstrations captured through pose estimation or motion c

Cited by 0SourceScholar
2025

Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities

ICCV 2025poster

Recent Vision-and-Language Navigation (VLN) advancements are promising, but their idealized assumptions about robot movement and control fail to reflect physically embodied deployment challenges. To bridge this gap, we introduce VLN-PE, a physically realistic VLN platform supporting humanoid, quadru…

2025

Rotation-Equivariant Robot Vision: A Perspective via Correspondence-Matching and Pre-training

IROS 2025

Correspondence matching is a fundamental and crucial task in robot vision. In recent years, deep learning-based keypoint matching techniques have shown outstanding performance in downstream tasks. Conventional learning-based correspondence matching methods rely on large datasets and a specific train

Cited by 0SourceScholar
2024

DVT: Decoupled Dual-Branch View Transformation for Monocular Bird’s Eye View Semantic Segmentation

IROS 2024poster

Monocular Bird’s Eye View (BEV) semantic segmentation is critical for autonomous driving for its inherent advantages in spatial representation and downstream tasks. However, it is challenging to simultaneously learn view transformation and pixel-wise classification. Previous works suffer from non-fl…

Cited by 0SourcecodeScholar
2024

Dive Deeper into Rectifying Homography for Stereo Camera Online Self-Calibration

ICRA 2024poster

Accurate estimation of stereo camera extrinsic parameters is crucial to guarantee the performance of stereo matching algorithms. In prior arts, the online self-calibration of stereo cameras has commonly been formulated as a specialized visual odometry problem, without taking into account the princip…

Cited by 8SourceScholar
2024

Enhanced Language-guided Robot Navigation with Panoramic Semantic Depth Perception and Cross-modal Fusion

IROS 2024poster

Integrating visual observation with linguistic instruction holds significant promise for enhancing robot navigation across unstructured environments and enriches the human-robot interaction experience. However, while panoramic RGB views furnish robots with extensive environmental visuals, current me…

Cited by 0SourcecodeScholar
2024

GenerOcc: Self-supervised Framework of Real-time 3D Occupancy Prediction for Monocular Generic Cameras

IROS 2024poster

In the context of 3D scene perception tasks, the significance of 3D occupancy prediction has been progressively growing, aiming to forecast the occupancy state of voxels in a discrete 3D space. However, existing methods typically exhibit several limitations, such as restricted adaptability to non-pi…

Cited by 0SourceScholar
2024

InstructDET: Diversifying Referring Object Detection with Generalized Instructions

ICLR 2024poster

We propose InstructDET, a data-centric method for referring object detection (ROD) that localizes target objects based on user instructions. While deriving from referring expressions (REC), the instructions we leverage are greatly diversified to encompass common user intentions related to object det…

2024

LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation

CVPR 2024poster

Referring video object segmentation (RVOS) aims to segment the target instance referred by a given text expression in a video clip. The text expression normally contains sophisticated description of the instance's appearance action and relation with others. It is therefore rather difficult for a RVO…

2024

MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer

NeurIPS 2024poster

Transferring visual-language knowledge from large-scale foundation models for video recognition has proved to be effective. To bridge the domain gap, additional parametric modules are added to capture the temporal information. However, zero-shot generalization diminishes with the increase in the num…

2024

Multimodal Evolutionary Encoder for Continuous Vision-Language Navigation

IROS 2024poster

Can multimodal encoder evolve when facing increasingly tough circumstances? Our work investigates this possibility in the context of continuous vision-language navigation (continuous VLN), which aims to navigate robots under linguistic supervision and visual feedback. We propose a multimodal evoluti…

Cited by 0SourcecodeScholar
2024

PICNN: A Pathway towards Interpretable Convolutional Neural Networks

AAAI 2024technical

Convolutional Neural Networks (CNNs) have exhibited great performance in discriminative feature learning for complex visual tasks. Besides discrimination power, interpretability is another important yet under-explored property for CNNs. One difficulty in the CNN interpretability is that filters and…

2024

Relaxing the Limitations of the Optimal Reciprocal Collision Avoidance Algorithm for Mobile Robots in Crowds

RA-L 2024

The Optimal Reciprocal Collision Avoidance (ORCA) algorithm is widely used for modeling agents in collision avoidance scenarios. However, suffering from limitations such as the improper reciprocal assumption that each agent is supposed to take half the responsibility for collision avoidance, the per

Cited by 5SourceScholar
2024

SG-RoadSeg: End-to-End Collision-Free Space Detection Sharing Encoder Representations Jointly Learned via Unsupervised Deep Stereo

ICRA 2024poster

Collision-free space detection is of utmost importance for autonomous robot perception and navigation. State-of-the-art (SoTA) approaches generally extract features from RGB images and an additional source or modality of 3-D information, such as depth or disparity images, using a pair of independent…

Cited by 2SourceScholar
2024

SNF-Feat: Semantic-Guided Negative-Sample-Free Representation Learning for Local Feature Extraction

IROS 2024poster

Local feature extraction constitutes a foundational module crucial for numerous downstream tasks of computer vision. Its primary challenge lies in the generation of discriminative feature representations. Prior methodologies have employed contrastive learning within their pipelines, yet have encount…

Cited by 0SourceScholar
2024

Vision-and-Language Navigation via Causal Learning

CVPR 2024poster

In the pursuit of robust and generalizable environment perception and language understanding the ubiquitous challenge of dataset bias continues to plague vision-and-language navigation (VLN) agents hindering their performance in unseen environments. This paper introduces the generalized cross-modal…

2023

A Dual Semantic-Aware Recurrent Global-Adaptive Network for Vision-and-Language Navigation

IJCAI 2023poster

Vision-and-Language Navigation (VLN) is a realistic but challenging task that requires an agent to locate the target region using verbal and visual cues. While significant advancements have been achieved recently, there are still two broad limitations: (1) The explicit information mining for signifi…

2023

FDLNet: Boosting Real-time Semantic Segmentation by Image-size Convolution via Frequency Domain Learning

ICRA 2023poster

This paper proposes a novel real-time semantic segmentation network via frequency domain learning, called FDLNet, which revisits the segmentation task from two critical perspectives: spatial structure description and multilevel feature fusion. We first devise an image-size convolution (IS-Conv) as a…

Cited by 6SourcecodeScholar
2023

Multiple Thinking Achieving Meta-Ability Decoupling for Object Navigation

ICML 2023poster

We propose a meta-ability decoupling (MAD) paradigm, which brings together various object navigation methods in an architecture system, allowing them to mutually enhance each other and evolve together. Based on the MAD paradigm, we design a multiple thinking (MT) model that leverages distinct thinki…

Cited by 10SourcePDFScholar
2023

RMRL: Robot Navigation in Crowd Environments With Risk Map-Based Deep Reinforcement Learning

RA-L 2023

Achieving safe and effective navigation in crowds is a crucial yet challenging problem. Recent work has mainly encoded the pedestrian-robot state pairs, which cannot fully capture the interactions among humans. Besides, existing work attempts to achieve “hard” collision avoidance, which may leave no

Cited by 20SourceScholar
2023

Search for or Navigate to? Dual Adaptive Thinking for Object Navigation

ICCV 2023poster

"Search for" or "Navigate to"? When we find a specific object in an unknown environment, the two choices always arise in our subconscious mind. Before we see the target, we search for the target based on prior experience. Once we have seen the target, we can navigate to it by remembering the target…

Cited by 21PDFScholar
2022

HoloSeg: An Efficient Holographic Segmentation Network for Real-time Scene Parsing

ICRA 2022poster

Real-time semantic segmentation is a crucial but challenging dense prediction task for scene parsing. However, the existing CNN-based methods commonly bias the model in favor of speed-boosting compromising spatial resolution due to business requirements and hardware constrains, which impedes the hig…

Cited by 8SourcecodeScholar
2019

Improving Learning-based Ego-motion Estimation with Homomorphism-based Losses and Drift Correction

IROS 2019poster

Visual odometry is an essential problem for mobile robots. Traditional methods for solving VO mostly utilize geometric optimization. While capable of achieving high accuracy, these methods require accurate sensor calibration and complicated parameter tuning to work well in practice. With the rise of…

Cited by 16SourceScholar
2018

Monocular Visual Odometry Scale Recovery Using Geometrical Constraint

ICRA 2018poster

Scale recovery is one of the essential problems for monocular visual odometry. The camera height is usually used as an absolute reference to recover the scale. In this case, the precision of scale recovery depends on the accuracy of the road region detection and road geometrical model calculation. I…

Cited by 31SourceScholar
2017

Scale Recovery for Monocular Visual Odometry Using Depth Estimated With Deep Convolutional Neural Fields

ICCV 2017poster

Scale recovery is one of the central problems for monocular visual odometry. Normally, road plane and camera height are specified as reference to recover the scale. The performances of these methods depend on the plane recognition and height measurement of camera. In this work, we propose a novel me…

Cited by 81PDFScholar