← Search

Dianxi Shi

20 accepted papers

2026

AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation

CVPR 2026

Monocular depth estimation is essential for applications such as robotics. The complementary characteristics of event and image modalities have inspired fusion-based methods for robust depth estimation. However, existing methods rely on convolutional or attention-based architectures, which either ha

Cited by 0SourceScholar
2026

Learning Reward–Cost Balance in Safe RL via Score-Based World Models

ICML 2026poster

Safe reinforcement learning (Safe RL) seeks to optimize long-term performance while ensuring adherence to safety constraints. However, most existing approaches address safety in a simplified manner, typically by linearly combining rewards and costs, which provides limited guidance when safety and pe…

Cited by 0SourceScholar
2026

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

AAAI 2026technical

While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this challenge, we

Cited by 0SourcePDFScholar
2026

Zero-shot Active Mapping via Fused 360-BEV Representations and Vision–Language Models

ICML 2026poster

Active mapping enables embodied agents to understand and interact in previously unseen environments. However, most methods struggle to achieve zero-shot generalization to large-scale scenes and lack support for language instructions. We propose a VLM-based active mapping method that achieves zero-sh…

Cited by 0SourceScholar
2025

Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios

AAAI 2025technical

Learning optimal policies in multi-agent cooperative settings with visual observations is significant and challenging. Agents must first perform state representation learning for their image observations and then learn policies in the abstracted state space. Aiming at this problem, we propose a nove…

Cited by 0SourcePDFScholar
2025

Enhancing Visual Localization with Cross-Domain Image Generation

ICML 2025poster

Visual localization aims to predict the absolute camera pose for a single query image. However, predominant methods focus on single-camera images and scenes with limited appearance variations, limiting their applicability to cross-domain scenes commonly encountered in real-world applications. Furthe…

2025

Gaussian Mixture Model for Graph Domain Adaptation

IJCAI 2025

Unsupervised domain adaptation (UDA) has been widely studied with the goal of transferring knowledge from a label-rich source domain to a related but unlabeled target domain. Most UDA techniques achieve this by reducing the feature discrepancies between the two domains to learn domain-invariant feat

Cited by 0SourcePDFScholar
2025

Hierarchy Coverage Path Planning With Proactive Extremum Prevention in Unknown Environments

RA-L 2025

The local extremum is a crucial factor that affects the efficiency of online coverage path planning (CPP). Most online CPP methods generate coverage motions point by point in unknown environments. However, these solutions ignore efficient global coverage and probably result in local extremum. This l

Cited by 0SourceScholar
2025

Multi-Agent Hierarchical Graph Attention Actor-Critic Reinforcement Learning

ICASSP 2025accepted

Multi-agent systems often face challenges such as elevated communication demands and intricate interactions. We propose an innovative hierarchical graph attention actor-critic reinforcement learning method to address the issues, which uses the hierarchical graph attention to capture the relationship…

Cited by 0SourceScholar
2025

UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block

IJCAI 2025

Depth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range and temporal resolution but face difficulties with sparse data. Combining event

Cited by 0SourcePDFScholar
2024

Crowd Perception Communication-Based Multi- Agent Path Finding With Imitation Learning

RA-L 2024

Deep reinforcement learning-based Multi-Agent Path Finding (MAPF) has gained significant attention due to its remarkable adaptability to environments. Existing methods primarily leverage multi-agent communication in a fully-decentralized framework to maintain scalability while enhancing information

Cited by 2SourceScholar
2024

Spatial-Aware Dynamic Lightweight Self-Supervised Monocular Depth Estimation

RA-L 2024

Self-supervised monocular depth estimation has attracted extensive attention in recent years. Lightweight depth estimation methods are crucial for resource-constrained edge devices. However, existing lightweight methods often encounter the challenge of limited representation capacity and increased c

Cited by 10SourceScholar
2024

Unified Single-Stage Transformer Network for Efficient RGB-T Tracking

IJCAI 2024poster

Most existing RGB-T tracking networks extract modality features in a separate manner, which lacks interaction and mutual guidance between modalities. This limits the network's ability to adapt to the diverse dual-modality appearances of targets and the dynamic relationships between the modalities. A…

2023

Collision-free Coverage Path Planning for the Variable-speed Curvature-constrained Robot

ICRA 2023poster

Dubins coverage has been extensively researched to address the coverage path planning (CPP) problem of a known environment for the curvature-constrained robot. However, its fixed-speed assumption prevents the robot from accelerating to reduce the time and limits its flexibility to avoid obstacles. T…

Cited by 2SourceScholar
2023

Gyro-Net: IMU Gyroscopes Random Errors Compensation Method Based on Deep Learning

RA-L 2023

To solve the problem of inaccurate orientation estimation after long-term operations of the Inertial Measurement Unit (IMU), we present a learning-based method (called Gyro-Net) to estimate and compensate for IMU gyroscope random errors. We firstly introduce a semi-dense network structure, which ext

Cited by 23SourceScholar
2023

Improved Event-Based Dense Depth Estimation via Optical Flow Compensation

ICRA 2023poster

Event cameras have the potential to overcome the limitations of classical computer vision in real-world applications. Depth estimation is a crucial step for high-level robotics tasks and has attracted much attention from the community. In this paper, we propose an event-based dense depth estimation…

Cited by 7SourceScholar
2023

NeRF-IBVS: Visual Servo Based on NeRF for Visual Localization and Navigation

NeurIPS 2023poster

Visual localization is a fundamental task in computer vision and robotics. Training existing visual localization methods requires a large number of posed images to generalize to novel views, while state-of-the-art methods generally require dense ground truth 3D labels for supervision. However, acqui…

Cited by 9SourcePDFScholar
2022

Self-supervised representations for multi-view reinforcement learning

UAI 2022poster

Learning policies from raw, pixel images are quite important for the real-world application of deep reinforcement learning (RL). Standard model-free RL algorithms focus on single-view settings and unify the representation learning and policy learning into an end-to-end training process. However, suc…

2019

FA-Harris: A Fast and Asynchronous Corner Detector for Event Cameras

IROS 2019poster

Recently, the emerging bio-inspired event cameras have demonstrated potentials for a wide range of robotic applications in dynamic environments. In this paper, we propose a novel fast and asynchronous event-based corner detection method which is called FA-Harris. FA-Harris consists of several compon…

Cited by 67SourceScholar