← Search

Zitian Wang

5 accepted papers

2025

Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs

ICCV 2025poster

Preference alignment has emerged as an effective strategy to enhance the performance of Multimodal Large Language Models (MLLMs) following supervised fine-tuning. While existing preference alignment methods predominantly target hallucination factors, they overlook the factors essential for multi-mod…

2024

Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection

ICRA 2024poster

Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird’s-eye view (BEV) representation space. However, our empirical findings indicate that previous methods have limitations in generating fusion BEV features free from cross-modal conflicts. T…

Cited by 12SourcecodeScholar
2023

Object as Query: Lifting Any 2D Object Detector to 3D Detection

ICCV 2023poster

3D object detection from multi-view images has drawn much attention over the past few years. Existing methods mainly establish 3D representations from multi-view images and adopt a dense detection head for object detection, or employ object queries distributed in 3D space to localize objects. In thi…

Cited by 50PDFcodeScholar
2022

Distribution-Aware Single-Stage Models for Multi-Person 3D Pose Estimation

CVPR 2022poster

In this paper, we present a novel Distribution-Aware Single-stage (DAS) model for tackling the challenging multi-person 3D pose estimation problem. Different from existing top-down and bottom-up methods, the proposed DAS model simultaneously localizes person positions and their corresponding body jo…

Cited by 50PDFScholar
2021

Lvio-Fusion: A Self-adaptive Multi-sensor Fusion SLAM Framework Using Actor-critic Method

IROS 2021poster

State estimation with sensors is essential for mobile robots. Due to different performance of sensors in different environments, how to fuse measurements of various sensors is a problem. In this paper, we propose a tightly coupled multi-sensor fusion framework, Lvio-Fusion, which fuses stereo camera…

Cited by 48SourcecodeScholar