← Search

ZEQUN JIE

37 accepted papers

2026

Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

CVPR 2026

Large vision-language models (LVLMs) have shown remarkable capabilities in visual-language understanding. Despite their success, LVLMs still suffer from generating hallucinations in complex generation tasks, leading to inconsistencies between visual inputs and generated content. To address this issu

Cited by 0SourceScholar
2026

M4V: Multimodal Mamba for Efficient Text-to-Video Generation

CVPR 2026

Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal space remains computationally demanding, particularly when employing Transformers, which incur quadratic complexity in sequ

Cited by 0SourceScholar
2026

X-SAM: From Segment Anything to Any Segmentation

AAAI 2026technical

Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although the Segment Anything Model (SAM) represents a significant advancement in visual-prompt-driven image segmentation, it exh

Cited by 0SourcePDFScholar
2025

CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting

ICCV 2025poster

Recent works in 3D representation learning and multimodal pre-training have made remarkable progress. However, typically multimodal 3D models are only capable of handling point clouds. Compared to the emerging 3D representation technique, 3D Gaussian Splatting (3DGS), the spatially sparse point clou…

Cited by 0SourcePDFScholar
2025

FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction

NeurIPS 2025poster

This work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVAR facilitates autoregressive learning with ground-truth prediction, enabling each step to independently produce plausibl…

Cited by 0SourcecodeScholar
2025

RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

ICCV 2025poster

Large Multimodal Models (LMMs) have demonstrated exceptional comprehension and interpretation capabilities in Autonomous Driving (AD) by incorporating large language models. Despite the advancements, current data-driven AD approaches tend to concentrate on a single dataset and specific tasks, neglec…

Cited by 0SourcePDFScholar
2024

3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance

ECCV 2024poster

"In this paper, we propose 3DSS-VLG, a weakly supervised approach for 3D Semantic Segmentation with 2D Vision-Language Guidance, an alternative approach that a 3D model predicts dense-embedding for each point which is co-embedded with both the aligned image and text spaces from the 2D vision-languag…

2024

AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement Learning

CVPR 2024poster

Powered by massive curated training data Segment Anything Model (SAM) has demonstrated its impressive generalization capabilities in open-world scenarios with the guidance of prompts. However the vanilla SAM is class-agnostic and heavily relies on user-provided prompts to segment objects of interest…

2024

InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

CVPR 2024poster

In this paper we present a novel paradigm to enhance the ability of object detector e.g. expanding categories or improving detection performance by training on syn- thetic dataset generated from diffusion models. Specifically we integrate an instance-level grounding head into a pre- trained generati…

Cited by 13SourcePDFScholar
2024

Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting Learning

AAAI 2024technical

Camera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction relies on reliable depth estimation for multi-camera images. However, existing approaches exhaustively predict depths fo…

2024

Investigating Compositional Challenges in Vision-Language Models for Visual Grounding

CVPR 2024highlight

Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performance gains contributed by large vision and language pre-training we find that state-of…

2024

Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models

NeurIPS 2024poster

Large Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the perception capabilities of LMMs. The current methods follow the paradigm of adaptin…

2024

Making Large Language Models Better Planners with Reasoning-Decision Alignment

ECCV 2024oral

"Data-driven approaches for autonomous driving (AD) have been widely adopted in the past decade but are confronted with dataset bias and uninterpretability. Inspired by the knowledge-driven nature of human driving, recent approaches explore the potential of large language models (LLMs) to improve un…

Cited by 12SourcePDFScholar
2023

AeDet: Azimuth-Invariant Multi-View 3D Object Detection

CVPR 2023poster

Recent LSS-based multi-view 3D object detection has made tremendous progress, by processing the features in Brid-Eye-View (BEV) via the convolutional detector. However, the typical convolution ignores the radial symmetry of the BEV features and increases the difficulty of the detector optimization.…

2023

Curriculum Multi-Negative Augmentation for Debiased Video Grounding

AAAI 2023technical

Video Grounding (VG) aims to locate the desired segment from a video given a sentence query. Recent studies have found that current VG models are prone to over-rely the groundtruth moment annotation distribution biases in the training set. To discourage the standard VG model's behavior of exploiting…

2023

MSMDFusion: Fusing LiDAR and Camera at Multiple Scales With Multi-Depth Seeds for 3D Object Detection

CVPR 2023poster

Fusing LiDAR and camera information is essential for accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features from two drastically different modalities. Recent approaches aim at e…

2022

Expansion and Shrinkage of Localization for Weakly-Supervised Semantic Segmentation

NeurIPS 2022accept

Generating precise class-aware pseudo ground-truths, a.k.a, class activation maps (CAMs), is essential for Weakly-Supervised Semantic Segmentation. The original CAM method usually produces incomplete and inaccurate localization maps. To tackle with this issue, this paper proposes an Expansion and Sh…

2022

MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes

ECCV 2022poster

"3D dense captioning is a recently-proposed novel task, where point clouds contain more geometric information than the 2D counterpart. However, it is also more challenging due to the higher complexity and wider variety of inter-object relations contained in point clouds. Existing methods only treat…

2022

PromptDet: Towards Open-Vocabulary Detection Using Uncurated Images

ECCV 2022poster

"The goal of this work is to establish a scalable pipeline for expanding an object detector towards novel/unseen categories, using zero manual annotations. To achieve that, we make the following four contributions: (i) in pursuit of generalisation, we propose a two-stage open-vocabulary object detec…

2020

Central Similarity Quantization for Efficient Image and Video Retrieval

CVPR 2020poster

Existing data-dependent hashing methods usually learn hash functions from pairwise or triplet data relationships, which only capture the data similarity locally, and often suffer from low learning efficiency and low collision rate. In this work, we propose a new global similarity metric, termed as c…

Cited by 400PDFcodeScholar
2020

MTL-NAS: Task-Agnostic Neural Architecture Search Towards General-Purpose Multi-Task Learning

CVPR 2020poster

We propose to incorporate neural architecture search (NAS) into general-purpose multi-task learning (GP-MTL). Existing NAS methods typically define different search spaces according to different tasks. In order to adapt to different task combinations (i.e., task sets), we disentangle the GP-MTL netw…

Cited by 109PDFcodeScholar
2020

NMS by Representative Region: Towards Crowded Pedestrian Detection by Proposal Pairing

CVPR 2020poster

Although significant progress has been made in pedestrian detection recently, pedestrian detection in crowded scenes is still challenging. The heavy occlusion between pedestrians imposes great challenges to the standard Non-Maximum Suppression (NMS). A relative low threshold of intersection over uni…

Cited by 202PDFScholar
2019

Geometry-Aware Distillation for Indoor Semantic Segmentation

CVPR 2019poster

It has been shown that jointly reasoning the 2D appearance and 3D information from RGB-D domains is beneficial to indoor scene semantic segmentation. However, most existing approaches require accurate depth map as input to segment the scene which severely limits their applications. In this paper, we…

Cited by 113PDFScholar
2018

Joint Task-Recursive Learning for Semantic Segmentation and Depth Estimation

ECCV 2018poster

In this paper, we propose a novel joint Task-Recursive Learning (TRL) framework for the closing-loop semantic segmentation and monocular depth estimation tasks. TRL can recursively refine the results of both tasks through serialized task-level interactions. In order to mutually-boost for each other,…

Cited by 261SourcePDFScholar
2018

Left-Right Comparative Recurrent Model for Stereo Matching

CVPR 2018poster

Leveraging the disparity information from both left and right views is crucial for stereo disparity estimation. Left-right consistency check is an effective way to enhance the disparity estimation by referring to the information from the opposite view. However, the conventional left-right consisten…

Cited by 115SourcePDFScholar
2018

Modeling Varying Camera-IMU Time Offset in Optimization-Based Visual-Inertial Odometry

ECCV 2018poster

Combining cameras and inertial measurement units (IMUs) has been proven effective in motion tracking, as these two sensing modalities offer complementary characteristics that are suitable for fusion. While most works focus on global-shutter cameras and synchronized sensor measurements, consumer-grad…

Cited by 25SourcePDFScholar
2018

Revisiting Dilated Convolution: A Simple Approach for Weakly- and Semi-Supervised Semantic Segmentation

CVPR 2018poster

Despite remarkable progress, weakly supervised segmentation methods are still inferior to their fully supervised counterparts. We obverse that the performance gap mainly comes from the inability of producing dense and integral pixel-level object localization for training images only with image-level…

Cited by 700SourcePDFScholar
2017

Deep Self-Taught Learning for Weakly Supervised Object Localization

CVPR 2017spotlight

Most existing weakly supervised localization (WSL) approaches learn detectors by finding positive bounding boxes based on features learned with image-level supervision. However, those features do not contain spatial location related information and usually provide poor-quality positive samples for t…

Cited by 249PDFScholar
2017

Predicting Scene Parsing and Motion Dynamics in the Future

NeurIPS 2017poster

It is important for intelligent systems, e.g. autonomous vehicles and robotics to anticipate the future in order to plan early and make decisions accordingly. Predicting the future scene parsing and motion dynamics helps the agents better understand the visual environment better as the former provid…

2017

Video Scene Parsing With Predictive Feature Learning

ICCV 2017poster

Video scene parsing is challenging due to the following two reasons: firstly, it is non-trivial to learn meaningful video representations for producing the temporally consistent labeling map; secondly, such a learning process becomes more difficult with insufficient labeled video training data. In t…

Cited by 156PDFScholar
2016

Reversible Recursive Instance-Level Object Segmentation

CVPR 2016poster

In this work, we propose a novel Reversible Recursive Instance-level Object Segmentation (R2-IOS) framework to address the challenging instance-level object segmentation task. R2-IOS consists of a reversible proposal refinement sub-network that predicts bounding box offsets for refining the object p…

Cited by 65PDFScholar
2016

Tree-Structured Reinforcement Learning for Sequential Object Localization

NeurIPS 2016poster

Existing object proposal algorithms usually search for possible object regions over multiple locations and scales \emph{ separately}, which ignore the interdependency among different objects and deviate from the human perception procedure. To incorporate global interdependency between objects into o…

Cited by 161SourcePDFScholar