← Search

Chengjian Feng

13 accepted papers

2026

LeapAlign: Post-training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories

CVPR 2026

This paper focuses on the alignment of flow-matching models with human preference. A promising way is fine-tuning by directly backpropagating reward signals through the differentiable generation process of flow matching. However, backpropagating through long trajectories results in prohibitive memor

Cited by 0SourcecodeScholar
2026

X-SAM: From Segment Anything to Any Segmentation

AAAI 2026technical

Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although the Segment Anything Model (SAM) represents a significant advancement in visual-prompt-driven image segmentation, it exh

Cited by 0SourcePDFScholar
2025

DisTime: Distribution-based Time Representation for Video Large Language Models

ICCV 2025poster

Despite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and limited temporally aware datasets. Existing methods for temporal expression either conflate time with text-based numeric…

2025

RoboTrom-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction

ICCV 2025poster

In language-guided visual navigation, agents locate target objects in unseen environments using natural language instructions. For reliable navigation in unfamiliar scenes, agents should possess strong perception, planning, and prediction capabilities. Additionally, when agents revisit previously ex…

Cited by 0SourcePDFScholar
2025

RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

ICCV 2025poster

Large Multimodal Models (LMMs) have demonstrated exceptional comprehension and interpretation capabilities in Autonomous Driving (AD) by incorporating large language models. Despite the advancements, current data-driven AD approaches tend to concentrate on a single dataset and specific tasks, neglec…

Cited by 0SourcePDFScholar
2025

RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

ICCV 2025poster

Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial interactions and managing data collection costs. To address these issues, we propose the multimodal robotic manipulation…

2025

RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case

ICCV 2025poster

Collecting real-world data for rare high-risk scenarios, long-tailed driving events, and complex interactions remains challenging, leading to poor performance of existing autonomous driving systems in these critical situations. In this paper, we propose RoboTron-Sim that improves real-world driving…

Cited by 0SourcePDFScholar
2024

InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

CVPR 2024poster

In this paper we present a novel paradigm to enhance the ability of object detector e.g. expanding categories or improving detection performance by training on syn- thetic dataset generated from diffusion models. Specifically we integrate an instance-level grounding head into a pre- trained generati…

Cited by 13SourcePDFScholar
2024

UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection

ECCV 2024poster

"Temporal Action Detection (TAD) focuses on detecting pre-defined actions, while Moment Retrieval (MR) aims to identify the events described by open-ended natural language within untrimmed videos. Despite that they focus on different events, we observe they have a significant connection. For instanc…

2023

AeDet: Azimuth-Invariant Multi-View 3D Object Detection

CVPR 2023poster

Recent LSS-based multi-view 3D object detection has made tremendous progress, by processing the features in Brid-Eye-View (BEV) via the convolutional detector. However, the typical convolution ignores the radial symmetry of the BEV features and increases the difficulty of the detector optimization.…

2022

PromptDet: Towards Open-Vocabulary Detection Using Uncurated Images

ECCV 2022poster

"The goal of this work is to establish a scalable pipeline for expanding an object detector towards novel/unseen categories, using zero manual annotations. To achieve that, we make the following four contributions: (i) in pursuit of generalisation, we propose a two-stage open-vocabulary object detec…

2021

TOOD: Task-Aligned One-Stage Object Detection

ICCV 2021poster

One-stage object detection is commonly implemented by optimizing two sub-tasks: object classification and localization, using heads with two parallel branches, which might lead to a certain level of spatial misalignment in predictions between the two tasks. In this work, we propose a Task-aligned On…

Cited by 1116PDFcodeScholar