← Search

Rui Tian

11 accepted papers

2026

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RL

CVPR 2026

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipeline to strengthen the image understanding and generation capabilities while unlo

Cited by 0SourcecodeScholar
2025

REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents

ICCV 2025poster

Commercial video generation models have exhibited realistic, high-fidelity results but are still restricted to limited access.One crucial obstacle for large-scale applications is the expensive training and inference cost.In this paper, we argue that videos contain significantly more redundant inform…

2025

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

NeurIPS 2025poster

We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More imp…

Cited by 0SourceScholar
2024

DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

NeurIPS 2024poster

Most large multimodal models (LMMs) are implemented by feeding visual tokens as a sequence into the first layer of a large language model (LLM). The resulting architecture is simple but significantly increases computation and memory costs, as it has to handle a large number of additional tokens in…

Cited by 13SourcePDFScholar
2024

ISP-Teacher:Image Signal Process with Disentanglement Regularization for Unsupervised Domain Adaptive Dark Object Detection

AAAI 2024technical

Object detection in dark conditions has always been a great challenge due to the complex formation process of low-light images. Currently, the mainstream methods usually adopt domain adaptation with Teacher-Student architecture to solve the dark object detection problem, and they imitate the dark co…

2023

ResFormer: Scaling ViTs With Multi-Resolution Training

CVPR 2023poster

Vision Transformers (ViTs) have achieved overwhelming success, yet they suffer from vulnerable resolution scalability, i.e., the performance drops drastically when presented with input resolutions that are unseen during training. We introduce, ResFormer, a framework that is built upon the seminal id…

2022

Accurate and Robust Object SLAM With 3D Quadric Landmark Reconstruction in Outdoors

RA-L 2022

Object-oriented SLAM is a popular technology in autonomous driving and robotics. In this letter, we propose a stereo visual SLAM with a robust quadric landmark representation method.The system consists of four components, including deep learning detection, quadric landmark initialization, object dat

Cited by 27SourceScholar
2022

Object-Aware SLAM Based on Efficient Quadric Initialization and Joint Data Association

RA-L 2022

Semantic simultaneous localization and mapping (SLAM) is a popular technology enabling indoor mobile robots to sufficiently perceive and interact with the environment. In this paper, we propose an object-aware semantic SLAM system, which consists of a quadric initialization method, an object-level d

Cited by 19SourceScholar
2022

SemLoc: Accurate and Robust Visual Localization with Semantic and Structural Constraints from Prior Maps

ICRA 2022poster

Semantic information and geometrical structures of a prior map can be leveraged in visual localization to bound drift errors and improve accuracy. In this paper, we propose SemLoc, a pure visual localization system, for accurate localization in a prior semantic map. To tightly couple semantic and st…

Cited by 9SourceScholar
2021

Accurate and Robust Scale Recovery for Monocular Visual Odometry Based on Plane Geometry

ICRA 2021poster

Scale ambiguity is a fundamental problem in monocular visual odometry. Typical solutions include loop closure detection and environment information mining. For applications like self-driving cars, loop closure is not always available, hence mining prior knowledge from the environment becomes a more…

Cited by 34SourceScholar