← Search

Shengju Qian

11 accepted papers

2026

AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer

ICLR 2026poster

The digital industry demands high-quality, diverse modular 3D assets, especially for user-generated content (UGC). In this work, we introduce AssetFormer, an autoregressive Transformer-based model designed to generate modular 3D assets from textual descriptions. Our pilot study leverages real-world…

Cited by 0SourcecodeScholar
2026

CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modeling

CVPR 2026

Open-world promptable 3D semantic segmentation remains brittle as semantics are inferred in the input sensor coordinates. Yet, humans, in contrast, interpret parts via functional roles in a canonical space -- wings extend laterally, handles protrude to the side, and legs support from below. Psychoph

Cited by 0SourcecodeScholar
2026

Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation

CVPR 2026

Multimodal autoregressive (AR) models, based on next-token prediction and transformer architecture, have demonstrated remarkable capabilities in various multimodal tasks including text-to-image (T2I) generation. Despite their strong performance in general T2I tasks, our research reveals that these m

Cited by 0SourceScholar
2025

MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation

CVPR 2025highlight

Recent advances in auto-regressive transformers have revolutionized generative modeling across different domains, from language processing to visual generation, demonstrating remarkable capabilities. However, applying these advances to 3D generation presents three key challenges: the unordered natur…

Cited by 1SourcePDFScholar
2024

LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

ICLR 2024oral

We present LongLoRA, an efficient fine-tuning approach that extends the context sizes of pre-trained large language models (LLMs), with limited computation cost. Typically, training LLMs with long context sizes is computationally expensive, requiring extensive training hours and GPU resources. For e…

2024

Prompt Highlighter: Interactive Control for Multi-Modal LLMs

CVPR 2024poster

This study targets a critical aspect of multi-modal LLMs' (LLMs&VLMs) inference: explicit controllable text generation. Multi-modal LLMs empower multi-modality understanding with the capability of semantic generation yet bring less explainability and heavier reliance on prompt contents due to their…

2024

Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

NeurIPS 2024spotlight

Multi-Modal Large Language Models (MLLMs) have demonstrated impressive performance in various VQA tasks. However, they often lack interpretability and struggle with complex visual inputs, especially when the resolution of the input image is high or when the interested region that could provide key i…

2023

On Efficient Transformer-Based Image Pre-training for Low-Level Vision

IJCAI 2023poster

Pre-training has marked numerous state of the arts in high-level computer vision, while few attempts have ever been made to investigate how pre-training acts in image processing systems. In this paper, we tailor transformer-based pre-training regimes that boost various low-level tasks. To comprehens…

2019

Aggregation via Separation: Boosting Facial Landmark Detector With Semi-Supervised Style Translation

ICCV 2019poster

Facial landmark detection, or face alignment, is a fundamental task that has been extensively studied. In this paper, we investigate a new perspective of facial landmark detection and demonstrate it leads to further notable improvement. Given that any face images can be factored into space of style…

Cited by 102PDFcodeScholar
2019

Make a Face: Towards Arbitrary High Fidelity Face Manipulation

ICCV 2019poster

Recent studies have shown remarkable success in face manipulation task with the advance of GANs and VAEs paradigms, but the outputs are sometimes limited to low-resolution and lack of diversity. In this work, we propose Additive Focal Variational Auto-encoder (AF-VAE), a novel approach that can arbi…

Cited by 86PDFScholar