← Search

Tianwei Lin

28 accepted papers

2026

H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation

AAAI 2026technical

Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data. Recent robotic foundation models often pre-train on cross-embodiment robot datasets to increase data scale, while they face significant limitations as the di

Cited by 0SourcePDFScholar
2026

OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis

ICLR 2026poster

Computed Tomography (CT) is one of the most widely used and diagnostically information-dense imaging modalities, covering critical organs such as the heart, lungs, liver, and colon. Clinical interpretation relies on both \textbf{slice-driven} local features (e.g., sub-centimeter nodules, lesion boun…

Cited by 0SourcecodeScholar
2026

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis

ICML 2026poster

The advancement of Medical Vision-Language Models (VLMs) for 3D Computed Tomography (CT) analysis is hindered by a misalignment between optimization objectives and clinical rigor. Current Reinforcement Learning (RL) paradigms rely on lexical proxy signals that induce ``\textbf{evaluation hallucinati…

Cited by 0SourceScholar
2026

SEM: Enhancing Spatial Understanding for Robust Robot Manipulation

ICRA 2026poster

A key challenge in robot manipulation lies in developing policy models with consistent spatial understanding—the ability to reason about 3D geometry, object relations, and robot state. Existing mainstream models take 2D images as input, without performing explicit 3D modeling, and thus lack spatial …

2026

Self-Improving Robot Policy with Compositional World Model

RSS 2026poster

Despite the sustained scaling on model capacity and data acquisition, Vision–Language–Action (VLA) models remain brittle in contact-rich and dynamic manipulation tasks, where minor execution deviations can compound into failures. While reinforcement learning (RL) offers a principled path to robustne…

Cited by 0SourceScholar
2026

TumorChain: Interleaved Multimodal Chain-of-Thought Reasoning for Traceable Clinical Tumor Analysis

ICLR 2026poster

Accurate tumor analysis is central to clinical radiology and precision oncology, where early detection, reliable lesion characterization, and pathology-level risk assessment directly guide diagnosis, staging, and treatment planning. Chain-of-Thought (CoT) reasoning is particularly critical in this s…

Cited by 0SourcecodeScholar
2026

Unified Personalized Understanding, Generating and Editing

CVPR 2026

Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a "one-size-fits-all" paradigm and struggle to model user-specific concepts (e.g., generate a photo of \texttt \<maeve> ) in a consis

Cited by 6SourceScholar
2025

BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence

CVPR 2025poster

In embodied intelligence systems, a key component is 3D perception algorithm, which enables agents to understand their surrounding environments. Previous algorithms primarily rely on point cloud, which, despite offering precise geometric information, still constrain perception performance due to inh…

2025

GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding

CVPR 2025poster

3D Semantic Occupancy Prediction is fundamental for spatial understanding, yet existing approaches face challenges in scalability and generalization due to their reliance on extensive labeled data and computationally intensive voxel-wise representations. In this paper, we introduce GaussTR, a novel…

2025

Generating Multimodal Driving Scenes via Next-Scene Prediction

CVPR 2025poster

Generative models in Autonomous Driving (AD) enable diverse scenario creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multi…

2025

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

ICML 2025spotlight

We present **HealthGPT**, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregressive paradigm. Our bootstrapping philosophy is to progressively adapt heterogeneous comprehension and generation kno…

2025

TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition

ACL 2025long

While Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) effectively address resource constraints during fine-tuning, their performance often falls short, especially in multidimensional task scenarios. To address this issue, one straightforward solution is to introduce ta…

2024

EDA: Evolving and Distinct Anchors for Multimodal Motion Prediction

AAAI 2024technical

Motion prediction is a crucial task in autonomous driving, and one of its major challenges lands in the multimodality of future behaviors. Many successful works have utilized mixture models which require identification of positive mixture components, and correspondingly fall into two main lines: pre…

2024

Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention Modulation

CVPR 2024poster

Recently diffusion-based methods like InstructPix2Pix (IP2P) have achieved effective instruction-based image editing requiring only natural language instructions from the user. However these methods often inadvertently alter unintended areas and struggle with multi-instruction editing resulting in c…

2024

Symphonize 3D Semantic Scene Completion with Contextual Instance Queries

CVPR 2024poster

3D Semantic Scene Completion (SSC) has emerged as a nascent and pivotal undertaking in autonomous driving aiming to predict the voxel occupancy within volumetric scenes. However prevailing methodologies primarily focus on voxel-wise feature aggregation while neglecting instance semantics and scene c…

2024

WidthFormer: Toward Efficient Transformer-based BEV View Transformation

IROS 2024poster

We present WidthFormer, a novel transformer-based module to compute Bird’s-Eye-View (BEV) representations from multi-view cameras for real-time autonomous-driving applications. WidthFormer is computationally efficient, robust and does not require any special engineering effort to deploy. We first in…

Cited by 3SourcecodeScholar
2023

AdaCM: Adaptive ColorMLP for Real-Time Universal Photo-Realistic Style Transfer

AAAI 2023technical

Photo-realistic style transfer aims at migrating the artistic style from an exemplar style image to a content image, producing a result image without spatial distortions or unrealistic artifacts. Impressive results have been achieved by recent deep models. However, deep neural network based methods…

Cited by 4SourcePDFScholar
2023

Master: Meta Style Transformer for Controllable Zero-Shot and Few-Shot Artistic Style Transfer

CVPR 2023poster

Transformer-based models achieve favorable performance in artistic style transfer recently thanks to its global receptive field and powerful multi-head/layer attention operations. Nevertheless, the over-paramerized multi-layer structure increases parameters significantly and thus presents a heavy bu…

Cited by 20SourcePDFScholar
2023

Planning-Oriented Autonomous Driving

CVPR 2023poster

Modern autonomous driving system is characterized as modular tasks in sequential order, i.e., perception, prediction, and planning. In order to perform a wide diversity of tasks and achieve advanced-level intelligence, contemporary approaches either deploy standalone models for individual tasks, or…

2022

Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model

CVPR 2022poster

To achieve disentangled image manipulation, previous works depend heavily on manual annotation. Meanwhile, the available manipulations are limited to a pre-defined set the models were trained for. We propose a novel framework, i.e., Predict, Prevent, and Evaluate (PPE), for disentangled text-driven…

Cited by 49PDFcodeScholar
2022

Towards Bidirectional Arbitrary Image Rescaling: Joint Optimization and Cycle Idempotence

CVPR 2022poster

Deep learning based single image super-resolution models have been widely studied and superb results are achieved in upscaling low-resolution images with fixed scale factor and downscaling degradation kernel. To improve real world applicability of such models, there are growing interests to develop…

Cited by 39PDFScholar
2021

AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer

ICCV 2021poster

Fast arbitrary neural style transfer has attracted widespread attention from academic, industrial and art communities due to its flexibility in enabling various applications. Existing solutions either attentively fuse deep style feature into deep content feature without considering feature distribut…

Cited by 444PDFcodeScholar
2021

Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer

CVPR 2021poster

Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize comple…

Cited by 119PDFcodeScholar
2021

Learning Semantic Person Image Generation by Region-Adaptive Normalization

CVPR 2021poster

Human pose transfer has received great attention due to its wide applications, yet is still a challenging task that is not well solved. Recent works have achieved great success to transfer the person image from the source to the target pose. However, most of them cannot well capture the semantic app…

Cited by 81PDFcodeScholar
2021

MVFNet: Multi-View Fusion Network for Efficient Video Recognition

AAAI 2021technical

Conventionally, spatiotemporal modeling network and its complexity are the two most concentrated research topics in video action recognition. Existing state-of-the-art methods have achieved excellent accuracy regardless of the complexity meanwhile efficient spatiotemporal modeling solutions are slig…

2021

Paint Transformer: Feed Forward Neural Painting With Stroke Prediction

ICCV 2021poster

Neural painting refers to the procedure of producing a series of strokes for a given image and non-photo-realistically recreating it using neural networks. While reinforcement learning (RL) based agents can generate a stroke sequence step by step for this task, it is not easy to train a stable RL ag…

Cited by 93PDFcodeScholar
2019

BMN: Boundary-Matching Network for Temporal Action Proposal Generation

ICCV 2019poster

Temporal action proposal generation is an challenging and promising task which aims to locate temporal regions in real-world videos where action or event may occur. Current bottom-up proposal generation methods can generate proposals with precise boundary, but cannot efficiently generate adequately…

Cited by 791PDFcodeScholar
2018

BSN: Boundary Sensitive Network for Temporal Action Proposal Generation

ECCV 2018poster

Temporal action proposal generation is an important yet challenging problem, since temporal proposals with rich action content are indispensable for analysing real-world videos with long duration and high proportion irrelevant content. This problem requires methods not only generating proposals with…