← Search

Zhiwu Lu

43 accepted papers

2026

Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation

ICLR 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) offers a robust mechanism for enhancing mathematical reasoning in large models. However, we identify a systematic lack of emphasis on more challenging questions in existing methods from both algorithmic and data perspectives, despite their import…

Cited by 0SourcecodeScholar
2026

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

CVPR 2026

In this work, we introduce LLaDA-V, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instruction tuning with masked diffusion models, representing a departure from the autoregressive paradigms dominant in current multimodal approaches. Built upon LLaDA, a repres

Cited by 0SourcecodeScholar
2026

PortraitRL: Reinforcement Learning for Personalized Portrait Pose Transfer with Multi-Objective Reward Modeling

ICML 2026poster

Portrait pose transfer (PPT) requires generative models to preserve fine-grained identity details while following complex pose and layout modification instructions. Existing methods often struggle with extensive data annotation requirements or employ optimization objectives that are suboptimal for a…

Cited by 0SourceScholar
2026

Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits

CVPR 2026

As social media platforms proliferate, users increasingly demand intuitive ways to create diverse, high-quality portrait collections. In this work, we introduce Portrait Collection Generation (PCG), a novel task that generates coherent portrait collections by editing a reference portrait image throu

Cited by 0SourceScholar
2026

When would Vision-Proprioception Policies Fail in Robotic Manipulation?

ICLR 2026poster

Proprioceptive information is critical for precise servo control by providing real-time robotic states. Its collaboration with vision is highly expected to enhance performances of the manipulation policy in complex tasks. However, recent studies have reported inconsistent observations on the general…

Cited by 0SourcecodeScholar
2025

Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval

AAAI 2025technical

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion stra…

Cited by 2SourcePDFScholar
2025

MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents

ICLR 2025poster

Recently, Role-Playing Agents (RPAs) have garnered increasing attention for their potential to deliver emotional value and facilitate sociological research. However, existing studies are primarily confined to the textual modality, unable to simulate humans' multimodal perceptual capabilities. To bri…

2024

FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding

NeurIPS 2024poster

Contrastive Language-Image Pre-training (CLIP) achieves impressive performance on tasks like image classification and image-text retrieval by learning on large-scale image-text datasets. However, CLIP struggles with dense prediction tasks due to the poor grasp of the fine-grained details. Although e…

Cited by 3SourcePDFScholar
2024

Image Retrieval with Composed Query by Multi-Scale Multi-Modal Fusion

ICASSP 2024accepted

Image retrieval with composed query (IR-CQ) is a challenging task since it aims to retrieve the target image according to a hybrid-modality query which consists of a reference image and a text modifier. Previous approaches mainly focus on designing various multi-modal fusion modules to fuse the hybr…

Cited by 0SourceScholar
2024

Multi-Level Contrastive Learning For Hybrid Cross-Modal Retrieval

ICASSP 2024accepted

Hybrid image retrieval is a significant task for a wide range of applications. In this scenario, the hybrid query for searching images consists of a reference image and a text modifier. The reference image provides a vital visual context and displays some semantic details, while the text modifier sp…

Cited by 0SourceScholar
2024

Progressive Image Synthesis from Semantics to Details with Denoising Diffusion GAN

ICASSP 2024accepted

Although denoising diffusion probabilistic models (DDPMs) have shown remarkable progress in image generation, they typically face two main challenges: the time-expensive sampling process and the semantically meaningless latent space, which are often addressed separately in previous works. In particu…

Cited by 0SourceScholar
2024

UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling

ICLR 2024poster

Large-scale vision-language pre-trained models have shown promising transferability to various downstream tasks. As the size of these foundation models and the number of downstream tasks grow, the standard full fine-tuning paradigm becomes unsustainable due to heavy computational and storage costs.…

2024

Unsupervised Continual Learning of Image Representation Via Rememory-Based Simsiam

ICASSP 2024accepted

Unsupervised continual learning (UCL) of image representation has garnered attention due to practical need. However, recent UCL methods focus on mitigating the catastrophic forgetting with a replay buffer (i.e., rehearsal-based strategy), which needs much extra storage. To overcome this drawback, we…

Cited by 0SourceScholar
2024

VDT: General-purpose Video Diffusion Transformers via Mask Modeling

ICLR 2024poster

This work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation. It features transformer blocks with modularized temporal and spatial attention modules to leverage the rich spatial-temporal representation inherited in transformers. A…

2022

BMU-MoCo: Bidirectional Momentum Update for Continual Video-Language Modeling

NeurIPS 2022accept

Video-language models suffer from forgetting old/learned knowledge when trained with streaming data. In this work, we thus propose a continual video-language modeling (CVLM) setting, where models are supposed to be sequentially trained on five widely-used video-text datasets with different data dist…

Cited by 5SourcePDFScholar
2022

COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

CVPR 2022poster

Large-scale single-stream pre-training has shown dramatic performance in image-text retrieval. Regrettably, it faces low inference efficiency due to heavy attention layers. Recently, two-stream methods like CLIP and ALIGN with high inference efficiency have also shown promising performance, however,…

Cited by 81PDFScholar
2022

Fine-Grained Analysis of Stability and Generalization for Modern Meta Learning Algorithms

NeurIPS 2022accept

The support/query episodic training strategy has been widely applied in modern meta learning algorithms. Supposing the $n$ training episodes and the test episodes are sampled independently from the same environment, previous work has derived a generalization bound of $O(1/\sqrt{n})$ for smooth non-c…

Cited by 8SourcePDFScholar
2022

LGDN: Language-Guided Denoising Network for Video-Language Modeling

NeurIPS 2022accept

Video-language modeling has attracted much attention with the rapid growth of web videos. Most existing methods assume that the video frames and text description are semantically correlated, and focus on video-language modeling at video level. However, this hypothesis often fails for two reasons: (1…

Cited by 14SourcePDFScholar
2022

Learning Versatile Neural Architectures by Propagating Network Codes

ICLR 2022poster

This work explores how to design a single neural network capable of adapting to multiple heterogeneous vision tasks, such as image segmentation, 3D detection, and video recognition. This goal is challenging because both network architecture search (NAS) spaces and methods in different tasks are inco…

2022

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

AAAI 2022technical

Simultaneous Localization and Mapping (SLAM) and Autonomous Driving are becoming increasingly more important in recent years. Point cloud-based large scale place recognition is the spine of them. While many models have been proposed and have achieved acceptable performance by learning short-range lo…

Cited by 78SourcePDFScholar
2022

Visual Prompt Tuning for Few-Shot Text Classification

COLING 2022main

Deploying large-scale pre-trained models in the prompt-tuning paradigm has demonstrated promising performance in few-shot learning. Particularly, vision-language pre-training models (VL-PTMs) have been intensively explored in various few-shot downstream tasks. However, most existing works only apply…

2021

A Global Occlusion-Aware Approach to Self-Supervised Monocular Visual Odometry

AAAI 2021technical

Self-Supervised monocular visual odometry (VO) is often cast into a view synthesis problem based on depth and camera pose estimation. One of the key challenges is to accurately and robustly estimate depth with occlusions and moving objects in the scene. Existing methods simply detect and mask out re…

Cited by 6SourcePDFScholar
2021

Contrastive prototype learning with augmented embeddings for few-shot learning

UAI 2021poster

Most recent few-shot learning (FSL) methods are based on meta-learning with episodic training. In each meta-training episode, a discriminative feature embedding and/or classifier are first constructed from a support set in an inner loop, and then evaluated in an outer loop using a query set for mode…

Cited by 43SourcePDFScholar
2021

Counterfactual VQA: A Cause-Effect Look at Language Bias

CVPR 2021poster

Recent VQA models may tend to rely on language bias as a shortcut and thus fail to sufficiently learn the multi-modal knowledge from both vision and language. In this paper, we investigate how to capture and mitigate language bias in VQA. Motivated by causal effects, we proposed a novel counterfactu…

Cited by 491PDFcodeScholar
2021

HR-NAS: Searching Efficient High-Resolution Neural Architectures With Lightweight Transformers

CVPR 2021poster

High-resolution representations (HR) are essential for dense prediction tasks such as segmentation, detection, and pose estimation. Learning HR representations is typically ignored in previous Neural Architecture Search (NAS) methods that focus on image classification. This work proposes a novel NAS…

Cited by 74PDFcodeScholar
2021

IEPT: Instance-Level and Episode-Level Pretext Tasks for Few-Shot Learning

ICLR 2021poster

The need of collecting large quantities of labeled training data for each new task has limited the usefulness of deep neural networks. Given data from a set of source tasks, this limitation can be overcome using two transfer learning approaches: few-shot learning (FSL) and self-supervised learning (…

2021

L2M-GAN: Learning To Manipulate Latent Space Semantics for Facial Attribute Editing

CVPR 2021poster

A deep facial attribute editing model strives to meet two requirements: (1) attribute correctness -- the target attribute should correctly appear on the edited face image; (2) irrelevance preservation -- any irrelevant information (e.g., identity) should not be changed after editing. Meeting both re…

Cited by 86PDFcodeScholar
2021

MELR: Meta-Learning via Modeling Episode-Level Relationships for Few-Shot Learning

ICLR 2021poster

Most recent few-shot learning (FSL) approaches are based on episodic training whereby each episode samples few training instances (shots) per class to imitate the test condition. However, this strict adhering to test condition has a negative side effect, that is, the trained model is susceptible to…

Cited by 133SourcePDFScholar
2021

Self-Supervised Video Representation Learning with Constrained Spatiotemporal Jigsaw

IJCAI 2021poster

This paper proposes a novel pretext task for self-supervised video representation learning by exploiting spatiotemporal continuity in videos. It is motivated by the fact that videos are spatiotemporal by nature and a representation learned by detecting spatiotemporal continuity/discontinuity is thus…

Cited by 24SourcePDFScholar
2020

Learning Depth-Guided Convolutions for Monocular 3D Object Detection

CVPR 2020poster

3D object detection from a single image without LiDAR is a challenging task due to the lack of accurate depth information. Conventional 2D convolutions are unsuitable for this task because they fail to capture local object and its scale information, which are vital for 3D object detection. To better…

Cited by 384PDFcodeScholar
2019

Large-Scale Few-Shot Learning: Knowledge Transfer With Class Hierarchy

CVPR 2019poster

Recently, large-scale few-shot learning (FSL) becomes topical. It is discovered that, for a large-scale FSL problem with 1,000 classes in the source domain, a strong baseline emerges, that is, simply training a deep feature embedding model using the aggregated source classes and performing nearest n…

Cited by 164PDFcodeScholar
2019

Recursive Visual Attention in Visual Dialog

CVPR 2019oral

Visual dialog is a challenging vision-language task, which requires the agent to answer multi-round questions about an image. It typically needs to address two major problems: (1) How to answer visually-grounded questions, which is the core challenge in visual question answering (VQA); (2) How to in…

Cited by 144PDFcodeScholar
2018

Domain-Invariant Projection Learning for Zero-Shot Recognition

NeurIPS 2018poster

Zero-shot learning (ZSL) aims to recognize unseen object classes without any training samples, which can be regarded as a form of transfer learning from seen classes to unseen ones. This is made possible by learning a projection between a feature space and a semantic space (e.g. attribute space). Ke…

Cited by 66SourcePDFScholar