← Search

Jiawei Zhao

22 accepted papers

2026

GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward Model

ICML 2026poster

Aligning Multimodal Large Language Models (MLLMs) with human preferences remains a fundamental challenge. While Generative Reward Models (GRMs) offer a promising reasoning-based alternative to scalar models, they are often hindered by severe position bias and prohibitively high computational overhea…

Cited by 0SourceScholar
2026

Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?

ICLR 2026poster

Reinforcement learning has been central to recent advances in large language model reasoning, but most algorithms rely on on-policy training that demands fresh rollouts at every update, limiting efficiency and scalability. Asynchronous RL systems alleviate this by decoupling rollout generation from…

Cited by 0SourcecodeScholar
2026

VCP-Attack: Visual-Contrastive Projection for Transferable Black-Box Targeted Attacks on Large Vision-Language Models

CVPR 2026

Large vision-language models (LVLMs) have achieved impressive performance across a variety of multimodal tasks, yet remain vulnerable to targeted adversarial attacks, particularly in black-box settings. In this paper, we propose VCP-Attack, a transferable targeted attack framework that combines stru

Cited by 0SourceScholar
2026

WMVLM: Evaluating Diffusion Model Image Watermarking via Vision-Language Models

ICML 2026poster

Digital watermarking is essential for securing generated images from diffusion models. Accurate watermark evaluation is critical for algorithm development, yet existing methods have significant limitations: they lack a unified framework for both residual and semantic watermarks, provide results with…

Cited by 0SourceScholar
2025

A Study on the Generation of Single Cell Droplets via the Combination of Lateral-Field Optoelectronic Tweezers and Electrowetting-on-Dielectric

IROS 2025

Microfluidic technology is currently a popular approach in the field of single-cell research, which is used to reveal the heterogeneity among cells. However, most of the existing microfluidic technologies for single-cell research lack the ability to control the microenvironment of single cells after

Cited by 0SourceScholar
2025

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

NeurIPS 2025poster

Reinforcement learning, such as PPO and GRPO, has powered recent breakthroughs in LLM reasoning. Scaling rollout to sample more prompts enables models to selectively use higher-quality data for training, which can stabilize RL training and improve model performance, but at the cost of significant co…

Cited by 0SourcecodeScholar
2025

From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications

ICML 2025poster

Large Language Models (LLMs) matrices can often be expressed in low-rank format with potential to relax memory and compute resource requirements. Unlike previous works which pivot around developing novel matrix decomposition algorithms, in this work we focus to study the emerging non-uniform low-ran…

Cited by 0SourcePDFScholar
2025

High-Precision Parallel Manipulation of Multi-Particle System Using Optoelectronic Tweezers

IROS 2025

This paper presents a multi-particle parallel manipulation optoelectronic tweezers system integrated with computer vision technology, enabling the parallel and precise manipulation of dozens of particles. This system significantly enhances manipulation efficiency while maintaining high precision. By

Cited by 0SourceScholar
2025

ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization

NeurIPS 2025poster

The optimal bit-width for achieving the best trade-off between quantized model size and accuracy has been a subject of ongoing debate. While some advocate for 4-bit quantization, others propose that 1.58-bit offers superior results. However, the lack of a cohesive framework for different bits has le…

Cited by 0SourceScholar
2025

SQL Injection Jailbreak: A Structural Disaster of Large Language Models

ACL 2025finding

Large Language Models (LLMs) are susceptible to jailbreak attacks that can induce them to generate harmful content.Previous jailbreak methods primarily exploited the internal properties or capabilities of LLMs, such as optimization-based jailbreak methods and methods that leveraged the model’s conte…

2024

Dynamic Adaptive Imaging System on Optoelectronic Tweezers Platform

ICRA 2024poster

Optoelectronic tweezers (OET) has shown great promise in various applications, especially in the precise manipulation of microparticles and microorganisms on a micron and nanometer scale. This technology significantly enhances the efficiency of single-cell sorting and the development of antibody-bas…

Cited by 2SourceScholar
2024

GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

ICML 2024oral

Training Large Language Models (LLMs) presents significant memory challenges, predominantly due to the growing size of weights and optimizer states. Common memory-reduction approaches, such as low-rank adaptation (LoRA), add a trainable low-rank matrix to the frozen pre-trained weight in each layer,…

2024

Mini-Sequence Transformers: Optimizing Intermediate Memory for Long Sequences Training

NeurIPS 2024poster

We introduce Mini-Sequence Transformer (MsT), a simple and effective methodology for highly efficient and accurate LLM training with extremely long sequences. MsT partitions input sequences and iteratively processes mini-sequences to reduce intermediate memory usage. Integrated with activation recom…

Cited by 0SourcePDFScholar
2024

S$^{2}$FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity

NeurIPS 2024poster

Current PEFT methods for LLMs can achieve high quality, efficient training, or scalable serving, but not all three simultaneously. To address this limitation, we investigate sparse fine-tuning and observe a remarkable improvement in generalization ability. Utilizing this key insight, we propose a…

Cited by 3SourcePDFScholar
2023

Microrobot Control Method Based on Movement of Field Free Point in Gradient Magnetic Field

IROS 2023poster

The untethered microrobots driven by multiple external physics fields have promising ability in minimally invasive disease treatments. One common type of the driving fields is gradient magnetic field, which can provide microrobots with adequate driving force in complicated environment. In this study…

Cited by 2SourceScholar
2023

Parallel Cell Array Patterning and Target Cell Lysis on an Optoelectronic Micro-Well Device

IROS 2023poster

This work presents a novel electrical method, implemented in the form of a microfluidic device, for cell arraying and target cell lysis. The microfluidic device contains a micro-well array on the photoconductive layer based on the optoelectronic tweezers (OET) method, where parallel cell manipulatio…

Cited by 1SourceScholar
2023

UMIRobot: An Open-{Software, Hardware} Low-Cost Robotic Manipulator for Education

IROS 2023poster

Robot teleoperation has been studied for the past 70 years and is relevant in many contexts, such as in the handling of hazardous materials and telesurgery. The COVID19 pandemic has rekindled interest in this topic, but the existing robotic education kits fall short of being suitable for teleoperate…

Cited by 1SourcecodeScholar
2021

Transformer-Based Dual Relation Graph for Multi-Label Image Recognition

ICCV 2021poster

The simultaneous recognition of multiple objects in one image remains a challenging task, spanning multiple events in the recognition field such as various object scales, inconsistent appearances, and confused inter-class relationships. Recent research efforts mainly resort to the statistic label co…

Cited by 120PDFcodeScholar
2020

Learning compositional functions via multiplicative weight updates

NeurIPS 2020poster

Compositionality is a basic structural feature of both biological and artificial neural networks. Learning compositional functions via gradient descent incurs well known problems like vanishing and exploding gradients, making careful learning rate tuning essential for real-world applications. This p…

2019

signSGD with Majority Vote is Communication Efficient and Fault Tolerant

ICLR 2019poster

Training neural networks on large datasets can be accelerated by distributing the workload over a network of machines. As datasets grow ever larger, networks of hundreds or thousands of machines become economically viable. The time cost of communicating gradients limits the effectiveness of using su…

Cited by 227SourcePDFScholar