← Search

Zhuowei Li

12 accepted papers

2026

Decoupling Vision and Language: Codebook Anchored Visual Adaptation

CVPR 2026

Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders often underperform in domain-specific visual tasks such as medical image diagnosis or fine-grained classification, where representation errors can cascad

Cited by 0SourceScholar
2026

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

ICLR 2026poster

While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to identify trustworthy responses, especially in complex tasks requiring multi-step reasoning. In this paper, we propose a…

Cited by 0SourcecodeScholar
2026

VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models

CVPR 2026

Advances in large reasoning models have shown strong performance on complex reasoning tasks by scaling test-time compute through extended inference-time thinking. However, recent studies observe that in vision-dependent tasks, extended textual reasoning at inference time can often degrade performanc

Cited by 0SourceScholar
2025

A Robotic System for Long-Term Personalized Automated Cultivation of Colorectal Cancer Organoids

RA-L 2025

Organoids are a class of popular three-dimensional in vitro models that recapitulate the structural, genetic, and functional characteristics of their native tissues, offering powerful tools for drug screening and precision medicine. While recent advances have integrated robotic automation into organ

Cited by 0SourceScholar
2025

MLLM-as-a-Judge for Image Safety without Human Labeling

CVPR 2025highlight

Image content safety has become a significant challenge with the rise of visual media on online platforms. Meanwhile, in the age of AI-generated content (AIGC), many image generation models are capable of producing harmful content, such as images containing sexual or violent material. Thus, it becom…

Cited by 2SourcePDFScholar
2025

Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing

CVPR 2025poster

Developing a face anti-spoofing model that meets the security requirements of clients worldwide is challenging due to the domain gap between training datasets and the diverse end-user test data. Moreover, for security and privacy reasons, it is undesirable for clients to share a large amount of thei…

Cited by 0SourcePDFScholar
2025

Show and Segment: Universal Medical Image Segmentation via In-Context Learning

CVPR 2025poster

Medical image segmentation remains challenging due to the vast diversity of anatomical structures, imaging modalities, and segmentation tasks. While deep learning has made significant advances, current approaches struggle to generalize as they require task-specific training or fine-tuning on unseen…

Cited by 0SourcePDFScholar
2025

The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information Steering

ICML 2025poster

Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded contents. In this paper, we investigate the internal dynamics of hallucination by examining the tokens logits rankings througho…

2024

A Retinex Structure-based Low-light Enhancement Model Guided by Spatial Consistency

ICRA 2024poster

Images captured by robotics under low-light conditions are often plagued by several challenges, including diminished contrast, increased noise, loss of fine details, and unnatural color reproduction. These factors can significantly hinder the performance of computer vision tasks such as object detec…

Cited by 11SourceScholar
2024

Voltage Regulation in Polymer Electrolyte Fuel Cell Systems Using Gaussian Process Model Predictive Control

IROS 2024poster

This study presents a novel approach using Gaussian process model predictive control (MPC) to stabilize the output voltage of a polymer electrolyte fuel cell (PEFC) by regulating hydrogen and airflow rates. Two Gaussian process models capture PEFC dynamics, accounting for constraints like hydrogen p…

Cited by 3SourceScholar
2023

DeFormer: Integrating Transformers with Deformable Models for 3D Shape Abstraction from a Single Image

ICCV 2023poster

Explicit 3D shape abstraction from a single 2D image is a long-standing problem in computer vision and graphics. By leveraging a set of primitives to represent the target shape, recent methods have achieved promising results. However, these methods either use a relatively larger number of primitives…

Cited by 8PDFScholar