← Search

Zihua Zhao

7 accepted papers

2026

Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models

ICLR 2026poster

Although reinforcement learning with verifiable rewards (RLVR) shows promise in improving the reasoning ability of large language models (LLMs), the scaling up dilemma remains due to the reliance on human-annotated labels especially for complex tasks. Recent self-rewarding methods provide a label-fr…

Cited by 0SourcecodeScholar
2025

Consensus Graph-Based Spectral Ensemble Clustering via Low-Rank Tensor Learning

ICASSP 2025accepted

Ensemble clustering using co-association matrices integrates multiple base clusterings but often overlooks interactions between crucial samples and base clusterings. This neglect can introduce noise and lead to information loss and instability. To address these issues, we propose the Consensus Graph…

Cited by 0SourceScholar
2025

Differential-informed Sample Selection Accelerates Multimodal Contrastive Learning

ICCV 2025poster

The remarkable success of contrastive-learning-based multimodal models has been greatly driven by training on ever-larger datasets with expensive compute consumption. Sample selection as an alternative efficient paradigm plays an important direction to accelerate the training process. However, recen…

2025

Multi-modal Medical Diagnosis via Large-small Model Collaboration

CVPR 2025poster

Recent advances in medical AI have shown a clear trend towards large models in healthcare. However, developing large models for multi-modal medical diagnosis remains challenging due to a lack of sufficient modal-complete medical data. Most existing multi-modal diagnostic models are relatively small…

Cited by 0SourcePDFScholar
2024

Exploring Training on Heterogeneous Data with Mixture of Low-rank Adapters

ICML 2024poster

Training a unified model to take multiple targets into account is a trend towards artificial general intelligence. However, how to efficiently mitigate the training conflicts among heterogeneous data collected from different domains or tasks remains under-explored. In this study, we explore to lever…

2024

Mitigating Noisy Correspondence by Geometrical Structure Consistency Learning

CVPR 2024poster

Noisy correspondence that refers to mismatches in cross-modal data pairs is prevalent on human-annotated or web-crawled datasets. Prior approaches to leverage such data mainly consider the application of uni-modal noisy label learning without amending the impact on both cross-modal and intra-modal g…

2024

Probabilistic Conformal Distillation for Enhancing Missing Modality Robustness

NeurIPS 2024poster

Multimodal models trained on modality-complete data are plagued with severe performance degradation when encountering modality-missing data. Prevalent cross-modal knowledge distillation-based methods precisely align the representation of modality-missing data and that of its modality-complete counte…