← Search

Zhang Zhang

30 accepted papers

2026

BaseReward: A Strong Baseline for Multimodal Reward Model

ICLR 2026poster

The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for achieving this goal, but a systematic guide for building state-of-the-art Multimodal Reward Models (MRMs) is currently l…

Cited by 0SourceScholar
2026

MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models

ICLR 2026poster

Unified Multimodal Large Language Models (U-MLLMs) have garnered considerable interest for their ability to seamlessly integrate generation and comprehension tasks. However, existing research lacks a unified evaluation standard, often relying on isolated benchmarks to assess these capabilities. More…

Cited by 0SourceScholar
2026

OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

ICML 2026poster

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and simple object manipulation, they often lack the systematic …

Cited by 0SourceScholar
2026

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

AAAI 2026technical

Although Vision Language Models (VLMs) have shown generalization in medical imaging, pathology presents unique challenges due to ultra-high resolution, complex tissue structures, and nuanced semantics. These factors make pathology VLMs prone to hallucinations, i.e., generating outputs inconsistent w

Cited by 0SourcePDFScholar
2026

Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner

AAAI 2026technical

Recent advances in vision-language models (VLMs) have enabled broad progress in the general medical field. However, pathology still remains a more challenging sub-domain, with current pathology-specific VLMs exhibiting limitations in both diagnostic accuracy and reasoning plausibility. Such shortcom

Cited by 0SourcePDFScholar
2026

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

ICLR 2026poster

Multimodal Reward Models (MRMs) play a crucial role in enhancing the performance of Multimodal Large Language Models (MLLMs). While recent advancements have primarily focused on improving the model structure and training data of MRMs, there has been limited exploration into the effectiveness of long…

Cited by 0SourcecodeScholar
2026

SimScale: Learning to Drive via Real-World Simulation at Scale

CVPR 2026

Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-distribution ones. However, such cases are underrepresented in real-world corpus collected by human experts. To complement for the lack of data diversity,

Cited by 0SourcecodeScholar
2026

The Geometric Origin of Grokking: Accelerating Generalization via Active Structural Reorganization

ICML 2026poster

Grokking, the phenomenon where models suddenly generalize long after overfitting training data, remains a puzzling challenge in neural network dynamics. Through mechanistic analysis, we find that this transition is fundamentally driven by a structural reorganization of token embeddings, with the ons…

Cited by 0SourceScholar
2025

Beyond Online Sampling: Bridging Offline-to-Online Alignment via Dynamic Data Transformation for LLMs

EMNLP 2025

While Direct Preference Optimization (DPO) eliminates complex reward modeling in aligning large language models (LLMs) with human preferences, its online variant faces significant efficiency bottlenecks due to costly real-time preference sampling and the reward model annotation. We propose a novel f

2025

DAA: Amplifying Unknown Discrepancy for Test-Time Discovery

NeurIPS 2025poster

Test-Time Discovery (TTD) addresses the critical challenge of identifying and adapting to novel classes during inference while maintaining performance on known classes, which is a capability essential for dynamic real-world environments such as healthcare and autonomous driving. Recent TTD methods a…

Cited by 0SourceScholar
2025

Integrating Expert Knowledge and Traffic Data for Lane-Changing Intention Prediction in Autonomous Vehicles

RA-L 2025

Accurate vehicle intention prediction is critical for autonomous driving safety in complex traffic environments. To address the interpretability limitations of data-driven methods while maintaining high accuracy, this letter proposes a knowledge-data co-learning framework featuring: (1) a knowledge-

Cited by 0SourceScholar
2025

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

ICML 2025poster

Existing efforts to align multimodal large language models (MLLMs) with human preferences have only achieved progress in narrow areas, such as hallucination reduction, but remain limited in practical applicability and generalizability. To this end, we introduce **MM-RLHF**, a dataset containing **12…

Cited by 13SourcePDFScholar
2025

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

ICLR 2025poster

Comprehensive evaluation of Multimodal Large Language Models (MLLMs) has recently garnered widespread attention in the research community. However, we observe that existing benchmarks present several common barriers that make it difficult to measure the significant challenges that models face in the…

Cited by 41SourcePDFScholar
2025

MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video conte…

Cited by 0SourceScholar
2024

Attribute-Guided Pedestrian Retrieval: Bridging Person Re-ID with Internal Attribute Variability

CVPR 2024poster

In various domains such as surveillance and smart retail pedestrian retrieval centering on person re-identification (Re-ID) plays a pivotal role. Existing Re-ID methodologies often overlook subtle internal attribute variations which are crucial for accurately identifying individuals with changing ap…

Cited by 9SourcePDFScholar
2024

CDUMA: An Adaptive Approach for Mitigating Confounder for MCQA

ICASSP 2024accepted

Multiple-choice question answering (MCQA) requires the model to select the correct answer from a set of candidate options when given a passage and a question. Previous research has achieved promising results with the assistance of Pre-trained Language Models(PrLMs). However, it has been observed tha…

Cited by 0SourceScholar
2024

Clear Up Confusion: Advancing Cross-Domain Few-Shot Relation Extraction through Relation-Aware Prompt Learning

NAACL 2024short

Cross-domain few-shot Relation Extraction (RE) aims to transfer knowledge from a source domain to a different target domain to address low-resource problems.Previous work utilized label descriptions and entity information to leverage the knowledge of the source domain.However, these models are prone…

Cited by 0SourcePDFScholar
2024

Fusion Makes Perfection: An Efficient Multi-Grained Matching Approach for Zero-Shot Relation Extraction

NAACL 2024short

Predicting unseen relations that cannot be observed during the training phase is a challenging task in relation extraction. Previous works have made progress by matching the semantics between input instances and label descriptions. However, fine-grained matching often requires laborious manual annot…

2023

AdaNPC: Exploring Non-Parametric Classifier for Test-Time Adaptation

ICML 2023poster

Many recent machine learning tasks focus to develop models that can generalize to unseen distributions. Domain generalization (DG) has become one of the key topics in various fields. Several literatures show that DG can be arbitrarily hard without exploiting target domain information. To address thi…

2023

Always the Best Fit: Adaptive Domain Gap Filling from Causal Perspective for Few-Shot Relation Extraction

EMNLP 2023short findings

Cross-domain Relation Extraction aims to transfer knowledge from a source domain to a different target domain to address low-resource challenges. However, the semantic gap caused by data bias between domains is a major challenge, especially in few-shot scenarios. Previous work has mainly focused on…

Cited by 0SourceScholar
2023

Free Lunch for Domain Adversarial Training: Environment Label Smoothing

ICLR 2023poster

A fundamental challenge for machine learning models is how to generalize learned models for out-of-distribution (OOD) data. Among various approaches, exploiting invariant features by Domain Adversarial Training (DAT) received widespread attention. Despite its success, we observe training instability…

2023

OneNet: Enhancing Time Series Forecasting Models under Concept Drift by Online Ensembling

NeurIPS 2023poster

Online updating of time series forecasting models aims to address the concept drifting problem by efficiently updating forecasting models based on streaming data. Many algorithms are designed for online time series forecasting, with some exploiting cross-variable dependency while others assume indep…

2022

Cross-Domain Cross-Set Few-Shot Learning via Learning Compact and Aligned Representations

ECCV 2022poster

"Few-shot learning (FSL) aims to recognize novel queries with only a few support samples through leveraging prior knowledge from a base dataset. In this paper, we consider the domain shift problem in FSL and aim to address the domain gap between the support set and the query set. Different from prev…

2022

Delving into Sample Loss Curve to Embrace Noisy and Imbalanced Data

AAAI 2022technical

Corrupted labels and class imbalance are commonly encountered in practically collected training data, which easily leads to over-fitting of deep neural networks (DNNs). Existing approaches alleviate these issues by adopting a sample re-weighting strategy, which is to re-weight sample by designing…

2019

Towards Rich Feature Discovery With Class Activation Maps Augmentation for Person Re-Identification

CVPR 2019poster

The fundamental challenge of small inter-person variation requires Person Re-Identification (Re-ID) models to capture sufficient fine-grained information. This paper proposes to discover diverse discriminative visual cues without extra assistance, e.g., pose estimation, human parsing. Specifically,…

Cited by 311PDFScholar
2018

Adversarially Occluded Samples for Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is the task of retrieving particular persons across different cameras. Despite its great progress in recent years, it is still confronted with challenges like pose variation, occlusion, and similar appearance among different persons. The large gap between training and…

Cited by 305SourcePDFScholar
2017

Learning Deep Context-Aware Features Over Body and Latent Parts for Person Re-Identification

CVPR 2017poster

Person Re-identification (ReID) is to identify the same person across different cameras. It is a challenging task due to the large variations in person pose, occlusion, background clutter, etc. How to extract powerful features is a fundamental problem in ReID and is still an open problem today. In t…

Cited by 829PDFScholar
2016

ReD-SFA: Relation Discovery Based Slow Feature Analysis for Trajectory Clustering

CVPR 2016poster

For spectral embedding/clustering, it is still an open problem on how to construct an relation graph to reflect the intrinsic structures in data. In this paper, we proposed an approach, named Relation Discovery based Slow Feature Analysis (ReD-SFA), for feature learning and graph construction simult…

Cited by 15PDFScholar