← Search

Mingyang Zhou

18 accepted papers

2026

Bayes-inspired Integration of Pretrained Priors and Few-Shot Evidence for Few-Shot Classification

ICML 2026poster

Few-shot classification aims to adapt a pretrained model to novel classes with limited examples. While current methods often heuristically combine pretrained knowledge and few-shot evidence, we seek a more principled understanding of their relationship. In this paper, we propose a Bayesian-inspired …

Cited by 0SourceScholar
2025

Adaptive Preference Arithmetic: A Personalized Agent with Adaptive Preference Arithmetic for Dynamic Preference Modeling

NeurIPS 2025poster

As large language models (LLMs) are increasingly used as personalized user assistants, effectively adapting to users' evolving preferences is critical for delivering high-quality personalized responses. While user preferences are often stable in content, their relative strengths shift over time due…

Cited by 0SourceScholar
2025

Denoising Diffusion Models are Good General Gaze Feature Learners

IJCAI 2025

Since the collection of labeled gaze data is laborious and time-consuming, methods which can learn generalizable features by leveraging large-scale available unlabeled data are desirable. In recent years, we have witnessed the tremendous capabilities of diffusion models in generating images as well

Cited by 0SourcePDFScholar
2025

Hierarchical Reward Modeling for Fault Localization in Large Code Repositories

EMNLP 2025

Large Language Models (LLMs) exhibit significant potential in complex software engineering tasks, however, their fault localization capabilities within repository are constrained by inherent limitations in max context length. Although Test-Time Scaling (TTS) can generate multiple candidate solutions

2025

IPSI: Enhancing Structural Inference with Automatically Learned Structural Priors

NeurIPS 2025poster

We propose IPSI, a general iterative framework for structural inference in interacting dynamical systems. It integrates a pretrained structural estimator and a joint inference module based on the Variational Autoencoder (VAE); these components are alternately updated to progressively refine the infe…

Cited by 0SourceScholar
2025

M2-TabFact: Multi-Document Multi-Modal Fact Verification with Visual and Textual Representations of Tabular Data

ACL 2025finding

Tabular data is used to store information in many real-world systems ranging from finance to healthcare. However, such structured data is often communicated to humans in visually interpretable formats (e.g. charts and textual paragraphs), making it imperative that fact-checking models should be able…

Cited by 0SourcePDFScholar
2025

Pretraining Context Compressor for Large Language Models with Embedding-Based Memory

ACL 2025long

Efficient processing of long contexts in large language models (LLMs) is essential for real-world applications like retrieval-augmented generation and in-context learning, especially in resource-constrained environments such as edge computing. This paper explores the embedding-based context compress…

Cited by 0SourcePDFScholar
2025

R-CHAR: A Metacognition-Driven Framework for Role-Playing in Large Language Models

EMNLP 2025

Role-playing capabilities in large language models (LLMs) often lack cognitive consistency in complex scenarios that require deep understanding and coherent reasoning. While recent reasoning models excel in math and coding tasks, they show limited effectiveness in open-ended role-playing scenarios.

Cited by 0SourcePDFScholar
2024

Aligning Large Language Models for Controllable Recommendations

ACL 2024long

Inspired by the exceptional general intelligence of Large Language Models (LLMs), researchers have begun to explore their application in pioneering the next generation of recommender systems — systems that are conversational, explainable, and controllable. However, existing literature primarily conc…

2024

Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning

ACL 2024findings

Advances in large vision-language models (LVLMs) have led to significant progress in generating natural language descriptions for visual contents. These powerful models are known for producing texts that are factually inconsistent with the visual input. While some efforts mitigate such inconsistenci…

2024

Modeling Personalized Retweeting Behaviors for Multi-Stage Cascade Popularity Prediction

IJCAI 2024poster

Predicting the size of message cascades is critical in various applications, such as online advertising and early detection of rumors. However, most existing deep learning approaches rely on cascade observation, which hinders accurate cascade prediction before message posting. Besides, these approac…

2024

Motif-oriented influence maximization for viral marketing in large-scale social networks

NeurIPS 2024poster

The influence maximization (IM) problem aims to identify a budgeted set of nodes with the highest potential to influence the largest number of users in a cascade model, a key challenge in viral marketing. Traditional \emph{IM} approaches consider each user/node independently as a potential target cu…

Cited by 0SourcePDFScholar
2023

Enhanced Chart Understanding via Visual Language Pre-training on Plot Table Pairs

ACL 2023findings

Building cross-model intelligence that can understand charts and communicate the salient information hidden behind them is an appealing challenge in the vision and language (V+L) community. The capability to uncover the underlined table data of chart figures is a critical key to automatic chart unde…

Cited by 0SourcePDFScholar
2023

Explainable Recommendation with Personalized Review Retrieval and Aspect Learning

ACL 2023long

Explainable recommendation is a technique that combines prediction and generation tasks to produce more persuasive results. Among these tasks, textual generation demands large amounts of data to achieve satisfactory accuracy. However, historical user reviews of items are often insufficient, making i…

2022

A Joint Learning Framework for Restaurant Survival Prediction and Explanation

EMNLP 2022main

The bloom of the Internet and the recent breakthroughs in deep learning techniques open a new door to AI for E-commence, with a trend of evolving from using a few financial factors such as liquidity and profitability to using more advanced AI techniques to process complex and multi-modal data. In th…

2022

Focus! Relevant and Sufficient Context Selection for News Image Captioning

EMNLP 2022finding

News Image Captioning requires describing an image by leveraging additional context derived from a news article. Previous works only coarsely leverage the article to extract the necessary context, which makes it challenging for models to identify relevant events and named entities. In our paper, we…

Cited by 11SourcePDFScholar
2022

Unsupervised Vision-and-Language Pre-Training via Retrieval-Based Multi-Granular Alignment

CVPR 2022oral

Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require pre-training on a large set of parallel image-text data, which is costly to collect, compared to image-only or text-only…

Cited by 41PDFScholar
2021

UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-Training

CVPR 2021poster

Vision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English languages, we introduce UC^2, the first machine translation-augmented framework for cross-lingual cross-modal representation…

Cited by 101PDFScholar