← Search

Yanru Wu

13 accepted papers

2026

CORE-MTL: Rethinking Gradient Balancing via Causal Orthogonal Representations

ICML 2026poster

Multi-task learning (MTL) aims to construct a joint model for multiple tasks by sharing a common representation across domains. To achieve this goal, existing optimization-centric methods either balance task gradients or modify the shared architecture. However, as these approaches remain agnostic to…

Cited by 0SourceScholar
2026

Learning Optimal Prompt Ensemble for Multi-source Visual Prompt Transfer

AAAI 2026technical

Prompt tuning has emerged as a lightweight strategy for adapting foundation models to downstream tasks, particularly for resource-constrained systems. As pre-trained prompts become valuable assets, combining multiple source prompts offers a promising approach to enhance generalization for new tasks

Cited by 0SourcePDFScholar
2026

TMT: Cross-domain Semantic Segmentation with Region-adaptive Transferability Estimation

ICASSP 2026poster

Recent advances in Vision Transformers (ViTs) have significantly advanced semantic segmentation performance. However, their adaptation to new target domains remains challenged by distribution shifts, which often disrupt global attention mechanisms. While existing global and patch-level adaptation me…

Cited by 0SourcePDFScholar
2025

A High-Dimensional Statistical Method for Optimizing Transfer Quantities in Multi-Source Transfer Learning

NeurIPS 2025poster

Multi-source transfer learning provides an effective solution to data scarcity in real-world supervised learning scenarios by leveraging multiple source tasks. In this field, existing works typically use all available samples from sources in training, which constrains their training efficiency and m…

Cited by 0SourcecodeScholar
2025

Exploiting Task Relationships in Continual Learning via Transferability-Aware Task Embeddings

NeurIPS 2025poster

Continual learning (CL) has been a critical topic in contemporary deep neural network applications, where higher levels of both forward and backward transfer are desirable for an effective CL performance. Existing CL strategies primarily focus on task models — either by regularizing model updates or…

Cited by 0SourcecodeScholar
2025

Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment

ICLR 2025spotlight

Many real-world user queries (e.g. *"How do to make egg fried rice?"*) could benefit from systems capable of generating responses with both textual steps with accompanying images, similar to a cookbook. Models designed to generate interleaved text and images face challenges in ensuring consistency w…

Cited by 8SourcePDFScholar
2025

Reinforced Domain Selection for Continuous Domain Adaptation

ICASSP 2025accepted

Continuous Domain Adaptation (CDA) effectively bridges significant domain shifts by progressively adapting from the source domain through intermediate domains to the target domain. However, selecting intermediate domains without explicit metadata remains a substantial challenge that has not been ext…

Cited by 0SourceScholar
2025

The Impact of Large Language Models in Academia: from Writing to Speaking

ACL 2025finding

Large language models (LLMs) are increasingly impacting human society, particularly in textual information. Based on more than 30,000 papers and 1,000 presentations from machine learning conferences, we examined and compared the words used in writing and speaking, representing the first large-scale…

Cited by 0SourcePDFScholar
2025

TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review

EMNLP 2025

While Large Language Models (LLMs) have shown significant potential in assisting peer review, current methods often struggle to generate thorough and insightful reviews while maintaining efficiency. In this paper, we propose TreeReview, a novel framework that models paper review as a hierarchical an

2025

pFedGPA: Diffusion-based Generative Parameter Aggregation for Personalized Federated Learning

AAAI 2025technical

Federated Learning (FL) offers a decentralized approach to model training, where data remains local and only model parameters are shared between the clients and the central server. Traditional methods, such as Federated Averaging (FedAvg), linearly aggregate these parameters which are usually traine…

Cited by 0SourcePDFScholar
2024

Balancing Speciality and Versatility: a Coarse to Fine Framework for Supervised Fine-tuning Large Language Model

ACL 2024findings

Aligned Large Language Models (LLMs) showcase remarkable versatility, capable of handling diverse real-world tasks. Meanwhile, aligned LLMs are also expected to exhibit speciality, excelling in specific applications. However, fine-tuning with extra data, a common practice to gain speciality, often l…

2024

H-ensemble: An Information Theoretic Approach to Reliable Few-Shot Multi-Source-Free Transfer

AAAI 2024technical

Multi-source transfer learning is an effective solution to data scarcity by utilizing multiple source tasks for the learning of the target task. However, access to source data and model details is limited in the era of commercial models, giving rise to the setting of multi-source-free (MSF) transfer…

Cited by 3SourcePDFScholar
2023

SDG-L: A Semiparametric Deep Gaussian Process based Framework for Battery Capacity Prediction

ICASSP 2023accepted

Lithium-ion batteries are becoming increasingly omnipresent in energy supply. However, the durability of energy storage using lithium-ion batteries is threatened by their dropping capacity with the growing number of charging/discharging cycles. An accurate capacity prediction is the key to ensure sy…

Cited by 0SourceScholar