← Search

Yun Li

47 accepted papers

2026

Learning Latent Proxies for Controllable Single-Image Relighting

CVPR 2026

Single-image relighting is highly under-constrained: small illumination changes can produce large, nonlinear variations in shading, shadows, and specularities, while geometry and materials remain unobserved. Existing diffusion-based approaches either rely on intrinsic- or G-buffer-based pipelines th

Cited by 0SourceScholar
2026

Learning from Comparison: Constrained Projection Policy Optimization for Pareto-Front Improvement

ICML 2026poster

Constrained multi-objective reinforcement learning aims to discover a diverse set of feasible trade-offs, yet scalarization and signed, normalized group-relative advantages can be brittle under objective-scale drift, near-ties, and feasibility scarcity. We propose constrained projection policy optim…

Cited by 0SourceScholar
2026

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

ICLR 2026poster

Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across diverse medical specialties, limiting their performance. Recent efforts introduce multi-agent collaboration frameworks insp…

Cited by 0SourceScholar
2025

Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models

ACL 2025long

Designing complex computer-aided design (CAD) models is often time-consuming due to challenges such as computational inefficiency and the difficulty of generating precise models. We propose a novel language-guided framework for industrial design automation to address these issues, integrating large…

2025

Better Process Supervision with Bi-directional Rewarding Signals

ACL 2025finding

Process supervision, i.e., evaluating each step, is critical for complex large language model (LLM) reasoning and test-time searching with increased inference compute. Existing approaches, represented by process reward models (PRMs), primarily focus on rewarding signals up to the current step, exhib…

2025

CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models

EMNLP 2025

Large Vision-Language Models (LVLMs) process multimodal inputs consisting of text tokens and vision tokens extracted from images or videos. Due to the rich visual information, a single image can generate thousands of vision tokens, leading to high computational costs during the prefilling stage and

2025

Collaborative Document Simplification Using Multi-Agent Systems

COLING 2025main

Research on text simplification has been ongoing for many years. However, the task of document simplification (DS) remains a significant challenge due to the need to consider complex factors such as technical terminology, metaphors, and overall coherence. In this work, we introduce a novel multi-age…

Cited by 3SourcePDFScholar
2025

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective

AAAI 2025technical

Recent Large Vision-Language Models (LVLMs) have shown promising reasoning capabilities on text-rich images from charts, tables, and documents. However, the abundant text within such images may increase the model's sensitivity to language. This raises the need to evaluate LVLM performance on cross-…

2025

DCASI: A Sequence-based Attack Investigation Method Using DTW Contrastive Learning

ICASSP 2025accepted

The stealth and persistence of APT attacks make investigation particularly challenging, further complicated by the diversity and volume of host logs. Existing methods, though effective, have limitations: 1) They rely heavily on manual processing and complex models that often fail to capture temporal…

Cited by 0SourceScholar
2025

Design of a Bio-Inspired Stiffness Controllable Continuum Robot for Object Grasping and Moving

RA-L 2025

Continuum robots (CRs) possess better compliance than rigid manipulators. However, existing CRs suffer from difficulties in manipulating objects for the conflicting needs of high stiffness and flexibility. This letter proposes an elephant trunk-inspired CR for grasping and moving objects. The CR fea

Cited by 2SourceScholar
2025

From Parameters to Performance: A Data-Driven Study on LLM Structure and Development

EMNLP 2025

Large language models (LLMs) have achieved remarkable success across various domains, driving significant technological advancements and innovations. Despite the rapid growth in model scale and capability, systematic, data-driven research on how structural configurations affect performance remains s

2025

Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News Detection

AAAI 2025technical

The questionable responses caused by knowledge hallucination may lead to LLMs' unstable ability in decision-making. However, it has never been investigated whether the LLMs' hallucination is possibly usable for generating negative reasoning to assist fake news detection. In this paper, we propose a…

Cited by 0SourcePDFScholar
2025

LIFTED: Multimodal Clinical Trial Outcome Prediction via Large Language Models and Mixture-of-Experts

EMNLP 2025

Clinical trials are pivotal yet costly processes, often spanning multiple years and requiring substantial expenses, motivating predictive models to identify likely-to-fail drugs early and save resources. Recent approaches leverage deep learning to integrate multimodal data for clinical outcome predi

Cited by 0SourcePDFScholar
2025

Learning Simultaneous Facial Canonical Correlation Representation for Face Hallucination

ICASSP 2025accepted

The low resolution (LR) problem is rather challenging in face analysis. Most existing face hallucination methods assume that LR face images have only one resolution, but multiple resolutions may be available from different sources. To solve this issue, we propose a novel simultaneous facial canonica…

Cited by 0SourceScholar
2025

MEGen: Generative Backdoor into Large Language Models via Model Editing

ACL 2025finding

Large language models (LLMs) have exhibited remarkable versatility and adaptability, while their widespread adoption across various applications also raises critical safety concerns.This paper focuses on the impact of backdoored LLMs. Traditional backdoor injection methods are primarily limited to y…

2025

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

ICML 2025poster

The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due to modality misalignment, where the models prioritize textual knowledge over visual input, leading to hallucinations th…

2025

Multi-PrefDrive: Optimizing Large Language Models for Autonomous Driving Through Multi-Preference Tuning

IROS 2025

This paper introduces Multi-PrefDrive, a framework that significantly enhances LLM-based autonomous driving through multidimensional preference tuning. Aligning LLMs with human driving preferences is crucial yet challenging, as driving scenarios involve complex decisions where multiple incorrect act

Cited by 1SourcecodeScholar
2025

PEARL: Parallel Speculative Decoding with Adaptive Draft Length

ICLR 2025poster

Speculative decoding (SD), where an extra draft model is employed to provide multiple **draft** tokens first and then the original target model verifies these tokens in parallel, has shown great power for LLM inference acceleration. However, existing SD methods suffer from the mutual waiting problem…

Cited by 0SourcePDFScholar
2025

Post-Hoc Watermarking for Robust Detection in Text Generated by Large Language Models

COLING 2025main

Research on text simplification has been ongoing for many years, yet document simplification remains a significant challenge due to the need to address complex factors such as technical terminology, metaphors, and overall coherence. In this work, we introduce a novel multi-agent framework AgentSimp…

2025

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

ICCV 2025poster

The use of Multimodal Large Language Models (MLLMs) as an end-to-end solution for Embodied AI and Autonomous Driving has become a prevailing trend. While MLLMs have been extensively studied for visual semantic understanding tasks, their ability to perform precise and quantitative spatial-temporal un…

Cited by 0SourcePDFScholar
2025

Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning

RA-L 2025

End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an crucial part of Embodied AI. Despite successes in applying Multimodal Large Language Models (MLLMs) for high-level traffic scene semantic understanding, it remains challenging to effectively tra

Cited by 23SourceScholar
2025

ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models

EMNLP 2025

Large Language Models (LLMs), constrained by limited context windows, often face significant performance degradation when reasoning over long contexts. To address this, Retrieval-Augmented Generation (RAG) retrieves and reasons over chunks but frequently sacrifices logical coherence due to its relia

Cited by 0SourcePDFScholar
2024

CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models

NeurIPS 2024poster

Artificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized healthcare. However, the trustworthiness of Med-LVLMs remains unverified, posing s…

2024

Calibrated Self-Rewarding Vision Language Models

NeurIPS 2024poster

Large Vision-Language Models (LVLMs) have made substantial progress by integrating pre-trained large language models (LLMs) and vision models through instruction tuning. Despite these advancements, LVLMs often exhibit the hallucination phenomenon, where generated text responses appear linguistically…

2024

Context-based and Diversity-driven Specificity in Compositional Zero-Shot Learning

CVPR 2024poster

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object pairs based on a limited set of observed examples. Current CZSL methodologies despite their advancements tend to neglect the distinct specificity levels present in attributes. For instance given images of sliced strawb…

Cited by 16SourcePDFScholar
2024

Exploring the Impact of Table-to-Text Methods on Augmenting LLM-based Question Answering with Domain Hybrid Data

NAACL 2024industry

Augmenting Large Language Models (LLMs) for Question Answering (QA) with domain specific data has attracted wide attention. However, domain data often exists in a hybrid format, including text and semi-structured tables, posing challenges for the seamless integration of information. Table-to-Text Ge…

Cited by 17SourcePDFScholar
2024

Learning Spectral Canonical ℱ-Correlation Representation for Face Super-Resolution

ICASSP 2024accepted

Face super-resolution (FSR) is a powerful technique for restoring high-resolution face images from the captured low-resolution ones with the assistance of prior information. Existing FSR methods based on explicit or implicit covariance matrices are difficult to reveal complex nonlinear relationships…

Cited by 0SourceScholar
2024

RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models

EMNLP 2024main

The recent emergence of Medical Large Vision Language Models (Med-LVLMs) has enhanced medical diagnosis. However, current Med-LVLMs frequently encounter factual issues, often generating responses that do not align with established medical facts. Retrieval-Augmented Generation (RAG), which utilizes e…

2024

STimage-1K4M: A histopathology image-gene expression dataset for spatial transcriptomics

NeurIPS 2024poster

Recent advances in multi-modal algorithms have driven and been driven by the increasing availability of large image-text datasets, leading to significant strides in various fields, including computational pathology. However, in most existing medical image-text datasets, the text typically provides h…

2023

Advancing Example Exploitation Can Alleviate Critical Challenges in Adversarial Training

ICCV 2023oral

Deep neural networks have achieved remarkable results across various tasks. However, they are susceptible to adversarial examples, which are generated by adding adversarial perturbations to original data. Adversarial training (AT) is the most effective defense mechanism against adversarial examples…

Cited by 10PDFcodeScholar
2023

Learning Supervised Covariation Projection Through General Covariance

ICASSP 2023accepted

Canonical correlation analysis (CCA) is a classical yet powerful tool for learning two-view feature representation in various fields. But, most CCA approaches are based on the conventional covariance measure, which makes them difficult to uncover the complicatedly nonlinear relationship between dist…

Cited by 0SourceScholar
2023

On the Identifiability and Interpretability of Gaussian Process Models

NeurIPS 2023poster

In this paper, we critically examine the prevalent practice of using additive mixtures of Mat\'ern kernels in single-output Gaussian process (GP) models and explore the properties of multiplicative mixtures of Mat\'ern kernels for multi-output GP models. For the single-output case, we derive a serie…

Cited by 7SourcePDFScholar
2023

ParaLS: Lexical Substitution via Pretrained Paraphraser

ACL 2023long

Lexical substitution (LS) aims at finding appropriate substitutes for a target word in a sentence. Recently, LS methods based on pretrained language models have made remarkable progress, generating potential substitutes for a target word through analysis of its contextual surroundings. However, thes…

2023

SCA: Streaming Cross-Attention Alignment For Echo Cancellation

ICASSP 2023accepted

End-to-End deep learning has shown promising results for speech enhancement tasks, such as noise suppression, dereverberation, and speech separation. However, most state-of-the-art methods for echo cancellation are either classical DSP-based or hybrid DSP-ML algorithms. Components such as the delay…

Cited by 0SourceScholar
2023

StereoVAE: A lightweight stereo-matching system using embedded GPUs

ICRA 2023poster

We propose a lightweight system for stereo-matching using embedded graphic processing units (GPUs). The proposed system overcomes the trade-off between accuracy and processing speed in stereo matching, thus further improving the matching accuracy while ensuring real-time processing. The basic idea i…

Cited by 8SourceScholar
2022

Learning Canonical F-Correlation Projection for Compact Multiview Representation

CVPR 2022poster

Canonical correlation analysis (CCA) matters in multiview representation learning. But, CCA and its most variants are essentially based on explicit or implicit covariance matrices. It means that they have no ability to model the nonlinear relationship among features due to intrinsic linearity of cov…

Cited by 11PDFScholar
2021

An Unsupervised Method for Building Sentence Simplification Corpora in Multiple Languages

EMNLP 2021finding

The availability of parallel sentence simplification (SS) is scarce for neural SS modelings. We propose an unsupervised method to build SS corpora from large-scale bilingual translation corpora, alleviating the need for SS supervised corpora. Our method is motivated by the following two findings: ne…

2021

Task Aligned Generative Meta-learning for Zero-shot Learning

AAAI 2021technical

Zero-shot learning (ZSL) refers to the problem of learning to classify instances from novel classes (unseen) that are absent in the training set (seen). Most ZSL methods infer the correlation between visual features and attributes to train the classifier for unseen classes. They may have a strong bi…

Cited by 47SourcePDFScholar
2020

A Visual-Pilot Deep Fusion for Target Speech Separation in Multitalker Noisy Environment

ICASSP 2020accepted

Separating the target speech in multi-talker noisy environment is a challenging problem for audio-only source separation algorithms. The major problem behind is that the separated speech from the same talker can switch among the outputs across consecutive segments, causing the talker permutation iss…

Cited by 5SourceScholar
2020

CooBa: Cross-project Bug Localization via Adversarial Transfer Learning

IJCAI 2020poster

Bug localization plays an important role in software quality control. Many supervised machine learning models have been developed based on historical bug-fix information. Despite being successful, these methods often require sufficient historical data (i.e., labels), which is not always available es…

Cited by 0SourcePDFScholar
2018

Low Resolution Face Recognition and Reconstruction Via Deep Canonical Correlation Analysis

ICASSP 2018accepted

Low-resolution (LR) face identification is always a challenge in computer vision. In this paper, we propose a new LR face recognition and reconstruction method using deep canonical correlation analysis (DCCA). Unlike linear CCA-based methods, our proposed method can learn flexible nonlinear represen…

Cited by 0SourceScholar
2016

Outlying sequence detection in large datasets: Comparison of universal hypothesis testing and clustering

ICASSP 2016accepted

Multiple observation sequences are collected, among which there is a small subset of outliers. A sequence is considered an outlier if the observations therein are generated by a mechanism different from that generating the observations in the majority of sequences. In the universal setting, the goal…

Cited by 0SourceScholar
2016

Stable dysphonia measures selection for Parkinson speech rehabilitation via diversity regularized ensemble

ICASSP 2016accepted

Vocal impairment is a common symptom for the vast majority of Parkinson's disease (PD) subjects. And it needs long term rehabilitation through personalized one-to-one periodic rehabilitation meetings with clinical speech experts. The significant challenge is that there are not enough experts to deli…

Cited by 0SourceScholar
2015

Universal outlier hypothesis testing: Application to anomaly detection

ICASSP 2015accepted

In outlier hypothesis testing, multiple observation sequences are collected, a small subset of which are outliers. Observations in an outlier sequence are generated by a mechanism different from that generating the observations in the majority of sequences. The goal is to best discern all the outlie…

Cited by 0SourceScholar