← Search

Lei YU

74 accepted papers

2026

Beyond Tie Points: Satellite Image Block Adjustment based on Dense Feature Consistency

CVPR 2026

Owing to the weak stereo geometry of satellite images, Planar Block Adjustment (PBA) is a predominant technique for correcting geometric distortions in satellite images, which treats elevation as a known constraint and primarily optimizes planar coordinates. Existing PBA methods mainly rely on expli

Cited by 0SourcecodeScholar
2026

Binary Message Passing for Generalizable Semi-Supervised Graph Anomaly Detection

AAAI 2026technical

Graph Neural Networks (GNNs) have achieved impressive performance in semi-supervised graph anomaly detection (GAD). While many GNN variants have been developed for this task, they largely focus on advanced message aggregation schemes, leaving the message routing aspect underexplored. We argue that t

Cited by 0SourcePDFScholar
2026

Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven Consistency

AAAI 2026technical

Super-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applicatio

Cited by 0SourcePDFScholar
2026

GenePheno: Interpretable Gene Knockout-Induced Phenotype Abnormality Prediction from Gene Sequences

AAAI 2026technical

Exploring how genetic sequences shape phenotypes is a fundamental challenge in biology and a key step toward scalable, hypothesis-driven experimentation. The task is complicated by the large modality gap between sequences and phenotypes, as well as the pleiotropic nature of gene–phenotype relationsh

Cited by 0SourcePDFScholar
2026

Hallucination Reduction with CASAL: Contrastive Activation Steering for Amortized Learning

ICLR 2026poster

Large Language Models (LLMs) exhibit impressive capabilities but often hallucinate, confidently providing incorrect answers instead of admitting ignorance. Prior work has shown that models encode linear representations of their own knowledge and that activation steering can reduce hallucinations. Th…

Cited by 0SourceScholar
2026

Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots

ICLR 2026poster

We **color-coded** the added changes to the **paper** and **Appendix** for the comfort of our reviewers. Multimodal Language Models (MMLMs) typically undergo post-training alignment to prevent harmful content generation. However, these alignment stages focus primarily on the *assistant* role, leavi…

Cited by 0SourceScholar
2026

PMGS: Reconstruction of Projectile Motion Across Large Spatiotemporal Spans via 3D Gaussian Splatting

AAAI 2026technical

Modeling complex rigid motion across large spatiotemporal spans remains an unresolved challenge in dynamic reconstruction. Existing paradigms are mainly confined to short-term, small-scale deformation and offer limited consideration for physical consistency. This study proposes PMGS, focusing on rec

Cited by 0SourcePDFScholar
2026

SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA

ICLR 2026poster

Federated fine-tuning is critical for improving the performance of large language models (LLMs) in handling domain-specific tasks while keeping training data decentralized and private. However, prior work has shown that clients' private data can actually be recovered via gradient inversion attacks.…

Cited by 0SourcecodeScholar
2026

SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imagery

CVPR 2026

While recent foundation models for remote sensing segmentation have shown notable progress, they still fall short in processing diverse multi-modal inputs, synergizing complementary prompt types, and leveraging semantic hierarchies. To address these limitations, we introduce SkySense-VITA, a unified

Cited by 0SourceScholar
2026

SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding

ICML 2026poster

Speculative decoding mitigates the memory-bound nature of LLM decoding by using a lightweight draft model to propose multiple tokens for parallel verification. However, its adoption has been limited by the lack of high-quality draft models and scalable training infrastructure. We introduce SpecForge…

Cited by 0SourceScholar
2026

VL-JEPA: Joint Embedding Predictive Architecture for Vision-language

ICLR 2026poster

We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the target texts. By learning in an abstract representation space, the model can focu…

Cited by 0SourceScholar
2025

An Efficient Residual-based Low-dose PET Reconstruction with Spatial-Frequency Integration

ICASSP 2025accepted

Positron emission tomography (PET) is a nuclear medical imaging technique where image quality depends on the dose of radionuclides administered to the patient. While standard-dose PET (SPET) offers high-quality imaging, it also poses radiation risks. If reconstructing low-dose PET (LPET) images can…

Cited by 0SourceScholar
2025

Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

EMNLP 2025

LLMs often adopt an assertive language style also when making false claims. Such ”overconfident hallucinations” mislead users and erode trust. Achieving the ability to express in language the actual degree of uncertainty around a claim is therefore of great importance. We find that ”verbal uncertain

Cited by 0SourcePDFScholar
2025

CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance

ICCV 2025poster

Semi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we p…

2025

CoGAP: A Personalized Federated Learning Method Using Collaborative Optimization for Medical Image Classification

ICASSP 2025accepted

Federated learning (FL) has been widely used in medical image processing to protect data privacy, but it has issues with data heterogeneity. Personalized federated learning have emerged to tackle these issues but often focuses too much on personalized models at the expense of global models. To addre…

Cited by 0SourceScholar
2025

Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts

ICML 2025poster

Diffusion models have transformed generative modeling but suffer from scalability limitations due to computational overhead and inflexible architectures that process all generative stages and tokens uniformly. In this work, we introduce Diff-MoE, a novel framework that combines Diffusion Transformer…

Cited by 0SourcePDFScholar
2025

Effective Diffusion Transformer Architecture for Image Super-Resolution

AAAI 2025technical

Recent advances indicate that diffusion model holds great promise in image super-resolution. While latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image…

2025

Emergence of a High-Dimensional Abstraction Phase in Language Transformers

ICLR 2025poster

A language model (LM) is a mapping from a linguistic context to an output token. However, much remains to be known about this mapping, including how its geometric properties relate to its function. We take a high-level geometric approach to its analysis, observing, across five pre-trained transforme…

2025

Exploring Scene Affinity for Semi-Supervised LiDAR Semantic Segmentation

CVPR 2025poster

This paper explores scene affinity (AIScene), namely intra-scene consistency and inter-scene correlation, for semi-supervised LiDAR semantic segmentation in driving scenes. Adopting teacher-student training, AIScene employs a teacher network to generate pseudo-labeled scenes from unlabeled data, whi…

2025

Geometric Signatures of Compositionality Across a Language Model’s Lifetime

ACL 2025long

By virtue of linguistic compositionality, few syntactic rules and a finite lexicon can generate an unbounded number of sentences. That is, language, though seemingly high-dimensional, can be explained using relatively few degrees of freedom. An open question is whether contemporary language models (…

Cited by 0SourcePDFScholar
2025

Holistic Large-Scale Scene Reconstruction via Mixed Gaussian Splatting

NeurIPS 2025poster

Recent advances in 3D Gaussian Splatting have shown remarkable potential for novel view synthesis. However, most existing large-scale scene reconstruction methods rely on the divide-and-conquer paradigm, which often leads to the loss of global scene information and requires complex parameter tuning…

Cited by 0SourcecodeScholar
2025

HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation

AAAI 2025technical

Feature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority…

Cited by 0SourcePDFScholar
2025

Intrinsic Test of Unlearning Using Parametric Knowledge Traces

EMNLP 2025

The task of “unlearning” certain concepts in large language models (LLMs) has gained attention for its role in mitigating harmful, private, or incorrect outputs. Current evaluations mostly rely on behavioral tests, without monitoring residual knowledge in model parameters, which can be adversarially

2025

MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals

ICASSP 2025accepted

Video-based physiology, exemplified by remote photoplethysmography (rPPG), extracts physiological signals such as pulse and respiration by analyzing subtle changes in video recordings. This non-contact, real-time monitoring method holds great potential for home settings. Despite the valuable contrib…

Cited by 0SourceScholar
2025

MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading

AAAI 2025technical

Lip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event fram…

2025

QUITO-X: A New Perspective on Context Compression from the Information Bottleneck Theory

EMNLP 2025

Generative large language models ( LLMs) have achieved remarkable success in various industrial applications, owing to their promising In-Context Learning capabilities. However, the issue of long context in complex tasks poses a significant barrier to their wider adoption, manifested in two main asp

Cited by 0SourcePDFScholar
2025

RaSS: Improving Denoising Diffusion Samplers with Reinforced Active Sampling Scheduler

CVPR 2025poster

Recent years have witnessed the great success of denoising diffusion samplers in improving the generative capability and sampling efficiency given a pre-trained diffusion model. However, most sampling schedulers in diffusion models lack the sampling dynamics and planning capability for future genera…

Cited by 0SourcePDFScholar
2025

Restricted Global-Aware Graph Filters Bridging GNNs and Transformer for Node Classification

NeurIPS 2025poster

Transformers have been widely regarded as a promising direction for breaking through the performance bottlenecks of Graph Neural Networks (GNNs), primarily due to their global receptive fields. However, a recent empirical study suggests that tuned classical GNNs can match or even outperform state-of…

Cited by 0SourceScholar
2025

Robust LLM safeguarding via refusal feature adversarial training

ICLR 2025poster

Large language models (LLMs) are vulnerable to adversarial attacks that can elicit harmful responses. Defending against such attacks remains challenging due to the opacity of jailbreaking mechanisms and the high computational cost of training LLMs robustly. We demonstrate that adversarial attacks sh…

Cited by 9SourcePDFScholar
2025

Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity

EMNLP 2025

In this paper, we introduce DiscoGP, a novel framework for extracting self-contained modular units, or sheaves, within neural language models (LMs). Sheaves extend the concept of functional circuits, a unit widely explored in interpretability research, by considering not only subsets of edges in an

2025

SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing

ICCV 2025poster

The multi-modal remote sensing foundation model (MM-RSFM) has significantly advanced various Earth observation tasks, such as urban planning, environmental monitoring, and natural disaster management. However, most existing approaches generally require the training of separate backbone networks for…

Cited by 0SourcePDFScholar
2025

Summary Factual Inconsistency Detection Based on LLMs Enhanced by Universal Information Extraction

ACL 2025finding

Automatic text summarization has a potential flaw that affects the factuality of summaries. Recently, Large Language Models (LLMs) have been introduced as detectors for factual inconsistencies in summaries. However, LLM-based methods rely on reasoning capabilities and face challenges in terms of eff…

Cited by 0SourcePDFScholar
2025

Text2Sql: Pure Fine-Tuning and Pure Knowledge Distillation

NAACL 2025industry

Text2Sql is a task that converts natural language questions into SQL queries. In previous research on LLM fine-tuning, researchers typically input both the entire database schema and the natural language question into the model. This approach has two issues: 1) the model’s context is limited when de…

Cited by 0SourcePDFScholar
2025

The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction

ACL 2025finding

Large language models (LLMs) excel on a variety of reasoning benchmarks, but previous studies suggest they sometimes struggle to generalize to unseen questions, potentially due to over-reliance on memorized training examples. However, the precise conditions under which LLMs switch between reasoning…

2024

EcoMatcher: Efficient Clustering Oriented Matcher for Detector-free Image Matching

ECCV 2024poster

"Detector-free local feature matching methods have demonstrated significant performance improvements since leveraging the power of Transformer architecture. The global receptive field allows for simultaneous interaction among all elements, proving particularly beneficial in regions with low texture…

Cited by 1SourcePDFScholar
2024

Enhancing Reinforcement Learning with Dense Rewards from Language Model Critic

EMNLP 2024main

Reinforcement learning (RL) can align language models with non-differentiable reward signals, such as human preferences. However, a major challenge arises from the sparsity of these reward signals - typically, there is only a single reward for an entire output. This sparsity of rewards can lead to i…

Cited by 9SourcePDFScholar
2024

FE-DeTr: Keypoint Detection and Tracking in Low-quality Image Frames with Events

ICRA 2024poster

Keypoint detection and tracking in traditional image frames are often compromised by image quality issues such as motion blur and extreme lighting conditions. Event cameras offer potential solutions to these challenges by virtue of their high temporal resolution and high dynamic range. However, they…

Cited by 4SourcecodeScholar
2024

Learning Quantized Adaptive Conditions for Diffusion Models

ECCV 2024poster

"The curvature of ODE trajectories in diffusion models hinders their ability to generate high-quality images in a few number of function evaluations (NFE). In this paper, we propose a novel and effective approach to reduce trajectory curvature by utilizing adaptive conditions. By employing a extreme…

Cited by 0SourcePDFScholar
2024

Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations

EMNLP 2024finding

State-of-the-art language models (LMs) sometimes generate that misalign with world knowledge. To explore the mechanistic causes of these hallucinations, we create diagnostic datasets with subject-relation queries and adapt interpretability methods to trace hallucinations through internal model repre…

2024

POA: Pre-training Once for Models of All Sizes

ECCV 2024poster

"Large-scale self-supervised pre-training has paved the way for one foundation model to handle many different vision tasks. Most pre-training methodologies train a single model of a certain size at one time. Nevertheless, various computation or storage constraints in real-world scenarios require sub…

2024

PQ-SAM: Post-training Quantization for Segment Anything Model

ECCV 2024poster

"Segment anything model (SAM) is a promising prompt-guided vision foundation model to segment objects of interest. However, the extensive computational requirements of SAM have limited its applicability in resource-constraint edge devices. Post-training quantization (PTQ) is an effective potential f…

Cited by 5SourcePDFScholar
2024

SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery

CVPR 2024poster

Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless these works primarily focus on a single modality without temporal and geo-context modeling hampering their capabilities for diverse tasks. In this study we pre…

Cited by 140SourcePDFScholar
2024

Toward Robust Keypoint Detection and Tracking: A Fusion Approach With Event-Aligned Image Features

RA-L 2024

Robust keypoint detection and tracking are crucial for various robotic tasks. However, conventional cameras struggle under rapid motion and lighting changes, hindering local and edge feature extraction essential for keypoint detection and tracking. Event cameras offer advantages in such scenarios du

Cited by 12SourceScholar
2023

A Natural Bias for Language Generation Models

ACL 2023short

After just a few hundred training updates, a standard probabilistic model for language generation has likely not yet learnt many semantic or syntactic rules of natural language, making it difficult to estimate the probability distribution over next tokens. Yet around this point, these models have id…

2023

Dynamic Coarse-To-Fine Learning for Oriented Tiny Object Detection

CVPR 2023poster

Detecting arbitrarily oriented tiny objects poses intense challenges to existing detectors, especially for label assignment. Despite the exploration of adaptive label assignment in recent oriented object detectors, the extreme geometry shape and limited feature of oriented tiny objects still induce…

2023

Generalizing Event-Based Motion Deblurring in Real-World Scenarios

ICCV 2023poster

Event-based motion deblurring has shown promising results by exploiting low-latency events. However, current approaches are limited in their practical usage, as they assume the same spatial resolution of inputs and specific blurriness distributions. This work addresses these limitations and aims to…

Cited by 29PDFcodeScholar
2023

High-Level Semantic Feature Matters Few-Shot Unsupervised Domain Adaptation

AAAI 2023technical

In few-shot unsupervised domain adaptation (FS-UDA), most existing methods followed the few-shot learning (FSL) methods to leverage the low-level local features (learned from conventional convolutional models, e.g., ResNet) for classification. However, the goal of FS-UDA and FSL are relevant yet dis…

Cited by 2SourcePDFScholar
2023

Learning from Training Dynamics: Identifying Mislabeled Data beyond Manually Designed Features

AAAI 2023technical

While mislabeled or ambiguously-labeled samples in the training set could negatively affect the performance of deep models, diagnosing the dataset and identifying mislabeled samples helps to improve the generalization power. Training dynamics, i.e., the traces left by iterations of optimization algo…

2023

Systematic word meta-sense extension

EMNLP 2023long main

The meaning of polysemous words often varies in a highly productive yet predictable way. Generalizing the regularity between conventional senses to derive novel word meaning is crucial for automated processing of non-literal language uses such as figurative expressions. We introduce a novel task cal…

Cited by 0SourcecodeScholar
2023

Word sense extension

ACL 2023long

Humans often make creative use of words to expressnovel senses. A long-standing effort in natural language processing hasbeen focusing on word sense disambiguation (WSD), but little has been explored about how the sense inventory of a word may be extended toward novel meanings. We present a paradigm…

2022

Deep Constrained Least Squares for Blind Image Super-Resolution

CVPR 2022poster

In this paper, we tackle the problem of blind image super-resolution(SR) with a reformulated degradation model and two novel modules. Following the common practices of blind SR, our method proposes to improve both the kernel estimation as well as the kernel-based high-resolution image restoration. T…

Cited by 131PDFcodeScholar
2022

Enabling Arbitrary Translation Objectives with Adaptive Tree Search

ICLR 2022poster

We introduce an adaptive tree search algorithm, which is a deterministic variant of Monte Carlo tree search, that can find high-scoring outputs under translation models that make no assumptions about the form or structure of the search objective. This algorithm enables the exploration of new kinds o…

Cited by 1SourcePDFScholar
2022

Prompt-Based Meta-Learning For Few-shot Text Classification

EMNLP 2022main

Few-shot Text Classification predicts the semantic label of a given text with a handful of supporting instances. Current meta-learning methods have achieved satisfying results in various few-shot situations. Still, they often require a large amount of data to construct many few-shot tasks for meta-t…

2022

RFLA: Gaussian Receptive Field Based Label Assignment for Tiny Object Detection

ECCV 2022poster

"Detecting tiny objects is one of the main obstacles hindering the development of object detection. The performance of generic object detectors tends to drastically deteriorate on tiny object detection tasks. In this paper, we point out that either box prior in the anchor-based detector or point pri…

2021

Event-Based Synthetic Aperture Imaging With a Hybrid Network

CVPR 2021poster

Synthetic aperture imaging (SAI) is able to achieve the see through effect by blurring out the off-focus foreground occlusions and reconstructing the in-focus occluded targets from multi-view images. However, very dense occlusions and extreme lighting conditions may bring significant disturbances to…

Cited by 39PDFcodeScholar
2021

Predicting emergent linguistic compositions through time: Syntactic frame extension via multimodal chaining

EMNLP 2021main

Natural language relies on a finite lexicon to express an unbounded set of emerging ideas. One result of this tension is the formation of new compositions, such that existing linguistic units can be combined with emerging items into novel expressions. We develop a framework that exploits the cogniti…

2020

A Mutual Information Maximization Perspective of Language Representation Learning

ICLR 2020spotlight

We show state-of-the-art word representation learning methods maximize an objective function that is a lower bound on the mutual information between different parts of a word sequence (i.e., a sentence). Our formulation provides an alternative perspective that unifies classical word embedding models…

Cited by 77SourceScholar
2020

Efficient Spatial-Temporal Normalization of SAE Representation for Event Camera

RA-L 2020

Event-based cameras are a new type of vision sensor that can encode spatial-temporal context in a pixel-level event stream. Its appealing properties offer great potential for applications requiring low processing latency and low power consumption. As an effective representation of events, the surfac

Cited by 13SourceScholar
2019

Variational Smoothing in Recurrent Neural Network Language Models

ICLR 2019poster

We present a new theoretical perspective of data noising in recurrent neural network language models (Xie et al., 2017). We show that each variant of data noising is an instance of Bayesian recurrent neural networks with a particular variational distribution (i.e., a mixture of Gaussians whose weig…

Cited by 3SourcePDFScholar