← Search

Yue Lu

34 accepted papers

2026

Attention Retention for Continual Learning with Vision Transformers

AAAI 2026technical

Continual learning (CL) empowers AI systems to progressively acquire knowledge from non-stationary data streams. However, catastrophic forgetting remains a critical challenge. In this work, we identify attention drift in Vision Transformers as a primary source of catastrophic forgetting, where the a

Cited by 0SourcePDFScholar
2026

DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models

CVPR 2026

Humans can robustly localize visual evidence and provide grounded answers even in noisy environments by identifying critical cues and then relating them to the full context in a bottom-up manner. Inspired by this, we propose DeepScan, a training-free framework that combines Hierarchical Scanning, Re

Cited by 0SourcecodeScholar
2026

Perceptual Flow Network for Visually Grounded Reasoning

ICML 2026poster

Despite the success of LVLMs, general optimization objectives (e.g., standard MLE) fail to constrain visual trajectories, leading to language bias and hallucination. To mitigate this, current methods introduce geometric priors from visual experts as additional supervision. However, we observe that s…

Cited by 0SourceScholar
2026

Web-CogReasoner: Towards Knowledge-Induced Cognitive Reasoning for Web Agents

ICLR 2026poster

Multimodal large-scale models have significantly advanced the development of web agents, enabling them to perceive and interact with the digital environment in a manner analogous to human cognition. In this paper, we argue that web agents must first acquire sufficient knowledge to engage in cognitiv…

Cited by 0SourcecodeScholar
2025

A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities

AISTATS 2025oral

A key property of neural networks is their capacity of adapting to data during training. Yet, our current mathematical understanding of feature learning and its relationship to generalization remain limited. In this work, we provide a random matrix analysis of how fully-connected two-layer neural ne…

Cited by 0SourceScholar
2025

BiMAC: Bidirectional Multimodal Alignment in Contrastive Learning

AAAI 2025technical

Achieving robust performance in vision-language tasks requires strong multimodal alignment, where textual and visual data interact seamlessly. Existing frameworks often combine contrastive learning with image captioning to unify visual and textual representations. However, reliance on global represe…

Cited by 0SourcePDFScholar
2025

MSA2: Multi-task Framework with Structure-aware and Style-adaptive Character Representation for Open-set Chinese Text Recognition

ICCV 2025poster

Most existing methods regard open-set Chinese text recognition (CTR) as a single-task problem, primarily focusing on prototype learning of linguistic components or glyphs to identify unseen characters. In contrast, humans identify characters by integrating multiple perspectives, including linguistic…

2025

Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual Learning

AAAI 2025technical

Visual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generato…

Cited by 0SourcePDFScholar
2025

Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering

NeurIPS 2025poster

Existing visual token pruning methods target prompt alignment and visual preservation with static strategies, overlooking the varying relative importance of these objectives across tasks, which leads to inconsistent performance. To address this, we derive the first closed-form error bound for visual…

Cited by 0SourceScholar
2024

Asymptotics of feature learning in two-layer networks after one gradient-step

ICML 2024spotlight

In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single gradient descent step. Leveraging the insight from (Ba et al., 2022), we model the trained network by a spiked Random Featur…

2024

Autoregressive 3D Shape Completion via Sphere-Guided Disentangled Representation

ICASSP 2024accepted

This paper introduces a novel 3D shape completion method based on sphere-guided disentangled representation. Utilizing an autoregressive transformer-based model, our approach efficiently constructs object completion distributions given incomplete point clouds. To enhance completion modeling, we prop…

Cited by 0SourceScholar
2024

Brush Your Text: Synthesize Any Scene Text on Images via Diffusion Model

AAAI 2024technical

Recently, diffusion-based image generation methods are credited for their remarkable text-to-image generation capabilities, while still facing challenges in accurately generating multilingual scene text images. To tackle this problem, we propose Diff-Text, which is a training-free scene text generat…

2024

Enhanced Deep Reinforcement Learning for Parcel Singulation in Non-Stationary Environments

ICASSP 2024accepted

In the rapidly expanding logistics sector, parcel singulation has emerged as a significant bottleneck. To address this, we propose an automated parcel singulator utilizing a sparse actuator array, which presents an optimal balance between cost and efficiency, albeit requiring a sophisticated control…

Cited by 0SourceScholar
2024

Enhancing Reinforcement Learning via Causally Correct Input Identification and Targeted Intervention

ICASSP 2024accepted

Causal confusion, characterized by the learning of spurious correlations, detrimentally affects the generalization and effectiveness of reinforcement learning (RL) algorithms, especially in environments without latent confounders often encountered in robot autonomous navigation tasks. This study add…

Cited by 0SourceScholar
2024

Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network

AAAI 2024technical

Scene text recognition is inherently a vision-language task. However, previous works have predominantly focused either on extracting more robust visual features or designing better language modeling. How to effectively and jointly model vision and language to mitigate heavy reliance on a single moda…

2024

Spotting the Unseen: Reciprocal Consensus Network Guided by Visual Archetypes

AAAI 2024technical

Humans often require only a few visual archetypes to spot novel objects. Based on this observation, we present a strategy rooted in ``spotting the unseen" by establishing dense correspondences between potential query image regions and a visual archetype, and we propose the Consensus Network (CoNet).…

2024

VME-Transformer: Enhancing Visual Memory Encoding for Navigation in Interactive Environments

RA-L 2024

The efficiency of a robotic system is primarily determined by its ability to navigate complex and interactive environments. In real-world scenarios, cluttered surroundings are common, requiring a robot to navigate diverse spaces and displace objects to pave a path towards its objective. Consequently

Cited by 14SourceScholar
2024

Visual Prompt Tuning in Null Space for Continual Learning

NeurIPS 2024poster

Existing prompt-tuning methods have demonstrated impressive performances in continual learning (CL), by selecting and updating relevant prompts in the vision-transformer models. On the contrary, this paper aims to learn each task by tuning the prompts in the direction orthogonal to the subspace span…

2023

A New Comprehensive Benchmark for Semi-Supervised Video Anomaly Detection and Anticipation

CVPR 2023poster

Semi-supervised video anomaly detection (VAD) is a critical task in the intelligent surveillance system. However, an essential type of anomaly in VAD named scene-dependent anomaly has not received the attention of researchers. Moreover, there is no research investigating anomaly anticipation, a more…

2023

Transformer Memory for Interactive Visual Navigation in Cluttered Environments

RA-L 2023

Substantial progress has been achieved in embodied visual navigation based on reinforcement learning (RL). These studies presume that the environment is stationary where all the obstacles are static. However, in real cluttered scenes, interactable objects (e.g. shoes and boxes) blocking the way of r

Cited by 20SourceScholar
2022

Precise Learning Curves and Higher-Order Scalings for Dot-product Kernel Regression

NeurIPS 2022accept

As modern machine learning models continue to advance the computational frontier, it has become increasingly important to develop precise estimates for expected performance improvements under different model and data scaling regimes. Currently, theoretical understanding of the learning curves that c…

Cited by 40SourcePDFScholar
2022

SGBANet: Semantic GAN and Balanced Attention Network for Arbitrarily Oriented Scene Text Recognition

ECCV 2022poster

"Scene text recognition is a challenging task due to the complex backgrounds and diverse variations of text instances. In this paper, we propose a novel Semantic GAN and Balanced Attention Network (SGBANet) to recognize the texts in scene images. The proposed method first generates the simple semant…

Cited by 32SourcePDFScholar
2022

UC-OWOD: Unknown-Classified Open World Object Detection

ECCV 2022poster

"Open World Object Detection (OWOD) is a challenging computer vision problem that requires detecting unknown objects and gradually learning the identified unknown classes. However, it cannot distinguish unknown instances as multiple unknown classes. In this work, we propose a novel OWOD problem call…

2021

DG-Font: Deformable Generative Networks for Unsupervised Font Generation

CVPR 2021poster

Font generation is a challenging problem especially for some writing systems that consist of a large number of characters and has attracted a lot of attention in recent years. However, existing methods for font generation are often in supervised learning. They require a large number of paired data,…

Cited by 150PDFcodeScholar
2020

Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimization

NeurIPS 2020poster

We consider a commonly studied supervised classification of a synthetic dataset whose labels are generated by feeding a one-layer non-linear neural network with random iid inputs. We study the generalization performances of standard classifiers in the high-dimensional regime where $\alpha=\frac{n}{d…

Cited by 71SourcePDFScholar
2020

The Role of Regularization in Classification of High-dimensional Noisy Gaussian Mixture

ICML 2020poster

We consider a high-dimensional mixture of two Gaussians in the noisy regime where even an oracle knowing the centers of the clusters misclassifies a small but finite fraction of the points. We provide a rigorous analysis of the generalization error of regularized convex classifiers, including ridge,…

Cited by 111SourcePDFScholar
2019

Generalized Approximate Survey Propagation for High-Dimensional Estimation

ICML 2019oral

In Generalized Linear Estimation (GLE) problems, we seek to estimate a signal that is observed through a linear transform followed by a component-wise, possibly nonlinear and noisy, channel. In the Bayesian optimal setting, Generalized Approximate Message Passing (GAMP) is known to achieve optimal p…

Cited by 13SourcePDFScholar