← Search

Yang Lu

58 accepted papers

2026

Boosting Multi-Domain Reasoning of LLMs via Curvature-Guided Policy Optimization

ICLR 2026poster

Multi-domain reinforcement learning (RL) for large language models (LLMs) involves highly intricate reward surfaces, posing significant challenges in finding parameters that excel across all domains. Recent empirical studies have further highlighted conflicts among domains, where gains in one capabi…

Cited by 0SourcecodeScholar
2026

Break the Tie: Learning Cluster-Customized Category Relationships for Categorical Data Clustering

AAAI 2026technical

Categorical attributes with qualitative values are ubiquitous in cluster analysis of real datasets. Unlike the Euclidean distance of numerical attributes, the categorical attributes lack well-defined relationships of their possible values (also called categories interchangeably), which hampers the e

Cited by 0SourcePDFScholar
2026

CARE: Class-Adaptive Expert Consensus for Reliable Learning with Long-Tailed Noisy Labels

ICML 2026poster

Learning from real-world data is frequently hindered by the compound challenge of long-tailed class distributions and noisy annotations. Existing methods partially address these issues but typically ignore the non-uniform impact of label noise across classes, resulting in ineffective correction for …

Cited by 0SourceScholar
2026

CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learning

CVPR 2026

Long-tailed distributions are common in real-world recognition tasks, where a few head classes have many samples while most tail classes have very few. Recently, fine-tuning foundation models for long-tailed learning has gained attention due to their excellent performance. However, most existing met

Cited by 0SourcecodeScholar
2026

Decision Boundary-aware Generation for Long-tailed Learning

CVPR 2026

Long-tailed data bias decision boundaries toward head classes and degrade tail class accuracy. Diffusion-based generative augmentation address this problem by generating additional data, while head-to-tail transfer further mitigate the generator bias inherit from long-tailed dataset. However, we sho

Cited by 0SourcecodeScholar
2026

Design of Bio-manta Based on Continuously Programmable Soft Pneumatic Actuators (CPSPAs)

RA-L 2026

Bionic autonomous underwater vehicles (AUVs) have been prevalent in underwater equipment for their excel features. Manta rays have been widely studied due to their highly efficient and maneuverable pectoral fin propulsion. Traditional single-hinge actuation methods have impeded bionic pectoral fin a

Cited by 0SourceScholar
2026

Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation

CVPR 2026

Open-vocabulary semantic segmentation (OVSS) aims to segment arbitrary category regions in images using open-vocabulary prompts, necessitating that existing methods possess pixel-level vision-language alignment capability. Typically, this capability involves computing the cosine similarity, ie, logi

Cited by 0SourcecodeScholar
2026

Fine-Tuning Impairs the Balancedness of Foundation Models in Long-tailed Personalized Federated Learning

CVPR 2026

Personalized federated learning (PFL) with foundation models has emerged as a promising paradigm enabling clients to adapt to heterogeneous data distributions. However, real-world scenarios often face the co-occurrence of non-IID data and long-tailed class distributions, presenting unique challenges

Cited by 0SourcecodeScholar
2026

GeoMind: Explicit Spatial Reasoning via Dual-Reference Geometric Modeling

IJCAI 2026

While Vision-Language Models (VLMs) excel at semantic understanding, they struggle to comprehend 3D spatial relationships from limited views. Their reliance on implicit geometric encoding often leads to severe hallucinations and inconsistencies in spatial reasoning tasks. To address this, we introdu

Cited by 0Scholar
2026

Joint Implicit and Explicit Language Learning for Pedestrian Attribute Recognition

AAAI 2026technical

Pedestrian attribute recognition (PAR) has received increasing attention due to its wide application in video surveillance and pedestrian analysis. Some text-enhanced methods tackle this task by converting attributes into language descriptions to facilitate interactive learning between attributes an

Cited by 0SourcePDFScholar
2026

SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning

CVPR 2026

In open-world semi-supervised learning (OWSSL), a model learns from labeled data and unlabeled data containing both known and novel classes. In practical OWSSL applications, models are expected to perform rigorous classification by directly selecting the most semantically relevant label from a candi

Cited by 0SourcecodeScholar
2026

Simultaneous Arrival Control for Distributed Multi-Robot Systems with Curvature and Constant-Speed Constraints

ICRA 2026poster

The simultaneous arrival of multiple mobile robots at their respective target points is crucial for cooperative tasks such as encirclement, interception, and disaster relief. Although the problem of simultaneous arrival is inherently complex, it becomes even more challenging in multi-robot systems w…

Cited by 0Scholar
2026

Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective

AAAI 2026technical

Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction capability, specifically this pixel-level multimodal alignment. Although existing m

Cited by 0SourcePDFScholar
2025

Asynchronous Federated Clustering with Unknown Number of Clusters

AAAI 2025technical

Federated Clustering (FC) is crucial to mining knowledge from unlabeled non-Independent Identically Distributed (non-IID) data provided by multiple clients while preserving their privacy. Most existing attempts learn cluster distributions at local clients, then securely pass the desensitized informa…

2025

MACA: Multi-Anchor Classification Approach for Unsupervised Domain Adaptation

ICASSP 2025accepted

Unsupervised Domain Adaptation for image classification aims to adapt models trained on a labeled source domain to an unlabeled target domain, improving target domain classification performance. However, previous UDA classification researches tend to assume the two domain distributions after domain…

Cited by 0SourceScholar
2025

MaskViM: Domain Generalized Semantic Segmentation with State Space Models

AAAI 2025technical

Domain Generalized Semantic Segmentation (DGSS) aims to utilize segmentation model training on known source domains to make predictions on unknown target domains. Currently, there are two network architectures: one based on Convolutional Neural Networks (CNNs) and the other based on Visual Transform…

Cited by 0SourcePDFScholar
2025

Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch

CVPR 2025poster

Federated Semi-Supervised Learning (FSSL) aims to leverage unlabeled data across clients with limited labeled data to train a global model with strong generalization ability. Most FSSL methods rely on consistency regularization with pseudo-labels, converting predictions from local or global models i…

2025

NeurIPT: Foundation Model for Neural Interfaces

NeurIPS 2025poster

Electroencephalography (EEG) has wide-ranging applications, from clinical diagnosis to brain-computer interfaces (BCIs). With the increasing volume and variety of EEG data, there has been growing interest in establishing foundation models (FMs) to scale up and generalize neural decoding. Despite sho…

Cited by 0SourceScholar
2025

PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt Relocation

ICCV 2025accepted

Visual prompt tuning (VPT), i.e., fine-tuning some lightweight prompt tokens, provides an efficient and effective approach for adapting pre-trained models to various downstream tasks. However, most prior art indiscriminately uses a fixed prompt distribution across different tasks, neglecting the imp…

2025

Progressive Data Dropout: An Embarrassingly Simple Approach to Train Faster

NeurIPS 2025poster

The success of the machine learning field has reliably depended on training on large datasets. While effective, this trend comes at an extraordinary cost. This is due to two deeply intertwined factors: the size of models and the size of datasets. While promising research efforts focus on reducing th…

Cited by 0SourcecodeScholar
2025

Unlocker: Disentangle the Deadlock of Learning between Label-noisy and Long-tailed Data

NeurIPS 2025poster

In real world, the observed label distribution of a dataset often mismatches its true distribution due to noisy labels. In this situation, noisy labels learning (NLL) methods directly integrated with long-tail learning (LTL) methods tend to fail due to a dilemma: NLL methods normally rely o…

Cited by 0SourceScholar
2025

Versatile Distributed Maneuvering With Generalized Formations Using Guiding Vector Fields

ICRA 2025

This paper presents a unified approach to realize versatile distributed maneuvering with generalized formations. Specifically, we decompose the robots' maneuvers into two independent components, i.e., interception and enclosing, which are parameterized by two independent virtual coordinates. Treatin

Cited by 0SourceScholar
2025

Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language Models

ACL 2025long

Large Vision-Language Models (LVLMs) have shown impressive progress by integrating visual perception with linguistic understanding to produce contextually grounded outputs. Despite these advancements achieved, LVLMs still suffer from the hallucination problem, e.g., they tend to produce content that…

Cited by 0SourcePDFScholar
2025

Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect Detection

CVPR 2025poster

As an important part of intelligent manufacturing, pixel-level surface defect detection (SDD) aims to locate defect areas through mask prediction. Previous methods adopt the image-independent static convolution to indiscriminately classify per-pixel features for mask prediction, which leads to subop…

2025

Weighted Density for The Win: Accurate Subspace Density Clustering

ICASSP 2025accepted

k-clustering typically struggles with the detection of irregular-distributed clusters due to the natural bias, while density clustering usually cannot well-adapt to different datasets and clustering tasks as it is not an oriented optimization process. This paper, therefore, proposes to perform densi…

Cited by 0SourceScholar
2025

You Are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-tailed Data

ICCV 2025poster

Data heterogeneity, stemming from local non-IID data and global long-tailed distributions, is a major challenge in federated learning (FL), leading to significant performance gaps compared to centralized learning. Previous research found that poor representations and biased classifiers are the main…

2025

YuLan-Mini: Pushing the Limits of Open Data-efficient Language Model

ACL 2025long

Due to the immense resource demands and the involved complex techniques, it is still challenging for successfully pre-training a large language models (LLMs) with state-of-the-art performance. In this paper, we explore the key bottlenecks and designs during pre-training, and make the following contr…

Cited by 0SourcePDFScholar
2024

CLIP-Guided Federated Learning on Heterogeneity and Long-Tailed Data

AAAI 2024technical

Federated learning (FL) provides a decentralized machine learning paradigm where a server collaborates with a group of clients to learn a global model without accessing the clients' data. User heterogeneity is a significant challenge for FL, which together with the class-distribution imbalance furth…

2024

Dynamically Anchored Prompting for Task-Imbalanced Continual Learning

IJCAI 2024poster

Existing continual learning literature relies heavily on a strong assumption that tasks arrive with a balanced data stream, which is often unrealistic in real-world applications. In this work, we explore task-imbalanced continual learning (TICL) scenarios where the distribution of task data is non-u…

2024

Feature Fusion from Head to Tail for Long-Tailed Visual Recognition

AAAI 2024technical

The imbalanced distribution of long-tailed data presents a considerable challenge for deep learning models, as it causes them to prioritize the accurate classification of head classes but largely disregard tail classes. The biased decision boundary caused by inadequate semantic information in tail c…

2024

Federated Learning with Extremely Noisy Clients via Negative Distillation

AAAI 2024technical

Federated learning (FL) has shown remarkable success in cooperatively training deep models, while typically struggling with noisy labels. Advanced works propose to tackle label noise by a re-weighting strategy with a strong assumption, i.e., mild label noise. However, it may be violated in many real…

2024

Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual Recognition

NeurIPS 2024poster

Long-tailed visual recognition has received increasing attention recently. Despite fine-tuning techniques represented by visual prompt tuning (VPT) achieving substantial performance improvement by leveraging pre-trained knowledge, models still exhibit unsatisfactory generalization performance on tai…

2024

Learning-Based Near-Optimal Motion Planning for Intelligent Vehicles With Uncertain Dynamics

RA-L 2024

Motion planning has been an important research topic in achieving safe and flexible maneuvers for intelligent vehicles. However, it remains challenging to realize efficient and optimal planning in the presence of uncertain model dynamics. In this paper, a sparse kernel-based reinforcement learning (

Cited by 6SourceScholar
2024

Proposal Distillation of Multi-Modal Feature Aggregation Network for Video Object Detection

ICASSP 2024accepted

Video object detection is a challenging task due to deteriorated object appearances. In order to bolster per-frame feature representations, one way is to aggregate features from relevant frames. However, relying exclusively on RGB modal for feature aggregation may limit the detection performance for…

Cited by 0SourceScholar
2024

Relationship Prompt Learning is Enough for Open-Vocabulary Semantic Segmentation

NeurIPS 2024poster

Open-vocabulary semantic segmentation (OVSS) aims to segment unseen classes without corresponding labels. Existing Vision-Language Model (VLM)-based methods leverage VLM's rich knowledge to enhance additional explicit segmentation-specific networks, yielding competitive results, but at the cost of e…

Cited by 0SourcePDFScholar
2024

Self-Training Domain Adaptation Via Weight Transmission Between Generators

ICASSP 2024accepted

Unsupervised domain adaptation (UDA) aims to transfer knowledge from the labeled source domain to the fully-unlabeled target domain, thus improving the classification performance of the target domain. Recently, self-training has shown its effectiveness on UDA. However, the feature space for generati…

Cited by 0SourceScholar
2024

Spatio-Temporal Correlation Learning for Multiple Object Tracking

ICASSP 2024accepted

Multi-object tracking (MOT) has gained remarkable progress in recent years, while due to the complexity of real-world environments, there are still many challenges that remain unsolved, such as object occlusion and deformation. To effectively alleviate this problem, we propose a simple yet effective…

Cited by 0SourceScholar
2024

Visual-Linguistic Representation Learning with Deep Cross-Modality Fusion for Referring Multi-Object Tracking

ICASSP 2024accepted

Referring multi-object tracking is a new rising research topic that aims at detecting and tracking the referred objects in a video sequence based on a natural language expression. Compared with traditional multi-object tracking, this setting guides object tracking with high-level semantic informatio…

Cited by 0SourceScholar
2023

Contrastive Domain Adaptation Via Delimitation Discriminator

ICASSP 2023accepted

Unsupervised domain adaptation aims to transfer the knowledge learned from the labeled source domain to the unlabeled target domain, thereby improving the classification performance of the target domain. Recent methods use contrastive learning to optimize this task, however, these methods only focus…

Cited by 5SourceScholar
2023

DDK: A Deep Koopman Approach for Longitudinal and Lateral Control of Autonomous Ground Vehicles

ICRA 2023poster

Autonomous driving has attracted lots of attention in recent years. For some tasks, e.g., trajectory prediction, motion planning, and trajectory tracking, an accurate vehicle model can reduce the difficulty of these tasks and improve task completion performance. Prior works focused on parameter esti…

Cited by 8SourceScholar
2023

Fast and Accurate Binary Neural Networks Based on Depth-Width Reshaping

AAAI 2023technical

Network binarization (i.e., binary neural networks, BNNs) can efficiently compress deep neural networks and accelerate model inference but cause severe accuracy degradation. Existing BNNs are mainly implemented based on the commonly used full-precision network backbones, and then the accuracy is imp…

2023

Label-Noise Learning with Intrinsically Long-Tailed Data

ICCV 2023poster

Label noise is one of the key factors that lead to the poor generalization of deep learning models. Existing label-noise learning methods usually assume that the ground-truth classes of the training data are balanced. However, the real-world data is often imbalanced, leading to the inconsistency bet…

Cited by 25PDFcodeScholar
2023

Learning to Reconnect Interrupted Trajectories for Weakly Supervised Multi-Object Tracking

ICASSP 2023accepted

Recently, some weakly supervised multi-object tracking (MOT) methods learn identity embedding features with pseudo identity labels rather than the high-cost manual ones. However, these pseudo identity labels may contain many false or missing identities, which adversely affect the optimization of tra…

Cited by 0SourceScholar
2023

Long-Tailed Visual Recognition via Self-Heterogeneous Integration With Knowledge Excavation

CVPR 2023poster

Deep neural networks have made huge progress in the last few decades. However, as the real-world data often exhibits a long-tailed distribution, vanilla deep models tend to be heavily biased toward the majority classes. To address this problem, state-of-the-art methods usually adopt a mixture of exp…

2023

Personalized Federated Learning on Long-Tailed Data via Adversarial Feature Augmentation

ICASSP 2023accepted

Personalized Federated Learning (PFL) aims to learn personalized models for each client based on the knowledge across all clients in a privacy-preserving manner. Existing PFL methods generally assume that the underlying global data across all clients are uniformly distributed without considering the…

Cited by 0SourceScholar
2022

Bounding Box Distribution Learning and Center Point Calibration for Robust Visual Tracking

ICASSP 2022accepted

Visual tracking aims at both robust target classification and accurate localization. However, the reliability of the target bounding box and classification score are not properly addressed by most existing trackers, resulting in inaccurate tracking performance. In this paper, we propose to learn bou…

Cited by 0SourceScholar
2022

Federated Learning on Heterogeneous and Long-Tailed Data via Classifier Re-Training with Federated Features

IJCAI 2022poster

Federated learning (FL) provides a privacy-preserving solution for distributed machine learning tasks. One challenging problem that severely damages the performance of FL models is the co-occurrence of data heterogeneity and long-tail distribution, which frequently appears in real FL applications. I…

2022

Multi-Focus Guided Semantic Aggregation for Video Object Detection

ICASSP 2022accepted

For the task of video object detection, it is useful to aggregate semantic information from supporting frames. However, existing methods only focus on the current frame during the semantic aggregation, called Single-Focus methods. They neglect semantic information among supporting frames and deterio…

Cited by 0SourceScholar
2021

Speech-Based Depression Prediction Using Encoder-Weight-Only Transfer Learning and a Large Corpus

ICASSP 2021accepted

Speech-based algorithms have gained interest for the management of behavioral health conditions such as depression. We explore a speech-based transfer learning approach that uses a lightweight encoder and that transfers only the encoder weights, enabling a simplified run-time model. Our study uses a…

Cited by 0SourceScholar
2018

DeepPINK: reproducible feature selection in deep neural networks

NeurIPS 2018poster

Deep learning has become increasingly popular in both supervised and unsupervised machine learning thanks to its outstanding empirical performance. However, because of their intrinsic complexity, most deep learning methods are largely treated as black box tools with little interpretability. Even tho…

2018

Learning Generative ConvNets via Multi-Grid Modeling and Sampling

CVPR 2018poster

This paper proposes a multi-grid method for learning energy-based generative ConvNet models of images. For each grid, we learn an energy-based probabilistic model where the energy function is defined by a bottom-up convolutional neural network (ConvNet or CNN). Learning such a model requires generat…

Cited by 93SourcePDFScholar