← Search

Lei huang

72 accepted papers

2026

ARNS: Adaptive Relation-Aware Negative Sampling with Curriculum Learning for Inductive Knowledge Graph Completion

AAAI 2026technical

Inductive knowledge graph completion (KGC) aims to predict missing links involving unseen entities, making it a particularly challenging task for knowledge representation learning. Traditional embedding-based methods often fall short in this setting due to their limited structural reasoning capabili

Cited by 0SourcePDFScholar
2026

Deep Global-sense Hard-negative Discriminative Generation Hashing for Cross-modal Retrieval

ICLR 2026poster

Hard negative generation (HNG) provides valuable signals for deep learning, but existing methods mostly rely on local correlations while neglecting the global geometry of the embedding space. This limitation often leads to weak discrimination, particularly in cross-modal hashing, which obtains compa…

Cited by 0SourceScholar
2026

Intra-class Distribution-guided Generative Hashing with Neighbor Refinement for Cross-modal Retrieval

CVPR 2026

Recent cross-modal hashing methods have introduced sample generation strategies to enrich training signals. Despite these advances, sample generation-driven hashing still faces two major challenges: (1) Interpolation-based methods adopt deterministic and class-independent generation that restricts s

Cited by 0SourcecodeScholar
2026

LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning

AAAI 2026technical

Joint multilingual instruction tuning is a widely adopted approach to improve the multilingual instruction-following ability and downstream performance of large language models (LLMs), but the resulting multilingual capability remains highly sensitive to the composition and selection of the training

Cited by 0SourcePDFScholar
2026

Polysemic Semantic Instance Network for Cross-Modal Hashing

AAAI 2026technical

Hashing techniques are widely adopted in large-scale cross-modal retrieval due to their efficiency and low storage cost. However, semantic ambiguities, including polysemy, multi-object images, and missing semantic descriptions, significantly degrade the accuracy of alignment and retrieval performanc

Cited by 0SourcePDFScholar
2026

QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpression

AAAI 2026technical

Large language models (LLMs) have shown promising capabilities in hardware description language (HDL) generation. However, existing approaches often rely on free-form natural language descriptions that are often ambiguous, redundant, and unstructured, which poses significant challenges for downstrea

Cited by 0SourcePDFScholar
2026

Steering Performance Optimization for Wheeled Mobile Robots in Granular Media Via DRFM: Enhancing Locomotion Precision and Energy Efficiency

ICRA 2026poster

Degraded steering performance and increased energy consumption present significant barriers to deploying wheeled mobile robots (WMRs) in granular media such as sand and lunar regolith. This study presents and experimentally validates a systematic optimization framework based on Dynamic Resistive For…

Cited by 0SourceScholar
2026

Steering Performance Optimization for Wheeled Mobile Robots in Granular Media via DRFM: Enhancing Locomotion Precision and Energy Efficiency

RA-L 2026

Degraded steering performance and increased energy consumption present significant barriers to deploying wheeled mobile robots (WMRs) in granular media such as sand and lunar regolith. This study presents and experimentally validates a systematic optimization framework based on Dynamic Resistive For

Cited by 0SourceScholar
2025

Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention Learning

ACL 2025long

Large language models (LLMs) are known to suffer from severe hallucination issues. One of the main causes lies in the knowledge misalignment between the pre-training stage and the supervised fine-tuning stage. The unfamiliar knowledge encountered during fine-tuning may encourage LLMs to generate fac…

2025

CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning

ACL 2025long

Current large language models (LLMs) often exhibit imbalanced multilingual capabilities due to their English-centric training corpora. To address this, existing fine-tuning approaches operating at the data-level (e.g., through data augmentation or distillation) typically introduce implicit cross-lin…

Cited by 0SourcePDFScholar
2025

Clustering Properties of Self-Supervised Learning

ICML 2025poster

Self-supervised learning (SSL) methods via joint embedding architectures have proven remarkably effective at capturing semantically rich representations with strong clustering properties, magically in the absence of label supervision. Despite this, few of them have explored leveraging these untapped…

Cited by 0SourcePDFScholar
2025

Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clustering

EMNLP 2025

The rapid growth of scientific literature demands efficient methods to organize and synthesize research findings. Existing taxonomy construction methods, leveraging unsupervised clustering or direct prompting of large language models (LLMs), often lack coherence and granularity. We propose a novel c

Cited by 0SourcePDFScholar
2025

From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems

EMNLP 2025

Research is a fundamental process driving the advancement of human civilization, yet it demands substantial time and effort from researchers. In recent years, the rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance rese

Cited by 0SourcePDFScholar
2025

Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced Optimization

ACL 2025long

Ensuring contextual faithfulness in retrieval-augmented large language models (LLMs) is crucial for building trustworthy information-seeking systems, particularly in long-form question-answering (LFQA) scenarios. In this work, we identify a salient correlation between LFQA faithfulness and retrieval…

2025

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models

EMNLP 2025

Large vision-language models (LVLMs) have demonstrated exceptional capabilities in understanding visual information with human languages but also exhibit an imbalance in multilingual capabilities. In this work, we delve into the multilingual working pattern of LVLMs and identify a salient correlatio

2025

Length Controlled Generation for Black-box LLMs

ACL 2025long

Large language models (LLMs) have demonstrated impressive instruction following capabilities, while still struggling to accurately manage the length of the generated text, which is a fundamental requirement in many real-world applications. Existing length control methods involve fine-tuning the para…

2025

One for All: Update Parameterized Knowledge Across Multiple Models with Once Edit

ACL 2025long

Large language models (LLMs) encode vast world knowledge but struggle to stay up-to-date, often leading to errors and hallucinations. Knowledge editing offers an efficient alternative to retraining, enabling targeted modifications by updating specific model parameters. However, existing methods prim…

Cited by 0SourcePDFScholar
2025

SLAM: Towards Efficient Multilingual Reasoning via Selective Language Alignment

COLING 2025main

Despite the significant improvements achieved by large language models (LLMs) in English reasoning tasks, these models continue to struggle with multilingual reasoning. Recent studies leverage a full-parameter and two-stage training paradigm to teach models to first understand non-English questions…

2025

Towards Global-Topology Relation Graph for Inductive Knowledge Graph Completion

AAAI 2025technical

Knowledge Graphs (KGs) are structured data presented as directed graphs. Due to the common issues of incompleteness and inaccuracy encountered during construction and maintenance, completing KGs becomes a critical task. Inductive Knowledge Graph Completion (KGC) excels at inferring patterns or model…

Cited by 0SourcePDFScholar
2025

Unveiling Entity-Level Unlearning for Large Language Models: A Comprehensive Analysis

COLING 2025main

Large language model unlearning has garnered increasing attention due to its potential to address security and privacy concerns, leading to extensive research in the field. However, existing studies have predominantly focused on instance-level unlearning, specifically targeting the removal of predef…

Cited by 1SourcePDFScholar
2024

A versatile informative diffusion model for single-cell ATAC-seq data generation and analysis

NeurIPS 2024poster

The rapid advancement of single-cell ATAC sequencing (scATAC-seq) technologies holds great promise for investigating the heterogeneity of epigenetic landscapes at the cellular level. The amplification process in scATAC-seq experiments often introduces noise due to dropout events, which results in ex…

Cited by 0SourcePDFScholar
2024

Advancing Large Language Model Attribution through Self-Improving

EMNLP 2024main

Teaching large language models (LLMs) to generate text with citations to evidence sources can mitigate hallucinations and enhance verifiability in information-seeking systems. However, improving this capability requires high-quality attribution data, which is costly and labor-intensive. Inspired by…

Cited by 6SourcePDFScholar
2024

CoR-GS: Sparse-View 3D Gaussian Splatting via Co-Regularization

ECCV 2024poster

"3D Gaussian Splatting (3DGS) creates a radiance field consisting of 3D Gaussians to represent a scene. With sparse training views, 3DGS easily suffers from overfitting, negatively impacting rendering. This paper introduces a new co-regularization perspective for improving sparse-view 3DGS. When tra…

2024

Discrete Modeling via Boundary Conditional Diffusion Processes

NeurIPS 2024poster

We present an novel framework for efficiently and effectively extending the powerful continuous diffusion processes to discrete modeling. Previous approaches have suffered from the discrepancy between discrete data and continuous modeling. Our study reveals that the absence of guidance from discrete…

Cited by 0SourcePDFScholar
2024

Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models

ACL 2024long

Though advanced in understanding visual information with human languages, Large Vision-Language Models (LVLMs) still suffer from multimodal hallucinations. A natural concern is that during multimodal interaction, the generated hallucinations could influence the LVLMs’ subsequent generation. Thus, we…

2024

Learning Fine-Grained Grounded Citations for Attributed Large Language Models

ACL 2024findings

Despite the impressive performance on information-seeking tasks, large language models (LLMs) still struggle with hallucinations. Attributed LLMs, which augment generated text with in-line citations, demonstrate potential in mitigating hallucinations and improving verifiability. However, current app…

2024

Modulate Your Spectrum in Self-Supervised Learning

ICLR 2024poster

Whitening loss offers a theoretical guarantee against feature collapse in self-supervised learning (SSL) with joint embedding architectures. Typically, it involves a hard whitening approach, transforming the embedding and applying loss to the whitened output. In this work, we introduce Spectral Tran…

2024

Robust Synthetic-to-Real Transfer for Stereo Matching

CVPR 2024poster

With advancements in domain generalized stereo matching networks models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However few studies have investigated the robustness after fine-tuning them in real-world scenarios during which the domain generalization ability ca…

2024

VIRL: Self-Supervised Visual Graph Inverse Reinforcement Learning

CoRL 2024poster

Learning dense reward functions from unlabeled videos for reinforcement learning exhibits scalability due to the vast diversity and quantity of video resources. Recent works use visual features or graph abstractions in videos to measure task progress as rewards, which either deteriorate in unseen do…

Cited by 0SourceScholar
2023

Bias Reduced Semidefinite Relaxation Method for Multistatic Localization in the Absence of Transmitter Position And Its Synchronization

ICASSP 2023accepted

This paper addresses the challenging problem of multistatic localization of a stationary object with a set of synchronized receivers, when the transmitter position is unknown and the synchronization with the transmitter is unavailable. Using the time delay measurements from the direct and indirect p…

Cited by 0SourceScholar
2023

MDM: Molecular Diffusion Model for 3D Molecule Generation

AAAI 2023technical

Molecule generation, especially generating 3D molecular geometries from scratch (i.e., 3D de novo generation), has become a fundamental task in drug design. Existing diffusion based 3D molecule generation methods could suffer from unsatisfactory performances, especially when generating large molecul…

2023

WHC: Weighted Hybrid Criterion for Filter Pruning on Convolutional Neural Networks

ICASSP 2023accepted

Filter pruning has attracted increasing attention in recent years for its capacity in compressing and accelerating convolutional neural networks. Various data-independent criteria, including norm-based and relationship-based ones, were proposed to prune the most unimportant filters. However, these s…

Cited by 0SourceScholar
2022

An Investigation into Whitening Loss for Self-supervised Learning

NeurIPS 2022accept

A desirable objective in self-supervised learning (SSL) is to avoid feature collapse. Whitening loss guarantees collapse avoidance by minimizing the distance between embeddings of positive pairs under the conditioning that the embeddings from different views are whitened. In this paper, we propose…

2022

Bi-Level Doubly Variational Learning for Energy-Based Latent Variable Models

CVPR 2022poster

Energy-based latent variable models (EBLVMs) are more expressive than conventional energy-based models. However, its potential on visual tasks are limited by its training process based on maximum likelihood estimate that requires sampling from two intractable distributions. In this paper, we propose…

Cited by 7PDFScholar
2022

Broadband Sound Source Localisation via Non-Synchronous Measurements for Service Robots: A Tensor Completion Approach

RA-L 2022

Constraint by the physical geometry, the lower and upper frequency bound and the scale of the scanning area of a microphone array are limited. Owing to its movable feature, for the service robots, achieving a wider working frequency range with a global view requires a virtually larger and denser arr

Cited by 16SourceScholar
2022

Delving Into the Estimation Shift of Batch Normalization in a Network

CVPR 2022poster

Batch normalization (BN) is a milestone technique in deep learning. It normalizes the activation using mini-batch statistics during training but the estimated population statistics during inference. This paper focuses on investigating the estimation of population statistics. We define the estimation…

Cited by 28PDFcodeScholar
2022

Revisiting Domain Generalized Stereo Matching Networks From a Feature Consistency Perspective

CVPR 2022poster

Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization ca…

Cited by 79PDFcodeScholar
2022

Semidefinite Relaxation Method for Moving Object Localization Using a Stationary Transmitter at Unknown Position

ICASSP 2022accepted

This paper addresses the multistatic localization of a moving object in position and velocity using time delay (TD) and Doppler frequency shift (DFS) measurements, where the position of the transmitter is unknown and has not yet been synchronized with the receivers. Based on the TD and DFS measureme…

Cited by 0SourceScholar
2022

Understanding the Failure of Batch Normalization for Transformers in NLP

NeurIPS 2022accept

Batch Normalization (BN) is a core and prevalent technique in accelerating the training of deep neural networks and improving the generalization on Computer Vision (CV) tasks. However, it fails to defend its position in Natural Language Processing (NLP), which is dominated by Layer Normalization (LN…

2021

CCT-Net: Category-Invariant Cross-Domain Transfer for Medical Single-to-Multiple Disease Diagnosis

ICCV 2021poster

A medical imaging model is usually explored for the diagnosis of a single disease. However, with the expanding demand for multi-disease diagnosis in clinical applications, multi-function solutions need to be investigated. Previous works proposed to either exploit different disease labels to conduct…

Cited by 11PDFScholar
2021

Group Whitening: Balancing Learning Efficiency and Representational Capacity

CVPR 2021poster

Batch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model's learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population…

Cited by 24PDFcodeScholar
2021

Many-to-One Distribution Learning and K-Nearest Neighbor Smoothing for Thoracic Disease Identification

AAAI 2021technical

Chest X-rays are an important and accessible clinical imaging tool for the detection of many thoracic diseases. Over the past decade, deep learning, with a focus on the convolutional neural network (CNN), has become the most powerful computer-aided diagnosis technology for improving disease identifi…

Cited by 14SourcePDFScholar
2021

Rot-Pro: Modeling Transitivity by Projection in Knowledge Graph Embedding

NeurIPS 2021poster

Knowledge graph embedding models learn the representations of entities and relations in the knowledge graphs for predicting missing links (relations) between entities. Their effectiveness are deeply affected by the ability of modeling and inferring different relation patterns such as symmetry, asymm…

2021

S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-Bit Neural Networks via Guided Distribution Calibration

CVPR 2021poster

Previous studies dominantly target at self-supervised learning on real-valued networks and have achieved many promising results. However, on the more challenging binary neural networks (BNNs), this task has not yet been fully explored in the community. In this paper, we focus on this more difficult…

Cited by 23PDFcodeScholar
2021

Slimmable Generative Adversarial Networks

AAAI 2021technical

Generative adversarial networks (GANs) have achieved remarkable progress in recent years, but the continuously growing scale of models make them challenging to deploy widely in practical applications. In particular, for real-time generation tasks, different devices require generators of different si…

2021

Visual-Textual Attentive Semantic Consistency for Medical Report Generation

ICCV 2021poster

Diagnosing diseases from medical radiographs and writing reports requires professional knowledge and is time-consuming. To address this, automatic medical report generation approaches have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding siz…

Cited by 24PDFScholar
2020

Layer-wise Conditioning Analysis in Exploring the Learning Dynamics of DNNs

ECCV 2020poster

Conditioning analysis uncovers the landscape of an optimization objective by exploring the spectrum of its curvature matrix. This has been well explored theoretically for linear models. We extend this analysis to deep neural networks (DNNs) in order to investigate their learning dynamics. To this en…

Cited by 12SourcePDFScholar
2020

On the Number of Linear Regions of Convolutional Neural Networks

ICML 2020poster

One fundamental problem in deep learning is understanding the outstanding performance of deep Neural Networks (NNs) in practice. One explanation for the superiority of NNs is that they can realize a large class of complicated functions, i.e., they have powerful expressivity. The expressivity of a Re…

Cited by 98SourcePDFScholar
2019

A Novel Deep Hashing Method with Top Similarity for Image Retrieval

ICASSP 2019accepted

Due to the advantages of retrieval speed and storage space, deep hashing methods have become a research hotspot in the field of large-scale image retrieval. Most of existing deep hashing methods pay close attention to similarity between images without images at the top of the ranking list similar to…

Cited by 0SourceScholar
2019

Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical Images

CVPR 2019poster

Medical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires imag…

Cited by 327PDFScholar
2019

Iterative Normalization: Beyond Standardization Towards Efficient Whitening

CVPR 2019poster

Batch Normalization (BN) is ubiquitously employed for accelerating neural network training and improving the generalization capability by performing standardization within mini-batches. Decorrelated Batch Normalization (DBN) further boosts the above effectiveness by whitening. However, DBN relies…

Cited by 183PDFcodeScholar
2017

Centered Weight Normalization in Accelerating Training of Deep Neural Networks

ICCV 2017poster

Training deep neural networks is difficult for the pathological curvature problem. Re-parameterization is an effective way to relieve the problem by learning the curvature approximately or constraining the solutions of weights with good properties for optimization. This paper proposes to re-paramete…

Cited by 88PDFcodeScholar
2016

Accurate asymptotic analysis for John's test in multichannel signal detection

ICASSP 2016accepted

John's test, which is also known as the locally most invariant test for sphericity of Gaussian variables, is one of the most frequently used methods in multichannel signal detection. The application of John's test requires closed-form and accurate formula to set threshold according to a prescribed f…

Cited by 0SourceScholar
2016

Iteratively reweighted tensor SVD for robust multi-dimensional harmonic retrieval

ICASSP 2016accepted

In this paper, parameter estimation for multi-dimensional sinusoids in additive impulsive noise is addressed. Our underlying idea is to minimize the ℓ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">p</sub> -norm of the residual error tensor, where 1 <;…

Cited by 6SourceScholar
2016

Least squares phase retrieval using feasible point pursuit

ICASSP 2016accepted

Phase retrieval has recently attracted renewed interest. It is revisited here through a new approach based on nonconvex quadratically constrained quadratic programming (QCQP). A least-squares (LS) formulation is adopted, and a recently developed non-convex QCQP approximation technique called feasibl…

Cited by 0SourceScholar
2016

Sparse recovery of multiple measurement vectors in impulsive noise: A smooth block successive minimization algorithm

ICASSP 2016accepted

This paper considers the sparse recovery problem of multiple measurement vector (MMV) model corrupted in impulsive noise. To ensure outlier-robust sparse recovery, we formulate an MMV problem that includes the generalized ℓ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.…

Cited by 2SourceScholar
2015

An improved cross-correlation approach to parameter estimation based on fractional Fourier transform for ISAR motion compensation

ICASSP 2015accepted

Motion compensation (MOCOMP) is a key procedure in inverse synthetic aperture radar (ISAR) imaging because the accuracy of estimated parameter has a strong influence on the imaging quality. Generally, the backscattered signal of a moving target is sampled in fast time dimension, which can be approxi…

Cited by 0SourceScholar
2015

Joint direction-of-arrival and frequency estimation without source enumeration

ICASSP 2015accepted

Joint estimation of the directions-of-arrival (DOAs) and frequencies of multiple signals is addressed in this paper. By constructing a set of joint diagonalization matrices, two cost functions that do not require a priori information of the source number are devised for DOA and frequency estimation…

Cited by 0SourceScholar
2015

Robust widely linear beamformer based on a projection constraint

ICASSP 2015accepted

For noncircular signals, optimal widely linear (WL) minimum variance distortionless response (MVDR) beamformer has a powerful performance by exploiting the noncircularity of the received signals. Though, the noncircularity rate can be estimated by the steering vector (SV) of the signal of interest (…

Cited by 0SourceScholar