← Search

Le zhang

42 accepted papers

2026

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

AAAI 2026technical

The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further

Cited by 0SourcePDFScholar
2026

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs

CVPR 2026

Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they are for. While existing benchmarks effectively evaluate the geometric perception capabilities of multimodal large language models (MLLMs), they fall s

Cited by 0SourceScholar
2026

Graph Meets Deep Unfolding: An Interpretable Mutual-benefit Multi-view Learning Network

AAAI 2026technical

Significant efforts have been focused on enhancing the utilization of multiple node features and topological structures in multi-view graph learning through explicit model-driven and implicit deep learning-based methodologies. The former excels in embedding prior knowledge, thereby offering theoreti

Cited by 0SourcePDFScholar
2026

WiTTA-Bench: Benchmarking Test-Time Adaptation for WiFi Sensing

CVPR 2026

WiFi sensing offers passive and privacy-preserving perception that complements vision-based sensing, but its performance degrades sharply under domain shifts caused by changes in environment, subjects, or hardware. This challenge is exacerbated in real-world deployments where source data are unavail

Cited by 0SourcecodeScholar
2025

Assessing and Learning Alignment of Unimodal Vision and Language Models

CVPR 2025highlight

How well are unimodal vision and language models aligned? While prior work has explored this question, their assessment methods do not directly translate to practical vision-language tasks. In this paper, we propose a direct assessment method, inspired by linear probing, to evaluate vision-language…

2025

CharacterBench: Benchmarking Character Customization of Large Language Models

AAAI 2025technical

Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs’ character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a si…

2025

Codar: Complex-valued Neural Network for Crossing-Floor Intrusion Detection via WiFi

ICASSP 2025accepted

WiFi systems offer enormous potential for device-free human intrusion detection. Current methods often require routers to be deployed in multiple adjacent rooms on the same floor, which is redundant and costly. To solve this, we introduce the first work on intrusion detection in the crossing-floor s…

Cited by 0SourceScholar
2025

Handling Spatial-Temporal Data Heterogeneity for Federated Continual Learning via Tail Anchor

CVPR 2025poster

Federated Continual Learning (FCL) allows each client to continually update its knowledge from task streams, enhancing the applicability of federated learning in real-world scenarios. However, FCL needs to address not only spatial data heterogeneity between clients but also temporal data heterogenei…

2025

Improving Acoustic Scene Classification in Low-Resource Conditions

ICASSP 2025accepted

Acoustic Scene Classification (ASC) identifies an environment based on an audio signal. This paper explores ASC in low-resource conditions and proposes a novel model, DS-FlexiNet, which combines depthwise separable convolutions from MobileNetV2 with ResNet-inspired residual connections for a balance…

Cited by 0SourceScholar
2025

MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices

ICCV 2025poster

Recent advancements in deep neural networks have driven significant progress in image enhancement (IE). However, deploying deep learning models on resource-constrained platforms, such as mobile devices, remains challenging due to high computation and memory demands. To address these challenges and f…

2025

REARANK: Reasoning Re-ranking Agent via Reinforcement Learning

EMNLP 2025

We present REARANK, a large language model (LLM)-based listwise reasoning rerank- ing agent. REARANK explicitly reasons be- fore reranking, significantly improving both performance and interpretability. Leveraging reinforcement learning and data augmentation, REARANK achieves substantial improvement

2025

Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level Tasks

CVPR 2025poster

Parameter-efficient fine-tuning (PEFT) adapts pre-trained models to new tasks by updating only a small subset of parameters, achieving efficiency but still facing significant inference costs driven by input token length. This challenge is even more pronounced in pixel-level tasks, which require long…

2025

Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning

CVPR 2025poster

Heterogeneous Federated Learning (HFL) has received widespread attention due to its adaptability to different models and data. The HFL approach utilizing auxiliary models for knowledge transfer enhances flexibility. However, existing frameworks face the challenges of aggregation bias and local over…

2025

ThermalGaussian: Thermal 3D Gaussian Splatting

ICLR 2025poster

Thermography is especially valuable for the military and other users of surveillance cameras. Some recent methods based on Neural Radiance Fields (NeRF) are proposed to reconstruct the thermal scenes in 3D from a set of thermal and RGB images. However, unlike NeRF, 3D Gaussian splatting (3DGS) preva…

2025

WiFi CSI Based Temporal Activity Detection via Dual Pyramid Network

AAAI 2025technical

We address the challenge of WiFi-based temporal activity detection and propose an efficient Dual Pyramid Network that integrates Temporal Signal Semantic Encoders and Local Sensitive Response Encoders. The Temporal Signal Semantic Encoder splits feature learning into high and low-frequency componen…

2024

Cognitive Virtual Sensing Technique for Feedforward Active Noise Control

ICASSP 2024accepted

The virtual sensing (VS) technique enables an active noise control (ANC) system to estimate the virtual error signal for control using remote monitoring microphones. However, instances where noise characteristics and primary paths exhibit variations lead to a noticeable decline in performance for th…

Cited by 0SourceScholar
2024

Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding

CVPR 2024poster

Vision-Language Models (VLMs) such as CLIP exhibit strong image-text comprehension abilities facilitating advances in several downstream tasks such as zero-shot image classification image-text retrieval and text-to-image generation. However the compositional reasoning abilities of existing VLMs rema…

Cited by 18SourcePDFScholar
2024

CorrMatch: Label Propagation via Correlation Matching for Semi-Supervised Semantic Segmentation

CVPR 2024poster

This paper presents a simple but performant semi-supervised semantic segmentation approach called CorrMatch. Previous approaches mostly employ complicated training strategies to leverage unlabeled data but overlook the role of correlation maps in modeling the relationships between pairs of locations…

2024

DGR: A General Graph Desmoothing Framework for Recommendation via Global and Local Perspectives

IJCAI 2024poster

Graph Convolutional Networks (GCNs) have become pivotal in recommendation systems for learning user and item embeddings by leveraging the user-item interaction graph's node information and topology. However, these models often face the famous over-smoothing issue, leading to indistinct user and item…

2024

Deep Feature Surgery: Towards Accurate and Efficient Multi-Exit Networks

ECCV 2024poster

"Multi-exit network is a promising architecture for efficient model inference by sharing backbone networks and weights among multiple exits. However, the gradient conflict of the shared weights results in sub-optimal accuracy. This paper introduces Deep Feature Surgery (), which consists of feature…

2024

Early Preparation Pays Off: New Classifier Pre-tuning for Class Incremental Semantic Segmentation

ECCV 2024poster

"Class incremental semantic segmentation aims to preserve old knowledge while learning new tasks, however, it is impeded by catastrophic forgetting and background shift issues. Prior works indicate the pivotal importance of initializing new classifiers and mainly focus on transferring knowledge from…

2024

Exploring the Best Practices of Query Expansion with Large Language Models

EMNLP 2024finding

Large Language Models (LLMs) are foundational in language technologies, particularly in information retrieval (IR). In this paper, we thoroughly explore the best practice of leveraging LLMs for query expansion. To this end, we introduce a training-free, straightforward yet effective framework called…

2024

MSA Generation with Seqs2Seqs Pretraining: Advancing Protein Structure Predictions

NeurIPS 2024poster

Deep learning models like AlphaFold2 have revolutionized protein structure prediction, achieving unprecedented accuracy. However, the dependence on robust multiple sequence alignments (MSAs) continues to pose a challenge, especially for proteins that lack a wealth of homologous sequences. To overcom…

2024

RCBEVDet: Radar-camera Fusion in Bird's Eye View for 3D Object Detection

CVPR 2024poster

Three-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However relying solely on cameras is difficult to achieve highly accurate and robust…

2023

Feature Modulation Transformer: Cross-Refinement of Global Representation via High-Frequency Prior for Image Super-Resolution

ICCV 2023poster

Transformer-based methods have exhibited remarkable potential in single image super-resolution (SISR) by effectively extracting long-range dependencies. However, most of the current research in this area has prioritized the design of transformer blocks to capture global information, while overlookin…

Cited by 79PDFcodeScholar
2023

MoqaGPT : Zero-Shot Multi-modal Open-domain Question Answering with Large Language Model

EMNLP 2023long findings

Multi-modal open-domain question answering typically requires evidence retrieval from databases across diverse modalities, such as images, tables, passages, etc. Even Large Language Models (LLMs) like GPT-4 fall short in this task. To enable LLMs to tackle the task in a zero-shot manner, we introduc…

Cited by 0SourcecodeScholar
2022

Mining Relations among Cross-Frame Affinities for Video Semantic Segmentation

ECCV 2022poster

"The essence of video semantic segmentation (VSS) is how to leverage temporal information for prediction. Previous efforts are mainly devoted to developing new techniques to calculate the cross-frame affinities such as optical flow and attention. Instead, this paper contributes from a different angl…

2022

Probing Simile Knowledge from Pre-trained Language Models

ACL 2022long

Simile interpretation (SI) and simile generation (SG) are challenging tasks for NLP because models require adequate world knowledge to produce predictions. Previous works have employed many hand-crafted resources to bring knowledge-related into models, which is time-consuming and labor-intensive. In…

2022

TreeMix: Compositional Constituency-based Data Augmentation for Natural Language Understanding

NAACL 2022long

Data augmentation is an effective approach to tackle over-fitting. Many previous works have proposed different data augmentations strategies for NLP, such as noise injection, word replacement, back-translation etc. Though effective, they missed one important characteristic of language–compositionali…

2021

A Multi-Stage Progressive Learning Strategy for Covid-19 Diagnosis Using Chest Computed Tomography with Imbalanced Data

ICASSP 2021accepted

In this paper, a multi-stage progressive learning strategy is investigated to train classifiers for COVID-19 Diagnosis using imbalanced Chest Computed Tomography Data acquired from patients infected with COVID-19 Pneumonia, Community Acquired Pneumonia (CAP) and from normal healthy subjects. In the…

Cited by 0SourceScholar
2021

Learning to Iteratively Solve Routing Problems with Dual-Aspect Collaborative Transformer

NeurIPS 2021poster

Recently, Transformer has become a prevailing deep architecture for solving vehicle routing problems (VRPs). However, it is less effective in learning improvement models for VRP because its positional encoding (PE) method is not suitable in representing VRP solutions. This paper presents a novel Dua…

2021

Two-Stream Convolution Augmented Transformer for Human Activity Recognition

AAAI 2021technical

Recognition of human activities is an important task due to its far-reaching applications such as healthcare system, context-aware applications, and security monitoring. Recently, WiFi based human activity recognition (HAR) is becoming ubiquitous due to its non-invasiveness. Existing WiFi-based HAR…

2020

Disentangling Human Error from Ground Truth in Segmentation of Medical Images

NeurIPS 2020poster

Recent years have seen increasing use of supervised learning methods for segmentation tasks. However, the predictive performance of these algorithms depends on the quality of labels. This problem is particularly pertinent in the medical image domain, where both the annotation cost and inter-observer…

2020

Identification of Essential Proteins Using A Novel Multi-Objective Optimization Method

ICASSP 2020accepted

Using graph theory to identify essential proteins is a hot topic at present. These methods are called network-based methods. However, the generalization ability of most network-based methods is not satisfactory. Hence, in this paper, we consider the identification of essential proteins as a multi-ob…

Cited by 0SourceScholar
2019

Contrast Prior and Fluid Pyramid Integration for RGBD Salient Object Detection

CVPR 2019poster

The large availability of depth sensors provides valuable complementary information for salient object detection (SOD) in RGBD images. However, due to the inherent difference between RGB and depth information, extracting features from the depth channel using ImageNet pre-trained backbone models and…

Cited by 451PDFScholar
2018

Crowd Counting With Deep Negative Correlation Learning

CVPR 2018poster

Deep convolutional networks (ConvNets) have achieved unprecedented performances on many computer vision tasks. However, their adaptations to crowd counting on single images are still in their infancy and suffer from severe over-fitting. Here we propose a new learning strategy to produce generalizabl…

2017

Analysis of keyword spotting performance across IARPA babel languages

ICASSP 2017accepted

With the completion of the IARPA Babel program, it is possible to systematically analyze the performance of speech recognition systems across a wide variety of languages. We select 16 languages from the dataset and compare performance using a deep neural network-based acoustic model. The focus is on…

Cited by 0SourceScholar
2017

Robust Visual Tracking Using Oblique Random Forests

CVPR 2017poster

Random forest has emerged as a powerful classification technique with promising results in various vision tasks including image classification, pose estimation and object detection. However, current techniques have shown little improvements in visual tracking as they mostly rely on piece wise orthog…

Cited by 99PDFcodeScholar
2017

The 2016 BBN Georgian telephone speech keyword spotting system

ICASSP 2017accepted

In this paper we describe the 2016 BBN conversational telephone speech keyword spotting system; the culmination of four years of research and development under the IARPA Babel program. The system was constructed in response to the NIST Open Keyword Search (OpenKWS) evaluation of 2016. We present our…

Cited by 0SourceScholar