← Search

Mao Ye

51 accepted papers

2026

CHAL: Causal-guided Hierarchical Anomaly-aware Learning for Moving Infrared Small Target Detection

CVPR 2026

Infrared small target detection is one highly special category of object detection, faced with tiny target imaging size and cluttered backgrounds. Currently, almost all existing methods are target-centered, directly learning the target features from backgrounds. However, due to weak target signals,

Cited by 0SourcecodeScholar
2026

Consistent Text-to-Image Generation via Scene De-Contextualization

ICLR 2026poster

Consistent text-to-image (T2I) generation seeks to produce identity-preserving images of the same subject across diverse scenes, yet it often fails due to a phenomenon called identity (ID) shift. Previous methods have tackled this issue, but typically rely on the unrealistic assumption of knowing al…

Cited by 0SourcecodeScholar
2026

Cross-domain Joint Learning with Prototype-guided Mixture-of-Experts for Infrared Moving Small Target Detection

AAAI 2026technical

Infrared small target detection often faces significant domain gaps across datasets due to varying sensors and scene distributions. Currently, most existing methods are typically based on single-domain learning (i.e., training and test are on the same dataset), requiring training separate detectors

Cited by 0SourcePDFScholar
2026

Domain-Auxiliary Infrared Moving Small Target Detection by Learning to Overlook Domain Discrepancy

AAAI 2026technical

Currently, almost all traditional infrared small target detection methods work on the assumption that training and test sets always belong to the same domain, and training samples are sufficient. However, in real applications, a new detection task could often have no sufficient training samples from

Cited by 0SourcePDFScholar
2026

Hierarchical Frequency-Guided Alignment Transformer for Compressed Video Quality Enhancement

AAAI 2026technical

During the video encoding process, the original spatial domain signal is first transformed into the frequency domain, followed by quantization and compression. As a result, the quality degradation in compressed videos primarily stems from distortions in the frequency domain information. However, exi

Cited by 0SourcePDFScholar
2025

Bayesian Test-Time Adaptation for Vision-Language Models

CVPR 2025poster

Test-time adaptation with pre-trained vision-language models, such as CLIP, aims to adapt the model to new, potentially out-of-distribution test data. Existing methods calculate the similarity between visual embedding and learnable class embeddings, which are initialized by text embeddings, for zer…

Cited by 0SourcePDFScholar
2025

Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing Data

CVPR 2025poster

Domain shift (the difference between source and target domains) poses a significant challenge in clinical applications, e.g., Diabetic Retinopathy (DR) grading. Despite considering certain clinical requirements, like source data privacy, conventional transfer methods are predominantly model-centered…

2025

FreeCap: Hybrid Calibration-Free Motion Capture in Open Environments

AAAI 2025technical

We propose a novel hybrid calibration-free method FreeCap to accurately capture global multi-person motions in open environments. Our system combines a single LiDAR with expandable moving cameras, allowing for flexible and precise motion estimation in a unified world coordinate. In particular, We in…

Cited by 0SourcePDFScholar
2025

Generative Data Mining with Longtail-Guided Diffusion

ICML 2025poster

It is difficult to anticipate the myriad challenges that a predictive model will encounter once deployed. Common practice entails a reactive, cyclical approach: model deployment, data mining, and retraining. We instead develop a proactive longtail discovery process by imagining additional data durin…

Cited by 0SourcePDFScholar
2025

High Dynamic Range Novel View Synthesis with Single Exposure

ICML 2025poster

High Dynamic Range Novel View Synthesis (HDR-NVS) aims to establish a 3D scene HDR model from Low Dynamic Range (LDR) imagery. Typically, multiple-exposure LDR images are employed to capture a wider range of brightness levels in a scene, as a single LDR image cannot represent both the brightest and…

2025

Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational Pathology

ICLR 2025poster

Histopathology Whole-Slide Images (WSIs) provide an important tool to assess cancer prognosis in computational pathology (CPATH). While existing survival analysis (SA) approaches have made exciting progress, they are generally limited to adopting highly-expressive network architectures and only coar…

2025

Motion Prior Knowledge Learning with Homogeneous Language Descriptions for Moving Infrared Small Target Detection

AAAI 2025technical

Different from traditional object detection, pure vision is not enough to infrared small target detection, due to small target size and weak background contrast. For promoting detection performance, more target representations are needed. Currently, motion representations have been proved to be one…

2025

Multimodal Causal Reasoning for UAV Object Detection

NeurIPS 2025poster

Unmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited inf…

Cited by 0SourceScholar
2025

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

CVPR 2025highlight

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality, human-centric video datasets presents a challenge to progress in…

Cited by 2SourcePDFScholar
2025

Proxy Denoising for Source-Free Domain Adaptation

ICLR 2025oral

Source-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to an unlabeled target domain with no access to the source data. Inspired by the success of large Vision-Language (ViL) models in many applications, the latest research has validated ViL's benefit for SFDA by using their p…

2025

Pseudo Visible Feature Fine-Grained Fusion for Thermal Object Detection

CVPR 2025poster

Thermal object detection is a critical task in various fields, such as surveillance and autonomous driving. Current state-of-the-art (SOTA) models always leverage a prior Thermal-To-Visible (T2V) translation model to obtain visible spectrum information, followed by a cross-modality aggregation modul…

2025

Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image Classification

AAAI 2025technical

Whole Slide Image (WSI) classification has very significant applications in clinical pathology, e.g., tumor identification and cancer diagnosis. Currently, most research attention is focused on Multiple Instance Learning (MIL) using static datasets. One of the most obvious weaknesses of these method…

2025

Self-Prompting Analogical Reasoning for UAV Object Detection

AAAI 2025technical

Unmanned Aerial Vehicle Object Detection (UAVOD) presents unique challenges due to varying altitudes, dynamic backgrounds, and the small size of objects. Traditional detection methods often struggle with these challenges, as they typically rely on visual feature only and fail to extract the semantic…

Cited by 0SourcePDFScholar
2025

Uncertainty-Guided Enhancement on Driving Perception System Via Foundation Models

ICRA 2025

Multimodal foundation models offer promising advancements for enhancing driving perception systems, but their high computational and financial costs pose challenges. We develop a method that leverages foundation models to refine predictions from existing driving perception modelssuch as enhancing ob

Cited by 4SourceScholar
2024

Cloud Object Detector Adaptation by Integrating Different Source Knowledge

NeurIPS 2024poster

We propose to explore an interesting and promising problem, Cloud Object Detector Adaptation (CODA), where the target domain leverages detections provided by a large cloud model to build a target detector. Despite with powerful generalization capability, the cloud model still cannot achieve error-fr…

Cited by 3SourcePDFScholar
2024

Source-Free Domain Adaptation with Frozen Multimodal Foundation Model

CVPR 2024poster

Source-Free Domain Adaptation (SFDA) aims to adapt a source model for a target domain with only access to unlabeled target training data and the source model pretrained on a supervised source domain. Relying on pseudo labeling and/or auxiliary supervision conventional methods are inevitably error-pr…

2023

Bit Allocation using Optimization

ICML 2023poster

In this paper, we consider the problem of bit allocation in Neural Video Compression (NVC). First, we reveal a fundamental relationship between bit allocation in NVC and Semi-Amortized Variational Inference (SAVI). Specifically, we show that SAVI with GoP (Group-of-Picture)-level likelihood is equiv…

2023

Efficient Transformer-based 3D Object Detection with Dynamic Token Halting

ICCV 2023poster

Balancing efficiency and accuracy is a long-standing problem for deploying deep learning models. The trade-off is even more important for real-time safety-critical systems like autonomous vehicles. In this paper, we propose an effective approach for accelerating transformer-based 3D object detectors…

Cited by 8PDFScholar
2023

Homeomorphism Alignment for Unsupervised Domain Adaptation

ICCV 2023poster

Existing unsupervised domain adaptation (UDA) methods rely on aligning the features from the source and target domains explicitly or implicitly in a common space (i.e., the domain invariant space). Explicit distribution matching ignores the discriminability of learned features, while the implicit co…

Cited by 14PDFcodeScholar
2023

Independent Feature Decomposition and Instance Alignment for Unsupervised Domain Adaptation

IJCAI 2023poster

Existing Unsupervised Domain Adaptation (UDA) methods typically attempt to perform knowledge transfer in a domain-invariant space explicitly or implicitly. In practice, however, the obtained features is often mixed with domain-specific information which causes performance degradation. To overcome th…

2022

A Multiscale Gradient-Backpropagation Optimization Framework for Deformable Convolution Based Compressed Video Enhancement

ICASSP 2022accepted

Deep learning based compressed video quality enhancement has raised lots of interest recently. To explore the information over multiple frames, deformable convolution has been used for temporal alignment. However, in the existing methods, the deformable convolution is used in a relatively naïve way,…

Cited by 1SourceScholar
2022

BOME! Bilevel Optimization Made Easy: A Simple First-Order Approach

NeurIPS 2022accept

Bilevel optimization (BO) is useful for solving a variety of important machine learning problems including but not limited to hyperparameter optimization, meta-learning, continual learning, and reinforcement learning. Conventional BO methods need to differentiate through the low-level optimization p…

Cited by 94SourcePDFScholar
2022

Diffusion-based Molecule Generation with Informative Prior Bridges

NeurIPS 2022accept

AI-based molecule generation provides a promising approach to a large area of biomedical sciences and engineering, such as antibody design, hydrolase engineering, or vaccine development. Because the molecules are governed by physical laws, a key challenge is to incorporate prior information into the…

Cited by 119SourcePDFScholar
2022

First Hitting Diffusion Models for Generating Manifold, Graph and Categorical Data

NeurIPS 2022accept

We propose a family of First Hitting Diffusion Models (FHDM), deep generative models that generate data with a diffusion process that terminates at a random first hitting time. This yields an extension of the standard fixed-time diffusion models that terminate at a pre-specified deterministic time.…

Cited by 29SourcePDFScholar
2022

Future gradient descent for adapting the temporal shifting data distribution in online recommendation systems

UAI 2022poster

One of the key challenges of learning an online recommendation model is the temporal domain shift, which causes the mismatch between the training and testing data distribution and hence domain generalization error. To overcome, we propose to learn a meta future gradient generator that forecasts the…

Cited by 8SourcePDFScholar
2022

KUNet: Imaging Knowledge-Inspired Single HDR Image Reconstruction

IJCAI 2022poster

Recently, with the rise of high dynamic range (HDR) display devices, there is a great demand to transfer traditional low dynamic range (LDR) images into HDR versions. The key to success is how to solve the many-to-many mapping problem. However, the existing approaches either do not consider constrai…

2022

MetaTeacher: Coordinating Multi-Model Domain Adaptation for Medical Image Classification

NeurIPS 2022accept

In medical image analysis, we often need to build an image recognition system for a target scenario with the access to small labeled data and abundant unlabeled data, as well as multiple related models pretrained on different source scenarios. This presents the combined challenges of multi-source-fr…

2022

Multi-Class 3D Object Detection with Single-Class Supervision

ICRA 2022poster

While multi-class 3D detectors are needed in many robotics applications, training them with fully labeled datasets can be expensive in labeling cost. An alternative approach is to have targeted single-class labels on disjoint data samples. In this paper, we are interested in training a multi-class 3…

Cited by 2SourceScholar
2022

Pareto navigation gradient descent: a first-order algorithm for optimization in pareto set

UAI 2022poster

Many modern machine learning applications, such as multi-task learning, require finding optimal model parameters to trade-off multiple objective functions that may conflict with each other. The notion of the Pareto set allows us to focus on the set of (often infinite number of) models that cannot be…

2022

Source-Free Object Detection by Learning To Overlook Domain Style

CVPR 2022oral

Source-free object detection (SFOD) needs to adapt a detector pre-trained on a labeled source domain to a target domain, with only unlabeled training data from the target domain. Existing SFOD methods typically adopt the pseudo labeling paradigm with model adaption alternating between predicting pse…

Cited by 68PDFcodeScholar
2021

MaxUp: Lightweight Adversarial Training With Data Augmentation Improves Neural Network Training

CVPR 2021poster

We propose MaxUp, an embarrassingly simple, highly effective technique for improving the generalization performance of machine learning models, especially deep neural networks. The idea is to generate a set of augmented data with some random perturbations or transforms, and minimize the maximum, or…

Cited by 82PDFcodeScholar
2021

Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision

AAAI 2021technical

We consider the post-training quantization problem, which discretizes the weights of pre-trained deep neural networks without re-training the model. We propose multipoint quantization, a quantization method that approximates a full-precision weight vector using a linear combination of multiple vecto…

2021

VCNet and Functional Targeted Regularization For Learning Causal Effects of Continuous Treatments

ICLR 2021oral

Motivated by the rising abundance of observational data with continuous treatments, we investigate the problem of estimating the average dose-response curve (ADRF). Available parametric methods are limited in their model space, and previous attempts in leveraging neural network to enhance model expr…

2021

argmax centroid

NeurIPS 2021poster

We propose a general method to construct centroid approximation for the distribution of maximum points of a random function (a.k.a. argmax distribution), which finds broad applications in machine learning. Our method optimizes a set of centroid points to compactly approximate the argmax distribution…

Cited by 0SourcePDFScholar
2020

Black-Box Certification with Randomized Smoothing: A Functional Optimization Based Framework

NeurIPS 2020poster

Randomized classifiers have been shown to provide a promising approach for achieving certified robustness against adversarial attacks in deep learning. However, most existing methods only leverage Gaussian smoothing noise and only work for $\ell_2$ perturbation. We propose a general framework of adv…

2020

Distribution-Aware Coordinate Representation for Human Pose Estimation

CVPR 2020poster

While being the de facto standard coordinate representation for human pose estimation, heatmap has not been investigated in-depth. This work fills this gap. For the first time, we find that the process of decoding the predicted heatmaps into the final joint coordinates in the original image space is…

Cited by 627PDFcodeScholar
2020

Go Wide, Then Narrow: Efficient Training of Deep Thin Networks

ICML 2020poster

For deploying a deep learning model into production, it needs to be both accurate and compact to meet the latency and memory constraints. This usually results in a network that is deep (to ensure performance) and yet thin (to improve computational efficiency). In this paper, we propose an efficient…

Cited by 23SourcePDFScholar
2020

Good Subnetworks Provably Exist: Pruning via Greedy Forward Selection

ICML 2020poster

Recent empirical works show that large deep neural networks are often highly redundant and one can find much smaller subnetworks without a significant drop of accuracy. However, most existing methods of network pruning are empirical and heuristic, leaving it open whether good subnetworks provably ex…

2020

Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough

NeurIPS 2020poster

Despite the great success of deep learning, recent works show that large deep neural networks are often highly redundant and can be significantly reduced in size. However, the theoretical question of how much we can prune a neural network given a specified tolerance of accuracy drop is still open. T…

2019

Fast Human Pose Estimation

CVPR 2019poster

Existing human pose estimation approaches often only consider how to improve the model generalisation performance, but putting aside the significant efficiency problem. This leads to the development of heavy models with poor scalability and cost-effectiveness in practical use. In this work, we inves…

Cited by 356PDFScholar
2015

3D Reconstruction in the Presence of Glasses by Acoustic and Stereo Fusion

CVPR 2015poster

We present a practical and inexpensive method to reconstruct 3D scenes that include piece-wise planar transparent objects. Our work is motivated by the need for automatically generating 3D models of interior scenes, in which glass structures are common. These large structures are often invisible to…

Cited by 44SourcePDFScholar