← Search

Ming Sun

61 accepted papers

2026

InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion Prior

CVPR 2026

Video inverse problems such as inpainting, deblurring and super-resolution are fundamental to streaming, telepresence, and AR/VR, where high perceptual quality must coexist with tight latency constraints. Diffusion-based priors currently deliver state-of-the-art reconstructions, but existing approac

Cited by 0SourceScholar
2026

SEPS: Semantic-Enhanced Patch Slimming Framework for Fine-Grained Cross-Modal Alignment

ICML 2026poster

Fine-grained cross-modal alignment is pivotal for multimodal reasoning yet remains limited by Semantic Sparsity Bias—a fundamental asymmetry where dense visual signals are under-represented by sparse textual captions. This disparity leads to the inadvertent suppression of contextually vital visual r…

Cited by 0SourceScholar
2026

SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models

ICML 2026poster

Evaluating large language models (LLMs) for software engineering has been limited by narrow task coverage, language bias, and insufficient alignment with real-world developer workflows. Existing benchmarks often focus on algorithmic problems or Python-centric bug fixing, leaving critical dimensions …

Cited by 0SourceScholar
2026

Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring

CVPR 2026

Classical video quality assessment (VQA) methods generate a numerical score to judge a video's perceived visual fidelity and clarity. Yet, a score fails to describe the video's complex quality dimensions (e.g., noise), restricting its applicability. Benefiting from the human-friendly linguistic outp

Cited by 0SourcecodeScholar
2026

ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restoration

CVPR 2026

Look-Up Table based methods have emerged as a promising direction for efficient image restoration tasks. Recent LUT-based methods focus on improving their performance by expanding the receptive field. However, they inevitably introduce extra computational and storage overhead, which hinders their de

Cited by 0SourcecodeScholar
2025

Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling

IJCAI 2025

Diffusion models have gained attention for their success in modeling complex distributions, achieving impressive perceptual quality in SR tasks. However, existing diffusion-based SR methods often suffer from high computational costs, requiring numerous iterative steps for training and inference. Exi

Cited by 0SourcePDFScholar
2025

Directional Source Separation for Robust Speech Recognition on Smart Glasses

ICASSP 2025accepted

Modern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this wor…

Cited by 15SourceScholar
2025

Divide-and-Conquer Variational Bayesian Inference for Multi-task Learning of High-resolution SAR Imagery

ICASSP 2025accepted

Conventional statistical-driven synthetic aperture radar (SAR) imaging algorithms can only encode a single and/or static prior, leading to that limited features can be accessed quantitatively. To this end, a novel multi-task learning framework is proposed by devising a divide-and-conquer variational…

Cited by 0SourceScholar
2025

Effective Integration of KAN for Keyword Spotting

ICASSP 2025accepted

Keyword spotting (KWS) is an important speech processing component for smart devices with voice assistance capability. In this paper, we investigate if Kolmogorov-Arnold Networks (KAN) can be used to enhance the performance of KWS. We explore various approaches to integrate KAN for a model architect…

Cited by 0SourceScholar
2025

EntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright Protection

CVPR 2025poster

High-quality open-source datasets are essential for advancing deep neural networks. However, the unauthorized commercial use of these datasets has raised significant concerns about copyright protection. One promising approach is backdoor watermark-based dataset ownership verification (BW-DOV), in wh…

2025

KVQ: Boosting Video Quality Assessment via Saliency-guided Local Perception

CVPR 2025poster

Video Quality Assessment (VQA), which intends to predict the perceptual quality of videos, has attracted increasing attention. Due to factors like motion blur or specific distortions, the quality of different regions in a video varies. Recognizing the region-wise local quality within a video is bene…

2025

Plug-and-Play Tri-Branch Invertible Block for Image Rescaling

AAAI 2025technical

High-resolution (HR) images are commonly downscaled to low-resolution (LR) to reduce bandwidth, followed by upscaling to restore their original details. Recent advancements in image rescaling algorithms have employed invertible neural networks (INNs) to create a unified framework for downscaling and…

2025

Preventing Latent Diffusion Model-Based Image Mimicry via Angle Shifting and Ensemble Learning

IJCAI 2025

The remarkable progress of Latent Diffusion Models (LDMs) in image generation has raised concerns about the potential for unauthorized image mimicry. To address these concerns, studies on adversarial attacks against LDMs have gained increasing attention in recent years. However, existing methods fac

2025

Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior

ICLR 2025poster

Recent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating…

2025

Translational Generative Retrieval via Potential Query Generation

ICASSP 2025accepted

Document retrieval aims to find documents related to the query from all candidate documents. Existing studies develop the Generative Retrieval approach, which assigns a unique DocID to each document, and then measures document-query relevance based on the probability of generating the expected DocID…

Cited by 0SourceScholar
2025

Visual Autoregressive Modeling for Image Super-Resolution

ICML 2025poster

Image Super-Resolution (ISR) has seen significant progress with the introduction of remarkable generative models. However, challenges such as the trade-off issues between fidelity and realism, as well as computational complexity, have also posed limitations on their application. Building upon the tr…

2025

Zero-Shot Cross-Domain Slot Filling with Retrieval Augmented In-Context Learning

ICASSP 2025accepted

Zero-shot cross-domain slot filling is becoming increasingly important due to its ability to generalize to new domains without the need for annotating domain-specific data, which aligns well with the requirements of industrial deployments. Recent advanced works deal with this task through question a…

Cited by 0SourceScholar
2024

A New Dataset and Framework for Real-World Blurred Images Super-Resolution

ECCV 2024poster

"Recent Blind Image Super-Resolution (BSR) methods have shown proficiency in general images. However, we find that the efficacy of recent methods obviously diminishes when employed on image data with blur, while image data with intentional blur constitute a substantial proportion of general data. To…

2024

AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition

ICASSP 2024accepted

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-chan…

Cited by 0SourceScholar
2024

CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality Enhancement

CVPR 2024poster

Recently numerous approaches have achieved notable success in compressed video quality enhancement (VQE). However these methods usually ignore the utilization of valuable coding priors inherently embedded in compressed videos such as motion vectors and residual frames which carry abundant temporal a…

2024

KVQ: Kwai Video Quality Assessment for Short-form Videos

CVPR 2024poster

Short-form UGC video platforms like Kwai and TikTok have been an emerging and irreplaceable mainstream media form thriving on user-friendly engagement and kaleidoscope creation etc. However the advancing content generation modes e.g. special effects and sophisticated processing workflows e.g. de-art…

2024

OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal

ECCV 2024poster

"Deep learning-based methods have shown remarkable performance in single JPEG artifacts removal task. However, existing methods tend to degrade on double JPEG images, which are prevalent in real-world scenarios. To address this issue, we propose Offset-Aware Partition Transformer for double JPEG art…

2024

PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild

CVPR 2024poster

Video quality assessment (VQA) is a challenging problem due to the numerous factors that can affect the perceptual quality of a video e.g. content attractiveness distortion type motion pattern and level. However annotating the Mean opinion score (MOS) for videos is expensive and time-consuming which…

Cited by 5SourcePDFScholar
2024

XPSR: Cross-modal Priors for Diffusion-based Image Super-Resolution

ECCV 2024poster

"Diffusion-based methods, endowed with a formidable generative prior, have received increasing attention in Image Super-Resolution (ISR) recently. However, as low-resolution (LR) images often undergo severe degradation, it is challenging for ISR models to perceive the semantic and degradation inform…

2023

Accelerating Monte Carlo Tree Search with Probability Tree State Abstraction

NeurIPS 2023poster

Monte Carlo Tree Search (MCTS) algorithms such as AlphaGo and MuZero have achieved superhuman performance in many challenging tasks. However, the computational complexity of MCTS-based algorithms is influenced by the size of the search space. To address this issue, we propose a novel probability tre…

Cited by 4SourcePDFScholar
2023

Disentangled Training with Adversarial Examples for Robust Small-Footprint Keyword Spotting

ICASSP 2023accepted

A keyword spotting (KWS) engine continuously running on the device is exposed to various speech signals that are usually unseen beforehand. It is a challenging problem to build a small-footprint and high-performing KWS model with robustness under different acoustic environments. In this paper, we ex…

Cited by 0SourceScholar
2023

Quality-Aware Pre-Trained Models for Blind Image Quality Assessment

CVPR 2023poster

Blind image quality assessment (BIQA) aims to automatically evaluate the perceived quality of a single image, whose performance has been improved by deep learning-based methods in recent years. However, the paucity of labeled data somewhat restrains deep learning-based BIQA methods from unleashing t…

Cited by 99SourcePDFScholar
2023

Reconstructed Convolution Module Based Look-Up Tables for Efficient Image Super-Resolution

ICCV 2023poster

Look-up table (LUT)-based methods have shown the great efficacy in single image super-resolution (SR) task. However, previous methods don't delve into the essential reason of restricted receptive field (RF) size in LUT, which is caused by the interaction of space and channel features in vanilla co…

Cited by 22PDFcodeScholar
2022

Federated Self-Supervised Learning for Acoustic Event Classification

ICASSP 2022accepted

Standard acoustic event classification (AEC) solutions require large-scale collection of data from client devices for model optimization. Federated learning (FL) is a compelling frame- work that decouples data collection and model training to enhance customer privacy. In this work, we investigate th…

Cited by 14SourceScholar
2022

Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology

ICASSP 2022accepted

Acoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is…

Cited by 0SourceScholar
2022

Wikitag: Wikipedia-Based Knowledge Embeddings Towards Improved Acoustic Event Classification

ICASSP 2022accepted

Acoustic event classification (AEC) is the task of determining whether certain events occur in an audio clip. Inspired by previous research [1], [2], [3] that embeddings from event labels can be leveraged to facilitate the learning of new detectors with no or limited audio samples, we introduce Wiki…

Cited by 0SourceScholar
2021

AutoSampling: Search for Effective Data Sampling Schedules

ICML 2021spotlight

Data sampling acts as a pivotal role in training deep learning models. However, an effective sampling schedule is difficult to learn due to its inherent high-dimension as a hyper-parameter. In this paper, we propose an AutoSampling method to automatically learn sampling schedules for model training,…

Cited by 8SourcePDFScholar
2021

BN-NAS: Neural Architecture Search With Batch Normalization

ICCV 2021poster

Model training and evaluation are two main time-consuming processes during neural architecture search (NAS). Although weight-sharing based methods have been proposed to reduce the number of trained networks, these methods still need to train the supernet for hundreds of epochs and evaluate thousands…

Cited by 45PDFcodeScholar
2021

Evolving Search Space for Neural Architecture Search

ICCV 2021poster

Automation of neural architecture design has been a coveted alternative to human experts. Various search methods have been proposed aiming to find the optimal architecture in the search space. One would expect the search results to improve when the search space grows larger since it would potentiall…

Cited by 55PDFcodeScholar
2021

GLiT: Neural Architecture Search for Global and Local Image Transformer

ICCV 2021poster

We introduce the first Neural Architecture Search (NAS) method to find a better transformer architecture for image recognition. Recently, transformers without CNN-based backbones are found to achieve impressive performance for image recognition. However, the transformer is designed for NLP tasks and…

Cited by 131PDFcodeScholar
2021

Inception Convolution With Efficient Dilation Search

CVPR 2021poster

As a variant of standard convolution, a dilated convolution can control effective receptive fields and handle large scale variance of objects without introducing additional computational costs. To fully explore the potential of dilated convolution, we proposed a new type of dilated convolution (refe…

Cited by 45PDFcodeScholar
2021

Multi-Task Self-Supervised Pre-Training for Music Classification

ICASSP 2021accepted

Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. B…

Cited by 0SourceScholar
2021

Unsupervised and Semi-Supervised Few-Shot Acoustic Event Classification

ICASSP 2021accepted

Few-shot Acoustic Event Classification (AEC) aims to learn a model to recognize novel acoustic events using very limited labeled data. Previous works utilize supervised pre-training as well as meta-learning approaches, which heavily rely on labeled data. Here, we study unsupervised and semi-supervis…

Cited by 0SourceScholar
2020

A Comparison of Pooling Methods on LSTM Models for Rare Acoustic Event Classification

ICASSP 2020accepted

Acoustic event classification (AEC) and acoustic event detection (AED) refer to the task of detecting whether specific target events occur in audios. As long short-term memory (LSTM) leads to state-of-the-art results in various speech related tasks, it is employed as a popular solution for AEC as we…

Cited by 0SourceScholar
2020

Efficient Transfer Learning via Joint Adaptation of Network Architecture and Weight

ECCV 2020poster

Transfer learning can boost the performance on the target task by leveraging the knowledge of the source domain. Recent works in neural architecture search (NAS), especially one-shot NAS, can aid transfer learning by establishing sufficient network search space. However, existing NAS methods tend to…

Cited by 7SourcePDFScholar
2020

Few-Shot Acoustic Event Detection Via Meta Learning

ICASSP 2020accepted

We study few-shot acoustic event detection (AED) in this paper. Few-shot learning enables detection of new events with very limited labeled data. Compared to other research areas like computer vision, few-shot learning for audio recognition has been under-studied. We formulate few-shot AED problem a…

Cited by 0SourceScholar
2020

Improving Auto-Augment via Augmentation-Wise Weight Sharing

NeurIPS 2020poster

The recent progress on automatically searching augmentation policies has boosted the performance substantially for various tasks. A key component of automatic augmentation search is the evaluation process for a particular augmentation policy, which is utilized to return reward and usually runs thous…

2020

Large-Scale Object Detection in the Wild From Imbalanced Multi-Labels

CVPR 2020oral

Training with more data has always been the most stable and effective way of improving performance in deep learn-ing era. As the largest object detection dataset so far, OpenImages brings great opportunities and challenges for object detection in general and sophisticated scenarios. However, owing t…

Cited by 75PDFScholar
2020

Powering One-shot Topological NAS with Stabilized Share-parameter Proxy

ECCV 2020poster

One-shot NAS method has attracted much interest from the research community due to its remarkable training efficiency and capacity to discover high performance models. However, the search spaces of previous one-shot based works usually relied on hand-craft design and were short for flexibility on th…

Cited by 21SourcePDFScholar
2020

Raw Waveform Based End-to-end Deep Convolutional Network for Spatial Localization of Multiple Acoustic Sources

ICASSP 2020accepted

In this paper, we present an end-to-end deep convolutional neural network operating on multi-channel raw audio data to localize multiple simultaneously active acoustic sources in space. Previously reported deep learning based approaches work well in localizing a single source directly from multi-cha…

Cited by 0SourceScholar
2019

Efficient Neural Architecture Transformation Search in Channel-Level for Object Detection

NeurIPS 2019poster

Recently, Neural Architecture Search has achieved great success in large-scale image classification. In contrast, there have been limited works focusing on architecture search for object detection, mainly because the costly ImageNet pretraining is always required for detectors. Training from scratch…

Cited by 67SourcePDFScholar
2019

Hierarchical Residual-pyramidal Model for Large Context Based Media Presence Detection

ICASSP 2019accepted

We study media presence detection, that is, learning to recognize if a sound segment (typically lasting for a few seconds) of a long recorded stream contains media (TV) sound. This problem is difficult because non-media sound sources can be quite diverse (e.g. human voicing, non-vocal sounds and non…

Cited by 0SourceScholar
2019

Improving Emotion Classification through Variational Inference of Latent Variables

ICASSP 2019accepted

Conventional models for emotion recognition from speech signal are trained in supervised fashion using speech utterances with emotion labels. In this study we hypothesize that speech signal depends on multiple latent variables including the emotional state, age, gender, and speech content. We propos…

Cited by 0SourceScholar
2019

POD: Practical Object Detection With Scale-Sensitive Network

ICCV 2019poster

Scale-sensitive object detection remains a challenging task, where most of the existing methods not learn it explicitly and not robust to scale variance. In addition, the most existing methods are less efficient during training or slow during inference, which are not friendly to real-time applicatio…

Cited by 27PDFScholar
2019

Semi-supervised Acoustic Event Detection Based on Tri-training

ICASSP 2019accepted

This paper presents our work of training acoustic event detection (AED) models using unlabeled dataset. Recent acoustic event detectors are based on large-scale neural networks, which are typically trained with huge amounts of labeled data. Labels for acoustic events are expensive to obtain, and rel…

Cited by 0SourceScholar
2018

Compact Generalized Non-local Network

NeurIPS 2018poster

The non-local module is designed for capturing long-range spatio-temporal dependencies in images and videos. Although having shown excellent performance, it lacks the mechanism to model the interactions between positions across channels, which are of vital importance in recognizing fine-grained obje…

2018

Monophone-Based Background Modeling for Two-Stage On-Device Wake Word Detection

ICASSP 2018accepted

Accurate on-device wake word detection is crucial to products with far-field voice control such as the Amazon Echo. It is quite challenging to build a wake word system with both low False Reject Rate (FRR) and low False Alarm Rate (FAR) in real scenarios where there are various types of background s…

Cited by 0SourceScholar
2018

Multi-Attention Multi-Class Constraint for Fine-grained Image Recognition

ECCV 2018poster

Attention-based learning for fine-grained image recognition remains a challenging task, where most of the existing methods treat each object part in isolation, while neglecting the correlations among them. In addition, the multi-stage or multi-scale mechanisms involved make the existing methods less…

Cited by 501SourcePDFScholar
2018

Time-Delayed Bottleneck Highway Networks Using a DFT Feature for Keyword Spotting

ICASSP 2018accepted

This paper presents a novel deep neural network (DNN) architecture with highway blocks (HWs) using a complex discrete Fourier transform (DFT) feature for keyword spotting. In our previous work, we showed that the feed-forward DNN with a time-delayed bottleneck layer (TDB-DNN) directly trained from t…

Cited by 0SourceScholar
2017

An empirical evaluation of zero resource acoustic unit discovery

ICASSP 2017accepted

Acoustic unit discovery (AUD) is a process of automatically identifying a categorical acoustic unit inventory from speech and producing corresponding acoustic unit tokenizations. AUD provides an important avenue for unsupervised acoustic model training in a zero resource setting where expert-provide…

Cited by 0SourceScholar
2016

Unsupervised user intent modeling by feature-enriched matrix factorization

ICASSP 2016accepted

Spoken language interfaces are being incorporated into various devices such as smart phones and TVs. However, dialogue systems may fail to respond correctly when users' request functionality is not supported by currently installed apps. This paper proposes a feature-enriched matrix factorization (MF…

Cited by 0SourceScholar