← Search

Ming-Ching Chang

21 accepted papers

2026

FAVE: A Structured Benchmark for Fine-Grained Audio-Visual Temporal Evaluation in Multimodal LLMs

CVPR 2026

Audio-visual large language models (AVLLMs) have made significant strides in understanding visual and auditory content. However, their ability to capture fine-grained temporal relationships between audio and visual streams remains insufficiently evaluated. To address this, we introduce FAVE (Fine-gr

Cited by 0SourceScholar
2026

FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPU

ICML 2026oral

Entropic optimal transport (EOT) via Sinkhorn iterations is widely used in modern machine learning, yet GPU solvers remain inefficient at scale. Tensorized implementations suffer quadratic HBM traffic from dense $n\times m$ interactions, while existing online backends avoid storing dense matrices bu…

Cited by 0SourceScholar
2026

How Bias Binds: Measuring Hidden Associations for Bias Control in Text-to-Image Compositions

AAAI 2026technical

Text-to-image generative models often exhibit bias related to sensitive attributes. However, current research tends to focus narrowly on single-object prompts with limited contextual diversity. In reality, each object or attribute within a prompt can contribute to bias. For example, the prompt ``an

Cited by 0SourcePDFScholar
2026

Partial Ring Scan: Revisiting Scan Order in Vision State Space Models

ICML 2026poster

State Space Models (SSMs) have emerged as efficient alternatives to attention for vision tasks, offering linear-time sequence processing with competitive accuracy. Vision SSMs, however, require serializing 2D images into 1D token sequences along a predefined scan order, a factor often overlooked. We…

Cited by 0SourceScholar
2025

E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFS

AAAI 2025technical

Deep neural network (DNN) models are increasingly popular in edge video analytic applications. However, the computeintensive nature of DNN models pose challenges for energyefficient inference on resource-constrained edge devices. Most existing solutions focus on optimizing DNN inference latency and…

Cited by 0SourcePDFScholar
2025

RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style Generation

IJCAI 2025

Arbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based fr

2024

A New Benchmark and Model for Challenging Image Manipulation Detection

AAAI 2024technical

The ability to detect manipulation in multimedia data is vital in digital forensics. Existing Image Manipulation Detection (IMD) methods are mainly based on detecting anomalous features arisen from image editing or double compression artifacts. All existing IMD techniques encounter challenges when i…

2024

Image Manipulation Detection With Implicit Neural Representation and Limited Supervision

ECCV 2024poster

"Image Manipulation Detection (IMD) is becoming increasingly important as tampering technologies advance. However, most state-of-the-art (SoTA) methods require high-quality training datasets featuring image- and pixel-level annotations. The effectiveness of these methods suffers when applied to mani…

Cited by 2SourcePDFScholar
2024

Improving Limited Supervised Foot Ulcer Segmentation Using Cross-Domain Augmentation Strategies

ICASSP 2024accepted

Diabetic foot ulcers pose health risks, including higher morbidity, mortality, and amputation rates. Monitoring wound areas is crucial for proper care, but manual segmentation is subjective due to complex wound features and background variation. Expert annotations are costly and time-intensive, thus…

Cited by 0SourceScholar
2024

MOTE-NAS: Multi-Objective Training-based Estimate for Efficient Neural Architecture Search

NeurIPS 2024poster

Neural Architecture Search (NAS) methods seek effective optimization toward performance metrics regarding model accuracy and generalization while facing challenges regarding search costs and GPU resources. Recent Neural Tangent Kernel (NTK) NAS methods achieve remarkable search efficiency based on a…

Cited by 0SourcePDFScholar
2024

Pushing the Limit of Fine-Tuning for Few-Shot Learning: Where Feature Reusing Meets Cross-Scale Attention

AAAI 2024technical

Due to the scarcity of training samples, Few-Shot Learning (FSL) poses a significant challenge to capture discriminative object features effectively. The combination of transfer learning and meta-learning has recently been explored by pre-training the backbone features using labeled base data and su…

Cited by 4SourcePDFScholar
2024

SMILEtrack: SiMIlarity LEarning for Occlusion-Aware Multiple Object Tracking

AAAI 2024technical

Despite recent progress in Multiple Object Tracking (MOT), several obstacles such as occlusions, similar objects, and complex scenes remain an open challenge. Meanwhile, a systematic study of the cost-performance tradeoff for the popular tracking-by-detection paradigm is still lacking. This paper in…

2022

Eyes Tell All: Irregular Pupil Shapes Reveal GAN-Generated Faces

ICASSP 2022accepted

Generative adversarial network (GAN) generated high-realistic human faces are visually challenging to discern from real ones. They have been used as profile images for fake social media accounts, which leads to high negative social impacts. In this work, we show that GAN-generated faces can be expos…

Cited by 0SourceScholar
2022

Improving Class Activation Map for Weakly Supervised Object Localization

ICASSP 2022accepted

We propose a Weakly Supervised Object Localization (WSOL) method that can locate an object within a given image using a pre-trained network learned with only class labels without location annotations. Most existing WSOL methods rely on thresholding a Class Activation Map (CAM) generated by the pre-t…

Cited by 0SourceScholar
2022

Text-Image De-Contextualization Detection Using Vision-Language Models

ICASSP 2022accepted

Text-image de-contextualization, which uses inconsistent image-text pairs, is an emerging form of misinformation and drawing increasing attention due to the great threat to information authenticity. With real content but semantic mismatch in multiple modalities, the detection of de-contextualization…

Cited by 0SourceScholar
2022

Vocbench: A Neural Vocoder Benchmark for Speech Synthesis

ICASSP 2022accepted

Neural vocoders, used for converting the spectral representations of an audio signal to the waveforms, are a commonly used component in speech synthesis pipelines. It focuses on synthesizing waveforms from low-dimensional representation, such as Mel-Spectrograms. In recent years, different approache…

Cited by 0SourceScholar
2020

Explainable and Efficient Sequential Correlation Network for 3D Single Person Concurrent Activity Detection

IROS 2020poster

We present the sequential correlation network (SCN) to improve concurrent activity detection. SCN combines a recurrent neural network and a correlation model hierarchically to model the complex correlations and temporal dynamics of concurrent activities. SCN has several advantages that enable effect…

Cited by 2SourceScholar
2018

Multi-Scale Structure-Aware Network for Human Pose Estimation

ECCV 2018poster

We develop a robust multi-scale structure-aware neural network for human pose estimation. This method improves the recent deep conv-deconv hourglass models with four key improvements: (1) multi-scale supervision to strengthen contextual feature learning in matching body keypoints by combining featur…

Cited by 382SourcePDFScholar
2017

Adaptive RNN Tree for Large-Scale Human Action Recognition

ICCV 2017poster

In this work, we present the RNN Tree (RNN-T), an adaptive learning framework for skeleton based human action recognition. Our method categorizes action classes and uses multiple Recurrent Neural Networks (RNNs) in a tree-like hierarchy. The RNNs in RNN-T are co-trained with the action category hier…

Cited by 140PDFScholar