← Search

Ming Wu

18 accepted papers

2026

Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents

ICML 2026poster

Consensus protocols form the backbone of distributed systems and blockchains, where implementation bugs can cause data corruption and financial losses. While LLM-based approaches show promise in code analysis, they struggle with deep protocol-level logic bugs involving complex state-dependent behavi…

Cited by 0SourceScholar
2026

BabyVision: Visual Reasoning Beyond Language

ICML 2026poster

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that …

Cited by 0SourceScholar
2025

MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecasting

CVPR 2025poster

Deep learning approaches for marine fog detection and forecasting have outperformed traditional methods, demonstrating significant scientific and practical importance. However, the limited availability of open-source datasets remains a major challenge. Existing datasets, often focused on a single re…

2025

Mind the Cost of Scaffold! Benign Clients May Even Become Accomplices of Backdoor Attack

ICCV 2025poster

By using a control variate to calibrate the local gradient of each client, Scaffold has been widely known as a powerful solution to mitigate the impact of data heterogeneity in Federated Learning. Although Scaffold achieves significant performance improvements, we show that this superiority is at th…

Cited by 0SourcePDFScholar
2025

Spatially Constrained and Deeply Learned Bilateral Structural Intensity-Depth Registration Autonomously Navigates a Flexible Endoscope

ICRA 2025

Endoscope tracking is commonly utilized to provide surgeons with in-body camera poses and visual fields during invasive procedures. The fundamental aspect of endoscopic navigation lies in precisely and continuously tracing the position and orientation of the endoscope within monocular endoscopic vid

Cited by 0SourceScholar
2024

Chat: Cascade Hole-Aware Transformers with Geometric Spatial Consistency for Accurate Monocular Endoscopic Depth Estimation

ICASSP 2024accepted

Monocular endoscopic depth estimation is essential for surgical navigation. Current deeply learned estimation methods still suffer from lack of real data labels and porous, artifacts (e.g., bubbles), illumination variations (e.g., specular highlight), and weak texture in endoscopic video images. Thi…

Cited by 0SourceScholar
2024

Deep Residual W-Unit Learning with Semantic Embedding for Automatic Pulmonary CT Artery-Vein Separation

ICASSP 2024accepted

Automatic segmentation of pulmonary arteries and veins in CT has great clinical significance. Because the growth range of a single vessel is vast, and the arteries and veins have barely identical intensity values on CT and grow very close to or even interleaved, accurate segmentation of them require…

Cited by 0SourceScholar
2024

Hierarchical Trajectory Deformation Algorithm With Hybrid Controller for Active Lower Limb Rehabilitation

RA-L 2024

Robot-aided active rehabilitation has shown to be an effective treatment approach for hemiplegic patients. This paper presents an active control framework for lower limb rehabilitation, combining an interaction layer with a hierarchical trajectory deformation algorithm (HTDA), and an assist-as-neede

Cited by 5SourceScholar
2024

Multi-Scale Representations by Varying Window Attention for Semantic Segmentation

ICLR 2024poster

Multi-scale learning is central to semantic segmentation. We visualize the effective receptive field (ERF) of canonical multi-scale representations and point out two risks learning them: \textit{scale inadequacy} and \textit{field inactivation}. A novel multi-scale learner, \textbf{varying window at…

2024

Privileged Prior Information Distillation for Image Matting

AAAI 2024technical

Performance of trimap-free image matting methods is limited when trying to decouple the deterministic and undetermined regions, especially in the scenes where foregrounds are semantically ambiguous, chromaless, or high transmittance. In this paper, we propose a novel framework named Privileged Prior…

Cited by 1SourcePDFScholar
2023

Doubly-Robust Self-Training

NeurIPS 2023poster

Self-training is a well-established technique in semi-supervised learning, which leverages unlabeled data by generating pseudo-labels and incorporating them with a limited labeled dataset for training. The effectiveness of self-training heavily relies on the accuracy of these pseudo-labels. In this…

2023

Pre-trained Language Models Can be Fully Zero-Shot Learners

ACL 2023long

How can we extend a pre-trained model to many language understanding tasks, without labeled or additional unlabeled data? Pre-trained language models (PLMs) have been effective for a wide range of NLP tasks. However, existing approaches either require fine-tuning on downstream labeled datasets or ma…

2023

SwiftAvatar: Efficient Auto-Creation of Parameterized Stylized Character on Arbitrary Avatar Engines

AAAI 2023technical

The creation of a parameterized stylized character involves careful selection of numerous parameters, also known as the "avatar vectors" that can be interpreted by the avatar engine. Existing unsupervised avatar vector estimation methods that auto-create avatars for users, however, often fail to wor…

2023

Weather2K: A Multivariate Spatio-Temporal Benchmark Dataset for Meteorological Forecasting Based on Real-Time Observation Data from Ground Weather Stations

AISTATS 2023poster

Weather forecasting is one of the cornerstones of meteorological work. In this paper, we present a new benchmark dataset named Weather2K, which aims to make up for the deficiencies of existing weather forecasting datasets in terms of real-time, reliability, and diversity, as well as the key bottlene…

2022

A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection

ICASSP 2022accepted

Polyphonic sound event localization and detection (SELD) aims at detecting types of sound events with corresponding temporal activities and spatial locations. In this paper, a trackwise ensemble event independent network with a novel data augmentation method is proposed. The proposed model is based…

Cited by 0SourceScholar
2022

Compressing Sentence Representation for Semantic Retrieval via Homomorphic Projective Distillation

ACL 2022findings

How to learn highly compact yet effective sentence representation? Pre-trained language models have been effective in many NLP tasks. However, these models are often huge and produce large sentence embeddings. Moreover, there is a big performance gap between large and small models. In this paper, we…

2021

A Diffusion FXLMS Algorithm for Multi-Channel Active Noise Control and Variable Spatial Smoothing

ICASSP 2021accepted

This paper studies the diffusion (Diff) control for multichannel ANC systems, where a group of controllers and error microphones are physically distributed at different locations within a large area. In this case, the conventional consensus agreement for controllers cannot be reached. To solve this…

Cited by 0SourceScholar
2021

Overfitting the Data: Compact Neural Video Delivery via Content-Aware Feature Modulation

ICCV 2021poster

Internet video delivery has undergone a tremendous explosion of growth over the past few years. However, the quality of video delivery system greatly depends on the Internet bandwidth. Deep Neural Networks (DNNs) are utilized to improve the quality of video delivery recently. These methods divide a…

Cited by 38PDFcodeScholar