← Search

Xiaotao Wang

10 accepted papers

2026

API: Adaptive Prototype Imputation for Incomplete Multimodal Sentiment Analysis

ICML 2026poster

Multimodal sentiment analysis aims to infer human emotions by integrating signals from diverse modalities. However, missing modalities are common in real-world applications due to sensor failure, data corruption, or privacy concerns. Existing approaches typically follow two main paradigms: recovery-…

Cited by 0SourceScholar
2026

Benchmarking the Scientific Mind: Toward Evaluation of Complex-Reasoning Biomedical VQA

ICML 2026poster

Despite progress of Multimodal Large Language Models (MLLMs) in biomedical visual question answering (VQA), existing benchmarks provide limited assessment of their scientific reasoning capabilities. Most datasets adopt single-image question construction and outcome-oriented evaluation, where correct…

Cited by 0SourceScholar
2024

Learning Real-World Image De-weathering with Imperfect Supervision

AAAI 2024technical

Real-world image de-weathering aims at removing various undesirable weather-related artifacts. Owing to the impossibility of capturing image pairs concurrently, existing real-world de-weathering datasets often exhibit inconsistent illumination, position, and textures between the ground-truth images…

2024

Self-Supervised High Dynamic Range Imaging with Multi-Exposure Images in Dynamic Scenes

ICLR 2024poster

Merging multi-exposure images is a common approach for obtaining high dynamic range (HDR) images, with the primary challenge being the avoidance of ghosting artifacts in dynamic scenes. Recent methods have proposed using deep neural networks for deghosting. However, the methods typically rely on suf…

2023

Beyond Image Borders: Learning Feature Extrapolation for Unbounded Image Composition

ICCV 2023poster

For improving image composition and aesthetic quality, most existing methods modulate the captured images by striking out redundant content near the image borders. However, such image cropping methods are limited in the range of image views. Some methods have been suggested to extrapolate the images…

Cited by 2PDFcodeScholar
2023

CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition

ICCV 2023poster

The co-occurrence signals (e.g., hand shape, facial expression, and lip pattern) play a critical role in Continuous Sign Language Recognition (CSLR). Compared to RGB data, skeleton data provide a more efficient and concise option, and lay a good foundation for the co-occurrence exploration in CSLR.…

Cited by 27PDFScholar
2023

Physics-Guided ISO-Dependent Sensor Noise Modeling for Extreme Low-Light Photography

CVPR 2023poster

Although deep neural networks have achieved astonishing performance in many vision tasks, existing learning-based methods are far inferior to the physical model-based solutions in extreme low-light sensor noise modeling. To tap the potential of learning-based sensor noise modeling, we investigate th…

2023

Self-supervised Learning to Bring Dual Reversed Rolling Shutter Images Alive

ICCV 2023poster

Modern consumer cameras usually employ the rolling shutter (RS) mechanism, where images are captured by scanning scenes row-by-row, yielding RS distortions for dynamic scenes. To correct RS distortions, existing methods adopt a fully supervised learning manner, where high framerate global shutter (G…

Cited by 5PDFScholar
2023

Spatially Adaptive Self-Supervised Learning for Real-World Image Denoising

CVPR 2023poster

Significant progress has been made in self-supervised image denoising (SSID) in the recent few years. However, most methods focus on dealing with spatially independent noise, and they have little practicality on real-world sRGB images with spatially correlated noise. Although pixel-shuffle downsampl…

2022

Deep Radial Embedding for Visual Sequence Learning

ECCV 2022poster

"Connectionist Temporal Classification (CTC) is a popular objective function in sequence recognition, which provides supervision for unsegmented sequence data through aligning sequence and its corresponding labeling iteratively. The blank class of CTC plays a crucial role in the alignment process an…

Cited by 20SourcePDFScholar