← Search

YingYing Wang

24 accepted papers

2026

Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

CVPR 2026

Pansharpening aims to generate high-resolution multi-spectral images by fusing the spatial detail of panchromatic images with the spectral richness of low-resolution MS data. However, most existing methods are evaluated under limited, low-resolution settings, limiting their generalization to real-wo

Cited by 0SourcecodeScholar
2026

MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement

AAAI 2026technical

Pan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the

Cited by 0SourcePDFScholar
2026

MobileFusion: Mobile-Friendly Infrared and Visible Image Fusion via Structural Re-parameterization

ICML 2026poster

Deep neural networks have recently advanced infrared and visible image fusion (IVIF), but most existing methods rely on sophisticated yet redundant designs, which hinder real-time deployment on mobile devices with limited compute and memory. In this paper, we present MobileFusion, an extremely light…

Cited by 0SourceScholar
2026

MovieGraph-ToM: Evaluating Long-Range Theory of Mind in Large Language Models via Implicit Social-Causal Graphs

AAAI 2026technical

The capacity for social reasoning, particularly Theory of Mind (ToM), is a foundational prerequisite for aligning Large Language Models (LLMs) with human values. However, current evaluations are predominantly confined to simplistic, short-text scenarios, obscuring their true capabilities and potenti

Cited by 0SourcePDFScholar
2026

RefChess: Monte-Carlo Move Selection for Zero-Shot Referring Image Segmentation

ICML 2026poster

Recent advances in zero-shot referring image segmentation (RIS), driven by foundation models such as SAM and CLIP, have improved cross-modal alignment between visual regions and natural language expressions. Nevertheless, selecting the correct segmentation proposal remains challenging, as existing m…

Cited by 0SourceScholar
2026

Self-supervised Multiplex Consensus Mamba for General Image Fusion

AAAI 2026technical

Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, gen

Cited by 0SourcePDFScholar
2026

Shaping Human–AI Collaboration in Education: Effects of AI-Assisted Decision-Making Paradigms and Human–AI Decision Consistency on Pre-Service Teachers’ Psychological States and Performance

AAAI 2026technical

Artificial intelligence is playing an increasingly important role in supporting decision-making, particularly in educational contexts, where it serves as a critical tool to assist teacher judgment and optimize instructional decisions. However, limited research has examined how different AI-assisted

Cited by 0SourcePDFScholar
2025

AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image Enhancement

AAAI 2025technical

Existing low-light image enhancement (LIE) methods have achieved noteworthy success in solving synthetic distortions, yet they often fall short in practical applications. The limitations arise from two inherent challenges in real-world LIE: 1) the collection of distorted/clean image pairs is often i…

Cited by 7SourcePDFScholar
2025

Accelerated Diffusion via High-Low Frequency Decomposition for Pan-Sharpening

AAAI 2025technical

Pan-sharpening aims to preserve the spectral information of the multi-spectral (MS) image while leveraging the high-frequency details from the guided high-resolution panchromatic (PAN) image to enhance its spatial resolution. The key challenge is how to preserve the spectral information from the MS…

Cited by 0SourcePDFScholar
2025

DPLUT: Unsupervised Low-light Image Enhancement with Lookup Tables and Diffusion Priors

AAAI 2025technical

Low-light image enhancement (LIE) aims at precisely and efficiently recovering an image degraded in poor illumination environments. Recent advanced LIE techniques are using deep neural networks, which require lots of low-normal light image pairs, network parameters, and computational resources. As a…

Cited by 6SourcePDFScholar
2025

FRN: Fractal-Based Recursive Spectral Reconstruction Network

NeurIPS 2025poster

Generating hyperspectral images (HSIs) from RGB images through spectral reconstruction can significantly reduce the cost of HSI acquisition. In this paper, we propose a Fractal-Based Recursive Spectral Reconstruction Network (FRN), which differs from existing paradigms that attempt to directly integ…

Cited by 0SourcecodeScholar
2025

Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables

NeurIPS 2025oral

Recently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicabil…

Cited by 0SourceScholar
2025

Physics-Informed LSTM for Shape and Contact Force Prediction of a Flexible Surgical Robot*

IROS 2025

Real-time morphological perception and precise end force feedback prediction of surgical robots constitute critical technical elements for ensuring safety and efficacy in complex interventional procedures such as Endoscopic Retrograde Cholangiopancreatography (ERCP). In this paper, we design a minia

Cited by 0SourceScholar
2025

Sp3ctralMamba: Physics-Driven Joint State Space Model for Hyperspectral Image Reconstruction

AAAI 2025technical

Hyperspectral image (HSI) reconstruction aims to restore the original 3D HSIs from the 2D hyperspectral snapshot compressive images (SCIs). The key to high-fidelity HSI reconstruction lies in designing refined spatial and spectral attention mechanisms, which are crucial for generating fine-grained r…

Cited by 0SourcePDFScholar
2024

Progressive High-Frequency Reconstruction for Pan-Sharpening with Implicit Neural Representation

AAAI 2024technical

Pan-sharpening aims to leverage the high-frequency signal of the panchromatic (PAN) image to enhance the resolution of its corresponding multi-spectral (MS) image. However, deep neural networks (DNNs) tend to prioritize learning the low-frequency components during the training process, which limits…

Cited by 11SourcePDFScholar
2023

Fine-Grained Private Knowledge Distillation

ICASSP 2023accepted

Knowledge distillation has emerged as a scalable and effective way for privacy-preserving machine learning. One remaining drawback is that it consumes privacy in a client-level manner. In order to attain fine-grained privacy accountant and improve utility, this work proposes a model-free reverse k-N…

Cited by 0SourceScholar
2023

Towards an Accurate Augmented-Reality-Assisted Orthopedic Surgical Robotic System Using Bidirectional Generalized Point Set Registration

IROS 2023poster

This paper presents a novel augmented reality (AR)-assisted orthopedic surgical robotic system based on Head-Mounted Display (HMD) devices. The proposed system can overlay the preoperative plans over the patient's anatomy and provide useful guidance for surgeons during interventions, with integrated…

Cited by 3SourceScholar
2022

A2DIO: Attention-Driven Deep Inertial Odometry for Pedestrian Localization based on 6D IMU

ICRA 2022poster

In this work, we propose A2DIO, a novel hybrid neural network model with a set of carefully designed attention mechanisms for pose invariant inertial odometry. The key idea is to extract both local and global features from the window of IMU measurements for velocity prediction. A2DIO leverages the c…

Cited by 20SourceScholar
2020

Boundary Content Graph Neural Network for Temporal Action Proposal Generation

ECCV 2020poster

Temporal action proposal generation plays an important role in video action understanding, which requires localizing high-quality action content precisely. However, generating temporal proposals with both precise boundaries and high-quality action content is extremely challenging. To address this is…

Cited by 213SourcePDFScholar
2019

3D Hand Shape and Pose Estimation From a Single RGB Image

CVPR 2019oral

This work addresses a novel and challenging problem of estimating the full 3D hand shape and pose from a single RGB image. Most current methods in 3D hand analysis from monocular RGB images only focus on estimating the 3D locations of hand keypoints, which cannot fully express the 3D shape of hand.…

Cited by 565PDFScholar