← Search

Yun Wang

40 accepted papers

2026

AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition

AAAI 2026technical

Automated scoring plays a crucial role in education by reducing the reliance on human raters and offering scalable and immediate evaluation of student work. While large language models (LLMs) have shown strong potential in this task, their use as end-to-end raters faces challenges such as low accura

Cited by 0SourcePDFScholar
2026

FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration

CVPR 2026

All-in-One Image Restoration (AIO-IR) aims to develop a unified model that can handle multiple degradations under complex conditions. However, existing methods often rely on task-specific designs or latent routing strategies, making it hard to adapt to real-world scenarios with various degradations.

Cited by 0SourcecodeScholar
2026

Investigating Advanced Reasoning of Large Language Models via Black-Box Interaction

ICML 2026poster

Existing tasks fall short in evaluating reasoning ability of Large Language Models (LLMs) in an interactive, unknown environment. This deficiency leads to the isolated assessment of deductive, inductive, and abductive reasoning, neglecting the integrated reasoning process that is indispensable for h…

Cited by 0SourceScholar
2026

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

ICML 2026poster

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 7K+ timestamped questions for diagnosing User-centric Continual Spatial intelligence in egocentric video streams. UCS-Bench targets a new problem that emphasizes dynamic spatial reasoning, long-term memory, …

Cited by 0SourceScholar
2026

M2FMoE: Multi-Resolution Multi-View Frequency Mixture-of-Experts for Extreme-Adaptive Time Series Forecasting

AAAI 2026technical

Forecasting time series with extreme events is critical yet challenging due to their high variance, irregular dynamics, and sparse but high-impact nature. While existing methods excel in modeling dominant regular patterns, their performance degrades significantly during extreme events, constituting

Cited by 0SourcePDFScholar
2026

RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation

AAAI 2026technical

Personalized image generation is crucial for improving the user experience, as it renders reference images into preferred ones according to user visual preferences. Although effective, existing methods face two main issues. First, existing methods treat all items in the user

Cited by 0SourcePDFScholar
2026

Scaling Speech Tokenizers with Diffusion Autoencoders

ICLR 2026poster

Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understanding and acoustics for reconstruction, and (2) achieving low bit rates and low token rates. We propose Speech Diffusion To…

Cited by 0SourceScholar
2025

Diff-LMM: Diffusion Teacher-Guided Spatio-Temporal Perception for Video Large Multimodal Models

IJCAI 2025

Dynamic spatio-temporal understanding is essential for video-based multimodal tasks, yet existing methods often struggle to capture fine-grained temporal and spatial relationships in long videos. Current approaches primarily rely on pre-trained CLIP encoders, which excel in semantic understanding bu

Cited by 0SourcePDFScholar
2025

DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label Supervision

AAAI 2025technical

Self-supervised stereo matching has drawn attention due to its ability to estimate disparity without needing ground-truth data. However, existing self-supervised stereo matching methods heavily rely on the photo-metric consistency assumption, which is vulnerable to natural disturbances, resulting in…

Cited by 0SourcePDFScholar
2025

Exploring Spectral Signatures of Chinese liquor using Machine Learning and SHapley Additive exPlanations

ICASSP 2025accepted

Chinese liquor holds great cultural and economic significance globally. The accurate classification of aroma types and alcohol content is crucial for quality control in Chinese liquor production. To address limitations such as subjectivity and sensor drift in current methods, this study introduces a…

Cited by 0SourceScholar
2025

FlowMamba: Learning Point Cloud Scene Flow with Global Motion Propagation

AAAI 2025technical

Scene flow methods based on deep learning have achieved impressive performance. However, current top-performing methods still struggle with ill-posed regions, such as extensive flat regions or occlusions, due to insufficient local evidence. In this paper, we propose a novel global-aware scene flow e…

Cited by 2SourcePDFScholar
2025

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

ICCV 2025poster

Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subject consistency due to the lack of fine-grained guidance and inter-frame interact…

Cited by 0SourcePDFScholar
2025

Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts

ICCV 2025poster

Recently, learning-based stereo matching networks have advanced significantly.However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets.Leveraging Vision Foundation Models (VFMs)…

2025

PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching

NeurIPS 2025poster

Temporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersion of users. Despite its importance, this task remains challenging due to the difficulty in modeling long-term temporal…

Cited by 0SourcecodeScholar
2024

Enhancing Cross-Subject fMRI-to-Video Decoding with Global-Local Functional Alignment

ECCV 2024poster

"Advancements in brain imaging enable the decoding of thoughts and intentions from neural activities. However, the fMRI-to-video decoding of brain signals across multiple subjects encounters challenges arising from structural and coding disparities among individual brains, further compounded by the…

2024

MAPSeg: Unified Unsupervised Domain Adaptation for Heterogeneous Medical Image Segmentation Based on 3D Masked Autoencoding and Pseudo-Labeling

CVPR 2024poster

Robust segmentation is critical for deriving quantitative measures from large-scale multi-center and longitudinal medical scans. Manually annotating medical scans however is expensive and labor-intensive and may not always be available in every domain. Unsupervised domain adaptation (UDA) is a well-…

2024

MinD-3D: Reconstruct High-quality 3D objects in Human Brain

ECCV 2024poster

"In this paper, we introduce Recon3DMind, an innovative task aimed at reconstructing 3D visuals from Functional Magnetic Resonance Imaging (fMRI) signals, marking a significant advancement in the fields of cognitive neuroscience and computer vision. To support this pioneering task, we present the fM…

2024

Multi-Task Learning for Front-End Text Processing in TTS

ICASSP 2024accepted

We propose a multi-task learning (MTL) model for jointly performing three tasks that are commonly solved in a text-to-speech (TTS) front-end: text normalization (TN), part-of-speech (POS) tagging, and homograph disambiguation (HD). Our framework utilizes a tree-like structure with a trunk that learn…

Cited by 0SourceScholar
2024

Neural Search Space in Gboard Decoder

EMNLP 2024industry

Gboard Decoder produces suggestions by looking for paths that best match input touch points on the context aware search space, which is backed by the language Finite State Transducers (FST). The language FST is currently an N-gram language model (LM). However, N-gram LMs, limited in context length,…

Cited by 0SourcePDFScholar
2024

NeuroPictor: Refining fMRI-to-Image Reconstruction via Multi-individual Pretraining and Multi-level Modulation

ECCV 2024poster

"Recent fMRI-to-image approaches mainly focused on associating fMRI signals with specific conditions of pre-trained diffusion models. These approaches, while producing high-quality images, capture only a limited aspect of the complex information in fMRI signals and offer little detailed control over…

2024

Research on bionic foldable wing for flapping wing micro air vehicle

ICRA 2024poster

This paper presents a bionic foldable wing that imitates the hind wing of ladybirds. Based on the folding mechanism of the hind wing of ladybirds and the theory of origami, the motion model of the bionic foldable wing is established, yield the motion law of the crease angles and the variation relati…

Cited by 0SourceScholar
2023

Exploring the Mutual Influence Between Self-Supervised Single-Frame and Multi-Frame Depth Estimation

RA-L 2023

Although both self-supervised single-frame and multi-frame depth estimation methods only require unlabeled monocular videos for training, the information they leverage varies because single-frame methods mainly rely on appearance-based features while multi-frame methods focus on geometric cues. Cons

Cited by 8SourcecodeScholar
2023

IHNet: Iterative Hierarchical Network Guided by High-Resolution Estimated Information for Scene Flow Estimation

ICCV 2023poster

Scene flow estimation, which predicts the 3D displacements of point clouds, is a fundamental task in autonomous driving. Most methods have adopted a coarse-to-fine structure to balance computational efficiency with accuracy, particularly when handling large displacements. However, inaccuracies in th…

Cited by 7PDFcodeScholar
2022

Conformer-Based Self-Supervised Learning For Non-Speech Audio Tasks

ICASSP 2022accepted

Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few works have comprehensively analyzed audio representation learning for non-speec…

Cited by 0SourceScholar
2022

Visual Attention-Based Self-Supervised Absolute Depth Estimation Using Geometric Priors in Autonomous Driving

RA-L 2022

Although existing monocular depth estimation methods have made great progress, predicting an accurate absolute depth map from a single image is still challenging due to the limited modeling capacity of networks and the scale ambiguity issue. In this paper, we introduce a fully Visual Attention-based

Cited by 27SourceScholar
2021

A Time-Domain Convolutional Recurrent Network for Packet Loss Concealment

ICASSP 2021accepted

Packet loss may affect a wide range of applications that use voice over IP (VoIP), e.g. video conferencing. In this paper, we investigate a time-domain convolutional recurrent network (CRN) for online packet loss concealment. The CRN comprises a convolutional encoder-decoder structure and long short…

Cited by 0SourceScholar
2021

Deep Wasserstein Graph Discriminant Learning for Graph Classification

AAAI 2021technical

Graph topological structures are crucial to distinguish different-class graphs. In this work, we propose a deep Wasserstein graph discriminant learning (WGDL) framework to learn discriminative embeddings of graphs in Wasserstein-metric (W-metric) matching space. In order to bypass the calculation of…

Cited by 19SourcePDFScholar
2021

Fast Hierarchy Preserving Graph Embedding via Subspace Constraints

ICASSP 2021accepted

Hierarchy preserving network embedding is a method that project nodes into feature space by preserving the hierarchy property of networks. Recently, researches on network representation have considerably profited from taking hierarchy into consideration. Among these works, SpaceNE <sup xmlns:mml="ht…

Cited by 0SourceScholar
2021

Wasserstein Coupled Graph Learning for Cross-Modal Retrieval

ICCV 2021poster

Graphs play an important role in cross-modal image-text understanding as they characterize the intrinsic structure which is robust and crucial for the measurement of cross-modal similarity. In this work, we propose a Wasserstein Coupled Graph Learning (WCGL) method to deal with the cross-modal retri…

Cited by 29PDFScholar
2020

Selective Attention Encoders by Syntactic Graph Convolutional Networks for Document Summarization

ICASSP 2020accepted

Abstractive text summarization is a challenging task, and one need to design a mechanism to effectively extract salient information from the source text and then generate a summary. A parsing process of the source text contains critical syntactic or semantic structures, which is useful to generate m…

Cited by 0SourceScholar
2019

A Comparison of Five Multiple Instance Learning Pooling Functions for Sound Event Detection with Weak Labeling

ICASSP 2019accepted

Sound event detection (SED) entails two subtasks: recognizing what types of sound events are present in an audio stream (audio tagging), and pinpointing their onset and offset times (localization). In the popular multiple instance learning (MIL) framework for SED with weak labeling, an important com…

Cited by 0SourceScholar
2018

A Light-Weight Multimodal Framework for Improved Environmental Audio Tagging

ICASSP 2018accepted

The lack of strong labels has severely limited the state-of-the-art fully supervised audio tagging systems to be scaled to larger dataset. Meanwhile, audio-visual learning models based on unlabeled videos have been successfully applied to audio tagging, but they are inevitably resource hungry and re…

Cited by 0SourceScholar
2017

A first attempt at polyphonic sound event detection using connectionist temporal classification

ICASSP 2017accepted

Sound event detection is the task of detecting the type, starting time, and ending time of sound events in audio streams. Recently, recurrent neural networks (RNNs) have become the mainstream solution for sound event detection. Because RNNs make a prediction at every frame, it is necessary to provid…

Cited by 0SourceScholar
2015

Semi-supervised training in low-resource ASR and KWS

ICASSP 2015accepted

In particular for “low resource” Keyword Search (KWS) and Speech-to-Text (STT) tasks, more untranscribed test data may be available than training data. Several approaches have been proposed to make this data useful during system development, even when initial systems have Word Error Rates (WER) abov…

Cited by 0SourceScholar