← Search

Yang Xiao

56 accepted papers

2026

AFT: AN EXEMPLAR-FREE CLASS INCREMENTAL LEARNING METHOD FOR ENVIRONMENTAL SOUND CLASSIFICATION

ICASSP 2026poster

As sounds carry rich information, environmental sound classification (ESC) is crucial for numerous applications such as rare wild animals detection. However, our world constantly changes, asking ESC models to adapt to new sounds periodically. The major challenge here is catastrophic forgetting, wher…

Cited by 0SourcePDFScholar
2026

Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal Transport

CVPR 2026

Semi-supervised learning (SSL) has achieved remarkable progress by leveraging both limited labeled data and abundant unlabeled data. However, unlabeled datasets often contain out-of-distribution (OOD) samples from unknown classes, which can lead to performance degradation in open-set SSL scenarios.

Cited by 0SourceScholar
2026

CoverPruneGS: Coverage-Preserving Structured Pruning for Hierarchical 3D Gaussian Splatting from Sparse-View Monocular Videos

ICML 2026poster

Reconstructing a complete yet compact 3DGS from sparse-view monocular long videos is challenging: hierarchical training with VFI can improve coverage, yet correlated pseudo views and repeated merging tend to accumulate near-duplicate Gaussians and exacerbate co-adaptation. To address this, we propos…

Cited by 0SourceScholar
2026

DEFANet: Dual-Path Edge-Target Collaboration with Frequency-Aware Enhancement for Infrared Small Target Detection

AAAI 2026technical

Infrared small target detection is challenging due to limited target size and low signal-to-noise ratio. Unlike common targets, infrared small targets contain a higher proportion of edge pixels and exhibit blurred boundaries due to diffraction and quantization artifacts, making boundaries uniquely v

Cited by 0SourcePDFScholar
2026

DeFB: Decomposed Feature Learning for Real-Time Multi-Person Eyeblink Detection in Untrimmed In-the-Wild Videos

AAAI 2026technical

Multi-person eyeblink detection in untrimmed in-the-wild videos is a recently emerged and challenging task. Due to its significant spatio-temporal fine-grained characteristics compared to general actions, we empirically find that general action detectors, though effective in general domains, struggl

Cited by 0SourcePDFScholar
2026

Environmental Sound Deepfake Detection Challenge: An Overview

ICASSP 2026poster

Recent progress in audio generation models has made it possible to create highly realistic and immersive soundscapes, which are now widely used in film and virtual-reality-related applications. However, these audio generators also raise concerns about potential misuse, such as producing deceptive au…

Cited by 0SourcePDFScholar
2026

Forget Many, Forget Right: Scalable and Precise Concept Unlearning in Diffusion Models

ICLR 2026poster

While multi-concept unlearning has shown progress, extending to large-scale scenarios remains difficult, as existing methods face three persistent challenges: **(i)** they often introduce conflicting weight updates, making some targets difficult to unlearn or causing degradation of generative capab…

Cited by 0SourceScholar
2026

InnovatorBench: Evaluating Agents’ Ability to Conduct Innovative AI Research

ICLR 2026poster

AI agents could accelerate scientific discovery by automating hypothesis formation, experiment design, coding, execution, and analysis, yet existing benchmarks probe narrow skills in simplified settings. To address this gap, we introduce InnovatorBench, a benchmark-platform pair for realistic, end-t…

Cited by 0SourcecodeScholar
2026

JUMP-Hand: Learning Joint-wise Uncertainty to Gate Mixture of View Experts for Multi-View 3D Hand Reconstruction

CVPR 2026

We propose JUMP-Hand, a novel multi-view 3D hand reconstruction method that explicitly models probabilistic joint-wise uncertainty as a gating mechanism for multi-view fusion. Existing approaches usually rely on naive pooling or implicit attention, overlooking that each hand joint exhibits varying v

Cited by 0SourcecodeScholar
2026

MOESCORE: MIXTURE-OF-EXPERTS-BASED TEXT-AUDIO RELEVANCE SCORE PREDICTION FOR TEXT-TO-AUDIO SYSTEM EVALUATION

ICASSP 2026poster

Recent advances in generative models have enabled modern Text-to-Audio (TTA) systems to synthesize audio with high perceptual quality. However, TTA systems often struggle to maintain semantic consistency with the input text, leading to mismatches in sound events, temporal tructures, or contextual re…

Cited by 0SourcePDFScholar
2026

MoEActok: A MoE-based Action Tokenizer for Vision-Language-Action Models

CVPR 2026

Recent works on vision-language-action (VLA) models have made great progress in exploring action tokenizers that convert continuous control signals into discrete tokens to align with LLM/VLM training paradigms.These approaches typically train a single tokenizer over entire manipulation trajectories,

Cited by 0SourcecodeScholar
2026

MoEG-HOI: Mixture of Expert Groups for One-Stage Hand-Object Interaction Motion Generation with Hand-Finger-Joint Semantic Guidance

AAAI 2026technical

In this paper, MoEG-HOI is proposed as a novel method for the challenging 3D hand-object interaction (HOI) motion generation task, by introducing Mixture-of-Experts (MoE) to this field for the first time. Almost all the mainstream approaches in HOI motion generation leverage diffusion model as its s

Cited by 0SourcePDFScholar
2026

Phys-Liquid: A Physics-Informed Dataset for Estimating 3D Geometry and Volume of Transparent Deformable Liquids

AAAI 2026technical

Estimating the geometric and volumetric properties of transparent deformable liquids is challenging due to optical complexities and dynamic surface deformations induced by container movements. Autonomous robots performing precise liquid manipulation tasks—such as dispensing, aspiration, and mixing—m

Cited by 0SourcePDFScholar
2026

SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling

AAAI 2026technical

Test-time compute scaling has emerged as a powerful paradigm for enhancing mathematical reasoning in large language models (LLMs) by allocating additional computational resources during inference. However, current methods employ uniform resource distribution across all reasoning sub-problems, creati

Cited by 0SourcePDFScholar
2026

Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic Event Classification

ICASSP 2026poster

Multimodal acoustic event classification plays a key role in audio-visual systems. Although combining audio and visual signals improves recognition, it is still difficult to align them over time and to reduce the effect of noise across modalities. Existing methods often treat audio and visual stream…

Cited by 0SourcePDFScholar
2026

Weak-to-Strong Generalization with Failure Trajectories

ICLR 2026poster

Weak-to-Strong generalization (W2SG) is a new trend to elicit the full capabilities of a strong model with supervision from a weak model. While existing W2SG studies focus on simple tasks like binary classification, we extend this paradigm to complex interactive decision-making environments. Speci…

Cited by 0SourcecodeScholar
2026

Your Language Model Secretly Contains Personality Subnetworks

ICLR 2026poster

Humans shift between different personas depending on social context. Large Language Models (LLMs) demonstrate a similar flexibility in adopting different personas and behaviors. Existing approaches, however, typically adapt such behavior through external knowledge such as prompting, retrieval-augmen…

Cited by 0SourcecodeScholar
2026

daVinci-Dev: Agent-native Mid-training for Software Engineering

ICML 2026oral

Recently, the frontier of Large Language Model (LLM) capabilities has shifted from single-turn code generation to agentic software engineering—a paradigm where models autonomously navigate, edit, and test complex repositories. While post-training methods have become the de facto approach for code ag…

Cited by 0SourceScholar
2025

AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning for Small-footprint Keyword Spotting

ACL 2025finding

Keyword spotting (KWS) offers a vital mechanism to identify spoken commands in voice-enabled systems, where user demands often shift, requiring models to learn new keywords continually over time. However, a major problem is catastrophic forgetting, where models lose their ability to recognize earlie…

Cited by 0SourcePDFScholar
2025

Boosting Vulnerability Detection of LLMs via Curriculum Preference Optimization with Synthetic Reasoning Data

ACL 2025finding

Large language models (LLMs) demonstrate considerable proficiency in numerous coding-related tasks; however, their capabilities in detecting software vulnerabilities remain limited. This limitation primarily stems from two factors: (1) the absence of reasoning data related to vulnerabilities, which…

2025

Exploring Text-Queried Sound Event Detection with Audio Source Separation

ICASSP 2025accepted

In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we firs…

Cited by 0SourceScholar
2025

LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling

NeurIPS 2025poster

Large language models (LLMs) have demonstrated remarkable reasoning capabilities through test-time scaling approaches, particularly when fine-tuned with chain-of-thought (CoT) data distilled from more powerful large reasoning models (LRMs). However, these reasoning chains often contain verbose eleme…

Cited by 0SourcecodeScholar
2025

M²N: A Progressive Macro-to-Micro 3D Modeling Scheme for Unveiling Drug-Target Affinity

AAAI 2025technical

Accurate drug-target affinity (DTA) prediction holds significant potential in the field of artificial intelligence (AI)-based drug discovery. However, existing methods primarily operate at a single scale, specifically at the macro (residue) scale for target proteins and the micro (atom) scale for dr…

Cited by 0SourcePDFScholar
2025

Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information Processing

ICCV 2025poster

The design of artificial neural networks (ANNs) is inspired by the structure of the human brain, and in turn, ANNs offer a potential means to interpret and understand brain signals. Existing methods primarily align brain signals with stimulus signals using Mean Squared Error (MSE), which focuses onl…

2025

PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space

NeurIPS 2025poster

3D human pose lifting from a single RGB image is a challenging task in 3D vision. Existing methods typically establish a direct joint-to-joint mapping from 2D to 3D poses based on 2D features. This formulation suffers from two fundamental limitations: inevitable error propagation from input predicte…

Cited by 0SourceScholar
2025

Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware Optimization

ICCV 2025poster

Text-to-image (T2I) diffusion models have achieved remarkable success in generating high-quality images from textual prompts. However, their ability to store vast amounts of knowledge raises concerns in scenarios where selective forgetting is necessary, such as removing copyrighted content, reducing…

2025

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States

ACL 2025long

As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks assess basic ToM abilities, they predominantly focus on stati…

2024

"Towards Dual Transparent Liquid Level Estimation in Biomedical Lab: Dataset, Methods and Practice"

ECCV 2024poster

"“Dual Transparent Liquid” refers to a liquid and its container, both being transparent. Accurately estimating the levels of such a liquid from arbitrary viewpoints is fundamental and crucial, especially in AI-guided autonomous biomedical laboratories for tasks like liquid dispensing, aspiration, an…

2024

A Label Disambiguation-Based Multimodal Massive Multiple Instance Learning Approach for Immune Repertoire Classification

AAAI 2024technical

One individual human’s immune repertoire consists of a huge set of adaptive immune receptors at a certain time point, representing the individual's adaptive immune state. Immune repertoire classification and associated receptor identification have the potential to make a transformative contribution…

2024

BEE-Net: Bridging Semantic and Instance with Gated Encoding and Edge Constraint for Efficient Panoptic Segmentation

ICRA 2024poster

Panoptic segmentation is a challenging perception task, which can help robots to comprehensively perceive the surrounding environment. In the task, we notice that semantic, instance, and panoptic have rich relations, however, which are rarely explored. In this work, we propose a novel panoptic, inst…

Cited by 0SourceScholar
2024

CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers

ICRA 2024poster

With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we prese…

Cited by 0SourceScholar
2024

NGP-RT: Fusing Multi-Level Hash Features with Lightweight Attention for Real-Time Novel View Synthesis

ECCV 2024poster

"This paper presents NGP-RT, a novel approach for enhancing the rendering speed of Instant-NGP to achieve real-time novel view synthesis. As a classic NeRF-based method, Instant-NGP stores implicit features in multi-level grids or hash tables and applies a shallow MLP to convert the implicit feature…

Cited by 0SourcePDFScholar
2024

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

NeurIPS 2024poster

The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i.e., AI4Science) once exclus…

2024

SAI3D: Segment Any Instance in 3D Scenes

CVPR 2024poster

Advancements in 3D instance segmentation have traditionally been tethered to the availability of annotated datasets limiting their application to a narrow spectrum of object categories. Recent efforts have sought to harness vision-language models like CLIP for open-set semantic reasoning yet these m…

2024

Towards Robust Evidence-Aware Fake News Detection via Improving Semantic Perception

COLING 2024main

Evidence-aware fake news detection aims to determine the veracity of a given news (i.e., claim) with external evidences. We find that existing methods lack sufficient semantic perception and are easily blinded by textual expressions. For example, they still make the same prediction after we flip the…

2023

A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation From a Single RGB Image

CVPR 2023poster

3D interacting hand pose estimation from a single RGB image is a challenging task, due to serious self-occlusion and inter-occlusion towards hands, confusing similar appearance patterns between 2 hands, ill-posed joint position mapping from 2D to 3D, etc.. To address these, we propose to extend A2J-…

2023

Real-Time Multi-Person Eyeblink Detection in the Wild for Untrimmed Video

CVPR 2023poster

Real-time eyeblink detection in the wild can widely serve for fatigue detection, face anti-spoofing, emotion analysis, etc. The existing research efforts generally focus on single-person cases towards trimmed video. However, multi-person scenario within untrimmed videos is also important for practic…

2023

Scoreformer: Score Fusion-Based Transformers for Weakly-Supervised Violence Detection

ICASSP 2023accepted

Violence detection is an application of anomaly detection, which is used to detect violence content in video clips. Using multimodal as input can improve the performance of violence detection. However, the existing MML Transformers-based fusion methods do not take into account the differences betwee…

Cited by 0SourceScholar
2022

Are All the Datasets in Benchmark Necessary? A Pilot Study of Dataset Evaluation for Text Classification

NAACL 2022long

In this paper, we ask the research question of whether all the datasets in the benchmark are necessary. We approach this by first characterizing the distinguishability of datasets when comparing different systems. Experiments on 9 datasets and 36 systems show that several existing benchmark datasets…

2022

C3P: Cross-Domain Pose Prior Propagation for Weakly Supervised 3D Human Pose Estimation

ECCV 2022poster

"This paper first proposes and solves weakly supervised 3D human pose estimation (HPE) problem in point cloud, via propagating the pose prior within unlabelled RGB-point cloud sequence to 3D domain. Our approach termed C3P does not require any labor-consuming 3D keypoint annotation for training. To…

2022

On the Robustness of Reading Comprehension Models to Entity Renaming

NAACL 2022long

We study the robustness of machine reading comprehension (MRC) models to entity renaming—do models make more wrong predictions when the same questions are asked about an entity whose name has been changed? Such failures imply that models overly rely on entity information to answer questions, and thu…

2022

Templates for 3D Object Pose Estimation Revisited: Generalization to New Objects and Robustness to Occlusions

CVPR 2022poster

We present a method that can recognize new objects and estimate their 3D pose in RGB images even under partial occlusions. Our method requires neither a training phase on these objects nor real images depicting them, only their CAD models. It relies on a small set of training objects to learn local…

Cited by 87PDFcodeScholar
2021

Re-ranking for image retrieval and transductive few-shot classification

NeurIPS 2021poster

In the problems of image retrieval and few-shot classification, the mainstream approaches focus on learning a better feature representation. However, directly tackling the distance or similarity measure between images could also be efficient. To this end, we revisit the idea of re-ranking the top-k…

Cited by 51SourcePDFScholar
2020

3DV: 3D Dynamic Voxel for Action Recognition in Depth Video

CVPR 2020poster

For depth-based 3D action recognition, one essential issue is to represent 3D motion pattern effectively and efficiently. To this end, 3D dynamic voxel (3DV) is proposed as a novel 3D motion representation manner. With 3D space voxelization, the key idea of 3DV is to encode the 3D motion information…

Cited by 126PDFcodeScholar
2020

Empirical Bayes Transductive Meta-Learning with Synthetic Gradients

ICLR 2020poster

We propose a meta-learning approach that learns from multiple tasks in a transductive setting, by leveraging the unlabeled query set in addition to the support set to generate a more powerful model for each task. To develop our framework, we revisit the empirical Bayes formulation for multi-task le…

Cited by 186SourceScholar
2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds

CVPR 2020oral

Towards 3D object tracking in point clouds, a novel point-to-box network termed P2B is proposed in an end-to-end learning manner. Our main idea is to first localize potential target centers in 3D search area embedded with target information. Then point-driven 3D target proposal and verification are…

Cited by 199PDFcodeScholar
2020

Pixel-Pair Occlusion Relationship Map (P2ORM): Formulation, Inference & Application

ECCV 2020poster

Inference & Application","We formalize concepts around geometric occlusion in 2D images (i.e., ignoring semantics), and propose a novel unified formulation of both occlusion boundaries and occlusion orientations via a pixel-pair occlusion relation. The former provides a way to generate large-scale a…

2019

A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth Image

ICCV 2019poster

For 3D hand and body pose estimation task in depth image, a novel anchor-based approach termed Anchor-to-Joint regression network (A2J) with the end-to-end learning ability is proposed. Within A2J, anchor points able to capture global-local spatial context information are densely set on depth image…

Cited by 221PDFcodeScholar
2018

Monocular Relative Depth Perception With Web Stereo Data Supervision

CVPR 2018poster

In this paper we study the problem of monocular relative depth perception in the wild. We introduce a simple yet effective method to automatically generate dense relative depth annotations from web stereo images, and propose a new dataset that consists of diverse images as well as corresponding dens…

Cited by 253SourcePDFScholar