← Search

Yan Zhao

31 accepted papers

2026

A Brain-Inspired Saliency Prediction Framework for Human-AI Cognitive Consistency in AIGC Content via Multi-Region Liquid Neurons

AAAI 2026technical

In recent years, human-AI cognitive consistency has emerged as a crucial perspective for evaluating the perceptual quality and interpretability of AIGC (Artificial Intelligence Generated Content). This paper proposes a biologically inspired saliency prediction framework that models six core regions

Cited by 0SourcePDFScholar
2026

D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos

AAAI 2026technical

Free-Viewpoint Video (FVV) enables immersive 3D experiences, but efficient compression of dynamic 3D representation remains a major challenge. Existing dynamic 3D Gaussian Splatting methods couple reconstruction with optimization-dependent compression and customized motion formats, limiting generali

Cited by 0SourcePDFScholar
2026

How Do Language Models Speak Languages? A Case Study on Unintended Code-Switching

ICML 2026poster

Unintended code-switching, which refers to the phenomenon where LLM unexpectedly switch languages, poses a fundamental challenge in the multilingual capabilities in LLMs. However, we still lack a mechanistic account of how this failure mode is implemented inside the model. For example, what internal…

Cited by 0SourceScholar
2026

Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising

ICML 2026poster

Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. However, the volume of time series data may vary significantly across domains due to low sampling rates and data regulations. To maximally create value from sparse data, th…

Cited by 0SourceScholar
2026

OmniZip: Learning a Unified and Lightweight Lossless Compressor for Multi-Modal Data

CVPR 2026

Lossless compression is essential for efficient data storage and transmission. Although learning-based lossless compressors achieve strong results, most of them are designed for a single modality, leading to redundant compressor deployments in multi-modal settings. Designing a unified multi-modal co

Cited by 0SourcecodeScholar
2026

TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile Data

ICLR 2026poster

Tactile sensing is crucial for embodied intelligence, providing fine-grained perception and control in complex environments. However, efficient tactile data compression, which is essential for real-time robotic applications under strict bandwidth constraints, remains underexplored. The inherent hete…

Cited by 0SourceScholar
2025

L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression

AAAI 2025technical

Learning-based probabilistic models can be combined with an entropy coder for data compression. However, due to the high complexity of learning-based models, their practical application as text compressors has been largely overlooked. To address this issue, our work focuses on a low-complexity desig…

2025

MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining

NeurIPS 2025poster

Data quality is a critical driver of large language model performance, yet existing model-based selection methods focus almost exclusively on English, neglecting other languages that are essential in the training mix for multilingual LLMs. We introduce MuRating, a scalable framework that transfers h…

Cited by 0SourceScholar
2025

SPOT-Trip: Dual-Preference Driven Out-of-Town Trip Recommendation

NeurIPS 2025poster

Out-of-town trip recommendation aims to generate a sequence of Points of Interest (POIs) for users traveling from their hometowns to previously unvisited regions based on personalized itineraries, e.g., origin, destination, and trip duration. Modeling the complex user preferences--which often exhibi…

Cited by 0SourceScholar
2025

Self-Supervised Learning and Image-Prompt Fusion for AIGC Image Quality Assessment

ICASSP 2025accepted

With the rapid advancement of artificial intelligence, the field of Artificial Intelligence Generated Content (AIGC) has seen significant growth. As AI-generated images (AIGIs) become increasingly prevalent, the AIGC image quality assessment(AIGCIQA) has gained critical importance. However, traditio…

Cited by 0SourceScholar
2024

Audio Prompt Tuning for Universal Sound Separation

ICASSP 2024accepted

Universal sound separation (USS) is a task to separate arbitrary sounds from an audio mixture. Existing USS systems are capable of separating arbitrary sources, given a few examples of the target sources as queries. However, separating arbitrary sounds with a single system is challenging, and the ro…

Cited by 0SourceScholar
2024

Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition

ICASSP 2024accepted

Cross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This p…

Cited by 0SourceScholar
2024

Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution Adaptation

ICASSP 2024accepted

In speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new s…

Cited by 0SourceScholar
2024

Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition

ICASSP 2024accepted

Swin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above in…

Cited by 0SourceScholar
2023

Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion Recognition

ICASSP 2023accepted

In this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and unlabeled testing (target) speech signals come from different co…

Cited by 0SourceScholar
2023

DualAfford: Learning Collaborative Visual Affordance for Dual-gripper Manipulation

ICLR 2023poster

It is essential yet challenging for future home-assistant robots to understand and manipulate diverse 3D objects in daily human environments. Towards building scalable systems that can perform diverse manipulation tasks over various 3D shapes, recent works have advocated and demonstrated promising r…

Cited by 16SourcePDFScholar
2023

Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions

NeurIPS 2023poster

Perceiving and manipulating 3D articulated objects in diverse environments is essential for home-assistant robots. Recent studies have shown that point-level affordance provides actionable priors for downstream manipulation tasks. However, existing works primarily focus on single-object scenarios wi…

Cited by 25SourcePDFScholar
2023

Leveraging SE(3) Equivariance for Learning 3D Geometric Shape Assembly

ICCV 2023poster

Shape assembly aims to reassemble parts (or fragments) into a complete object, which is a common task in our daily life. Different from the semantic part assembly (e.g., assembling a chair's semantic parts like legs into a whole chair), geometric part assembly (e.g., assembling bowl fragments into a…

Cited by 21PDFcodeScholar
2022

Adaptive Logit Adjustment Loss for Long-Tailed Visual Recognition

AAAI 2022technical

Data in the real world tends to exhibit a long-tailed label distribution, which poses great challenges for the training of neural networks in visual recognition. Existing methods tackle this problem mainly from the perspective of data quantity, i.e., the number of samples in each class. To be specif…

Cited by 67SourcePDFScholar
2022

VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated Objects

ICLR 2022poster

Perceiving and manipulating 3D articulated objects (e.g., cabinets, doors) in human environments is an important yet challenging task for future home-assistant robots. The space of 3D articulated objects is exceptionally rich in their myriad semantic categories, diverse shape geometry, and complicat…

Cited by 104SourcePDFScholar
2020

VALID: A Comprehensive Virtual Aerial Image Dataset

ICRA 2020poster

Aerial imagery plays an important role in land-use planning, population analysis, precision agriculture, and unmanned aerial vehicle tasks. However, existing aerial image datasets generally suffer from the problem of inaccurate labeling, single ground truth type, and few category numbers. In this wo…

Cited by 46SourceScholar
2019

An Image Coding Approach Based on Mixture-of-experts Regression Using Epanechnikov Kernel

ICASSP 2019accepted

In this paper, we propose an optimal modeling framework for image compression using EMM (Epanechnikov Mixture Model). Epanechnikov Kernel and its correlated statistics are basement of our Epanechnikov Mixture Regression (EMR). In our scheme, the stochastic processes of the pixel values are modelled…

Cited by 0SourceScholar
2019

Video-based, Occlusion-robust Multi-view Stereo Using Inner-boundary Depths of Textureless Areas

ICASSP 2019accepted

Occlusions and poor textures are two main problems in multi-view stereo reconstruction. This paper presents a video-based solution to address both challenges in depth estimation. We focus on reconstructing accurate inner boundaries of visible textureless areas, particularly for occluded background,…

Cited by 0SourceScholar
2018

Late Reverberation Suppression Using Recurrent Neural Networks with Long Short-Term Memory

ICASSP 2018accepted

Human speech is usually distorted by room reverberation. These corruptions degrade speech quality and intelligibility, especially under a long reverberation time, and they also pose a serious problem for many speech-related applications such as automatic speech recognition. In this paper, we propose…

Cited by 0SourceScholar