← Search

Yan Song

66 accepted papers

2026

Active Reasoning Vision-Language Model via Sequential Experimental Design

ICML 2026poster

Visual perception in modern Vision-Language Models (VLM) is constrained by a fundamental perceptual bandwidth bottleneck: a broad field-of-view inevitably sacrifices the fine-grained details necessary for complex reasoning. Inspired by the classical paradigms of active vision and information foragin…

Cited by 0SourceScholar
2026

Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel

CVPR 2026

Large vision foundation models, such as SAM2, have achieved remarkable performance in video object segmentation and tracking (VOST). However, their effectiveness is hindered by significant computational overhead. While model pruning is a widely used strategy to address this issue, traditional static

Cited by 0SourceScholar
2026

GO-PRE:Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction

ICML 2026poster

Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing methods rely on surrogate signals—such as parameter uncertainty or geometric heuristics—which are often misaligned with the ultimate goal: the fidelity o…

Cited by 0SourceScholar
2025

Efficient Reinforcement Learning with Large Language Model Priors

ICLR 2025poster

In sequential decision-making (SDM) tasks, methods like reinforcement learning (RL) and heuristic search have made notable advances in specific cases. However, they often require extensive exploration and face challenges in generalizing across diverse environments due to their limited grasp of the u…

Cited by 4SourcePDFScholar
2025

ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation

NeurIPS 2025poster

Vision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncert…

Cited by 0SourceScholar
2025

Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments

IROS 2025

Anomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects i

Cited by 0SourcecodeScholar
2025

PNP-RKD: A Positive-Negative Pair based Relational Knowledge Distillation Method for Cross-Domain Speaker Verification

ICASSP 2025accepted

Existing deep embedding learning based speaker verification (SV) methods suffer from performance degradation under domain shift conditions. This can be alleviated through unsupervised domain adaptation (UDA) techniques. While UDA improves global statistical consistency across domains, discriminative…

Cited by 0SourceScholar
2025

Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection

ICASSP 2025accepted

A significant challenge in sound event detection (SED) is the effective utilization of unlabeled data, given the limited availability of labeled data due to high annotation costs. Semi-supervised algorithms rely on labeled data to learn from unlabeled data, and the performance is constrained by the…

Cited by 0SourceScholar
2025

ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning

NeurIPS 2025poster

Recent research on Reasoning of Large Language Models (LLMs) has sought to further enhance their performance by integrating meta-thinking—enabling models to monitor, evaluate, and control their reasoning processes for more adaptive and effective problem-solving. However, current single-agent work la…

Cited by 0SourcecodeScholar
2025

Reinforcement Learning from Imperfect Corrective Actions and Proxy Rewards

ICLR 2025poster

In practice, reinforcement learning (RL) agents are often trained with a possibly imperfect proxy reward function, which may lead to a human-agent alignment issue (i.e., the learned policy either converges to non-optimal performance with low cumulative rewards, or achieves high cumulative rewards bu…

Cited by 1SourcePDFScholar
2025

ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning

NeurIPS 2025poster

Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs…

Cited by 0SourcecodeScholar
2025

medIKAL: Integrating Knowledge Graphs as Assistants of LLMs for Enhanced Clinical Diagnosis on EMRs

COLING 2025main

Electronic Medical Records (EMRs), while integral to modern healthcare, present challenges for clinical reasoning and diagnosis due to their complexity and information redundancy. To address this, we proposed medIKAL (Integrating Knowledge Graphs as Assistants of LLMs), a framework that combines Lar…

2024

AI-Olympics: Exploring the Generalization of Agents through Open Competitions

IJCAI 2024poster

Between 2021 and 2023, AI-Olympics---a series of online AI competitions, was hosted by the online evaluation platform Jidi in collaboration with the IJCAI committee. In these competitions, an agent is required to accomplish diverse sports tasks in a two-dimensional continuous world, while competing…

Cited by 2SourcePDFScholar
2024

Aspect-based Sentiment Analysis with Context Denoising

NAACL 2024findings

Given a sentence and a particular aspect term, aspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity towards this aspect term, which provides fine-grained analysis on sentiment understanding and it has attracted much attention in recent years. In order to achieve a good perfo…

2024

Bootstrapping Large Language Models for Radiology Report Generation

AAAI 2024technical

Radiology report generation (RRG) aims to automatically generate a free-text description from a specific clinical radiograph, e.g., chest X-Ray images. Existing approaches tend to perform RRG with specific models trained on the public yet limited data from scratch, where they often lead to inferior…

2024

Challenging Large Language Models with New Tasks: A Study on their Adaptability and Robustness

ACL 2024findings

Recent progress in large language models (LLMs) has marked a notable milestone in the field of artificial intelligence. The conventional evaluation of LLMs primarily relies on existing tasks and benchmarks, raising concerns about test set contamination and the genuine comprehension abilities of LLMs…

2024

ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences

ACL 2024long

Recently, the increasing demand for superior medical services has highlighted the discrepancies in the medical infrastructure. With big data, especially texts, forming the foundation of medical services, there is an exigent need for effective natural language processing (NLP) solutions tailored to t…

2024

Dialogue Summarization with Mixture of Experts based on Large Language Models

ACL 2024long

Dialogue summarization is an important task that requires to generate highlights for a conversation from different aspects (e.g., content of various speakers). While several studies successfully employ large language models (LLMs) and achieve satisfying results, they are limited by using one model a…

2024

Improving Radiology Report Generation with D2-Net: When Diffusion Meets Discriminator

ICASSP 2024accepted

Radiology report generation (RRG) aims to automatically provide observations and insight into a patient’s condition based on radiology images, which is able to greatly reduce the workload of physicians on the premise of ensuring the quality of medical treatment. Existing works leverage the Transform…

Cited by 0SourceScholar
2024

Learning Multimodal Contrast with Cross-modal Memory and Reinforced Contrast Recognition

ACL 2024findings

In many practical scenarios, contents from different modalities are not semantically aligned; for instance, visual and textual information may conflict with each other, resulting in non-compositional expression effects such as irony or humor. Effective modeling and smooth integration of multimodal i…

2024

Meta Representation Learning Method for Robust Speaker Verification in Unseen Domains

ICASSP 2024accepted

This paper presents a meta representation learning method for robust speaker verification (SV) in unseen domains. It is known that the existing embedding learning based SV systems may suffer from domain mismatch issues. To address this, we propose an episodic training procedure to compensate domain…

Cited by 0SourceScholar
2024

Norface: Improving Facial Expression Analysis by Identity Normalization

ECCV 2024poster

"Facial Expression Analysis remains a challenging task due to unexpected task-irrelevant noise, such as identity, head pose, and background. To address this issue, this paper proposes a novel framework, called Norface, that is unified for both Action Unit (AU) analysis and Facial Emotion Recognition…

2024

Prompting Few-shot Multi-hop Question Generation via Comprehending Type-aware Semantics

NAACL 2024findings

Given several documents, multi-hop question generation (MQG) is a task aims to generate complicated questions that require reasoning over multiple pieces of these documents to find the answer. To perform this task, existing studies focus on designing advanced architectures to locate essential keywor…

2024

RESEMO: A Benchmark Chinese Dataset for Studying Responsive Emotion from Social Media Content

ACL 2024findings

On social media platforms, users’ emotions are triggered when they encounter particular content from other users,where such emotions are different from those that spontaneously emerged, owing to the “responsive” nature. Analyzing the aforementioned responsive emotions from user interactions is a tas…

Cited by 0SourcePDFScholar
2023

AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer

ICASSP 2023accepted

In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have recently shown promise on DCASE2022 challenge task4 where t…

Cited by 0SourceScholar
2023

An Effective Anomalous Sound Detection Method Based on Representation Learning with Simulated Anomalies

ICASSP 2023accepted

In this paper, we propose an effective anomalous sound detection (ASD) method based on representation learning with simulated anomalies. Recently, ASD systems have used Outlier Exposure (OE) strategy to achieve promising performance in DCASE challenges. These exploit deep Convolutional Neural Networ…

Cited by 0SourceScholar
2023

End-to-end Aspect-based Sentiment Analysis with Combinatory Categorial Grammar

ACL 2023findings

End-to-end Aspect-based Sentiment Analysis (EASA) is a natural language processing (NLP) task that involves extracting aspect terms and identifying the sentiments for them, which provides a fine-grained level of text analysis and thus requires a deep understanding of the running text. Many previous…

2023

Improving Image Captioning via Predicting Structured Concepts

EMNLP 2023long main

Having the difficulty of solving the semantic gap between images and texts for the image captioning task, conventional studies in this area paid some attention to treating semantic concepts as a bridge between the two modalities and improved captioning performance accordingly. Although promising res…

Cited by 0SourceScholar
2023

Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection

ICASSP 2023accepted

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together wi…

Cited by 28SourceScholar
2023

Learning Semantic Relationship Among Instances for Image-Text Matching

CVPR 2023poster

Image-text matching, a bridge connecting image and language, is an important task, which generally learns a holistic cross-modal embedding to achieve a high-quality semantic alignment between the two modalities. However, previous studies only focus on capturing fragment-level relation within a sampl…

2023

Neural Episodic Control with State Abstraction

ICLR 2023top-25%

Existing Deep Reinforcement Learning (DRL) algorithms suffer from sample inefficiency. Generally, episodic control-based approaches are solutions that leverage highly rewarded past experiences to improve sample efficiency of DRL algorithms. However, previous episodic control-based approaches fail to…

Cited by 14SourcePDFScholar
2023

Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to…

Cited by 0SourceScholar
2023

Text Style Transfer with Contrastive Transfer Pattern Mining

ACL 2023long

Text style transfer (TST) is an important task in natural language generation, which aims to alter the stylistic attributes (e.g., sentiment) of a sentence and keep its semantic meaning unchanged. Most existing studies mainly focus on the transformation between styles, yet ignore that this transform…

2022

Domain Robust Deep Embedding Learning for Speaker Recognition

ICASSP 2022accepted

This paper presents a domain robust deep embedding learning method for speaker verification (SV) tasks. Most recent methods utilize deep neural networks (DNN) to learn compact and discriminative speaker embeddings from large-scale labeled datasets such as VoxCeleb and the NIST SRE corpus. Despite th…

Cited by 0SourceScholar
2022

Enhancing Structure-aware Encoder with Extremely Limited Data for Graph-based Dependency Parsing

COLING 2022main

Dependency parsing is an important fundamental natural language processing task which analyzes the syntactic structure of an input sentence by illustrating the syntactic relations between words. To improve dependency parsing, leveraging existing dependency parsers and extra data (e.g., through semi-…

2022

Frontend Attributes Disentanglement for Speech Emotion Recognition

ICASSP 2022accepted

Speech emotion recognition (SER) with limited size dataset is a challenging task, since a spoken utterance contains various disturbing attributes besides emotion, including speaker, content, and language. However, due to a close relationship between speaker and emotion attributes, simply fine-tuning…

Cited by 0SourceScholar
2022

Improving English-Arabic Transliteration with Phonemic Memories

EMNLP 2022finding

Transliteration is an important task in natural language processing (NLP) which aims to convert a name in the source language to the target language without changing its pronunciation. Particularly, transliteration from English to Arabic is highly needed in many applications, especially in countries…

2022

Improving Relation Extraction through Syntax-induced Pre-training with Dependency Masking

ACL 2022findings

Relation extraction (RE) is an important natural language processing task that predicts the relation between two given entities, where a good understanding of the contextual information is essential to achieve an outstanding model performance. Among different types of contextual information, the aut…

2022

Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift

ICASSP 2022accepted

In this paper, a self-supervised representation learning method is proposed for anomalous sound detection (ASD). ASD has received much research attention in recent DCASE challenges. It aims to identify whether a sound emitted from a machine is anomalous or not, given only normal sound data. This is…

Cited by 0SourceScholar
2021

An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification

ICASSP 2021accepted

In this paper, we present an effective end-to-end deep embedding learning method based on Dense-Residual networks, which combine the advantages of a densely connected convolutional network (DenseNet) and a residual network (ResNet), for speaker verification (SV). Unlike a model ensemble strategy whi…

Cited by 0SourceScholar
2021

An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection

ICASSP 2021accepted

This paper presents an improved mean teacher (MT) based method for large-scale weakly labeled semi-supervised sound event detection (SED), by focusing on learning a better student model. Two main improvements are proposed based on the authors’ previous perturbation based MT method. Firstly, an event…

Cited by 26SourceScholar
2021

Aspect-based Sentiment Analysis with Type-aware Graph Convolutional Networks and Layer Ensemble

NAACL 2021long

It is popular that neural graph-based models are applied in existing aspect-based sentiment analysis (ABSA) studies for utilizing word relations through dependency parses to facilitate the task with better semantic guidance for analyzing context and aspect words. However, most of these studies only…

2021

Cross-modal Memory Networks for Radiology Report Generation

ACL 2021long

Medical imaging plays a significant role in clinical practice of medical diagnosis, where the text reports of the images are essential in understanding them and facilitating later treatments. By generating the reports automatically, it is beneficial to help lighten the burden of radiologists and sig…

2021

Dependency-driven Relation Extraction with Attentive Graph Convolutional Networks

ACL 2021long

Syntactic information, especially dependency trees, has been widely used by existing studies to improve relation extraction with better semantic guidance for analyzing the context information associated with the given entities. However, most existing studies suffer from the noise in the dependency t…

2021

Improving Arabic Diacritization with Regularized Decoding and Adversarial Training

ACL 2021short

Arabic diacritization is a fundamental task for Arabic language processing. Previous studies have demonstrated that automatically generated knowledge can be helpful to this task. However, these studies regard the auto-generated knowledge instances as gold references, which limits their effectiveness…

2021

Improving Federated Learning for Aspect-based Sentiment Analysis via Topic Memories

EMNLP 2021main

Aspect-based sentiment analysis (ABSA) predicts the sentiment polarity towards a particular aspect term in a sentence, which is an important task in real-world applications. To perform ABSA, the trained model is required to have a good understanding of the contextual information, especially the part…

2021

Taming Pre-trained Language Models with N-gram Representations for Low-Resource Domain Adaptation

ACL 2021long

Large pre-trained models such as BERT are known to improve different downstream NLP tasks, even when such a model is trained on a generic domain. Moreover, recent studies have shown that when large domain-specific corpora are available, continued pre-training on domain-specific data can further impr…

2020

An Online Speaker-aware Speech Separation Approach Based on Time-domain Representation

ICASSP 2020accepted

Despite the significant progress of deep learning based speech separation methods, it remains challenging to extract and track the speech from target speakers, especially in a single-channel multiple speaker situation. Previously, the authors proposed a source-aware context network to exploit the te…

Cited by 0SourceScholar
2020

Joint Aspect Extraction and Sentiment Analysis with Directional Graph Convolutional Networks

COLING 2020main

End-to-end aspect-based sentiment analysis (EASA) consists of two sub-tasks: the first extracts the aspect terms in a sentence and the second predicts the sentiment polarities for such terms. For EASA, compared to pipeline and multi-task approaches, joint aspect extraction and sentiment analysis pro…

2020

Joint Chinese Word Segmentation and Part-of-speech Tagging via Multi-channel Attention of Character N-grams

COLING 2020main

Chinese word segmentation (CWS) and part-of-speech (POS) tagging are two fundamental tasks for Chinese language processing. Previous studies have demonstrated that jointly performing them can be an effective one-step solution to both tasks and this joint task can benefit from a good modeling of cont…

2020

Meet Changes with Constancy: Learning Invariance in Multi-Source Translation

COLING 2020main

Multi-source neural machine translation aims to translate from parallel sources of information (e.g. languages, images, etc.) to a single target language, which has shown better performance than most one-to-one systems. Despite the remarkable success of existing models, they usually neglect the fact…

2020

Summarizing Medical Conversations via Identifying Important Utterances

COLING 2020main

Summarization is an important natural language processing (NLP) task in identifying key information from text. For conversations, the summarization systems need to extract salient contents from spontaneous utterances by multiple speakers. In a special task-oriented scenario, namely medical conversat…

2020

Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection

ICASSP 2020accepted

Weakly labeled semi-supervised learning methods have recently drawn increasing attention from the research community for sound event detection tasks. Due to the weakness of the labelling, neural networks are often designed to perform sound event detection (SED) and audio tagging (AT) at the same tim…

Cited by 0SourceScholar
2019

A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification

ICASSP 2019accepted

Recently, an attention based convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) has achieved state-of-the-art performance for audio tagging (AT) and sound event detection (SED) tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challen…

Cited by 0SourceScholar
2019

Topic Detection in Conversational Telephone Speech Using CNN with Multi-stream Inputs

ICASSP 2019accepted

Topic detection for conversational telephone speech (CTS) is addressed in this paper. The low accuracy of automatic speech recognition (ASR) will cause severe performance deterioration for topic detection. To make up for this, we adopt two ASR systems, HMM-BiLSTM and CTC systems, to provide compleme…

Cited by 0SourceScholar
2018

Source-Aware Context Network for Single-Channel Multi-Speaker Speech Separation

ICASSP 2018accepted

Deep learning based approaches have achieved promising performance in speaker-dependent single-channel multispeaker speech separation. However, partly due to the label permutation problem, they may encounter difficulties in speaker-independent conditions. Recent methods address this problem by some…

Cited by 0SourceScholar
2016

Compact convolutional neural network transfer learning for small-scale image classification

ICASSP 2016accepted

Transfer learning methods have demonstrated state-of-the-art performance on various small-scale image classification tasks. This is generally achieved by exploiting the information from an ImageNet convolution neural network (ImageNet CNN). However, the transferred CNN model is generally with high c…

Cited by 0SourceScholar
2015

Improved language identification using deep bottleneck network

ICASSP 2015accepted

Effective representation plays an important role in automatic spoken language identification (LID). Recently, several representations that employ a pre-trained deep neural network (DNN) as the front-end feature extractor, have achieved state-of-the-art performance. However the performance is still f…

Cited by 0SourceScholar