← Search

Yan Zhou

28 accepted papers

2026

Online Velocity Estimation of a Robotic Fish Using Artificial Lateral Line System with Velocity-Decoupling Sensing Ability

ICRA 2026poster

The robotic fish has attracted widespread research interest over the past few decades, due to its outstanding agility and environmental friendliness. And the sensing ability of underwater environments is crucial for the robotic fish to accomplish various underwater tasks. Inspired by the lateral lin…

Cited by 0SourceScholar
2026

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation

CVPR 2026

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal interaction and limited modal diversity for comprehensive world knowledge represe

Cited by 0SourcecodeScholar
2025

Boosting Lightweight Camouflaged Object Detection with Multi-Scale Context and Boundary Awareness

ICASSP 2025accepted

To adapt to the resource-limited environment, this study introduces the lightweight boundary-aware camouflaged object detection(COD) network LMABnet. We enhance the feature representation capability of the lightweight network through a multi-scale feature fusion architecture, while effectively avoid…

Cited by 0SourceScholar
2025

Can Multimodal Large Language Models Understand Spatial Relations?

ACL 2025long

Spatial relation reasoning is a crucial task for multimodal large language models (MLLMs) to understand the objective world. However, current benchmarks have issues like relying on bounding boxes, ignoring perspective substitutions, or allowing questions to be answered using only the model’s prior k…

2025

Certainty-guided Reasoning and Refinement Network for Camouflaged Object Detection

ICASSP 2025accepted

Camouflaged object detection (COD), which aims to segment objects that are highly similar to their background, is a valuable yet challenging task. Due to the interference of clutter and noise in the background, existing methods often struggle to avoid misleading and accurately segment the camouflage…

Cited by 0SourceScholar
2025

Cypher-RI: Reinforcement Learning for Integrating Schema Selection into Cypher Generation

NeurIPS 2025poster

The increasing utilization of graph databases across various fields stems from their capacity to represent intricate interconnections. Nonetheless, exploiting the full capabilities of graph databases continues to be a significant hurdle, largely because of the inherent difficulty in translating natu…

Cited by 0SourceScholar
2025

Joint Edge and Regional Depth Enhancement Network for Camouflaged Object Detection

ICASSP 2025accepted

Camouflaged object detection (COD) is a task of identifying and locating target objects that are camouflaged, masked, or confused. Research claims that depth cues can provide effective object location cues. However, depth images often contain noise interference, which may negatively affect object re…

Cited by 0SourceScholar
2025

LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

ACL 2025long

Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction. Recent advancements have showcased the potential of building intelligent spoken chatbots based on large language models (LLMs). In this paper, we introduce LLaMA-Omni 2, a s…

2025

LLaMA-Omni: Seamless Speech Interaction with Large Language Models

ICLR 2025poster

Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how to build speech interaction models based on open-source LLMs.…

2025

Multi-scale Re-weighted Attention Feature Fusion for Non-Intrusive Load Monitoring

ICASSP 2025accepted

Non-Intrusive Load Monitoring (NILM) addresses the challenge of disaggregating total energy consumption into individual appliance usage, which is essential for enhancing energy efficiency and managing smart grids. Existing methods often overlook the impact of window sizes on the separation of applia…

Cited by 0SourceScholar
2025

N-ForGOT: Towards Not-forgetting and Generalization of Open Temporal Graph Learning

ICLR 2025poster

Temporal Graph Neural Networks (TGNNs) lay emphasis on capturing node interactions over time but often overlook evolution in node classes and dynamic data distributions triggered by the continuous emergence of new class labels, known as the open-set problem. This problem poses challenges for existin…

Cited by 0SourcePDFScholar
2024

CTC-based Non-autoregressive Textless Speech-to-Speech Translation

ACL 2024findings

Direct speech-to-speech translation (S2ST) has achieved impressive translation quality, but it often faces the challenge of slow decoding due to the considerable length of speech sequences. Recently, some research has turned to non-autoregressive (NAR) models to expedite decoding, yet the translatio…

2024

Deep Reinforcement Learning for Modelling Protein Complexes

ICLR 2024poster

Structure prediction of large protein complexes (a.k.a., protein multimer mod- elling, PMM) can be achieved through the one-by-one assembly using provided dimer structures and predicted docking paths. However, existing PMM methods struggle with vast search spaces and generalization challenges: (1) T…

Cited by 1SourcePDFScholar
2024

Fast Graph Sharpness-Aware Minimization for Enhancing and Accelerating Few-Shot Node Classification

NeurIPS 2024poster

Graph Neural Networks (GNNs) have shown superior performance in node classification. However, GNNs perform poorly in the Few-Shot Node Classification (FSNC) task that requires robust generalization to make accurate predictions for unseen classes with limited labels. To tackle the challenge, we propo…

2024

Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction

EMNLP 2024main

Automated red teaming is an effective method for identifying misaligned behaviors in large language models (LLMs). Existing approaches, however, often focus primarily on improving attack success rates while overlooking the need for comprehensive test case coverage. Additionally, most of these method…

2024

Using AI Uncertainty Quantification to Improve Human Decision-Making

ICML 2024poster

AI Uncertainty Quantification (UQ) has the potential to improve human decision-making beyond AI predictions alone by providing additional probabilistic information to users. The majority of past research on AI and human decision-making has concentrated on model explainability and interpretability, w…

Cited by 12SourcePDFScholar
2023

DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation

NeurIPS 2023poster

Direct speech-to-speech translation (S2ST) translates speech from one language into another using a single model. However, due to the presence of linguistic and acoustic diversity, the target speech follows a complex multimodal distribution, posing challenges to achieving both high-quality translati…

2023

Dichotomous Image Segmentation with Frequency Priors

IJCAI 2023poster

Dichotomous image segmentation (DIS) has a wide range of real-world applications and gained increasing research attention in recent years. In this paper, we propose to tackle DIS with informative frequency priors. Our model, called FP-DIS, stems from the fact that prior knowledge in the frequency do…

2023

Hierarchical Spatial-Temporal Transformer with Motion Trajectory for Individual Action and Group Activity Recognition

ICASSP 2023accepted

Group activity recognition, which aims to simultaneously understand individual action and group activity in video clips, plays a fundamental role in video analysis. In this paper, we propose a novel reasoning network, Hierarchical Spatial-Temporal Transformer termed HSTT, for individual action and g…

Cited by 0SourceScholar
2023

QAP: A Quantum-Inspired Adaptive-Priority-Learning Model for Multimodal Emotion Recognition

ACL 2023findings

Multimodal emotion recognition for video has gained considerable attention in recent years, in which three modalities (i.e., textual, visual and acoustic) are involved. Due to the diverse levels of informational content related to emotion, three modalities typically possess varying degrees of contri…

Cited by 16SourcePDFScholar
2023

TrojanSQL: SQL Injection against Natural Language Interface to Database

EMNLP 2023long main

The technology of text-to-SQL has significantly enhanced the efficiency of accessing and manipulating databases. However, limited research has been conducted to study its vulnerabilities emerging from malicious user interaction. By proposing TrojanSQL, a backdoor-based SQL injection framework for t…

Cited by 0SourceScholar
2022

AMOA: Global Acoustic Feature Enhanced Modal-Order-Aware Network for Multimodal Sentiment Analysis

COLING 2022main

In recent years, multimodal sentiment analysis (MSA) has attracted more and more interest, which aims to predict the sentiment polarity expressed in a video. Existing methods typically 1) treat three modal features (textual, acoustic, visual) equally, without distinguishing the importance of differe…

Cited by 25SourcePDFScholar
2022

Evolutionary Neural Architecture Design of Liquid State Machine for Image Classification

ICASSP 2022accepted

As a recurrent spiking neural network, liquid state machine (LSM) has attracted more and more attention in neuromorphic computing due to its biological plausibility, computation power, and hardware implementation. However, the neural architecture of LSM, such as hidden neuron number, synaptic densit…

Cited by 0SourceScholar
2021

An Adaptive Hybrid Framework for Cross-domain Aspect-based Sentiment Analysis

AAAI 2021technical

Cross-domain aspect-based sentiment analysis aims to utilize the useful knowledge in a source domain to extract aspect terms and predict their sentiment polarities in a target domain. Recently, methods based on adversarial training have been applied to this task and achieved promising results. In su…

Cited by 35SourcePDFScholar
2021

Does Explainable Artificial Intelligence Improve Human Decision-Making?

AAAI 2021technical

Explainable AI provides insights to users into the why for model predictions, offering potential for users to better understand and trust a model, and to recognize and correct AI predictions that are incorrect. Prior research on human and explainable AI interactions has focused on measures such as i…

Cited by 164SourcePDFScholar
2020

The Compressed Nested Array for Underdetermined DOA Estimation by Fourth-order Difference Coarrays

ICASSP 2020accepted

In this paper, a new sparse array structure, which further improves the degrees of freedom (DOFs) and enhanced the DOA estimation performance, for the fourth-order cumulant based direction of arrival (DOA) estimation is proposed. The new-formed array is hole-free and can achieve a large consecutive…

Cited by 0SourceScholar