← Search

Xiangyu Kong

14 accepted papers

2026

Explainable Depression Assessment from Face Videos by Weakly Supervised Learning

AAAI 2026technical

Existing video-based automatic depression assessment (ADA) approaches frequently achieve video-level depression assessment by aggregating features or predictions of individual frames or equal-length segments within the given video. While their performances have been largely enhanced by recent advanc

Cited by 0SourcePDFScholar
2026

MECAP-R1: EMOTION-AWARE POLICY WITH REINFORCEMENT LEARNING FOR MULTIMODAL EMOTION CAPTIONING

ICASSP 2026oral

Speech Emotion Captioning (SEC) has emerged as a notable research direction. The inherent complexity of emotional content in human speech makes it challenging for traditional discrete classification methods to provide an adequate representation. Consequently, utilizing natural language to describe s…

Cited by 0SourcePDFScholar
2025

Convex MPC With Unreachable Setpoint for a Class of Affine System

RA-L 2025

We propose a convex model predictive control (MPC) scheme for a class of affine input systems to reduce the dependence on terminal components and improve real-time control capability. Artificial reference variables are introduced to handle unreachable references, and the terminal set constraint is r

Cited by 0SourceScholar
2025

Differential High Order Control Barrier Function-Based Safe Reinforcement Learning

RA-L 2025

Safe reinforcement learning (RL) aims to learn policy while also ensuring the safety constraints. An increasingly common approach is to design a safety filter based on control barrier function (CBF) or high order control barrier function (HOCBF) for the RL policy. A quadratic programming (QP) is the

Cited by 2SourceScholar
2025

Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework

ICASSP 2025accepted

Multimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this issue, we present the Cross-Modal Alignment, Reconstruction, and Refinement (CM-ARR) framework, an innovative approach…

Cited by 0SourceScholar
2025

PerReactor: Offline Personalised Multiple Appropriate Facial Reaction Generation

AAAI 2025technical

In dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strateg…

2024

Adaptive High-Order Control Barrier Function-Based Iterative LQR for Real Time Safety-Critical Motion Planning

RA-L 2024

This letter proposes an adaptive high-order control barrier function-based iterative linear quadratic regulator (AHOCBF-ILQR) algorithm for real time safety-critical motion planning. Firstly, we propose a HOCBF-ILQR method, where a HOCBF-based controller is designed as a safety filter of ILQR to gua

Cited by 8SourceScholar
2024

CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making Agents

ICLR 2024spotlight

The generalization of decision-making agents encompasses two fundamental elements: learning from past experiences and reasoning in novel contexts. However, the predominant emphasis in most interactive environments is on learning, often at the expense of complexity in reasoning. In this paper, we int…

2024

Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy

NeurIPS 2024poster

Diplomacy is one of the most sophisticated activities in human society, involving complex interactions among multiple parties that require skills in social reasoning, negotiation, and long-term strategic planning. Previous AI agents have demonstrated their ability to handle multi-step games and larg…

2024

Semi-Supervised Volumetric Medical Image Segmentation via Class Prototype Guided Distribution-Aligned Representation Learning

ICASSP 2024accepted

We present SemiCRL, a novel framework for volumetric medical image segmentation that formulates an innovative contrastive learning methodology in a semi-supervised learning setting. We leverage the pseudo-labels generated in semi-supervised learning to guide the selection of negative samples for our…

Cited by 0SourceScholar
2023

Conflict-Based Cross-View Consistency for Semi-Supervised Semantic Segmentation

CVPR 2023poster

Semi-supervised semantic segmentation (SSS) has recently gained increasing research interest as it can reduce the requirement for large-scale fully-annotated training data. The current methods often suffer from the confirmation bias from the pseudo-labelling process, which can be alleviated by the c…

2023

Dasformer: Deep Alternating Spectrogram Transformer For Multi/Single-Channel Speech Separation

ICASSP 2023accepted

For the task of speech separation, previous study usually treats multi-channel and single-channel scenarios as two research tracks with specialized solutions developed respectively. Instead, we propose a simple and unified architecture - DasFormer (Deep alternating spectrogram transFormer) to handle…

Cited by 0SourceScholar