← Search

Kun Yang

28 accepted papers

2026

GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward Model

ICML 2026poster

Aligning Multimodal Large Language Models (MLLMs) with human preferences remains a fundamental challenge. While Generative Reward Models (GRMs) offer a promising reasoning-based alternative to scalar models, they are often hindered by severe position bias and prohibitively high computational overhea…

Cited by 0SourceScholar
2026

LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment

CVPR 2026

We present LoD-Loc v3, a novel method for generalized aerial visual localization in dense urban environments. While prior work LoD-Loc v2 achieves localization through semantic building silhouette alignment with low-detail city models, it suffers from two key limitations: poor cross-scene generaliza

Cited by 0SourcecodeScholar
2026

MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement

AAAI 2026technical

Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers pro

Cited by 0SourcePDFScholar
2026

Planar-Sector LOS Guidance for Interception of Agile Targets with Lifting-Wing Quadcopters

ICRA 2026poster

This paper proposes a Planar-Sector Line-of-Sight (PS-LOS) guidance law and an accompanying control stack for lifting-wing quadcopters, enabling robust image-based interception of agile targets. The PS-LOS relaxes conventional conical constraints, preserving maneuverability while reducing aerodynami…

2026

Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs

ICLR 2026poster

Multimodal Large Language Models (MLLMs) pose critical safety challenges, as they are susceptible not only to adversarial attacks such as jailbreaking but also to inadvertently generating harmful content for benign users. While internal safety alignment via Supervised Fine-Tuning (SFT) and Reinforce…

Cited by 0SourceScholar
2026

Teach to Reason Safely: Policy-Guided Safety Tuning for MLRMs

ICLR 2026poster

Multimodal Large Reasoning Models (MLRMs) have exhibited remarkable capabilities in complex multimodal tasks. However, our findings reveal a critical trade-off: reasoning-based models are more prone to generating harmful content, leading to degradation in safety performance. This paper presents a la…

Cited by 0SourceScholar
2025

Adaptive Fine-Grained Feature Mining and RoI Feature Interaction Network for Small Object Detection in Aerial Images

ICASSP 2025accepted

Object detection in drone view images remains challenging due to small-scale objects distributed non-uniformly and exhibiting weak semantic features. Additionally, external factors such as lighting conditions, viewing angles, and background clutter further complicate the detection of small objects.…

Cited by 0SourceScholar
2025

Autonomous Suturing Method for Robot-Assisted Minimally Invasive Surgery

IROS 2025

Robot-assisted minimally invasive surgery is widely used because of its superior postoperative recovery outcomes. However, the workload for surgeons remains high. The development of autonomous suturing capabilities in surgical robots is poised to significantly reduce surgeon workload. In this study,

Cited by 0SourceScholar
2025

Multi-Agent Decision Transformer for Power Control in Wireless Networks

ICASSP 2025accepted

This paper introduces a novel offline approach to power control in wireless networks using a multi-agent reinforcement learning (MARL) framework. We develop a multi-agent decision transformer method to optimize performance metrics including sum-rate or packet delay. In this distributed method, each…

Cited by 0SourceScholar
2025

NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on Thermodynamics

CVPR 2025poster

Thermal infrared imaging enables a non-invasive measurement of the surface temperature of objects with all-weather applicability. Leveraging such techniques for 3D reconstruction can accurately reflect the temperature distribution of a scene, thereby supporting applications such as building monitori…

Cited by 1SourcePDFScholar
2025

Parameter-free and Accessible Prompt Learning to Enhance Adversarial Robustness for Pre-trained Vision-Language Models

NAACL 2025long

Large pre-trained Vision-Language Models (VLMs) have revolutionized both computer vision and natural language processing. Despite their success, adversarial examples can still mislead VLMs into producing incorrect results. This work focuses on boosting the adversarial robustness of VLMs by searching…

Cited by 0SourcePDFScholar
2024

A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing Modalities

AAAI 2024technical

Multimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of modality completeness. However, many inevitable factors in real-world scenarios lead to uncertain missing modalities, which invalidate the fixed multimodal fusion…

Cited by 13SourcePDFScholar
2024

Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities

CVPR 2024poster

Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However in real-world applications some practical factors cause uncertain modality missingness which drastically degrades the model's…

Cited by 15SourcePDFScholar
2024

ERMVP: Communication-Efficient and Collaboration-Robust Multi-Vehicle Perception in Challenging Environments

CVPR 2024poster

Collaborative perception enhances perception performance by enabling autonomous vehicles to exchange complementary information. Despite its potential to revolutionize the mobile industry challenges in various environments such as communication bandwidth limitations localization errors and informatio…

2024

Efficient Prompt Optimization Through the Lens of Best Arm Identification

NeurIPS 2024poster

The remarkable instruction-following capability of large language models (LLMs) has sparked a growing interest in automatically finding good prompts, i.e., prompt optimization. Most existing works follow the scheme of selecting from a pre-generated pool of candidate prompts. However, these designs m…

Cited by 7SourcePDFScholar
2024

Robust Emotion Recognition in Context Debiasing

CVPR 2024poster

Context-aware emotion recognition (CAER) has recently boosted the practical applications of affective computing techniques in unconstrained environments. Mainstream CAER methods invariably extract ensemble representations from diverse contexts and subject-centred characteristics to perceive the targ…

Cited by 23SourcePDFScholar
2024

Towards Multimodal Sentiment Analysis Debiasing via Bias Purification

ECCV 2024poster

"Multimodal Sentiment Analysis (MSA) aims to understand human intentions by integrating emotion-related clues from diverse modalities, such as visual, language, and audio. Unfortunately, the current MSA task invariably suffers from unplanned dataset biases, particularly multimodal utterance-level la…

Cited by 18SourcePDFScholar
2024

Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models

NeurIPS 2024poster

The in-context learning (ICL) capability of pre-trained models based on the transformer architecture has received growing interest in recent years. While theoretical understanding has been obtained for ICL in reinforcement learning (RL), the previous results are largely confined to the single-agent…

Cited by 1SourcePDFScholar
2023

A Novel Efficient Multi-View Traffic-Related Object Detection Framework

ICASSP 2023accepted

With the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Acc…

Cited by 0SourceScholar
2023

AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception

ICCV 2023poster

Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIs…

Cited by 54PDFcodeScholar
2023

How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent Perception

NeurIPS 2023poster

Multi-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay,…

2023

Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception

ICCV 2023poster

Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this em…

Cited by 68PDFcodeScholar
2017

Point Set Registration With Global-Local Correspondence and Transformation Estimation

ICCV 2017poster

We present a new point set registration method with global-local correspondence and transformation estimation (GL-CATE). The geometric structures of point sets are exploited by combining the global feature, the point-to-point Euclidean distance, with the local feature, the shape distance (SD) which…

Cited by 50PDFScholar