← Search

Kun Qian

36 accepted papers

2026

Bridging the Modality Reliability Gap in Drug-Target Interaction Prediction via a Confidence-aware Multimodal Fusion Framework

AAAI 2026technical

With the rapid advancement of deep learning, drug target interaction (DTI) prediction has seen substantial performance enhancements. However, existing methodologies face a critical, yet unaddressed challenge, i.e., the Modality Reliability Gap. Such a gap arises from the unpredictable variance in t

Cited by 0SourcePDFScholar
2026

SFT Doesn’t Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs

ICLR 2026poster

Supervised Fine-Tuning (SFT) on domain-specific datasets is a common approach to adapt Large Language Models (LLMs) to specialized tasks but is often believed to degrade their general capabilities. In this work, we revisit this trade-off and present both empirical and theoretical insights. First, we…

Cited by 0SourceScholar
2026

Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning

ICLR 2026poster

Large Language Models (LLMs) have recently demonstrated strong potential in generating ‘believable human-like’ behavior in web environments. Prior work has explored augmenting training data with LLM-synthesized rationales and applying supervised fine-tuning (SFT) to enhance reasoning ability, which…

Cited by 0SourcecodeScholar
2025

Bottom-Up Synthesis of Knowledge-Grounded Task-Oriented Dialogues with Iteratively Self-Refined Prompts

NAACL 2025short

Training conversational question-answering (QA) systems demands a substantial amount of in-domain data, which is often scarce in practice. A common solution to this challenge is to generate synthetic data. Traditional methods typically follow a top-down approach, where a large language model (LLM) g…

Cited by 0SourcePDFScholar
2025

Bridging the Data Provenance Gap Across Text, Speech, and Video

ICLR 2025poster

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and first-of-its-kind longitudinal audit across modalities --- pop…

Cited by 1SourcePDFScholar
2025

Evaluation and Incident Prevention in an Enterprise AI Assistant

AAAI 2025technical

Enterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a comprehensive framework for monitoring, benchmarking, and continuously improving such complex, multi-component systems und…

Cited by 0SourcePDFScholar
2025

LBSNet: Lightweight Joint Boundary Detection and Semantic Segmentation for Transparent and Reflective Objects

RA-L 2025

Accurate visual detection of transparent and reflective objects remains a challenging issue for mobile manipulators. For the most common depth cameras and LiDAR sensors, the distinctive optical attributes inherent in both transparent and reflective objects pose a significant challenge. To address th

Cited by 5SourceScholar
2025

LiLoc: Lifelong Localization Using Adaptive Submap Joining and Egocentric Factor Graph

ICRA 2025

This paper proposes a versatile graph-based lifelong localization framework using LiDAR, LiLoc, which enhances its timeliness by maintaining a single central session while improves the accuracy through multi-modal factors between the central and subsidiary sessions. First, an adaptive submap joining

Cited by 3SourcecodeScholar
2025

Radio Frequency Ray Tracing with Neural Object Representation for Enhanced RF Modeling

CVPR 2025poster

Radio frequency (RF) propagation modeling poses unique electromagnetic simulation challenges. While recent neural representations have shown success in visible spectrum rendering, the fundamentally different scales and physics of RF signals require novel modeling paradigms. In this paper, we introdu…

Cited by 0SourcePDFScholar
2024

Consent in Crisis: The Rapid Decline of the AI Data Commons

NeurIPS 2024poster

General-purpose artificial intelligence (AI) systems are built on massive swathes of public web data, assembled into corpora such as C4, RefinedWeb, and Dolma. To our knowledge, we conduct the first, large-scale, longitudinal audit of the consent protocols for the web domains underlying AI training…

Cited by 36SourceScholar
2024

DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and Rewriting

EMNLP 2024main

Coherence in writing, an aspect that L2 English learners often struggle with, is crucial in assessing L2 English writing. Existing automated writing evaluation systems primarily use basic surface linguistic features to detect coherence in writing. However, little effort has been made to correct the…

2024

Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation

ICASSP 2024accepted

Medical image segmentation is an important task in modern analysis of medical images. Current methods tend to extract either local features with convolutions or global features with Transformers. However, few of them are able to effectively fuse global and local features to facilitate segmentation.…

Cited by 0SourceScholar
2024

FOTS: A Fast Optical Tactile Simulator for Sim2Real Learning of Tactile-Motor Robot Manipulation Skills

RA-L 2024

Simulation is a widely used tool in robotics to reduce hardware consumption and gather large-scale data. Despite previous efforts to simulate optical tactile sensors, there remain challenges in efficiently synthesizing images and replicating marker motion under different contact loads. In this work,

Cited by 21SourcecodeScholar
2024

Scalable Network and Adaptive Refinement Module for 6D Pose Estimation of Diverse Industrial Components*

IROS 2024poster

The estimation of the 6D pose of industrial components is essential for smart manufacturing. Especially for complex units that require intensive manual operations, such as a concentrator photovoltaics solar panel, accurate spatial localization provides visual aids for industrial automation. In this…

Cited by 0SourceScholar
2024

Time Sensitive Knowledge Editing through Efficient Finetuning

ACL 2024short

Large Language Models (LLMs) have demonstrated impressive capability in different tasks and are bringing transformative changes to many domains. However, keeping the knowledge in LLMs up-to-date remains a challenge once pretraining is complete. It is thus essential to design effective methods to bot…

2024

VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation

EMNLP 2024finding

As large language models achieve impressive scores on traditional benchmarks, an increasing number of researchers are becoming concerned about benchmark data leakage during pre-training, commonly known as the data contamination problem. To ensure fair evaluation, recent benchmarks release only the t…

2023

Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning Analysis

ICASSP 2023accepted

Translating mental health recognition from clinical research into real-world application requires extensive data, yet existing emotion datasets are impoverished in terms of daily mental health monitoring, especially when aiming for self-reported anxiety and depression recognition. We introduce the J…

Cited by 0SourceScholar
2023

Data-Driven Adaptive Iterative Learning Control of a Compliant Rehabilitation Robot for Repetitive Ankle Training

RA-L 2023

This letter investigates the repetitive range of motion (ROM) training control for a compliant ankle rehabilitation robot (CARR). The CARR utilizes four pneumatic muscle (PM) actuators to manipulate the ankle with three rational degree-of-freedoms (DoFs) and soft human-robot interaction, but the str

Cited by 21SourceScholar
2023

Federated Intelligent Terminals Facilitate Stuttering Monitoring

ICASSP 2023accepted

Stuttering is a complicated language disorder. The most common form of stuttering is developmental stuttering, which begins in childhood. Early monitoring and intervention are essential for the treatment of children with stuttering. Automatic speech recognition technology has shown its great potenti…

Cited by 0SourceScholar
2023

GVGNet: Gaze-Directed Visual Grounding for Learning Under-Specified Object Referring Intention

RA-L 2023

Referring Expression Comprehension (REC) and Referring Expression Segmentation (RES) enable robots to infer human's object referring intention through natural languages. In this letter, Gaze-directed Visual Grounding Network (GVGNet) is proposed to disambiguate human's under-specified object referri

Cited by 12SourceScholar
2023

KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords Learning

EMNLP 2023long main

In task-oriented dialogs (TOD), reinforcement learning (RL) algorithms train a model to directly optimize response for task-related metrics. However, RL often needs to perform exploration, which can be time-consuming due to the slow auto-regressive sequence generation process. We investigate an appr…

Cited by 0SourcecodeScholar
2023

Knowledge Transfer for on-Device Speech Emotion Recognition With Neural Structured Learning

ICASSP 2023accepted

Speech emotion recognition (SER) has been a popular research topic in human-computer interaction (HCI). As edge devices are rapidly springing up, applying SER to edge devices is promising for a huge number of HCI applications. Although deep learning has been investigated to improve the performance o…

Cited by 0SourceScholar
2022

A Glance-and-Gaze Network for Respiratory Sound Classification

ICASSP 2022accepted

A plethora of great successes has been achieved by the existing convolutional neural networks (CNN) for respiratory sound classification. Nevertheless, simultaneously capturing both the local and global features can never be an easy task due to the limitation of a CNN’s structure. In this contributi…

Cited by 0SourceScholar
2022

An Overview of the FIRST ICASSP Special Session on Computer Audition for Healthcare

ICASSP 2022accepted

Audio has been increasingly used as a novel digital phenotype that carries important information of the subject’s health status. We can find tremendous efforts given to this young and promising field, i.e., computer audition for healthcare (CA4H), whereas the application scenarios have not been full…

Cited by 0SourceScholar
2022

Database Search Results Disambiguation for Task-Oriented Dialog Systems

NAACL 2022long

As task-oriented dialog systems are becoming increasingly popular in our lives, more realistic tasks have been proposed and explored. However, new practical challenges arise. For instance, current dialog systems cannot effectively handle multiplesearch results when querying a database, due to the la…

Cited by 20SourcePDFScholar
2022

MMDF: Multi-Modal Deep Feature Based Place Recognition of Mobile Robots With Applications on Cross-Scene Navigation

RA-L 2022

Although the navigation of robots in urban environments has achieved great performance, there is still a problem of insufficient robustness in cross-scene (ground, water surface) navigation applications. An intuitive idea is to introduce multi-modal complementary data to improve the robustness of th

Cited by 16SourceScholar
2021

A Student-Teacher Architecture for Dialog Domain Adaptation Under the Meta-Learning Setting

AAAI 2021technical

Numerous new dialog domains are being created every day while collecting data for these domains is extremely costly since it involves human interactions. Therefore, it is essential to develop algorithms that can adapt to different domains efficiently when building data-driven dialog models. Most rec…

Cited by 7SourcePDFScholar
2021

Multi-Robot Dynamical Source Seeking in Unknown Environments

ICRA 2021poster

This paper presents an algorithmic framework for the distributed on-line source seeking, termed as DoSS, with a multi-robot system in an unknown dynamical environment. Our algorithm, building on a novel concept called dummy confidence upper bound (D-UCB), integrates both estimation of the unknown en…

Cited by 19SourceScholar
2021

Robust Iterative Learning Control for Pneumatic Muscle with State Constraint and Model Uncertainty

ICRA 2021poster

In this paper, we propose a novel iterative learning control (ILC) scheme for precise state tracking of pneumatic muscle (PM) actuators. Two critical issues are considered in our scheme: 1) state constraints on PM position and velocity; 2) uncertainties of the PM model. Based on the three-element fo…

Cited by 2SourceScholar
2021

Robust Multimodal Vehicle Detection in Foggy Weather Using Complementary Lidar and Radar Signals

CVPR 2021poster

Vehicle detection with visual sensors like lidar and camera is one of the critical functions enabling autonomous driving. While they generate fine-grained point clouds or high-resolution images with rich information in good weather conditions, they fail in adverse weather (e.g., fog) where opaque pa…

Cited by 193PDFcodeScholar
2021

Two-stream 2D/3D Residual Networks for Learning Robot Manipulations from Human Demonstration Videos

ICRA 2021poster

Learning manipulation skills from observing human demonstration videos is a promising aspect for intelligent robotic systems. Recent advances in video to command provide an end-to-end approach to translate a video into robot plans. However, the general video captioning methods focus more on the unde…

Cited by 9SourceScholar
2020

Social-STGCNN: A Social Spatio-Temporal Graph Convolutional Neural Network for Human Trajectory Prediction

CVPR 2020poster

Better machine understanding of pedestrian behaviors enables faster progress in modeling interactions between agents such as autonomous vehicles and humans. Pedestrian trajectories are not only influenced by the pedestrian itself but also by interaction with surrounding objects. Previous methods mod…

Cited by 1032PDFcodeScholar
2016

Wavelet features for classification of vote snore sounds

ICASSP 2016accepted

Location and form of the upper airway obstruction is essential for a targeted therapy of obstructive sleep apnea (OSA). Utilizing snore sounds (SnS) to reveal the pathological characters of OSA patients has been the subject of scientific research for several decades. Fewer studies exist on the evalu…

Cited by 0SourceScholar