← Search

Kun Huang

30 accepted papers

2026

CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning

ICML 2026poster

Mobile Agents can autonomously execute user instructions, which requires hybrid-capabilities reasoning, including screen summary, subtask planning, action decision and action function. However, existing agents struggle to achieve both decoupled enhancement and balanced integration of these capabilit…

Cited by 0SourceScholar
2026

MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning

ICLR 2026poster

The Chain of Action-Planning Thoughts (CoaT) paradigm has been shown to improve the reasoning performance of VLM-based mobile agents in GUI tasks. However, the scarcity of diverse CoaT trajectories limits the expressiveness and generalization ability of such agents. While self-training is commonly e…

Cited by 0SourceScholar
2025

MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity Recognition

EMNLP 2025

Grounded Multimodal Named Entity Recognition (GMNER), which aims to extract textual entities, their types, and corresponding visual regions from image-text data, has become a critical task in multimodal information extraction. However, existing methods face two major challenges. First, they fail to

2024

Adaptive Data Augmentation for Aspect Sentiment Quad Prediction

ICASSP 2024accepted

Aspect sentiment quad prediction (ASQP) aims to predict the quad sentiment elements for a given sentence, which is a critical task in the field of aspect-based sentiment analysis. However, the data imbalance issue has not received sufficient attention in ASQP task. In this paper, we divide the issue…

Cited by 0SourceScholar
2024

Model-Based Label-to-Image Diffusion for Semi-Supervised Choroidal Vessel Segmentation

ICASSP 2024accepted

Current successful choroidal vessel segmentation methods rely on large amounts of voxel-level annotations on the 3D optical coherence tomography images, which are hard and time-consuming. Semi-supervised learning solves this issue by enabling model learning from both unlabeled data and a limited amo…

Cited by 0SourceScholar
2023

Enhancing Personalized Dialogue Generation with Contrastive Latent Variables: Combining Sparse and Dense Persona

ACL 2023long

The personalized dialogue explores the consistent relationship between dialogue generation and personality. Existing personalized dialogue agents model persona profiles from three resources: sparse or dense persona descriptions and dialogue histories. However, sparse structured persona attributes ar…

2023

Guiding Dialogue Agents to Complex Semantic Targets by Dynamically Completing Knowledge Graph

ACL 2023findings

In the target-oriented dialogue, the representation and achievement of targets are two interrelated essential issues. In current approaches, the target is typically supposed to be a single object represented as a word, which makes it relatively easy to achieve the target through dialogue with the he…

2023

Learning Sparse Group Models Through Boolean Relaxation

ICLR 2023top-25%

We introduce an efficient algorithmic framework for learning sparse group models formulated as the natural convex relaxation of a cardinality-constrained program with Boolean variables. We provide theoretical techniques to characterize the equivalent condition when the relaxation achieves the exact…

Cited by 0SourcePDFScholar
2023

MTGP: Multi-turn Target-oriented Dialogue Guided by Generative Global Path with Flexible Turns

ACL 2023findings

Target-oriented dialogue guides the dialogue to a target quickly and smoothly. The latest approaches focus on global planning, which plans toward the target before the conversation instead of adopting a greedy strategy during the conversation. However, the global plan in existing works is fixed to c…

2023

Point Clouds Outlier Removal Method Based on Improved Mahalanobis and Completion

RA-L 2023

Point clouds have been regarded as a representative format for 3D visualization of real-world objects or scenes. However, point clouds acquired from depth cameras or laser scanning devices commonly contain outliers. Outlier removal performance will directly affect the downstream applications. Existi

Cited by 11SourceScholar
2023

Self-Supervised Boundary Point Prediction Task for Point Cloud Domain Adaptation

RA-L 2023

Unsupervised domain adaptation (UDA) could significantly improve the cross-domain performance of current supervised 3D deep learning methods and have a widespread application prospect. However, the domain gap between source domain and target domain renders the UDA problem highly challenging. In this

Cited by 10SourceScholar
2023

Towards Efficient Pre-Trained Language Model via Feature Correlation Distillation

NeurIPS 2023poster

Knowledge Distillation (KD) has emerged as a promising approach for compressing large Pre-trained Language Models (PLMs). The performance of KD relies on how to effectively formulate and transfer the knowledge from the teacher model to the student model. Prior arts mainly focus on directly aligning…

Cited by 4SourcePDFScholar
2023

Uncertainty Guided Label Denoising for Document-level Distant Relation Extraction

ACL 2023long

Document-level relation extraction (DocRE) aims to infer complex semantic relations among entities in a document. Distant supervision (DS) is able to generate massive auto-labeled data, which can improve DocRE performance. Recent works leverage pseudo labels generated by the pre-denoising model to r…

2022

Accurate Calibration of Multi-Perspective Cameras from a Generalization of the Hand-Eye Constraint

ICRA 2022poster

Multi-perspective cameras are quickly gaining importance in many applications such as smart vehicles and virtual or augmented reality. However, a large system size or absence of overlap in neighbouring fields-of-view often complicate their calibration. We present a novel solution which relies on the…

Cited by 10SourcecodeScholar
2022

Aligning Recommendation and Conversation via Dual Imitation

EMNLP 2022main

Human conversations of recommendation naturally involve the shift of interests which can align the recommendation actions and conversation process to make accurate recommendations with rich explanations. However, existing conversational recommendation systems (CRS) ignore the advantage of user inter…

Cited by 8SourcePDFScholar
2022

CR-GIS: Improving Conversational Recommendation via Goal-aware Interest Sequence Modeling

COLING 2022main

Conversational recommendation systems (CRS) aim to determine a goal item by sequentially tracking users’ interests through multi-turn conversation. In CRS, implicit patterns of user interest sequence guide the smooth transition of dialog utterances to the goal item. However, with the convenient expl…

Cited by 7SourcePDFScholar
2022

Deep 360° Optical Flow Estimation Based on Multi-Projection Fusion

ECCV 2022poster

"Optical flow computation is essential in the early stages of the video processing pipeline. This paper focuses on a less explored problem in this area, the 360° optical flow estimation using deep neural networks to support the increasingly popular VR applications. To address the distortions of pan…

Cited by 24SourcePDFScholar
2022

Know Thyself: Transferable Visual Control Policies Through Robot-Awareness

ICLR 2022poster

Training visual control policies from scratch on a new robot typically requires generating large amounts of robot-specific data. How might we leverage data previously collected on another robot to reduce or even completely remove this need for robot-specific data? We propose a "robot-aware control"…

2022

TopKG: Target-oriented Dialog via Global Planning on Knowledge Graph

COLING 2022main

Target-oriented dialog aims to reach a global target through multi-turn conversation. The key to the task is the global planning towards the target, which flexibly guides the dialog concerning the context. However, existing target-oriented dialog works take a local and greedy strategy for response g…

2022

Training Robots to Evaluate Robots: Example-Based Interactive Reward Functions for Policy Learning

CoRL 2022oral

Physical interactions can often help reveal information that is not readily apparent. For example, we may tug at a table leg to evaluate whether it is built well, or turn a water bottle upside down to check that it is watertight. We propose to train robots to acquire such interactive behaviors autom…

Cited by 4SourcecodeScholar
2021

B-splines for Purely Vision-based Localization and Mapping on Non-holonomic Ground Vehicles

ICRA 2021poster

Purely vision-based localization and mapping is a cost-effective and thus attractive solution to localization and mapping on smart ground vehicles. However, the accuracy and especially robustness of vision-only solutions remain rivalled by more expensive, lidar-based multi-sensor alternatives. We sh…

Cited by 8SourceScholar
2021

Fast Projection onto the Capped Simplex with Applications to Sparse Regression in Bioinformatics

NeurIPS 2021poster

We consider the problem of projecting a vector onto the so-called k-capped simplex, which is a hyper-cube cut by a hyperplane. For an n-dimensional input vector with bounded elements, we found that a simple algorithm based on Newton's method is able to solve the projection problem to high precision…

Cited by 11SourcePDFScholar
2021

Fusing RGBD Tracking and Segmentation Tree Sampling for Multi-Hypothesis Volumetric Segmentation

ICRA 2021poster

Despite rapid progress in scene segmentation in recent years, 3D segmentation methods are still limited when there is severe occlusion. The key challenge is estimating the segment boundaries of (partially) occluded objects, which are inherently ambiguous when considering only a single frame. In this…

Cited by 2SourcecodeScholar
2021

Transfer Learning via Optimal Transportation for Integrative Cancer Patient Stratification

IJCAI 2021poster

The Stratification of early-stage cancer patients for the prediction of clinical outcome is a challenging task since cancer is associated with various molecular aberrations. A single biomarker often cannot provide sufficient information to stratify early-stage patients effectively. Understanding the…

Cited by 5SourcePDFScholar
2020

Reliable frame-to-frame motion estimation for vehicle-mounted surround-view camera systems

ICRA 2020poster

Modern vehicles are often equipped with a surround-view multi-camera system. The current interest in autonomous driving invites the investigation of how to use such systems for a reliable estimation of relative vehicle displacement. Existing camera pose algorithms either work for a single camera, ma…

Cited by 12SourceScholar
2019

Motion Estimation of Non-Holonomic Ground Vehicles From a Single Feature Correspondence Measured Over N Views

CVPR 2019poster

The planar motion of ground vehicles is often non-holonomic, which enables a solution of the two-view relative pose problem from a single point feature correspondence. Man-made environments such as underground parking lots are however dominated by line features. Inspired by the planar tri-focal tens…

Cited by 16PDFScholar
2019

PPGNet: Learning Point-Pair Graph for Line Segment Detection

CVPR 2019poster

In this paper, we present a novel framework to detect line segments in man-made environments. Specifically, we propose to describe junctions, line segments and relationships between them with a simple graph, which is more structured and informative than end-point representation used in existing line…

Cited by 110PDFcodeScholar
2018

Learning to Parse Wireframes in Images of Man-Made Environments

CVPR 2018poster

In this paper, we propose a learning-based approach to the task of automatically extracting a "wireframe" representation for images of cluttered man-made environments. The wireframe contains all salient straight lines and their junctions of the scene that encode efficiently and accurately large-scal…