← Search

Kun LI

83 accepted papers

2026

CG-Floor: Centroid-Guided Diffusion for Large-Scale Floorplan Generation

CVPR 2026

Large-scale floorplan generation is critical for virtual space planning and architectural simulation. Although existing methods have shown success in generating small-scale floorplans with simple room shapes, they struggle to handle complex room connections and irregular room shapes that arise in la

Cited by 0SourceScholar
2026

Can Molecular Evolution Mechanism Enhance Molecular Representation?

AAAI 2026technical

Molecular evolution is the process of simulating the natural evolution of molecules in chemical space to explore potential molecular structures and properties. The relationships between similar molecules are often described through transformations such as adding, deleting, and modifying atoms and ch

Cited by 0SourcePDFScholar
2026

CoMA-SLAM: Collaborative Multi-Agent Gaussian SLAM with Geometric Consistency

AAAI 2026technical

Although Gaussian scene representation has achieved remarkable success in tracking and mapping, most existing methods are confined to single-agent systems. Current multi-agent solutions typically rely on centralized architectures, which struggle to account for communication bandwidth constraints. Fu

Cited by 0SourcePDFScholar
2026

Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction

ICML 2026poster

Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due to severe depth ambiguity and complex scene geometry. Existing monocular crowd reconstruction methods typically rely on single-plane assumptions, leading to unreliable metric scale and spat…

Cited by 0SourceScholar
2026

DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax Constraints

AAAI 2026technical

Dual-lens video inpainting aims to simultaneously restore missing or corrupted contents in videos captured by each lens of binocular systems. Although preliminary explorations have been conducted, existing methods still face two key challenges: limited exploitation of long-range reference informatio

Cited by 0SourcePDFScholar
2026

InterCoser: Interactive 3D Character Creation with Disentangled Fine-Grained Features

AAAI 2026technical

This paper aims to interactively generate and edit disentangled 3D characters based on precise user instructions. Existing methods generate and edit 3D characters via rough and simple editing guidance and entangled representations, making it difficult to achieve precise and comprehensive control ove

Cited by 0SourcePDFScholar
2026

MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models

CVPR 2026

Post-training quantization (PTQ) with computational equivalence for Large Language Models (LLMs) have demonstrated remarkable advances, however, their application to Multimodal Large Language Models (MLLMs) presents substantial challenges. In this paper, we analyze SmoothQuant as a case study and id

Cited by 0SourcecodeScholar
2026

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

CVPR 2026

Predicting enzyme kinetic parameters quantifies how efficiently an enzyme catalyzes a specific substrate under defined biochemical conditions. Canonical parameters such as the turnover number (k_\text cat ), Michaelis constant (K_\text m ), and inhibition constant (K_\text i ) depend jointly on the

Cited by 0SourceScholar
2026

OrionEdit: Bridging Reference and Source Images for Generalized Cross-Image Editing

CVPR 2026

Multimodal image synthesis has made significant progress, yet most editing methods still rely on textual instructions, which are less direct than visual guidance. Recently, a new paradigm edits one image using another as reference, enabling more intuitive manipulation through visual exemplars. We fo

Cited by 0SourcecodeScholar
2026

PCEvo: Path-Consistent Molecular Representation via Virtual Evolutionary

IJCAI 2026

Molecular representation learning aims to learn vector embeddings that capture molecular structure and geometry, thereby enabling property prediction and downstream scientific applications. In many AI for science tasks, labeled data are expensive to obtain and therefore limited in availability. Unde

Cited by 0Scholar
2026

Query-Guided Spatial–Temporal–Frequency Interaction for Music Audio–Visual Question Answering

ICLR 2026poster

Audio–Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natural language questions. Inspired by recent advances in Video QA, many existing AVQA approaches primarily focus on visual…

Cited by 0SourcecodeScholar
2026

Sequence-Free for Compound Protein Interaction Prediction

AAAI 2026technical

The prediction of compound–protein interactions (CPIs) is crucial for drug discovery. Most existing CPI prediction models rely on protein sequence information as input. However, in early-stage drug development, particularly in phenotype-driven studies or compound-response analyses, proteins are oft

Cited by 0SourcePDFScholar
2026

Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation

CVPR 2026

Co-manipulation requires multiple humans to synchronize their motions with a shared object while ensuring reasonable interactions, maintaining natural poses, and preserving stable states. However, most existing motion generation approaches are designed for single-character scenarios or fail to accou

Cited by 0SourceScholar
2026

VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video Reasoning

CVPR 2026

Reinforcement learning (RL) has emerged as an effective approach for improving video reasoning in multimodal large language models (MLLMs). However, existing methods remain inefficient for two reasons. First, training data are typically organized by task formats rather than underlying reasoning abil

Cited by 0SourcecodeScholar
2026

VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLM) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "Visual Prompts" (VP) like bounding box

Cited by 0SourcePDFScholar
2025

A Kinematic Constrained Batch Informed Trees Algorithm With Varied Density Sampling for Mobile Robot Path Planning

RA-L 2025

we proposed a novel Kinematic Batch Informed Trees algorithm (K-BIT*) to solve problems of the low efficiency, poor geometric smoothness and local optimum when conducting path planning for mobile robots. A variable density sampling strategy is designed which can automatically adjust the searching ra

Cited by 3SourceScholar
2025

AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment

EMNLP 2025

Multimodal Large Language Models (MLLMs) are increasingly applied in Personalized Image Aesthetic Assessment (PIAA) as a scalable alternative to expert evaluations. However, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education. In this work

Cited by 0SourcePDFScholar
2025

BVINet: Unlocking Blind Video Inpainting with Zero Annotations

ICCV 2025poster

Video inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusing primarily on the "how to inpaint". This reliance necessitates manual annotation of the corrupted regions using binary…

Cited by 0SourcePDFScholar
2025

BotSim: LLM-Powered Malicious Social Botnet Simulation

AAAI 2025technical

Social media platforms like X(Twitter) and Reddit are vital to global communication. However, advancements in Large Language Model (LLM) technology give rise to social media bots with unprecedented intelligence. These bots adeptly simulate human profiles, conversations, and interactions, disseminati…

2025

CODE: COllaborative Visual-UWB SLAM for Online Large-Scale Metric DEnse Mapping

IROS 2025

This paper presents a novel collaborative online dense mapping system for multiple Unmanned Aerial Vehicles (UAVs). The system confers two primary benefits: it facilitates simultaneous UAVs co-localization and real-time dense map reconstruction, and it recovers the metric scale even in GNSS-denied c

Cited by 0SourceScholar
2025

Capture the Key in Reasoning to Enhance CoT Distillation Generalization

ACL 2025long

As Large Language Models (LLMs) scale up and gain powerful Chain-of-Thoughts (CoTs) reasoning abilities, practical resource constraints drive efforts to distill these capabilities into more compact Smaller Language Models (SLMs). We find that CoTs consist mainly of simple reasoning forms, with a sma…

2025

Cluster-ALIV: Aerial LiDAR-Inertia-Visual Dense Reconstruction for Cluster UAV

RA-L 2025

Unmanned aerial vehicles (UAVs) equipped with LiDAR, camera, and Inertial Measurement Unit sensors are increasingly utilized for real-time dense reconstruction in large-scale rescue operations and environmental monitoring, among others. However, achieving algorithmic robustness remains challenging d

Cited by 2SourceScholar
2025

DORA: Dynamic Optimization Prompt for Continuous Reflection of LLM-based Agent

COLING 2025main

Autonomous agents powered by large language models (LLMs) hold significant potential across various domains. The Reflection framework is designed to help agents learn from past mistakes in complex tasks. While previous research has shown that reflection can enhance performance, our investigation rev…

2025

Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains

ACL 2025long

Knowledge Graphs (KGs) can serve as reliable knowledge sources for question answering (QA) due to their structured representation of knowledge. Existing research on the utilization of KG for large language models (LLMs) prevalently relies on subgraph retriever or iterative prompting, overlooking the…

Cited by 0SourcePDFScholar
2025

Dynamic Simulation Framework for Disinformation Dissemination and Correction With Social Bots

EMNLP 2025

In the “human-bot symbiotic” information ecosystem, social bots play key roles in spreading and correcting disinformation. Understanding their influence is essential for risk control and better governance. However, current studies often rely on simplistic user and network modeling, overlook the dyna

2025

Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning

EMNLP 2025

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos partially relevant to a given query. The core challenge lies in learning robust query-video alignment against spurious semantic correlations arising from inherent data uncertainty: 1) query ambiguity, where the query incompl

Cited by 0SourcePDFScholar
2025

Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution

ACL 2025finding

Improving context faithfulness in large language models is essential for developing trustworthy retrieval augmented generation systems and mitigating hallucinations, especially in long-form question answering (LFQA) tasks or scenarios involving knowledge conflicts. Existing methods either intervene…

Cited by 0SourcePDFScholar
2025

Identifying Bots on Social Media through Coordinated Group Perception

ICASSP 2025accepted

Identifying bots on social media has become a crucial and challenging task for regulating online discourse. Existing detection methods primarily focus on individual account-level information, identifying potential threats by detecting inconsistencies between genuine humans and anomalous bots in pers…

Cited by 0SourceScholar
2025

MMAD: Multi-label Micro-Action Detection in Videos

ICCV 2025poster

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising applications in human emotion analysis. In real-world scenari…

2025

Patch-level Sounding Object Tracking for Audio-Visual Question Answering

AAAI 2025technical

Answering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding objects related to the question along the timeline. In this paper, we present a new Patch-level Sounding Object Tracking (PS…

Cited by 6SourcePDFScholar
2025

Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition

AAAI 2025technical

Micro-Action Recognition (MAR) has gained increasing attention due to its crucial role as a form of non-verbal communication in social interactions, with promising potential for applications in human communication and emotion analysis. However, current approaches often overlook the inherent ambiguit…

2025

RAG-Zeval: Enhancing RAG Responses Evaluator through End-to-End Reasoning and Ranking-Based Reinforcement Learning

EMNLP 2025

Robust evaluation is critical for deploying trustworthy retrieval-augmented generation (RAG) systems. However, current LLM-based evaluation frameworks predominantly rely on directly prompting resource-intensive models with complex multi-stage prompts, underutilizing models’ reasoning capabilities an

2025

RESCUE: Crowd Evacuation Simulation via Controlling SDM-United Characters

ICCV 2025poster

Crowd evacuation simulation is critical for enhancing public safety, and demanded for realistic virtual environments. Current mainstream evacuation models overlook the complex human behaviors that occur during evacuation, such as pedestrian collisions, interpersonal interactions, and variations in b…

Cited by 0SourcePDFScholar
2025

Realm: Real-Time Line-of-Sight Maintenance in Multi-Robot Navigation with Unknown Obstacles

ICRA 2025

Multi-robot navigation in complex environments relies on inter-robot communication and mutual observation for situational awareness. This paper studies the multi-robot navigation problem in unknown environments with line-ofsight (LoS) connectivity constraints. While previous works are limited to kno

Cited by 8SourcecodeScholar
2025

Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition

ICASSP 2025accepted

Speech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain analysis, overlooking the valuable envelope structures of the frequency domain that are equally important for robust emo…

Cited by 0SourceScholar
2025

ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning

AAAI 2025technical

Zero-shot learning (ZSL) aims to recognize unseen classes by transferring semantic knowledge from seen classes to unseen ones, guided by semantic information. To this end, existing works have demonstrated remarkable performance by utilizing global visual features from Convolutional Neural Networks (…

2024

Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers

EMNLP 2024main

Query rewriting is a crucial technique for passage retrieval in open-domain conversational question answering (CQA). It decontexualizes conversational queries into self-contained questions suitable for off-the-shelf retrievers. Existing methods attempt to incorporate retriever’s preference during th…

Cited by 1SourcePDFScholar
2024

AutoFusion: Autonomous Visual Geolocation and Online Dense Reconstruction for UAV Cluster

ICRA 2024poster

Real-time dense reconstruction using Unmanned Aerial Vehicle (UAV) is becoming increasingly popular in large-scale rescue and environmental monitoring tasks. However, due to the energy constraints of a single UAV, the efficiency can be greatly improved through the collaboration of multi-UAVs. Nevert…

Cited by 0SourceScholar
2024

Beyond Read-Only: Crafting a Comprehensive Chinese Text-to-SQL Dataset for Database Manipulation and Query

NAACL 2024findings

Text-to-SQL aims to convert natural language into structured query language, which is a challenging task. Current research focuses mainly on read operations and ignores other aspects of database operations such as create, update, and delete operations. The benchmark datasets as well as models that h…

2024

CMDFusion: Bidirectional Fusion Network With Cross-Modality Knowledge Distillation for LiDAR Semantic Segmentation

RA-L 2024

2D RGB images and 3D LIDAR point clouds provide complementary knowledge for the perception system of autonomous vehicles. Several 2D and 3D fusion methods have been explored for the LIDAR semantic segmentation task, but they suffer from different problems. 2D-to-3D fusion methods require strictly pa

Cited by 21SourcecodeScholar
2024

Contrastive Learning Drug Response Models from Natural Language Supervision

IJCAI 2024poster

Deep learning-based drug response prediction (DRP) methods can accelerate the drug discovery process and reduce research and development costs. Despite their high accuracy, generating regression-aware representations remains challenging for mainstream approaches. For instance, the representations ar…

2024

EulerMormer: Robust Eulerian Motion Magnification via Dynamic Filtering within Transformer

AAAI 2024technical

Video Motion Magnification (VMM) aims to break the resolution limit of human visual perception capability and reveal the imperceptible minor motion that contains valuable information in the macroscopic domain. However, challenges arise in this task due to photon noise inevitably introduced by photog…

2024

Frequency Decoupling for Motion Magnification via Multi-Level Isomorphic Architecture

CVPR 2024poster

Video Motion Magnification (VMM) aims to reveal subtle and imperceptible motion information of objects in the macroscopic world. Prior methods directly model the motion field from the Eulerian perspective by Representation Learning that separates shape and texture or Multi-domain Learning from phase…

2024

Improve Student’s Reasoning Generalizability through Cascading Decomposed CoTs Distillation

EMNLP 2024main

Large language models (LLMs) exhibit enhanced reasoning at larger scales, driving efforts to distill these capabilities into smaller models via teacher-student learning.Previous works simply fine-tune student models on teachers’ generated Chain-of-Thoughts (CoTs) data. Although these methods enhance…

2024

Joint2Human: High-Quality 3D Human Generation via Compact Spherical Embedding of 3D Joints

CVPR 2024poster

3D human generation is increasingly significant in various applications. However the direct use of 2D generative methods in 3D generation often results in losing local details while methods that reconstruct geometry from generated images struggle with global view consistency. In this work we introdu…

Cited by 6SourcePDFScholar
2024

KeDuSR: Real-World Dual-Lens Super-Resolution via Kernel-Free Matching

AAAI 2024technical

Dual-lens super-resolution (SR) is a practical scenario for reference (Ref) based SR by utilizing the telephoto image (Ref) to assist the super-resolution of the low-resolution wide-angle image (LR input). Different from general RefSR, the Ref in dual-lens SR only covers the overlapped field of view…

2024

LPSNet: End-to-End Human Pose and Shape Estimation with Lensless Imaging

CVPR 2024poster

Human pose and shape (HPS) estimation with lensless imaging is not only beneficial to privacy protection but also can be used in covert surveillance scenarios due to the small size and simple structure of this device. However this task presents significant challenges due to the inherent ambiguity of…

Cited by 1SourcePDFScholar
2024

Solving Motion Planning Tasks with a Scalable Generative Model

ECCV 2024poster

"As autonomous driving systems being deployed to millions of vehicles, there is a pressing need of improving the system’s scalability, safety and reducing the engineering cost. A realistic, scalable, and practical simulator of the driving world is highly desired. In this paper, we present an efficie…

2024

Zero-shot Learning for Preclinical Drug Screening

IJCAI 2024poster

Conventional deep learning methods typically employ supervised learning for drug response prediction (DRP). This entails dependence on labeled response data from drugs for model training. However, practical applications in the preclinical drug screening phase demand that DRP models predict responses…

2023

Bidirectionally Deformable Motion Modulation For Video-based Human Pose Transfer

ICCV 2023poster

Video-based human pose transfer is a video-to-video generation task that animates a plain source human image based on a series of target human poses. Considering the difficulties in transferring highly structural patterns on the garments and discontinuous poses, existing methods often generate unsat…

Cited by 26PDFcodeScholar
2023

CT-GAT: Cross-Task Generative Adversarial Attack based on Transferability

EMNLP 2023long main

Neural network models are vulnerable to adversarial examples, and adversarial transferability further increases the risk of adversarial attacks. Current methods based on transferability often rely on substitute models, which can be impractical and costly in real-world scenarios due to the unavailab…

Cited by 0SourcecodeScholar
2023

Crowd3D: Towards Hundreds of People Reconstruction From a Single Image

CVPR 2023poster

Image-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of large number of people, large variations in human scale, and…

Cited by 13SourcePDFScholar
2023

Learning Semantic-Aware Disentangled Representation for Flexible 3D Human Body Editing

CVPR 2023poster

3D human body representation learning has received increasing attention in recent years. However, existing works cannot flexibly, controllably and accurately represent human bodies, limited by coarse semantics and unsatisfactory representation capability, particularly in the absence of supervised da…

Cited by 8SourcePDFScholar
2023

Narrator: Towards Natural Control of Human-Scene Interaction Generation via Relationship Reasoning

ICCV 2023poster

Naturally controllable human-scene interaction (HSI) generation has an important role in various fields, such as VR/AR content creation and human-centered AI. However, existing methods are unnatural and unintuitive in their controllability, which heavily limits their application in practice. Therefo…

Cited by 9PDFScholar
2023

SGP-TOD: Building Task Bots Effortlessly via Schema-Guided LLM Prompting

EMNLP 2023long findings

Building and maintaining end-to-end task bots using minimal human effort is a long-standing challenge in dialog research. In this work, we introduce SGP-TOD, Schema-Guided Prompting for building Task-Oriented Dialog systems effortlessly based on large language models (LLMs). Utilizing the predefined…

Cited by 0SourceScholar
2022

AFDetV2: Rethinking the Necessity of the Second Stage for Object Detection from Point Clouds

AAAI 2022technical

There have been two streams in the 3D detection from point clouds: single-stage methods and two-stage methods. While the former is more computationally efficient, the latter usually provides better detection accuracy. By carefully examining the two-stage approaches, we have found that if appropriate…

2022

FOF: Learning Fourier Occupancy Field for Monocular Real-time Human Reconstruction

NeurIPS 2022accept

The advent of deep learning has led to significant progress in monocular human reconstruction. However, existing representations, such as parametric models, voxel grids, meshes and implicit neural representations, have difficulties achieving high-quality results and real-time speed at the same time.…

Cited by 39SourcePDFScholar
2022

High-Fidelity Human Avatars From a Single RGB Camera

CVPR 2022poster

In this paper, we propose a coarse-to-fine framework to reconstruct a personalized high-fidelity human avatar from a monocular video. To deal with the misalignment problem caused by the changed poses and shapes in different frames, we design a dynamic surface network to recover pose-dependent surfac…

Cited by 40PDFScholar
2021

Cross-MPI: Cross-Scale Stereo for Image Super-Resolution Using Multiplane Images

CVPR 2021poster

Various combinations of cameras enrich computational photography, among which reference-based superresolution (RefSR) plays a critical role in multiscale imaging systems. However, existing RefSR approaches fail to accomplish high-fidelity super-resolution under a large resolution gap, e.g., 8x upsca…

Cited by 31PDFScholar
2021

Implicit Transformer Network for Screen Content Image Continuous Super-Resolution

NeurIPS 2021poster

Nowadays, there is an explosive growth of screen contents due to the wide application of screen sharing, remote cooperation, and online education. To match the limited terminal bandwidth, high-resolution (HR) screen contents may be downsampled and compressed. At the receiver side, the super-resolu…

2021

MDANet: Multi-Modal Deep Aggregation Network for Depth Completion

ICRA 2021poster

Depth completion aims to recover the dense depth map from sparse depth data and RGB image respectively. However, due to the huge difference between the multi-modal signal input, vanilla convolutional neural network and simple fusion strategy cannot extract features from sparse data and aggregate mul…

Cited by 17SourcecodeScholar
2021

Robust SRIF-based LiDAR-IMU Localization for Autonomous Vehicles

ICRA 2021poster

We present a tightly-coupled multi-sensor fusion architecture for autonomous vehicle applications, which achieves centimetre-level accuracy and high robustness in various scenarios. In order to realize robust and accurate point-cloud feature matching we propose a novel method for extracting structur…

Cited by 5SourceScholar
2020

4D Association Graph for Realtime Multi-Person Motion Capture Using Multiple Video Cameras

CVPR 2020oral

his paper contributes a novel realtime multi-person motion capture algorithm using multiview video inputs. Due to the heavy occlusions and closely interacting motions in each view, joint optimization on the multiview images and multiple temporal frames is indispensable, which brings up the essential…

Cited by 104PDFcodeScholar
2020

Constituency Lattice Encoding for Aspect Term Extraction

COLING 2020main

One of the remaining challenges for aspect term extraction in sentiment analysis resides in the extraction of phrase-level aspect terms, which is non-trivial to determine the boundaries of such terms. In this paper, we aim to address this issue by incorporating the span annotations of constituents o…

2020

Rebalancing Expanding EV Sharing Systems with Deep Reinforcement Learning

IJCAI 2020poster

Electric Vehicle (EV) sharing systems have recently experienced unprecedented growth across the world. One of the key challenges in their operation is vehicle rebalancing, i.e., repositioning the EVs across stations to better satisfy future user demand. This is particularly challenging in the shared…

2020

Vision Global Localization with Semantic Segmentation and Interest Feature Points

IROS 2020poster

In this work, we present a vision-only global localization architecture for autonomous vehicle applications, and achieves centimeter-level accuracy and high robustness in various scenarios. We first apply pixel-wise segmentation to the front-view mono camera and extract the semantic features, e.g. p…

Cited by 5SourceScholar
2018

Image Alignment via Multi-Model Geometric Fitting and Hierarchical Homography Estimation

ICASSP 2018accepted

It is challenging to achieve accurate alignment for building images containing multiple planes. We propose a multi-model geometric fitting and hierarchical homography estimation method to improve the alignment performance for building images. We first extract scale-invariant feature transform (SIFT)…

Cited by 0SourceScholar
2018

Image-Based PM2.5 Estimation and its Application on Depth Estimation

ICASSP 2018accepted

Air pollution is still a big threat to human health particularly for developing countries. It is highly demanding to measure air quality with daily-used devices such as smartphones. On the other hand, it is difficult to estimate the scene depth under the foul weather using traditional vision-based m…

Cited by 0SourceScholar
2018

Inverse Reinforcement Learning via Function Approximation for Clinical Motion Analysis

ICRA 2018poster

This paper introduces a new method for inverse reinforcement learning in large state spaces, where the learned reward function can be used to control high-dimensional robot systems and analyze complex human movement. To avoid solving the computationally expensive reinforcement learning problems in r…

Cited by 15SourceScholar
2018

Unsupervised Discovery of an Extended Phoneme Set in L2 English Speech for Mispronunciation Detection and Diagnosis

ICASSP 2018accepted

Second language (L2) speech is often labelled with the native, phoneme categories. Hence, we often observe segments for which it is difficult, if not impossible, to decide on a categorical phoneme label. We refer to these segments as “non-categorical” phoneme units. Existing approaches to mispronunc…

Cited by 0SourceScholar
2017

Clinical patient tracking in the presence of transient and permanent occlusions via geodesic feature

ICRA 2017poster

This paper develops a method to use RGB-D cameras to track the motions of a human spinal cord injury patient undergoing spinal stimulation and physical rehabilitation. Because clinicians must remain close to the patient during training sessions, the patient is usually under permanent and transient o…

Cited by 0SourceScholar
2015

Voice conversion using deep Bidirectional Long Short-Term Memory based Recurrent Neural Networks

ICASSP 2015accepted

This paper investigates the use of Deep Bidirectional Long Short-Term Memory based Recurrent Neural Networks (DBLSTM-RNNs) for voice conversion. Temporal correlations across speech frames are not directly modeled in frame-based methods using conventional Deep Neural Networks (DNNs), which results in…

Cited by 0SourceScholar