← Search

Yu Chen

102 accepted papers

2026

ANALYTIC INCREMENTAL LEARNING FOR SOUND SOURCE LOCALIZATION WITH IMBALANCE RECTIFICATION

ICASSP 2026poster

Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA) distributions, and inter-task imbalance induced by cross-task skews…

Cited by 0SourcePDFScholar
2026

ASSESS: A Semantic and Structural Evaluation Framework for Statement Similarity

ICLR 2026poster

Despite significant strides in statement autoformalization, a critical gap remains in the development of automated evaluation metrics capable of assessing formal translation quality. Existing metrics often fail to balance semantic and structural information: string-based methods neglect semantics, w…

Cited by 0SourcecodeScholar
2026

AV-SSAN: Audio-Visual Selective DOA Estimation Through Explicit Multi-Band Semantic-Spatial Alignment

AAAI 2026technical

Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize specific target sources. To address these limitations, we intr

Cited by 0SourcePDFScholar
2026

Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations

AAAI 2026technical

Automated essay scoring (AES) is a challenging task in cross-prompt settings due to the diversity of scoring criteria. While previous studies have focused on the output of large language models (LLMs) to improve scoring accuracy, we believe activations from intermediate layers may also provide valua

Cited by 0SourcePDFScholar
2026

Adversarial Game-Theoretic Algorithm for Dexterous Grasp Synthesis

ICRA 2026poster

For many complex tasks, multi-finger robot hands are poised to revolutionize how we interact with the world, but reliably grasping objects remains a significant challenge. We focus on the problem of synthesizing grasps for multi-finger robot hands that, given an target object's geometry and pose, co…

2026

BEYOND VISUAL REALISM: TOWARD RELIABLE FINANCIAL TIME SERIES GENERATION

ICASSP 2026poster

Generative models for financial time series often create data that look realistic and even reproduce stylized facts such as fat tails or volatility clustering. However, these apparent successes break down under trading backtests: models like GANs or WGAN-GP frequently collapse, yielding extreme and…

Cited by 0SourcePDFScholar
2026

Finite-Time Convergence Analysis of ODE-based Generative Models for Stochastic Interpolants

ICLR 2026poster

Stochastic interpolants offer a robust framework for continuously transforming samples between arbitrary data distributions via ordinary or stochastic differential equations (ODEs/SDEs), holding significant promise for generative modeling. While previous studies have analyzed the finite-time converg…

Cited by 0SourceScholar
2026

FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation

CVPR 2026

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photorealistic, they are typically sparse and discrete. Conversely, synthetic data sc

Cited by 0SourcecodeScholar
2026

Memoria-Bench: A Comprehensive Benchmark for Evaluating Memory in Long-Horizon Autonomous Agents

ICML 2026poster

Memory is a core capability of autonomous agents, yet existing benchmarks evaluate it primarily in constrained settings such as short dialogues or synthetic tasks, failing to reflect realistic agent deployments. We present \textbf{Memoria-Bench}, a benchmark for evaluating agent memory grounded in c…

Cited by 0SourceScholar
2026

NavMoE: Hybrid Model and Learning-Based Traversability Estimation for Local Navigation Via Mixture of Experts

ICRA 2026poster

This paper explores traversability estimation for robot navigation. A key bottleneck in traversability estimation lies in efficiently achieving reliable and robust predictions while accurately encoding both geometric and semantic information across diverse environments. We introduce Navigation via M…

2026

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modaliti…

Cited by 0SourcecodeScholar
2026

PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text Alignment

ICML 2026poster

While CLIP has achieved strong performance across vision–language tasks, fine-grained image–text alignment remains challenging. Recent efforts improve textual granularity by leveraging long, detailed descriptions and replacing CLIP’s text encoder with LLM, but often overlook the visual-side bottlene…

Cited by 0SourceScholar
2026

PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching

ICML 2026poster

Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without external supervision. However, current methods rely on heuristic intrinsic rewards, which often lack a well-defined theo…

Cited by 0SourceScholar
2026

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

CVPR 2026

Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. We introduce NUMINA, a training-free identify-then-guide framework for improved numerical alignment. NUMINA identifies prompt-layout inconsi

Cited by 0SourcecodeScholar
2025

A Chaotic Dynamics Framework Inspired by Dorsal Stream for Event Signal Processing

ICML 2025poster

Event cameras are bio-inspired vision sensors that encode visual information with high dynamic range, high temporal resolution, and low latency. Current state-of-the-art event stream processing methods rely on end-to-end deep learning techniques. However, these models are heavily dependent on data s…

Cited by 0SourcePDFScholar
2025

A Closer Look at Transformers for Time Series Forecasting: Understanding Why They Work and Where They Struggle

ICML 2025poster

Time-series forecasting is crucial across various domains, including finance, healthcare, and energy. Transformer models, originally developed for natural language processing, have demonstrated significant potential in addressing challenges associated with time-series data. These models utilize diff…

Cited by 0SourcePDFScholar
2025

A Pioneering Neural Network Method for Efficient and Robust Fuel Sloshing Simulation in Aircraft

AAAI 2025technical

Simulating fuel sloshing within aircraft tanks during flight is crucial for aircraft safety research. Traditional methods based on Navier-Stokes equations are computationally expensive. In this paper, we treat fluid motion as point cloud transformation and propose the first neural network method spe…

Cited by 1SourcePDFScholar
2025

ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data

NeurIPS 2025poster

Autoformalization, the automatic translation of mathematical content from natural language into machine-verifiable formal languages, has seen significant progress driven by advances in large language models (LLMs). Nonetheless, a primary barrier to further improvements is the limited availability of…

Cited by 0SourcecodeScholar
2025

An Interactive Hands-Free Controller for a Riding Ballbot to Enable Simple Shared Control Tasks

ICRA 2025

Our team developed a riding ballbot (called PURE) that is dynamically stable, omnidirectional, and driven by lean-to-steer control. A hands-free admittance control scheme (HACS) was previously integrated to allow riders with different torso functions to control the robot's movements via torso leanin

Cited by 1SourceScholar
2025

Confidence-Aware With Prototype Alignment for Partial Multi-label Learning

NeurIPS 2025poster

Label prototype learning has emerged as an effective paradigm in Partial Multi-Label Learning (PML), providing a distinctive framework for modeling structured representations of label semantics while naturally filtering noise through prototype-based label confidence estimation. However, existing pro…

Cited by 0SourceScholar
2025

Deep Gaussian from Motion: Exploring 3D Geometric Foundation Models for Gaussian Splatting

NeurIPS 2025poster

Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photorealistic images. However, the prerequisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. Although previous methods can reconstruct a few unpos…

Cited by 0SourceScholar
2025

Developing Chatbots for Sustainability: Experiential Learning in an Undergraduate Business Course

AAAI 2025technical

This paper presents an experiential learning pedagogy that teaches undergraduate business management information systems students hands-on AI skills through the lens of sustainability. The learning modules aim to empower undergraduate business students to gain interest and confidence in AI knowledge…

Cited by 0SourcePDFScholar
2025

FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network

NeurIPS 2025poster

The Transformer architecture serves as the foundation of modern AI systems, powering recent advances in Large Language Models (LLMs) and Large Multimodal Models (LMMs). Central to these models, attention mechanisms capture contextual dependencies via token interactions. Beyond inference, attention h…

Cited by 0SourceScholar
2025

H-PCC: Point Cloud Compression With Hybrid Mode Selection and Content Adaptive Down-Sampling

RA-L 2025

LiDAR sensors are integral to autonomous driving and augmented reality applications, providing essential depth information. However, managing the substantial volume of LiDAR point cloud data is crucial for practical application, necessitating efficient compression algorithms. Similar to other data c

Cited by 5SourceScholar
2025

Image Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual Inference

CVPR 2025poster

Existing full-reference image quality assessment (FR-IQA) methods often fail to capture the complex causal mechanisms that underlie human perceptual responses to image distortions, limiting their ability to generalize across diverse scenarios. In this paper, we propose an FR-IQA method based on abdu…

Cited by 0SourcePDFScholar
2025

Online Synthesis of Control Barrier Functions with Local Occupancy Grid Maps for Safe Navigation in Unknown Environments

IROS 2025

Control Barrier Functions (CBFs) have emerged as an effective and non-invasive safety filter for ensuring the safety of autonomous systems in dynamic environments with formal guarantees. However, most existing works on CBF synthesis focus on fully known settings. Synthesizing CBFs online based on pe

Cited by 0SourceScholar
2025

Propagative Distance Optimization for Motion Planning

ICRA 2025

This paper focuses on the motion planning problem for serial articulated robots with revolute joints under kinematic constraints. Many motion planners leverage iterative local optimization methods but are often trapped in local minima due to non-convexity of the problem. A key reason for the non-con

Cited by 1SourceScholar
2025

Pseudo-Label Reconstruction for Partial Multi-Label Learning

IJCAI 2025

In Partial Multi-Label Learning (PML), each instance is associated with a candidate label set containing multiple relevant labels along with other false positive labels. Currently, most PML methods directly extract instance correlation from instance features while ignoring the candidate labels, whic

Cited by 0SourcePDFScholar
2025

STGE-Former: Spatial-Temporal Graph-Enhanced Transformer for EEG-Based Major Depressive Disorder Detection

ICASSP 2025accepted

Applying deep learning techniques to Electroencephalogram (EEG) data has shown great potential in the field of depression detection. However, existing EEG-based depression detection models face challenges: they struggle to capture the complex spatiotemporal dependencies and the complementary nature…

Cited by 0SourceScholar
2025

Scaling and Taming Adversarial Training with Synthetic Data

ICCV 2025poster

Despite the success of adversarial training on small datasets, applying it to large-scale datasets like ImageNet remains challenging. Previous attempts using synthetic data show limited improvements. This work investigates the impact of synthetic data scaling, model scaling, and training strategies…

Cited by 0SourcePDFScholar
2025

Steering Large Language Models for Vulnerability Detection

ICASSP 2025accepted

Vulnerability detection remains a critical challenge in the field of security. Many existing approaches extract code representations for vulnerability detection. However, these methods often focus on the overall semantics of the code, neglecting to specifically target vulnerability-related semantics…

Cited by 0SourceScholar
2025

uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs

ICLR 2025spotlight

In this paper, we present a novel algorithm, `uniINF`, for the Heavy-Tailed Multi-Armed Bandits (HTMAB) problem, demonstrating robustness and adaptability in both stochastic and adversarial environments. Unlike the stochastic MAB setting where loss distributions are stationary with time, our study e…

Cited by 0SourcePDFScholar
2024

DOGS: Distributed-Oriented Gaussian Splatting for Large-Scale 3D Reconstruction Via Gaussian Consensus

NeurIPS 2024poster

The recent advances in 3D Gaussian Splatting (3DGS) show promising results on the novel view synthesis (NVS) task. With its superior rendering performance and high-fidelity rendering quality, 3DGS is excelling at its previous NeRF counterparts. The most recent 3DGS method focuses either on improving…

2024

FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods

ICLR 2024poster

This paper introduces the Fair Fairness Benchmark (FFB), a benchmarking framework for in-processing group fairness methods. Ensuring fairness in machine learning is important for ethical compliance. However, there exist challenges in comparing and developing fairness methods due to inconsistencies i…

2024

Generating Handwritten Mathematical Expressions From Symbol Graphs: An End-to-End Pipeline

CVPR 2024poster

In this paper we explore a novel challenging generation task i.e. Handwritten Mathematical Expression Generation (HMEG) from symbolic sequences. Since symbolic sequences are naturally graph-structured data we formulate HMEG as a graph-to-image (G2I) generation problem. Unlike the generation of natur…

2024

Graph-Propagation-Based Kinematic Algorithm for In-Pipe Truss Structure Robots

RA-L 2024

Robots designed for in-pipe navigation, inspection, and repair require flexibility for intricate pipeline traversal and the strength to carry payloads. However, conventional wheeled in-pipe robots face challenges in simultaneously achieving both substantial flexibility and payload-carrying capacity.

Cited by 0SourceScholar
2024

LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

NAACL 2024long

Today’s large language models (LLMs) typically train on short text segments (e.g., <4K tokens) due to the quadratic complexity of their Transformer architectures. As a result, their performance suffers drastically on inputs longer than those encountered during training, substantially limiting their…

2024

LOCSELECT: Target Speaker Localization with an Auditory Selective Hearing Mechanism

ICASSP 2024accepted

The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of speakers. In this paper, we present a target speaker localization a…

Cited by 0SourceScholar
2024

Multi-Constellation-Inspired Single-Shot Global LiDAR Localization

AAAI 2024technical

Global localization is a challenging task for intelligent robots, as its accuracy directly contributes to the performance of downstream navigation and planning tasks. However, existing literature focus more on the place retrieval and the success rate of localization, with limited attention given to…

2024

PKAD: Pretrained Knowledge is All You Need to Detect and Mitigate Textual Backdoor Attacks

EMNLP 2024finding

In textual backdoor attacks, attackers insert poisoned samples with triggered inputs and target labels into training datasets to manipulate model behavior, threatening the model’s security and reliability. Current defense methods can generally be categorized into inference-time and training-time one…

2024

Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation

ICML 2024poster

In the realm of reinforcement learning (RL), accounting for risk is crucial for making decisions under uncertainty, particularly in applications where safety and reliability are paramount. In this paper, we introduce a general framework on Risk-Sensitive Distributional Reinforcement Learning (RS-Dis…

Cited by 5SourcePDFScholar
2024

Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback

ICLR 2024poster

Risk-sensitive reinforcement learning (RL) aims to optimize policies that balance the expected reward and risk. In this paper, we present a novel risk-sensitive RL framework that employs an Iterated Conditional Value-at-Risk (CVaR) objective under both linear and general function approximations, enr…

Cited by 3SourcePDFScholar
2024

Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation

ICML 2024poster

This work pioneers regret analysis of risk-sensitive reinforcement learning in partially observable environments with hindsight observation, addressing a gap in theoretical exploration. We introduce a novel formulation that integrates hindsight observations into a Partially Observable Markov Decisio…

Cited by 1SourcePDFScholar
2024

Reflected Schrödinger Bridge for Constrained Generative Modeling

UAI 2024poster

Diffusion models have become the go-to method for large-scale generative models in real-world applications. These applications often involve data distributions confined within bounded domains, typically requiring ad-hoc thresholding techniques for boundary enforcement. Reflected diffusion models aim…

2024

Safe and Individualized Motion Planning for Upper-limb Exoskeleton Robots Using Human Demonstration and Interactive Learning

ICRA 2024poster

A typical application of upper-limb exoskeleton robots is deployment in rehabilitation training, helping patients to regain manipulative abilities. However, as the patient is not always capable of following the robot, safety issues may arise during the training. Due to the bias in different patients…

Cited by 1SourceScholar
2024

Spatial-Temporal Interaction Decoding Transformer for Unsupervised Multivariate Time Series Anomaly Detection

ICASSP 2024accepted

Time series data consists of a temporal dimension and features associated with each timestamp. Anomaly detection in this context necessitates the consideration of both temporal and spatial features. However, existing work focuses on separately addressing temporal and spatial features, neglecting the…

Cited by 0SourceScholar
2024

Tunable Stiffness Glove for Tremor Suppression Based on 3D Printed Structured Fabrics

IROS 2024poster

Tremors, which are prevalent symptoms in both Parkinson’s disease (PD) and essential tremor (ET), substantially diminish the quality of life for those affected. Traditional treatments, including pharmaceutical medications and invasive surgical procedures, often come with limitations and side effects…

Cited by 1SourceScholar
2024

Variational Schrödinger Diffusion Models

ICML 2024poster

Schrödinger bridge (SB) has emerged as the go-to method for optimizing transportation plans in diffusion models. However, SB requires estimating the intractable forward score functions, inevitably resulting in the (costly) implicit training loss based on simulated trajectories. To improve the scalab…

Cited by 9SourcePDFScholar
2023

AdaSfM: From Coarse Global to Fine Incremental Adaptive Structure from Motion

ICRA 2023poster

Despite the impressive results achieved by many existing Structure from Motion (SfM) approaches, there is still a need to improve the robustness, accuracy, and efficiency on large-scale scenes with many outlier matches and sparse view graphs. In this paper, we propose AdaSfM: a coarse-to-fine adapti…

Cited by 10SourceScholar
2023

Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality

EMNLP 2023long main

Contrastively trained vision-language models have achieved remarkable progress in vision and language representation learning. However, recent research has highlighted severe limitations of these models in their ability to perform compositional reasoning over objects, attributes, and relations. Scen…

Cited by 0SourceScholar
2023

Improving Noisy Student Training on Non-Target Domain Data for Automatic Speech Recognition

ICASSP 2023accepted

Noisy Student Training (NST) has recently demonstrated extremely strong performance in Automatic Speech Recognition (ASR). In this paper, we propose a data selection strategy named LM Filter to improve the performance of NST on non-target domain data in ASR tasks. Hypotheses with and without a Langu…

Cited by 0SourceScholar
2023

Inference and sampling of point processes from diffusion excursions

UAI 2023poster

Point processes often have a natural interpretation with respect to a continuous process. We propose a point process construction that describes arrival time observations in terms of the state of a latent diffusion process. In this framework, we relate the return times of a diffusion in a continuous…

Cited by 3SourcePDFScholar
2023

Information theoretic clustering via divergence maximization among clusters

UAI 2023poster

Information-theoretic clustering is one of the most promising and principled approaches to finding clusters with minimal apriori assumptions. The key criterion therein is to maximize the mutual information between the data points and their cluster labels. Such an approach, however, does not explicit…

2023

Joint Semantic and Strategy Matching for Persuasive Dialogue

EMNLP 2023long findings

Persuasive dialogue aims to persuade users to achieve some targets by conversations. While previous persuasion models have achieved notable successes, they mostly base themselves on utterance semantic matching, and an important aspect has been ignored, that is, the strategy of the conversations, for…

Cited by 0SourceScholar
2023

MMVC: Learned Multi-Mode Video Compression With Block-Based Prediction Mode Selection and Density-Adaptive Entropy Coding

CVPR 2023poster

Learning-based video compression has been extensively studied over the past years, but it still has limitations in adapting to various motion patterns and entropy models. In this paper, we propose multi-mode video compression (MMVC), a block wise mode ensemble deep video compression framework that s…

2023

MixPAVE: Mix-Prompt Tuning for Few-shot Product Attribute Value Extraction

ACL 2023findings

The task of product attribute value extraction is to identify values of an attribute from product information. Product attributes are important features, which help improve online shopping experience of customers, such as product search, recommendation and comparison. Most existing works only focus…

Cited by 31SourcePDFScholar
2023

Multi-Modal Learning and Relaxation of Physical Conflict for an Exoskeleton Robot with Proprioceptive Perception

ICRA 2023poster

Exoskeleton robots provide assistive forces to suit the human subject via physical human-robot interaction. During the closely-coupled interaction, a mismatch between the wearer and the robot may result in physical conflict, which could affect assistance efficiency or even compromise safety. Therefo…

Cited by 5SourceScholar
2023

Provably Convergent Schrödinger Bridge with Applications to Probabilistic Time Series Imputation

ICML 2023poster

The Schrödinger bridge problem (SBP) is gaining increasing attention in generative modeling and showing promising potential even in comparison with the score-based generative models (SGMs). SBP can be interpreted as an entropy-regularized optimal transport problem, which conducts projections onto ev…

2023

Search for Efficient Deep Visual-Inertial Odometry Through Neural Architecture Search

ICASSP 2023accepted

Recent deep learning based visual-inertial odometry (VIO) systems achieve impressive performance in various applications and challenging scenarios. However, it is difficult to deploy such VIO models directly on energy-constrained mobile platforms in real-time due to the extensive complexity of exist…

Cited by 0SourceScholar
2023

Two-Stage Trajectory-Tracking Control of Cable-Driven Upper-Limb Exoskeleton Robots with Series Elastic Actuators: A Simple, Accurate, and Force-Sensorless Method

IROS 2023poster

The advantages of cable-driven exoskeleton robots with series elastic actuators can be summarized in twofold: 1) the inertia of the robot joint is relatively low, which is more friendly for human-robot interaction; 2) the elastic element is tolerant to impacts and hence provides structural safety. A…

Cited by 1SourceScholar
2022

An End-to-End Deep Learning Framework For Multiple Audio Source Separation And Localization

ICASSP 2022accepted

Sound source separation and localization for situational awareness enables a wide range of applications such as hearing enhancement and audio beam-forming. We present an end-to-end deep learning framework to separate and localize multiple audio sources from the mixture of multi-channels. The propose…

Cited by 0SourceScholar
2022

Efficient Deep Visual and Inertial Odometry with Adaptive Visual Modality Selection

ECCV 2022poster

"In recent years, deep learning-based approaches for visual-inertial odometry (VIO) have shown remarkable performance outperforming traditional geometric methods. Yet, all existing methods use both the visual and inertial measurements for every pose estimation incurring potential computational redun…

2022

Estimating transfer entropy under long ranged dependencies

UAI 2022poster

Estimating Transfer Entropy (TE) between time series is a highly impactful problem in fields such as finance and neuroscience. The well-known nearest neighbor estimator of TE potentially fails if temporal dependencies are noisy and long ranged, primarily because it estimates TE indirectly relying o…

2022

Hierarchical Learning and Control for In-Hand Micromanipulation Using Multiple Laser-Driven Micro-Tools

IROS 2022poster

Laser-driven micro-tools are formulated by treating highly-focused laser beams as actuators, to control the tool's motion to contact then manipulate a micro object, which allows it to manipulate opaque micro objects, or large cells without causing photodamage. However, most existing laser-driven too…

Cited by 1SourceScholar
2022

Leveraging Structural Information to Improve Point Line Visual-Inertial Odometry

RA-L 2022

Leveraging line features to improve the accuracy of the SLAM system has been studied in many works. However, making full use of the characteristics of different line features (parallel, non-parallel) to improve the SLAM system is rarely mentioned. In this paper, we designed a VIO system based on poi

Cited by 36SourcecodeScholar
2022

Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation

ICML 2022spotlight

We study reinforcement learning with linear function approximation where the transition probability and reward functions are linear with respect to a feature mapping $\boldsymbol{\phi}(s,a)$. Specifically, we consider the episodic inhomogeneous linear Markov Decision Process (MDP), and propose a nov…

Cited by 38SourcePDFScholar
2021

A CNN Based Vision-Proprioception Fusion Method for Robust UGV Terrain Classification

RA-L 2021

The ability for ground vehicles to identify terrain types and characteristics can help provide more accurate localization and information-rich mapping solutions. Previous studies have shown the possibility of classifying terrain types based on proprioceptive sensors that monitor wheel-terrain intera

Cited by 28SourceScholar
2021

Retrieval-Augmented Generation for Code Summarization via Hybrid GNN

ICLR 2021spotlight

Source code summarization aims to generate natural language summaries from structured code snippets for better understanding code functionalities. However, automatic code summarization is challenging due to the complexity of the source code and the language gap between the source code and natural la…

2021

Semi-supervised Vein Segmentation of Ultrasound Images for Autonomous Venipuncture

IROS 2021poster

Venipuncture is an indispensable procedure for both diagnosis and treatment. In this paper, unlike existing solutions that fully or partially rely on professional assistance, a compact robotic system integrating both novel hardware and software developments is introduced. The hardware consists of a…

Cited by 7SourceScholar
2021

The LOB Recreation Model: Predicting the Limit Order Book from TAQ History Using an Ordinary Differential Equation Recurrent Neural Network

AAAI 2021technical

In an order-driven financial market, the price of a financial asset is discovered through the interaction of orders - requests to buy or sell at a particular price - that are posted to the public limit order book (LOB). Therefore, LOB data is extremely valuable for modelling market dynamics. However…

2020

GraphFlow: Exploiting Conversation Flow with Graph Neural Networks for Conversational Machine Comprehension

IJCAI 2020poster

Conversational machine comprehension (MC) has proven significantly more challenging compared to traditional MC since it requires better utilization of conversation history. However, most existing approaches do not effectively capture conversation history and thus have trouble handling questions invo…

2020

Iterative Deep Graph Learning for Graph Neural Networks: Better and Robust Node Embeddings

NeurIPS 2020poster

In this paper, we propose an end-to-end graph learning framework, namely \textbf{I}terative \textbf{D}eep \textbf{G}raph \textbf{L}earning (\alg), for jointly and iteratively learning graph structure and graph embedding. The key rationale of \alg is to learn a better graph structure based on better…

2020

Reinforcement Learning Based Graph-to-Sequence Model for Natural Question Generation

ICLR 2020poster

Natural question generation (QG) aims to generate questions from a passage and an answer. Previous works on QG either (i) ignore the rich structure information hidden in text, (ii) solely rely on cross-entropy loss that leads to issues like exposure bias and inconsistency between train/test measurem…

Cited by 221SourcecodeScholar
2019

$β^3$-IRT: A New Item Response Model and its Applications

AISTATS 2019poster

Item Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the $\beta^3$-IRT model, which models continuous responses and can generate a much enriched family o…

2019

Completely Blind Image Quality Assessment Using Latent Quality Factor from Image Local Structure Representation

ICASSP 2019accepted

Although opinion-unaware (OA) blind image quality assessment (BIQA) is the most difficult task, it is still very attractive because of its great potential for good generalization capability and practical usage. In this paper, a novel OA-BIQA algorithm is proposed. This algorithm is based on the cons…

Cited by 0SourceScholar
2019

Fast and Incremental Loop Closure Detection Using Proximity Graphs

IROS 2019poster

Visual loop closure detection, which can be considered as an image retrieval task, is an important problem in SLAM (Simultaneous Localization and Mapping) systems. The frequently used bag-of-words (BoW) models can achieve high precision and moderate recall. However, the requirement for lower time co…

Cited by 51SourcecodeScholar
2019

Fisher Efficient Inference of Intractable Models

NeurIPS 2019poster

Maximum Likelihood Estimators (MLE) has many good properties. For example, the asymptotic variance of MLE solution attains equality of the asymptotic Cram{\'e}r-Rao lower bound (efficiency bound), which is the minimum possible variance for an unbiased estimator. However, obtaining such MLE solution…

2019

Multisource Surveillance Video Coding by Exploiting 3D and 2D Knolwedge

ICASSP 2019accepted

The rapidly increasing surveillance video data has challenged the existing video coding standards. Even though knowledge based video coding scheme proposed for moving objects so far has achieved high efficiency, it does not take full advantages of local information and highly relies on the accuracy…

Cited by 0SourceScholar
2019

Sell-corpus: an Open Source Multiple Accented Chinese-english Speech Corpus for L2 English Learning Assessment

ICASSP 2019accepted

We present SELL-CORPUS, a multiple accented speech corpus for L2 English learning in China, aiming at the potential research of multiple accented acoustic model, mispronunciation detection and pronunciation assessment for future nationwide oral English tests. Our corpus contains 31.6 hour speech rec…

Cited by 0SourceScholar
2018

FSRNet: End-to-End Learning Face Super-Resolution With Facial Priors

CVPR 2018poster

Face Super-Resolution (SR) is a domain-specific superresolution problem. The facial prior knowledge can be leveraged to better super-resolve face images. We present a novel deep end-to-end trainable Face Super-Resolution Network (FSRNet), which makes use of the geometry prior, i.e., facial landmark…

2017

Adversarial PoseNet: A Structure-Aware Convolutional Network for Human Pose Estimation

ICCV 2017poster

For human pose estimation in monocular images, joint occlusions and overlapping upon human bodies often result in deviated pose predictions. Under these circumstances, bi- ologically implausible pose predictions may be produced. In contrast, human vision is able to predict poses by exploiting geomet…

Cited by 461PDFScholar
2016

Integration of machine learning and human learning for training optimization in robust linear regression

ICASSP 2016accepted

In this paper machine learning and human learning are applied jointly to optimize the training of linear regression. Human learning is exploited to label extra training data so as to resolve problems such as insufficient training and over-fitting. Considering the inevitable human errors in labeling,…

Cited by 0SourceScholar