← Search

YAN LIU

106 accepted papers

2026

A Lightweight Physics-Informed Neural Network for Sim-To-Real of Biped Robot

ICRA 2026poster

In this paper, we present a low-cost, easy-to-implement sim-to-real framework for biped locomotion that narrows the reality gap using only simulation data, without motion-capture or additional real-world measurements. First, a walking policy for the BRUCE robot is trained in Isaac Gym via reinforcem…

Cited by 0SourceScholar
2026

A Lightweight Physics-Informed Neural Network for Sim-to-Real of Biped Robot

RA-L 2026

In this paper, we present a low-cost, easy-to-implement sim-to-real framework for biped locomotion that narrows the reality gap using only simulation data, without motion-capture or additional real-world measurements. First, a walking policy for the BRUCE robot is trained in Isaac Gym via reinforcem

Cited by 0SourceScholar
2026

Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning

ICLR 2026poster

While scaling individual Large Language Models (LLMs) has delivered remarkable progress, the next frontier lies in scaling collaboration through multi-agent systems (MAS). However, purely autonomous MAS remain ``closed-world'' systems, constrained by the static knowledge horizon of pre-trained model…

Cited by 0SourcecodeScholar
2026

AerialExtreMatch: A Benchmark for Extreme-View Image Matching and Localization

RA-L 2026

Image matching serves as a core component for UAV localization guided by satellite imagery. However, this task remains highly challenging due to the extreme viewpoint discrepancies between low-altitude UAV images and nadir-view satellite maps. Existing datasets primarily focus on ground-level or hig

Cited by 0SourcecodeScholar
2026

Capturability as Controlled-Invariant Sets: Recursive Feasibility for Variable-Stepping Time S2S NMPC

RA-L 2026

Capturability characterizes a safe region of states for humanoid walking and is most commonly constructed by analyzing the one-dimensional divergent component of motion (DCM) of the center of mass. In this work, by exploiting the mathematical structure of the step-to-step (S2S) dynamics, we characte

Cited by 0SourceScholar
2026

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation

CVPR 2026

Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real-world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diffusion models. Current approaches leverage commonsense reasoning capability of la

Cited by 0SourcecodeScholar
2026

DroneDINO: Towards Heterogeneous Routed Mixture of Experts for Drone-based Unified Object Detection

ICML 2026oral

Recently, the rapid development of low-altitude aerial applications has driven the need for drone-based unified detectors. In contrast to task-specific detectors that suffer from poor scalability across diverse scenarios, existing unified detectors leverage the Mixture-of-Experts (MoE) architecture …

Cited by 0SourceScholar
2026

Dynamic-Static Decomposition for Novel View Synthesis of Dynamic Scenes with Spiking Neurons

CVPR 2026

Novel view synthesis for dynamic scenes remains challenging due to complex motion variations. Recent methods represent dynamic and static regions with separate Gaussians to improve efficiency and accuracy, but inaccurate assignment of static and dynamic Gaussian primitives still limits performance.

Cited by 0SourceScholar
2026

From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression Comprehension

AAAI 2026technical

Recent advances in Referring Expression Comprehension (REC) have been largely driven by supervised learning on curated datasets, where each expression is assumed to refer to exactly one known object. However, such assumptions rarely hold in real-world scenarios, where expressions can refer to multip

Cited by 0SourcePDFScholar
2026

Gait-Adaptive Perceptive Humanoid Locomotion With Real-Time Under-Base Terrain Reconstruction

RA-L 2026

For full-size humanoid robots, reliable locomotion on complex terrains—such as long staircases—remains challenging, even with recent advances in reinforcement-learning-based control. In such settings, limited perception, ambiguous terrain cues, and insufficient adaptation of gait timing can cause ev

Cited by 9SourcecodeScholar
2026

Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMs

AAAI 2026technical

Large Language Models (LLMs) have demonstrated remarkable proficiency in diverse tasks. This success raises a fundamental question in machine composition: Can symbolic music be considered a special form of language that can be jointly modeled with natural language for composition tasks? Recent studi

Cited by 0SourcePDFScholar
2026

Music Atelier: Exploring the Knowing–Doing Gap in LLM Creativity via Symbolic Music Composition

IJCAI 2026

The knowing--doing gap, the mismatch between ideas articulated during model reasoning and the realized creative artifact, remains a fundamental challenge in creative AI and persists in LLM-based artistic creation. This paper responds to this gap by introducing a systematic and interpretable framewor

Cited by 0Scholar
2026

On the Tension Between Optimality and Adversarial Robustness in Policy Optimization

ICLR 2026poster

Achieving optimality and adversarial robustness in deep reinforcement learning has long been regarded as conflicting goals. Nonetheless, recent theoretical insights presented in CAR suggest a potential alignment, raising the important question of how to realize this in practice. This paper first ide…

Cited by 0SourceScholar
2026

PINFDiT: Energy-Based Physics-Informed Diffusion Transformers for General-purpose Time Series Tasks

ICLR 2026poster

Time series analysis underpins scientific advances. While specialized models have advanced various time series tasks, scientific domains face unique challenges: limited samples with complex physical dynamics, missing observations, multi-resolution sampling, and requirements for physical consistency.…

Cited by 0SourceScholar
2026

Physics-Aware Spatiotemporal Causal Graph Network for Forecasting with Limited Data

ICML 2026poster

Spatiotemporal models have drawn significant interest recently due to their widespread applicability across many domains. These models are often made more practically useful by incorporating beneficial inductive biases, such as laws or symmetries from domain-relevant physics equations. This "physics…

Cited by 0SourceScholar
2026

PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localization

CVPR 2026

We present PiLoT, a unified framework that tackles UAV-based ego and target geo-localization. Conventional approaches rely on decoupled pipelines that fuse GNSS and Visual-Inertial Odometry (VIO) for ego-pose estimation, and active sensors like laser rangefinders for target localization. However, th

Cited by 0SourcecodeScholar
2026

Position: Beyond Prediction: Toward Verifiable Physiological Waveform Reasoning with Foundation Models and Agentic LLMs

ICML 2026poster

Physiological waveforms (e.g., ECG, PPG, EEG) encode clinically meaningful information in fine-grained morphology, precise timing, and cross-channel dynamics, yet most machine learning systems still treat them as generic time series and optimize end-to-end prediction. In this position paper, **we ar…

Cited by 0SourceScholar
2026

RAGFort: Dual-Path Defense Against Proprietary Knowledge Base Extraction in Retrieval-Augmented Generation

AAAI 2026technical

Retrieval-Augmented Generation (RAG) systems deployed over proprietary knowledge bases face growing threats from reconstruction attacks that aggregate model responses to replicate knowledge bases. Such attacks exploit both intra-class and inter-class paths—progressively extracting fine-grained knowl

Cited by 0SourcePDFScholar
2026

Self-Refine Learning in LLM Multi-Agent Systems for Legal Norm Cognition and Compliance

IJCAI 2026

As large language models (LLMs) increasingly serve as autonomous agents in social simulations, ensuring their ability to understand and comply with legal norms is essential. Yet, current LLM agents frequently exhibit reward hacking (RH) behaviors by optimizing metrics at the expense of norm adherenc

Cited by 0Scholar
2026

``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based Retrieval

ICML 2026poster

Large language models (LLMs) have been serving as effective backbones for retrieval systems, including Retrieval-Augmentation-Generation (RAG), Dense Information Retriever (IR), and Agent Memory Retrieval. Recent studies have demonstrated that such LLM-based Retrieval (LLMR) is vulnerable to adversa…

Cited by 0SourceScholar
2026

eRetinexGS: Retinex Modeling for Low-Light Scene Enhancement via Event Streams and 3D Gaussian Splatting

CVPR 2026

Perception under low illumination remains a major challenge for computer vision systems, as RGB sensors often fail to capture sufficient structural and color information in extremely dark environments. Event cameras, with their high dynamic range and temporal resolution, provide complementary cues t

Cited by 0SourceScholar
2025

A Large-scale Dataset and Benchmark for Commuting Origin-Destination Flow Generation

ICLR 2025poster

Commuting Origin-Destination~(OD) flows are critical inputs for urban planning and transportation, providing crucial information about the population residing in one region and working in another within an interested area. Due to the high cost of data collection, researchers have developed physical…

Cited by 0SourcePDFScholar
2025

A Unified Spatiotemporal Frequency Graph Neural Network for fMRI-based Brain Functional Connectivity Analysis

ICASSP 2025accepted

Analyzing functional connectivity patterns from resting-state functional magnetic resonance imaging (fMRI) requires unraveling its interrelations across spatial, temporal, and frequency domains. To comprehensively analyze four-dimensional (4D) fMRI data, we propose the Spatiotemporal Frequency Graph…

Cited by 0SourceScholar
2025

Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning

AAAI 2025technical

The capability of the reward model (RM) is crucial for the success of Reinforcement Learning from Human Feedback (RLHF) in aligning with human preferences. However, as training progresses, the output space distribution of the policy model shifts. The RM, initially trained on responses sampled from t…

Cited by 0SourcePDFScholar
2025

CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation

IJCAI 2025

Generative artificial intelligence in music has made significant strides, yet it still falls short of the substantial achievements seen in natural language processing, primarily due to the limited availability of music data. Knowledge-informed approaches have been shown to enhance the performance of

Cited by 0SourcePDFScholar
2025

Compose with Me: Collaborative Music Inpainter for Symbolic Music Infilling

AAAI 2025technical

The field of music generation has seen a surge of interest from both academia and industry, with innovative platforms such as Suno, Udio, and SkyMusic earning widespread recognition. However, the challenge of music infilling—modifying specific music segments without reconstructing the entire piece—r…

2025

E-NeMF: Event-based Neural Motion Field for Novel Space-time View Synthesis of Dynamic Scenes

ICCV 2025poster

Synthesizing novel space-time views from a monocular video is a highly ill-posed problem, and its effectiveness relies on accurately reconstructing motion and appearance of the dynamic scene.Frame-based methods for novel space-time view synthesis in dynamic scenes rely on simplistic motion assumptio…

Cited by 0SourcePDFScholar
2025

EDyGS: Event Enhanced Dynamic 3D Radiance Fields from Blurry Monocular Video

IJCAI 2025

The task of generating novel views in dynamic scenes plays a critical role in the 3D vision domain. Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS) have shown great promise in this domain but struggle with motion blur, which often arises in real-world scenarios due to camera or objec

2025

EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving

NeurIPS 2025poster

We introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 18…

Cited by 0SourceScholar
2025

FLIQA-AD: a Fusion Model with Large Language Model for Better Diagnose and MMSE Prediction of Alzheimer’s Disease

NAACL 2025short

Tracking a patient’s cognitive status early in the onset of the disease provides an opportunity to diagnose and intervene in Alzheimer’s disease (AD). However, relying solely on magnetic resonance imaging (MRI) images with traditional classification and regression models may not fully extract finer-…

Cited by 0SourcePDFScholar
2025

HighMATH: Evaluating Math Reasoning of Large Language Models in Breadth and Depth

EMNLP 2025

With the rapid development of large language models (LLMs) in math reasoning, the accuracy of models on existing math benchmarks has gradually approached 90% or even higher. More challenging math benchmarks are hence urgently in need to satisfy the increasing evaluation demands. To bridge this gap,

2025

Improving Generated and Retrieved Knowledge Combination Through Zero-shot Generation

ICASSP 2025accepted

Open-domain Question Answering (QA) has garnered substantial interest by combining the advantages of faithfully retrieved passages and relevant passages generated through Large Language Models (LLMs). However, there is a lack of definitive labels available to pair these sources of knowledge. In orde…

Cited by 0SourceScholar
2025

Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference Modeling

ACL 2025long

In Reinforcement Learning from Human Feedback (RLHF), the reward model (RM) evaluates the response quality based on the given context and assigns a reward. It plays a crucial role in aligning RLHF with human preferences. Although the current RM training paradigm concatenates the context and response…

Cited by 0SourcePDFScholar
2025

Mixture of Knowledge Minigraph Agents for Literature Review Generation

AAAI 2025technical

Literature reviews play a crucial role in scientific research for understanding the current state of research, identifying gaps, and guiding future studies on specific topics. However, the process of conducting a comprehensive literature review is yet time-consuming. This paper proposes a novel fram…

Cited by 0SourcePDFScholar
2025

Multi-Modality Test-Time Adaptation for Semantic Segmentation in Robotic Perception

ICRA 2025

Test-Time Adaptation (TTA) adjusts pre-trained models in unlabeled unseen environments during the test phase, making it more practical for robotic applications. However, the constant changes of the physical world create significant domain gaps between the received data during robot deployment and th

Cited by 0SourceScholar
2025

PsyAdvisor: A Plug-and-Play Strategy Advice Planner with Proactive Questioning in Psychological Conversations

ACL 2025long

Proactive questioning is essential in psychological conversations as it helps uncover deeper issues and unspoken concerns. Current psychological LLMs are constrained by passive response mechanisms, limiting their capacity to deploy proactive strategies for psychological counseling. To bridge this ga…

2025

SERENA: A Unified Stochastic Recursive Variance Reduced Gradient Framework for Riemannian Non-Convex Optimization

ICML 2025poster

Recently, the expansion of Variance Reduction (VR) to Riemannian stochastic non-convex optimization has attracted increasing interest. Inspired by recursive momentum, we first introduce Stochastic Recursive Variance Reduced Gradient (SRVRG) algorithm and further present Stochastic Recursive Gradient…

Cited by 0SourcePDFScholar
2025

Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models

NeurIPS 2025poster

With the growing power of text-to-image diffusion models, their potential to generate harmful or biased content has become a pressing concern, motivating the development of concept erasure techniques. Existing approaches, whether relying on retraining or not, frequently compromise the generative cap…

Cited by 0SourcecodeScholar
2024

Active Sequential Posterior Estimation for Sample-Efficient Simulation-Based Inference

NeurIPS 2024poster

Computer simulations have long presented the exciting possibility of scientific insight into complex real-world processes. Despite the power of modern computing, however, it remains challenging to systematically perform inference under simulation models. This has led to the rise of simulation-based…

2024

An Empirical Examination of Balancing Strategy for Counterfactual Estimation on Time Series

ICML 2024poster

Counterfactual estimation from observations represents a critical endeavor in numerous application fields, such as healthcare and finance, with the primary challenge being the mitigation of treatment bias. The balancing strategy aimed at reducing covariate disparities between different treatment gro…

Cited by 2SourcePDFScholar
2024

Beyond Mimicking Under-Represented Emotions: Deep Data Augmentation with Emotional Subspace Constraints for EEG-Based Emotion Recognition

AAAI 2024technical

In recent years, using Electroencephalography (EEG) to recognize emotions has garnered considerable attention. Despite advancements, limited EEG data restricts its potential. Thus, Generative Adversarial Networks (GANs) are proposed to mimic the observed distributions and generate EEG data. However,…

Cited by 14SourcePDFScholar
2024

CodeM: Less Data Yields More Versatility via Ability Matrix

ACL 2024findings

In the era of code large language models (code LLMs), data engineering plays a pivotal role during the instruction fine-tuning phase. To train a versatile model, previous efforts devote tremendous efforts into crafting instruction data covering all the downstream scenarios. Nonetheless, this will in…

2024

D3ETR: Decoder Distillation for Detection Transformer

IJCAI 2024poster

Although various knowledge distillation (KD) methods for CNN-based detectors have been proven effective in improving small students, build- ing baselines and recipes for DETR-based detec- tors remains a challenge. This paper concentrates on the transformer decoder of DETR-based detec- tors and explo…

Cited by 18SourcePDFScholar
2024

GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting

AAAI 2024technical

Time series forecasting is an essential area of machine learning with a wide range of real-world applications. Most of the previous forecasting models aim to capture dynamic characteristics from uni-modal numerical historical data. Although extra knowledge can boost the time series forecasting perfo…

Cited by 55SourcePDFScholar
2024

Improving Long Text Understanding with Knowledge Distilled from Summarization Model

ICASSP 2024accepted

Long text understanding is important yet challenging for natural language processing. A long article or document usually contains many redundant words that are not pertinent to its gist and sometimes can be regarded as noise. With recent advances of abstractive summarization, we propose our Gist Det…

Cited by 0SourceScholar
2024

LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin

ACL 2024long

Supervised fine-tuning (SFT) is a crucial step for large language models (LLMs), enabling them to align with human instructions and enhance their capabilities in downstream tasks. Substantially increasing instruction data is a direct solution to align the model with a broader range of downstream tas…

2024

MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music

IJCAI 2024poster

The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music. However, due to semantic gaps between Music Information Retrieval (MIR) algorithms and human understanding, discrepanc…

2024

Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization

EMNLP 2024finding

Large language models (LLMs) have revolutionized the role of AI, yet pose potential social risks. To steer LLMs towards human preference, alignment technologies have been introduced and gained increasing attention. Nevertheless, existing methods heavily rely on high-quality positive-negative trainin…

2024

Physics-Constrained Comprehensive Optical Neural Networks

NeurIPS 2024poster

With the advantages of low latency, low power consumption, and high parallelism, optical neural networks (ONN) offer a promising solution for time-sensitive and resource-limited artificial intelligence applications. However, the performance of the ONN model is often diminished by the gap between the…

Cited by 1SourcePDFScholar
2024

Position: TrustLLM: Trustworthiness in Large Language Models

ICML 2024poster

Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLM…

Cited by 95SourcePDFScholar
2024

Responding to the Call: Exploring Automatic Music Composition Using a Knowledge-Enhanced Model

AAAI 2024technical

Call-and-response is a musical technique that enriches the creativity of music, crafting coherent musical ideas that mirror the back-and-forth nature of human dialogue with distinct musical characteristics. Although this technique is integral to numerous musical compositions, it remains largely unch…

2024

Spiking NeRF: Representing the Real-World Geometry by a Discontinuous Representation

AAAI 2024technical

A crucial reason for the success of existing NeRF-based methods is to build a neural density field for the geometry representation via multiple perceptron layers (MLPs). MLPs are continuous functions, however, real geometry or density field is frequently discontinuous at the interface between the ai…

2024

StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback

ACL 2024long

The advancement of large language models (LLMs) has significantly propelled the field of code generation. Previous work integrated reinforcement learning (RL) with compiler feedback for exploring the output space of LLMs to enhance code generation quality. However, the lengthy code generated by LLMs…

2024

TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting

ICLR 2024poster

The past decade has witnessed significant advances in time series modeling with deep learning. While achieving state-of-the-art results, the best-performing architectures vary highly across applications and domains. Meanwhile, for natural language processing, the Generative Pre-trained Transformer (…

2024

TextGenSHAP: Scalable Post-Hoc Explanations in Text Generation with Long Documents

ACL 2024findings

Large language models (LLMs) have attracted great interest in many real-world applications; however, their “black-box” nature necessitates scalable and faithful explanations. Shapley values have matured as an explainability method for deep learning, but extending them to LLMs is difficult due to lon…

Cited by 5SourcePDFScholar
2024

The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models

ICLR 2024poster

Pre-trained Language models (PLMs) have been acknowledged to contain harmful information, such as social biases, which may cause negative social impacts or even bring catastrophic results in application. Previous works on this problem mainly focused on using black-box methods such as probing to dete…

Cited by 15SourcePDFScholar
2023

A Novel Efficient Multi-View Traffic-Related Object Detection Framework

ICASSP 2023accepted

With the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Acc…

Cited by 0SourceScholar
2023

DSRM: Boost Textual Adversarial Training with Distribution Shift Risk Minimization

ACL 2023long

Adversarial training is one of the best-performing methods in improving the robustness of deep language models. However, robust models come at the cost of high time consumption, as they require multi-step gradient ascents or word substitutions to obtain adversarial samples. In addition, these genera…

2023

Estimating Treatment Effects from Irregular Time Series Observations with Hidden Confounders

AAAI 2023technical

Causal analysis for time series data, in particular estimating individualized treatment effect (ITE), is a key task in many real world applications, such as finance, retail, healthcare, etc. Real world time series, i.e., large-scale irregular or sparse and intermittent time series, raise significan…

Cited by 20SourcePDFScholar
2023

GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular Video

ICCV 2023poster

3D human pose estimation has been researched for decades with promising fruits. 3D human pose lifting is one of the promising research directions toward the task where both estimated pose and ground truth pose data are used for training. Existing pose lifting works mainly focus on improving the perf…

Cited by 80PDFcodeScholar
2023

Hierarchical Gaussian Mixture based Task Generative Model for Robust Meta-Learning

NeurIPS 2023poster

Meta-learning enables quick adaptation of machine learning models to new tasks with limited data. While tasks could come from varying distributions in reality, most of the existing meta-learning methods consider both training and testing tasks as from the same uni-component distribution, overlooking…

Cited by 1SourcePDFScholar
2023

SVGformer: Representation Learning for Continuous Vector Graphics Using Transformers

CVPR 2023poster

Advances in representation learning have led to great success in understanding and generating data in various domains. However, in modeling vector graphics data, the pure data-driven approach often yields unsatisfactory results in downstream tasks as existing deep learning methods often require the…

Cited by 9SourcePDFScholar
2023

Uncovering and Categorizing Social Biases in Text-to-SQL

ACL 2023long

Large pre-trained language models are acknowledged to carry social bias towards different demographics, which can further amplify existing stereotypes in our society and cause even more harm. Text-to-SQL is an important task, models of which are mainly adopted by administrative industries, where unf…

2023

Uncovering and Quantifying Social Biases in Code Generation

NeurIPS 2023poster

With the popularity of automatic code generation tools, such as Copilot, the study of the potential hazards of these tools is gaining importance. In this work, we explore the social bias problem in pre-trained code generation models. We propose a new paradigm to construct code prompts and successful…

Cited by 19SourcePDFScholar
2022

Counterfactual Neural Temporal Point Process for Estimating Causal Influence of Misinformation on Social Media

NeurIPS 2022accept

Recent years have witnessed the rise of misinformation campaigns that spread specific narratives on social media to manipulate public opinions on different areas, such as politics and healthcare. Consequently, an effective and efficient automatic methodology to estimate the influence of the misinfor…

Cited by 24SourcePDFScholar
2022

EGCN: An Ensemble-based Learning Framework for Exploring Effective Skeleton-based Rehabilitation Exercise Assessment

IJCAI 2022poster

Recently, some skeleton-based physical therapy systems have been attempted to automatically evaluate the correctness or quality of an exercise performed by rehabilitation subjects. However, in terms of algorithms and evaluation criteria, the task remains not fully explored regarding making full use…

2022

I-SEA: Importance Sampling and Expected Alignment-Based Deep Distance Metric Learning for Time Series Analysis and Embedding

AAAI 2022technical

Learning effective embeddings for potentially irregularly sampled time-series, evolving at different time scales, is fundamental for machine learning tasks such as classification and clustering. Task-dependent embeddings rely on similarities between data samples to learn effective geometries. Howeve…

2022

MPII: Multi-Level Mutual Promotion for Inference and Interpretation

ACL 2022long

In order to better understand the rationale behind model behavior, recent works have exploited providing interpretation to support the inference prediction. However, existing methods tend to provide human-unfriendly interpretation, and are prone to sub-optimal performance due to one-side promotion,…

2022

Non-Autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition

ICASSP 2022accepted

Non-autoregressive (NAR) transformer models have been studied intensively in automatic speech recognition (ASR), and many NAR transformer models is to use the causal mask to limit token dependencies. However, the causal mask is designed for the left-to-right decoding process of the non-parallel auto…

Cited by 0SourceScholar
2022

Physics-Informed Long-Sequence Forecasting From Multi-Resolution Spatiotemporal Data

IJCAI 2022poster

Spatiotemporal data aggregated over regions or time windows at various resolutions demonstrate heterogeneous patterns and dynamics in each resolution. Meanwhile, the multi-resolution characteristic provides rich contextual information, which is critical for effective long-sequence forecasting. The i…

2022

Sparse Interaction Additive Networks via Feature Interaction Detection and Sparse Selection

NeurIPS 2022accept

There is currently a large gap in performance between the statistically rigorous methods like linear regression or additive splines and the powerful deep methods using neural networks. Previous works attempting to close this gap have failed to fully consider the exponentially growing number of feat…

2021

Depth Privileged Object Detection in Indoor Scenes via Deformation Hallucination

AAAI 2021technical

RGB-D object detection has achieved significant advance, because depth provides complementary geometric information to RGB images. Considering depth images are unavailable in some scenarios, we focus on depth privileged object detection in indoor scenes, where the depth images are only available in…

Cited by 7SourcePDFScholar
2021

Mixed Supervised Object Detection by Transferring Mask Prior and Semantic Similarity

NeurIPS 2021poster

Object detection has achieved promising success, but requires large-scale fully-annotated data, which is time-consuming and labor-extensive. Therefore, we consider object detection with mixed supervision, which learns novel object categories using weak annotations with the help of full annotations o…

2021

Multimodal Fusion via Teacher-Student Network for Indoor Action Recognition

AAAI 2021technical

Indoor action recognition plays an important role in modern society, such as intelligent healthcare in large mobile cabin hospitals. With the wide usage of depth sensors like Kinect, multimodal information including skeleton and RGB modalities brings a promising way to improve the performance. Howev…

2021

Physics-aware Spatiotemporal Modules with Auxiliary Tasks for Meta-Learning

IJCAI 2021poster

Modeling the dynamics of real-world physical systems is critical for spatiotemporal prediction tasks, but challenging when data is limited. The scarcity of real-world data and the difficulty in reproducing the data distribution hinder directly applying meta-learning techniques. Although the knowledg…

Cited by 10SourcePDFScholar
2021

VigDet: Knowledge Informed Neural Temporal Point Process for Coordination Detection on Social Media

NeurIPS 2021poster

Recent years have witnessed an increasing use of coordinated accounts on social media, operated by misinformation campaigns to influence public opinion and manipulate social outcomes. Consequently, there is an urgent need to develop an effective methodology for coordinated group detection to combat…

Cited by 34SourcePDFScholar
2020

Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction Detection

ICLR 2020poster

Recommendation is a prevalent application of machine learning that affects many users; therefore, it is important for recommender models to be accurate and interpretable. In this work, we propose a method to both interpret and augment the predictions of black-box recommender systems. In particular,…

Cited by 75SourcecodeScholar
2020

How does This Interaction Affect Me? Interpretable Attribution for Feature Interactions

NeurIPS 2020poster

Machine learning transparency calls for interpretable explanations of how inputs relate to predictions. Feature attribution is a way to analyze the impact of features on predictions. Feature interactions are the contextual dependence between features that jointly impact predictions. There are a numb…

2020

Multi-agent Trajectory Prediction with Fuzzy Query Attention

NeurIPS 2020poster

Trajectory prediction for scenes with multiple agents and entities is a challenging problem in numerous domains such as traffic prediction, pedestrian tracking and path planning. We present a general architecture to address this challenge which models the crucial inductive biases of motion, namely,…

2020

NoiseRank: Unsupervised Label Noise Reduction with Dependence Models

ECCV 2020poster

Label noise is increasingly prevalent in datasets acquired from noisy channels. Existing approaches that detect and remove label noise generally rely on some form of supervision, which is not scalable and error-prone. In this paper, we propose NoiseRank, for unsupervised label noise reduction using…

Cited by 43SourcePDFScholar
2020

Semi-Supervised Crowd Counting via Self-Training on Surrogate Tasks

ECCV 2020poster

Most existing crowd counting systems rely on the availability of the object location annotation which can be expensive to obtain. To reduce the annotation cost, one attractive solution is to leverage a large number of unlabeled images to build a crowd counting model in semi-supervised fashion. This…

Cited by 96SourcePDFScholar
2019

Adaptive Gradient Methods with Dynamic Bound of Learning Rate

ICLR 2019poster

Adaptive optimization methods such as AdaGrad, RMSprop and Adam have been proposed to achieve a rapid training process with an element-wise scaling term on learning rates. Though prevailing, they are observed to generalize poorly compared with SGD or even fail to converge due to unstable and extreme…

2018

Automatically Inferring Data Quality for Spatiotemporal Forecasting

ICLR 2018poster

Spatiotemporal forecasting has become an increasingly important prediction task in machine learning and statistics due to its vast applications, such as climate modeling, traffic prediction, video caching predictions, and so on. While numerous studies have been conducted, most existing works assume…

Cited by 9SourcePDFScholar
2018

Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting

ICLR 2018poster

Spatiotemporal forecasting has various applications in neuroscience, climate and transportation domain. Traffic forecasting is one canonical example of such learning task. The task is challenging due to (1) complex spatial dependency on road networks, (2) non-linear temporal dynamics with changing r…

2018

Hierarchical Deep Generative Models for Multi-Rate Multivariate Time Series

ICML 2018oral

Multi-Rate Multivariate Time Series (MR-MTS) are the multivariate time series observations which come with various sampling rates and encode multiple temporal dependencies. State-space models such as Kalman filters and deep learning models such as deep Markov models are mainly designed for time seri…

Cited by 59SourcePDFScholar
2018

Neural Interaction Transparency (NIT): Disentangling Learned Interactions for Improved Interpretability

NeurIPS 2018poster

Neural networks are known to model statistical interactions, but they entangle the interactions at intermediate hidden layers for shared representation learning. We propose a framework, Neural Interaction Transparency (NIT), that disentangles the shared learning across different interactions to obta…

Cited by 85SourcePDFScholar
2017

Variational Recurrent Adversarial Deep Domain Adaptation

ICLR 2017poster

We study the problem of learning domain invariant representations for time series data while transferring the complex temporal latent dependencies between the domains. Our model termed as Variational Recurrent Adversarial Deep Domain Adaptation (VRADA) is built atop a variational recurrent neural ne…

Cited by 185SourceScholar
2016

SPALS: Fast Alternating Least Squares via Implicit Leverage Scores Sampling

NeurIPS 2016poster

Tensor CANDECOMP/PARAFAC (CP) decomposition is a powerful but computationally challenging tool in modern data analytics. In this paper, we show ways of sampling intermediate steps of alternating minimization algorithms for computing low rank tensor CP decompositions, leading to the sparse alternatin…

2015

Accelerated Online Low Rank Tensor Learning for Multivariate Spatiotemporal Streams

ICML 2015poster

Low-rank tensor learning has many applications in machine learning. A series of batch learning algorithms have achieved great successes. However, in many emerging applications, such as climate data analysis, we are confronted with large-scale tensor streams, which poses significant challenges to exi…

Cited by 79SourcePDFScholar
2015

Functional Subspace Clustering with Application to Time Series

ICML 2015poster

Functional data, where samples are random functions, are increasingly common and important in a variety of applications, such as health care and traffic analysis. They are naturally high dimensional and lie along complex manifolds. These properties warrant use of the subspace assumption, but most st…

Cited by 40SourcePDFScholar
2015

HawkesTopic: A Joint Model for Network Inference and Topic Modeling from Text-Based Cascades

ICML 2015poster

Understanding the diffusion of information in social network and social media requires modeling the text diffusion process. In this work, we develop the HawkesTopic model (HTM) for analyzing text-based cascades, such as "retweeting a post" or "publishing a follow-up blog post". HTM combines Hawkes p…

Cited by 123SourcePDFScholar