← Search

Li Li

95 accepted papers

2026

ARDiff: Anisotropic Residual Diffusion for Heterogeneous Graph Learning

AAAI 2026technical

Learning representations on graphs is foundational for many downstream tasks, and its synergy with diffusion models has emerged as a promising direction. However, diffusion-based methods for heterogeneous graphs remain underexplored, confronting two principal challenges: (1) The presence of noise an

Cited by 0SourcePDFScholar
2026

Charts Are Not Images: On the Challenges of Scientific Chart Editing

ICLR 2026poster

Generative models, such as diffusion and autoregressive approaches, have demonstrated impressive capabilities in editing natural images. However, applying these tools to scientific charts rests on a flawed assumption: a chart is not merely an arrangement of pixels but a visual representation of stru…

Cited by 0SourcecodeScholar
2026

CoAct-1: Computer-using Multi-agent System with Coding Actions

ICLR 2026poster

Autonomous agents that operate computers via Graphical User Interfaces (GUIs) often struggle with efficiency and reliability on complex, long-horizon tasks. While augmenting these agents with planners can improve task decomposition, they remain constrained by the inherent limitations of performing a…

Cited by 0SourcecodeScholar
2026

CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs

ICLR 2026poster

Chain-of-Thought (CoT) prompting has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing implementations, such as in-context learning and fine-tuning, remain costly and inefficient. To improve CoT reasoning at a lower cost, and in…

Cited by 0SourceScholar
2026

Content-Adaptive Hierarchical Hyperprior for Neural Video Coding

CVPR 2026

While neural video codecs (NVCs) have recently demonstrated superior performance over traditional codecs through end-to-end learning, existing approaches primarily focus on architectural enhancements and coding module design, with limited exploration into optimizing hierarchical structures--specific

Cited by 0SourceScholar
2026

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

ICLR 2026poster

Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities. However, the intrinsic heterogeneity of diverse modalities presents substantial challenges to achieve effective cross-modal collaboration and integration. To address this…

Cited by 0SourcecodeScholar
2026

Developmental Federated Tuning: A Cognitive-Inspired Paradigm for Efficient LLM Adaptation

ICLR 2026poster

Federated fine-tuning enables Large Language Models (LLMs) to adapt to downstream tasks while preserving data privacy, but its resource-intensive nature severely limits deployment on edge devices. In this paper, we introduce Developmental Federated Tuning (DevFT), a resource-efficient approach inspi…

Cited by 0SourceScholar
2026

Diffusion-Enhanced Tree Planning for Autonomous Driving

RA-L 2026

In highly interactive urban driving, decision making is often naturally multi-stage, and decisions at different stages can lead to different reactions from surrounding vehicles. This calls for stage-wise evaluation and selection. Tree-based planning naturally supports multi-stage search and evaluati

Cited by 0SourceScholar
2026

Disentangling Length Bias in Preference Learning via Response-Conditioned Modeling

ICLR 2026poster

Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward model and employing a reinforcement learning algorithm to maximize the reward model's scores. However, these reward model…

Cited by 0SourceScholar
2026

Don't Reinvent the Wheel, Just Realign the Spokes: Resource-Efficient Federated Fine-Tuning via Rank-Wise Expert Assembly

ICML 2026spotlight

Federated fine-tuning presents a promising avenue for adapting Large Language Models (LLMs) to downstream tasks while preserving data privacy. However, the prohibitive computational and communication overhead of LLM adaptation inhibits its deployment on resource-constrained edge devices. In this pap…

Cited by 0SourceScholar
2026

EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model

ICLR 2026poster

Subject-driven generation is a critical task in creative AI; yet current state-of-the-art methods present a stark trade-off. They either rely on computationally expensive, per-subject fine-tuning, sacrificing efficiency and zero-shot capability, or employ feed-forward architectures built on diffusio…

Cited by 0SourcecodeScholar
2026

Mitigating Hallucinations in Large Language Models via Causal Reasoning

AAAI 2026technical

Large language models (LLMs) exhibit logically inconsistent hallucinations that appear coherent yet violate reasoning principles, with recent research suggesting an inverse relationship between causal reasoning capabilities and such hallucinations. However, existing reasoning approaches in LLMs, suc

Cited by 0SourcePDFScholar
2026

Perceptual Neural Video Compression with Color Separation and Rank Chain

CVPR 2026

Neural video compression (NVC) has achieved significant progress in recent years. The state-of-the-art (SOTA) NVC schemes, exemplified by the Deep Conditional Video Coding series, have focused on pursuing higher fidelity (e.g., PSNR), but lack sufficient exploitation of deep networks' advantages for

Cited by 0SourcecodeScholar
2026

Real-Time Neural Video Compression with Unified Intra and Inter Coding

CVPR 2026

Neural video compression (NVC) technologies have advanced rapidly in recent years, yielding state-of-the-art schemes such as DCVC-RT that offer superior compression efficiency to H.266/VVC and real-time encoding/decoding capabilities. Nonetheless, existing NVC schemes have several limitations, inclu

Cited by 0SourcecodeScholar
2026

Transform-Free Feature Coding via Entropy-Constrained Vector Quantization

AAAI 2026technical

Feature coding has recently emerged as a key technique for efficient transmission of intermediate representations in distributed AI systems. Existing approaches largely follow a transform-based pipeline inherited from image and video coding, where the transform module is used to remove spatial struc

Cited by 0SourcePDFScholar
2026

``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based Retrieval

ICML 2026poster

Large language models (LLMs) have been serving as effective backbones for retrieval systems, including Retrieval-Augmentation-Generation (RAG), Dense Information Retriever (IR), and Agent Memory Retrieval. Recent studies have demonstrated that such LLM-based Retrieval (LLMR) is vulnerable to adversa…

Cited by 0SourceScholar
2025

AD-LLM: Benchmarking Large Language Models for Anomaly Detection

ACL 2025finding

Anomaly detection (AD) is an important machine learning task with many real-world uses, including fraud detection, medical diagnosis, and industrial monitoring. Within natural language processing (NLP), AD helps detect issues like spam, misinformation, and unusual user activity. Although large langu…

2025

BEV-PolyNet: BEV-Based Polygonal End to End Parking Slot Detection Framework

RA-L 2025

Accurate parking slot detection is crucial for autonomous parking and intelligent driving, directly impacting safety and efficiency. However, most existing methods rely on AVM images, making them susceptible to image distortion and vehicle occlusion. Additionally, many approaches independently regre

Cited by 0SourceScholar
2025

DropoutGS: Dropping Out Gaussians for Better Sparse-view Rendering

CVPR 2025poster

Although 3D Gaussian Splatting (3DGS) has demonstrated promising results in novel view synthesis, its performance degrades dramatically with sparse inputs and generates undesirable artifacts. As the number of training views decreases, the novel view synthesis task degrades to a highly under-determin…

Cited by 0SourcePDFScholar
2025

Dur360BEV: A Real-World 360-Degree Single Camera Dataset and Benchmark for Bird-Eye View Mapping in Autonomous Driving

ICRA 2025

We present Dur360BEV, a novel spherical camera autonomous driving dataset equipped with a high-resolution 128-channel 3D LiDAR and a RTK-refined GNSS/INS system, along with a benchmark architecture designed to generate Bird-Eye-View (BEV) maps using only a single spherical camera. This dataset and b

Cited by 4SourceScholar
2025

Exploring the Efficacy of Multi-Agent Reinforcement Learning for Autonomous Cyber Defence: A CAGE Challenge 4 Perspective

AAAI 2025technical

As cyber threats become increasingly automated and sophisticated, novel solutions must be introduced to improve defence of enterprise networks. Deep Reinforcement Learning (DRL) has demonstrated potential in mitigating these advanced threats. Single DRL Agents have proven utility toward execution of…

2025

Few-Shot Domain Adaptation for Learned Image Compression

AAAI 2025technical

Learned image compression (LIC) has achieved state-of-the-art rate-distortion performance, deemed promising for next-generation image compression techniques. However, pre-trained LIC models usually suffer from significant performance degradation when applied to out-of-training-domain images, implyin…

Cited by 0SourcePDFScholar
2025

Flexible Group Count Enables Hassle-Free Structured Pruning

CVPR 2025poster

Densely structured pruning methods -- which generate pruned models in a fully dense format, allowing immediate compression benefits without additional demands -- are evolving owing to their practical significance. Traditional techniques in this domain mainly revolve around coarser granularities, suc…

Cited by 0SourcePDFScholar
2025

G-Depth: An Efficient Graph Method for Robust Depth Completion

ICASSP 2025accepted

Depth completion has played a vital role in enhancing depth perception in various applications such as autonomous driving, robotics, and 3D reconstruction. It is a critical task that involves generating a dense depth map from an RGB image and an aligned sparse depth map. However, many existing studi…

Cited by 0SourceScholar
2025

JPG-SLAM: Joint Point-Gaussian Splatting Representation for Dense Dynamic SLAM

ICRA 2025

This paper presents a simultaneous localization and mapping (SLAM) system to provide accurate pose estimation and dynamic scene reconstruction. Our approach proposes a Joint Point-Gaussian Splatting representation, which fully integrates the robustness of isotropic feature points in pose estimation

Cited by 3SourceScholar
2025

Learned Image Compression with Hierarchical Progressive Context Modeling

ICCV 2025poster

Context modeling is essential in learned image compression for accurately estimating the distribution of latents. While recent advanced methods have expanded context modeling capacity, they still struggle to efficiently exploit long-range dependency and diverse context information across different c…

2025

Leveraging Spatial Invariance to Boost Adversarial Transferability

ICCV 2025poster

Adversarial examples, crafted with imperceptible perturbations, reveal a significant vulnerability of Deep Neural Networks (DNNs). More critically, the transferability of adversarial examples allows attackers to induce unreasonable predictions without requiring knowledge about the target model. DNNs…

2025

LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem

EMNLP 2025

Backdoor attacks are powerful and effective, but distributing LLMs without a proven track record like ‘meta-llama‘ or ‘qwen‘ rarely gains community traction. We identify LoRA sharing as a unique scenario where users are more willing to try unendorsed assets, since such shared LoRAs allow them to enj

2025

Treble Counterfactual VLMs: A Causal Approach to Hallucination

EMNLP 2025

Vision-Language Models (VLMs) excel at tasks such as image captioning and visual question answering but frequently produce hallucinated outputs that deviate from the actual visual input or prompt. While prior work links hallucination to biases in data or representation, their causal origins remain u

2025

Tree of Agents: Improving Long-Context Capabilities of Large Language Models through Multi-Perspective Reasoning

EMNLP 2025

Large language models (LLMs) face persistent challenges when handling long-context tasks, most notably the “lost in the middle” issue, where information located in the middle of a long input tends to be underutilized. Some existing methods that reduce input have the risk of discarding key informatio

2025

Unsupervised Kernel-based Multi-view Feature Selection with Robust Self-representation and Binary Hashing

AAAI 2025technical

Unsupervised multi-view feature selection involves selecting a subset of crucial features across diverse views to diminish feature dimensionality without leveraging label information. While numerous studies have delved into this area, current solutions predominantly rely on linear multi-view data or…

Cited by 0SourcePDFScholar
2025

VETA-DiT: Variance-Equalized and Temporally Adaptive Quantization for Efficient 4-bit Diffusion Transformers

NeurIPS 2025poster

Diffusion Transformers (DiTs) have recently demonstrated remarkable performance in visual generation tasks, surpassing traditional U-Net-based diffusion models by significantly improving image and video generation quality and scalability. However, the large model size and iterative denoising process…

Cited by 0SourcecodeScholar
2024

Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition

ICASSP 2024accepted

How to effectively use limited supervised accent data to improve the accented ASR is of paramount importance. In this work, we propose an accent-specific quantization for joint unsupervised and supervised training (AQ-JUST) of end-to-end ASR models to address this issue. Specifically, two variants o…

Cited by 0SourceScholar
2024

Bayesian Deep Predictive Coding for Snake-like Robotic Control in Unknown Terrains

IROS 2024

Effectively modeling the spatio-temporal interactions both internally and externally is a challenge in controlling multi-linked snake robots. This paper presents an effective method based on deep predictive coding: SnakeFormer, to address the aforementioned issue. The main contributions include: 1)

Cited by 0SourceScholar
2024

Darkshot: Lighting Dark Images with Low-Compute and High-Quality

ICASSP 2024accepted

Nighttime photography encounters escalating challenges in extremely low-light conditions, primarily attributable to the ultra-low signal-to-noise ratio. For real-world deployment, a practical solution must not only produce visually appealing results but also require minimal computation. However, mos…

Cited by 0SourceScholar
2024

Domain-Wise Invariant Learning for Panoptic Scene Graph Generation

ICASSP 2024accepted

Panoptic Scene Graph Generation (PSG) involves the detection of objects and the prediction of their corresponding relationships (predicates). However, the presence of biased predicate annotations poses a significant challenge for PSG models, as it hinders their ability to establish a clear decision…

Cited by 0SourceScholar
2024

Evaluating Step-by-Step Reasoning through Symbolic Verification

NAACL 2024findings

Pre-trained language models (LMs) have shown remarkable reasoning performance using explanations or chain-of-thoughts (CoT)) for in-context learning. On the other hand, these reasoning tasks are usually presumed to be more approachable for symbolic programming. To understand the mechanism of reasoni…

2024

FedGCS: A Generative Framework for Efficient Client Selection in Federated Learning via Gradient-based Optimization

IJCAI 2024poster

Federated Learning faces significant challenges in statistical and system heterogeneity, along with high energy consumption, necessitating efficient client selection strategies. Traditional approaches, including heuristic and learning-based methods, fall short of addressing these complexities holist…

2024

Forest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis

ECCV 2024poster

"Synthesizing realistic 3D indoor scenes is a challenging task that traditionally relies on manual arrangement and annotation by expert designers. Recent advances in autoregressive models have automated this process, but they often lack semantic understanding of the relationships and hierarchies pre…

Cited by 7SourcePDFScholar
2024

GuessKT: Improving Knowledge Tracing via Considering Guess Behaviors

ICASSP 2024accepted

Knowledge tracing (KT) aims to predict students’ responses to given questions based on their historical question-answering interactions. Recent studies have proposed multiple types of KT models, mainly relying on learners’ feedback to capture the evolution of their knowledge states. However, these m…

Cited by 0SourceScholar
2024

How to Configure Good In-Context Sequence for Visual Question Answering

CVPR 2024poster

Inspired by the success of Large Language Models in dealing with new tasks via In-Context Learning (ICL) in NLP researchers have also developed Large Vision-Language Models (LVLMs) with ICL capabilities. However when implementing ICL using these LVLMs researchers usually resort to the simplest way l…

2024

HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-Tuning

NeurIPS 2024oral

Adapting Large Language Models (LLMs) to new tasks through fine-tuning has been made more efficient by the introduction of Parameter-Efficient Fine-Tuning (PEFT) techniques, such as LoRA. However, these methods often underperform compared to full fine-tuning, particularly in scenarios involving comp…

2024

Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation

CVPR 2024poster

As a new embodied vision task Instance ImageGoal Navigation (IIN) aims to navigate to a specified object depicted by a goal image in an unexplored environment. The main challenge of this task lies in identifying the target object from different viewpoints while rejecting similar distractors. Existin…

2024

MGS-SLAM: Monocular Sparse Tracking and Gaussian Mapping With Depth Smooth Regularization

RA-L 2024

This letter introduces a novel framework for dense Visual Simultaneous Localization and Mapping (VSLAM) based on Gaussian Splatting. Recently, SLAM based on Gaussian Splatting has shown promising results. However, in monocular scenarios, the Gaussian maps reconstructed lack geometric accuracy and ex

Cited by 22SourceScholar
2024

Mixing Left and Right-Hand Driving Data in a Hierarchical Framework With LLM Generation

RA-L 2024

Data-driven trajectory prediction is critical in autonomous vehicles, which requires high-quality data. However, discussions about the compatibility of data collected from different countries remain limited, with a typical issue being the different driving rules in various countries. Therefore, we p

Cited by 5SourceScholar
2024

Offline and Online Optical Flow Enhancement for Deep Video Compression

AAAI 2024technical

Video compression relies heavily on exploiting the temporal redundancy between video frames, which is usually achieved by estimating and using the motion information. The motion information is represented as optical flows in most of the existing deep video compression networks. Indeed, these network…

Cited by 20SourcePDFScholar
2024

Panoptic Scene Graph Generation with Semantics-Prototype Learning

AAAI 2024technical

Panoptic Scene Graph Generation (PSG) parses objects and predicts their relationships (predicate) to connect human language and visual scenes. However, different language preferences of annotators and semantic overlaps between predicates lead to biased predicate annotations in the dataset, i.e. diff…

2024

Ranking-based Client Imitation Selection for Efficient Federated Learning

ICML 2024poster

Federated Learning (FL) enables multiple devices to collaboratively train a shared model while ensuring data privacy. The selection of participating devices in each training round critically affects both the model performance and training efficiency, especially given the vast heterogeneity in traini…

Cited by 3SourcePDFScholar
2024

SANet: Small but Accurate Detector for Aerial Flying Object

ICRA 2024poster

This paper proposes SANet, a small but accurate detector for aerial flying objects. The detector introduces an attention module into the feature extraction module (FEM) for enhancing the accuracy. This FEM with fewer convolutional kernel channels can reduce the parameters, speed up the inference tim…

Cited by 1SourceScholar
2024

SRECT: Machine-Specific Spatial-Resolution Enhancement in Computed Tomography

ICASSP 2024accepted

Computed Tomography (CT) is an advanced imaging technology. To obtain high-resolution (HR) CT images from low-resolution (LR) sinograms, we present a deep-learning (DL) based CT super-resolution (SR) method.The proposed method combines a SR model in the sinogram domain and the iterative framework in…

Cited by 0SourceScholar
2023

ADMNet: Anti-Drone Real-Time Detection and Monitoring

IROS 2023poster

We propose a lightweight, effective, and efficient anti-drone network, namely ADMNet, for visually detecting and monitoring unfriendly drones with a constrained view field, flying against a complex environment. We merge an SPP module to the first head of YOLOv4 to improve accuracy and perform networ…

Cited by 4SourceScholar
2023

CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI Detection

NeurIPS 2023poster

Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories. A strong zero-shot HOI detector is supposed to be not only capable of discriminating novel interactions but also robust to positional distribution discrepancy between seen and unseen categories w…

Cited by 21SourcePDFScholar
2023

Cyclic-Bootstrap Labeling for Weakly Supervised Object Detection

ICCV 2023poster

Recent progress in weakly supervised object detection is featured by a combination of multiple instance detection networks (MIDN) and ordinal online refinement. However, with only image-level annotation, MIDN inevitably assigns high scores to some unexpected region proposals when generating pseudo l…

Cited by 12PDFcodeScholar
2023

HandNeRF: Neural Radiance Fields for Animatable Interacting Hands

CVPR 2023poster

We propose a novel framework to reconstruct accurate appearance and geometry with neural radiance fields (NeRF) for interacting hands, enabling the rendering of photo-realistic images and videos for gesture animation from arbitrary views. Given multi-view images of a single hand or interacting hands…

Cited by 29SourcePDFScholar
2023

Less Is More: Reducing Task and Model Complexity for 3D Point Cloud Semantic Segmentation

CVPR 2023poster

Whilst the availability of 3D LiDAR point cloud data has significantly grown in recent years, annotation remains expensive and time-consuming, leading to a demand for semi-supervised semantic segmentation methods with application domains such as autonomous driving. Existing work very often employs r…

2023

Neural Characteristic Function Learning for Conditional Image Generation

ICCV 2023poster

The emergence of conditional generative adversarial networks (cGANs) has revolutionised the way we approach and control the generation, by means of adversarially learning joint distributions of data and auxiliary information. Despite the success, cGANs have been consistently put under scrutiny due t…

Cited by 7PDFcodeScholar
2023

Reinforcement Learning Based Multi-Layer Bayesian Control for Snake Robots in Cluttered Scenes

IROS 2023poster

The majority of current research on reinforcement learning (RL) for snake robot control do not sufficiently account for the spatial and temporal dependencies within the robot or its interaction with its environment during movement. To address this issue, we propose an RL based multi-layer Bayesian m…

Cited by 2SourceScholar
2023

SimFIR: A Simple Framework for Fisheye Image Rectification with Self-supervised Representation Learning

ICCV 2023poster

In fisheye images, rich distinct distortion patterns are regularly distributed in the image plane. These distortion patterns are independent of the visual content and provide informative cues for rectification. To make the best of such rectification cues, we introduce SimFIR, a simple framework for…

Cited by 31PDFScholar
2023

VPGTrans: Transfer Visual Prompt Generator across LLMs

NeurIPS 2023poster

Since developing a new multimodal LLM (MLLM) by pre-training on tremendous image-text pairs from scratch can be exceedingly resource-consuming, connecting an existing LLM with a comparatively lightweight visual prompt generator (VPG) becomes a feasible paradigm. However, further tuning the VPG compo…

2022

An Information Fusion Approach to Learning with Instance-Dependent Label Noise

ICLR 2022poster

Instance-dependent label noise (IDN) widely exists in real-world datasets and usually misleads the training of deep neural networks. Noise transition matrix (NTM) (i.e., the probability that clean labels flip into noisy labels) is used to characterize the label noise and can be adopted to bridge the…

Cited by 45SourcePDFScholar
2022

Attentionpit: Soft Permutation Invariant Training for Audio Source Separation with Attention Mechanism

ICASSP 2022accepted

Permutation invariant training (PIT) has recently attracted attention as a framework to achieve end-to-end time-domain audio source separation. Its goal is to train a separation network that takes a mixture signal as input and produces the J underlying source signals. Since the order of the output s…

Cited by 0SourceScholar
2022

CODE-MVP: Learning to Represent Source Code from Multiple Views with Contrastive Pre-Training

NAACL 2022findings

Recent years have witnessed increasing interest in code representation learning, which aims to represent the semantics of source code into distributed vectors. Currently, various works have been proposed to represent the complex semantics of source code from different views, including plain text, Ab…

2022

DeepMLE: A Robust Deep Maximum Likelihood Estimator for Two-view Structure from Motion

IROS 2022poster

Two-view structure from motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM (vSLAM). Many existing end-to-end learning-based methods usually formulate it as a brute regression problem. However, the inadequate utilization of traditional geometry model makes the model not robust in un…

Cited by 10SourceScholar
2022

EXACT: Scalable Graph Neural Networks Training via Extreme Activation Compression

ICLR 2022poster

Training Graph Neural Networks (GNNs) on large graphs is a fundamental challenge due to the high memory usage, which is mainly occupied by activations (e.g., node embeddings). Previous works usually focus on reducing the number of nodes retained in memory. In parallel, unlike what has been developed…

Cited by 67SourcePDFScholar
2022

Enhancing Feedback Steering Controllers for Autonomous Vehicles With Deep Monte Carlo Tree Search

RA-L 2022

Steering control is a vital function for autonomous vehicles, whose performance largely determines the driving safety. The widely used feedback steering controllers degrade significantly when vehicles drive at high speeds. This is mainly because these controllers can not effectively exploit nonlinea

Cited by 8SourceScholar
2022

FedDC: Federated Learning With Non-IID Data via Local Drift Decoupling and Correction

CVPR 2022poster

Federated learning (FL) allows multiple clients to collectively train a high-performance global model without sharing their private data. However, the key challenge in federated learning is that the clients have significant statistical heterogeneity among their local data distributions, which would…

Cited by 346PDFcodeScholar
2022

HBP: An Efficient Block Permutation Solver Using Hungarian Algorithm and Spectrogram Inpainting for Multichannel Audio Source Separation

ICASSP 2022accepted

This paper proposes a method called "Hungarian Block Permutation (HBP)" to solve the block permutation problem in frequency-domain multichannel audio source separation. Many methods for frequency-domain multichannel audio source separation are designed to simultaneously solve frequency-wise source s…

Cited by 0SourceScholar
2022

Investigation And Comparison of Optimization Methods for Variational Autoencoder-Based Underdetermined Multichannel Source Separation

ICASSP 2022accepted

In this paper, we investigate two algorithms for variational autoencoder (VAE)-based underdetermined multichannel source separation. We previously extended the multichannel VAE (MVAE) method for determined multichannel source separation and proposed the generalized MVAE (GMVAE) method for underdeter…

Cited by 0SourceScholar
2022

Modeling Hierarchical Syntax Structure with Triplet Position for Source Code Summarization

ACL 2022long

Automatic code summarization, which aims to describe the source code in natural language, has become an essential task in software maintenance. Our fellow researchers have attempted to achieve such a purpose through various machine learning-based approaches. One key challenge keeping these approache…

2022

Optimization of Compressive Light Field Display in Dual-Guided Learning

ICASSP 2022accepted

Glass-free compressive light field (CLF) display gains much attention due to their compatibility in holographic-like and three-dimensional (3D) demonstration. Opposite to other analogous devices, CLF display can provide binocular and motion parallaxes by stacking multiple liquid crystal screens with…

Cited by 0SourceScholar
2022

Optimization-Based Maneuver Planning for a Tractor-Trailer Vehicle in a Curvy Tunnel: A Weak Reliance on Sampling and Search

RA-L 2022

This study is focused on the maneuver planning problem for a tractor-trailer vehicle in a curvy and tiny tunnel. Due to the curse of dimensionality, the prevalent sampling-and- search-based planners used to handle a rigid-body vehicle well become less efficient when the trailer number grows or when

Cited by 32SourceScholar
2022

Table2Graph: Transforming Tabular Data to Unified Weighted Graph

IJCAI 2022poster

Learning useful interactions between input features is crucial for tabular data modeling. Recent efforts start to explicitly model the feature interactions with graph, where each feature is treated as an individual node. However, the existing graph construction methods either heuristically formula…

Cited by 26SourcePDFScholar
2021

Dirichlet Energy Constrained Learning for Deep Graph Neural Networks

NeurIPS 2021poster

Graph neural networks (GNNs) integrate deep architectures and topological structure modeling in an effective way. However, the performance of existing GNNs would decrease significantly when they stack many layers, because of the over-smoothing issue. Node embeddings tend to converge to similar vecto…

Cited by 142SourcePDFScholar
2021

Fast and Accurate Neural Machine Translation with Translation Memory

ACL 2021long

It is generally believed that a translation memory (TM) should be beneficial for machine translation tasks. Unfortunately, existing wisdom demonstrates the superiority of TM-based neural machine translation (NMT) only on the TM-specialized translation tasks rather than general tasks, with a non-negl…

Cited by 63SourcePDFScholar
2021

Joint Optimization for Full-Duplex Cellular Communications Via Intelligent Reflecting Surface

ICASSP 2021accepted

The implementation of full-duplex (FD) theoretically doubles the spectral efficiency of cellular communications. We propose a multiuser FD cellular network relying on an intelligent reflecting surface (IRS). The IRS is deployed to cover a dead zone while suppressing user-side self-interference (SI)…

Cited by 0SourceScholar
2021

Scaling Symbolic Methods using Gradients for Neural Model Explanation

ICLR 2021poster

Symbolic techniques based on Satisfiability Modulo Theory (SMT) solvers have been proposed for analyzing and verifying neural network properties, but their usage has been fairly limited owing to their poor scalability with larger networks. In this work, we propose a technique for combining gradient-…

2021

SepNet: A Deep Separation Matrix Prediction Network for Multichannel Audio Source Separation

ICASSP 2021accepted

In this paper, we propose SepNet, a deep neural network (DNN) designed to predict separation matrices from multichannel observations. One well-known approach to blind source separation (BSS) involves independent component analysis (ICA). A recently developed method called independent low-rank matrix…

Cited by 2SourceScholar
2021

Teacher-Student Learning for Low-Latency Online Speech Enhancement Using Wave-U-Net

ICASSP 2021accepted

In this paper, we propose a low-latency online extension of wave-U-net for single-channel speech enhancement, which utilizes teacher-student learning to reduce the system latency while keeping the enhancement performance high. Wave-U-net is a recently proposed end-to-end source separation method, wh…

Cited by 28SourceScholar
2021

Towards understanding retrosynthesis by energy-based models

NeurIPS 2021poster

Retrosynthesis is the process of identifying a set of reactants to synthesize a target molecule. It is of vital importance to material design and drug discovery. Existing machine learning approaches based on language models and graph neural networks have achieved encouraging results. However, the in…

Cited by 47SourcePDFScholar
2019

Coordinated multi-robot planning while preserving individual privacy

ICRA 2019poster

We consider the problem of multiple robots that must cooperate within a shared environment, but which wish to limit the information they disclose during their coordination efforts. Specifically, we examine the problems of privacy-preserving rendezvous and persistent monitoring. In the former, the ro…

Cited by 24SourceScholar
2019

Fast MVAE: Joint Separation and Classification of Mixed Sources Based on Multichannel Variational Autoencoder with Auxiliary Classifier

ICASSP 2019accepted

This paper proposes an alternative algorithm for the multi-channel variational autoencoder (MVAE), a recently proposed multichannel source separation approach. While MVAE is notable for its impressive source separation performance, its convergence-guaranteed optimization algorithm and the fact that…

Cited by 0SourceScholar
2019

Joint Separation and Dereverberation of Reverberant Mixtures with Multichannel Variational Autoencoder

ICASSP 2019accepted

In this paper, we deal with a multichannel source separation problem under a highly reverberant condition. The multichannel variational autoencoder (MVAE) is a recently proposed source separation method that employs the decoder distribution of a conditional VAE (CVAE) as the generative model for the…

Cited by 0SourceScholar