← Search

Jun Yang

62 accepted papers

2026

ActiShade: Activating Overshadowed Knowledge to Guide Multi-Hop Reasoning in Large Language Models

AAAI 2026technical

In multi-hop reasoning, multi-round retrieval-augmented generation (RAG) methods typically rely on LLM-generated content as the retrieval query. However, these approaches are inherently vulnerable to knowledge overshadowing—a phenomenon where critical information is overshadowed during generation.

Cited by 0SourcePDFScholar
2026

An Efficient Learning-Based Task Planning Approach Using a Bio-Inspired Action Context-Free Grammar for Bimanual Manipulation

ICRA 2026poster

Task and Motion Planning (TAMP) frameworks for bimanual robots are limited by the combinatorial explosion at the task planning level, which can negatively affect human-robot interaction. This work introduces BAG-Learn Planning, an efficient learning-based task planning approach that combines a Bio-I…

Cited by 0Scholar
2026

BroRL: Scaling Reinforcement Learning via Broadened Exploration

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key ingredient for unlocking complex reasoning capabilities in large language models. Recent work ProRL \citep{liu2025prorl} has shown promise in scaling RL by increasing the number of training steps. However, performance plateau…

Cited by 0SourceScholar
2026

CBF-Based Hierarchical Quadratic Programs with Guaranteed Feasibility for Safety-Critical Systems (I)

ICRA 2026poster

Control Barrier Function (CBF) based quadratic programs (QPs) have become an effective method for enforcing safety in safety-critical systems and robotics. However, these methods often suffer from infeasibility or overly conservative relaxations when handling multiple constraints, potentially compro…

Cited by 0Scholar
2026

ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications

AAAI 2026technical

While Large Language Models (LLMs) demonstrate immense potential for automating integrated circuit (IC) development, their practical deployment is fundamentally limited by restricted context windows. Existing context-extension methods struggle to achieve effective semantic modeling and thorough mult

Cited by 0SourcePDFScholar
2026

EXVERUS: Verus Proof Repair via Counterexample Reasoning

ICML 2026poster

Large Language Models (LLMs) have shown promising results in automating formal verification. However, existing approaches treat proof generation as a static, end-to-end prediction over source code, relying on limited verifier feedback and lacking access to concrete program behaviors. We present EXVE…

Cited by 0SourceScholar
2026

FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification

AAAI 2026technical

We introduce FIXME, the first end-to-end and large-scale benchmark for evaluating Large Language Models (LLMs) in hardware design functional verification (FV). Comprising 747 tasks derived from real-world hardware designs, FIXME spans five core FV sub-sets: specification comprehension, reference mod

Cited by 0SourcePDFScholar
2026

Guarding Force: Safety-Critical Compliant Control for Robot-Environment Interaction

RA-L 2026

In this letter, we propose a safety-critical compliant control strategy designed to strictly enforce interaction force constraints during the physical interaction of robots with environments. The interaction force constraint is interpreted as a new force-constrained control barrier function (FC-CBF)

Cited by 2SourceScholar
2026

PRED-MPPI: Disturbance-Preview and Efficient MPPI for Robust Quadrotor Tracking with Hardware Validation

ICRA 2026poster

We propose PRED-MPPI, the first MPPI variant that seamlessly integrates real-time disturbance preview and adaptive discretization for quadrotor tracking control under significant model inaccuracies and time-varying disturbances. Unlike prior MPPI variants (e.g., mathcal{L}_1-MPPI, DA-MPPI), which as…

Cited by 0codeScholar
2026

Perception-Control Coupled Visual Servoing for Textureless Objects Using Keypoint-Based EKF

ICRA 2026poster

Visual servoing is fundamental to robotic applications, enabling precise positioning and control. However, applying it to textureless objects remains a challenge due to the absence of reliable visual features. Moreover, adverse visual conditions, such as occlusions, often corrupt visual feedback, le…

2026

Robust High-Precision Trajectory Planning for Payload Transportation in Overhead Cranes: A Disturbance-Aware Approach

RA-L 2026

The coordinated motion of the trolley and hoisting rope improves crane flexibility but poses challenges in precise trajectory conversion and tracking due to disturbances and inaccessible low-level controllers. This letter proposes a disturbance-aware high-precision trajectory planning method integra

Cited by 1SourceScholar
2026

Safety-Critical Steering Control for Rubber-Tired Container Gantry Cranes: A State-Interlocked CBF Approach

RA-L 2026

The rubber-tired container gantry crane (RTG) is a type of heavy-duty lifting equipment commonly used in container yards, which is driven by two-side rubber tires and steered via differential drive. While moving along the desired path, the RTG must remain centered of the lane with restricted heading

Cited by 0SourceScholar
2025

A Dynamic Knowledge Update-Driven Model with Large Language Models for Fake News Detection

IJCAI 2025

As the Internet and social media evolve rapidly, distinguishing credible news from a vast amount of complex information poses a significant challenge. Due to the suddenness and instability of news events, the authenticity labels of news can potentially shift as events develop, making it crucial for

Cited by 0SourcePDFScholar
2025

A New Variable-Gain Sliding Mode Filter and Its Application to Velocity Filtering

ICRA 2025

This paper proposes a new variable gain sliding mode filter augmented by variable windowing for achieving smooth and reactive response over a broad range of input frequencies. The proposed filter can be seen as a synergistic combination of Kikuuwe et al.'s [1] sliding mode filter with varying gain a

Cited by 0SourceScholar
2025

CompKBQA: Component-wise Task Decomposition for Knowledge Base Question Answering

EMNLP 2025

Knowledge Base Question Answering (KBQA) aims to extract accurate answers from the Knowledge Base (KB). Traditional Semantic Parsing (SP)-based methods are widely used but struggle with complex queries. Recently, large language models (LLMs) have shown promise in improving KBQA performance. However,

2025

DA-MPPI: Disturbance-Aware Model Predictive Path Integral via active disturbance estimation and compensation

IROS 2025

Model Predictive Path Integral (MPPI) controllers are drawing increasing attention for their ability to efficiently handle complex systems by leveraging GPU acceleration while with flexible prediction models and cost functions. However, their performance generally degrades with low-quality predictio

Cited by 2SourceScholar
2025

DeepPreNet: A Deep Learning Pre-Processing Method for Speech Distortion Correction in Parametric Array Loudspeaker

ICASSP 2025accepted

The parametric array loudspeaker produces highly directional sound via a nonlinear process in air, which also introduces inherent baseband distortions. However, conventional recursive modulators employed to compensate for nonlinearity demand substantially increased bandwidth and are not optimized fo…

Cited by 2SourceScholar
2025

Episodic Novelty Through Temporal Distance

ICLR 2025poster

Exploration in sparse reward environments remains a significant challenge in reinforcement learning, particularly in Contextual Markov Decision Processes (CMDPs), where environments differ across episodes. Existing episodic intrinsic motivation methods for CMDPs primarily rely on count-based approac…

Cited by 0SourcePDFScholar
2025

External Memory Matters: Generalizable Object-Action Memory for Retrieval-Augmented Long-Term Video Understanding

IJCAI 2025

Long video understanding with Large Language Models (LLMs) enables the description of objects that are not explicitly present in the training data. However, continuous changes in known objects and the emergence of new ones require up-to-date knowledge of objects and their dynamics for effective unde

Cited by 0SourcePDFScholar
2025

Flexible Active Safety Motion Control for Robotic Obstacle Avoidance: A CBF-Guided MPC Approach

RA-L 2025

A flexible active safety motion (FASM) control approach is proposed for collision avoidance in robot manipulators. The key feature is the use of control barrier functions (CBFs) to design flexible CBF-guided safety criteria (CBFSC) with dynamically optimized decay rates, providing both flexibility a

Cited by 21SourceScholar
2025

MemFreezing: A Novel Adversarial Attack on Temporal Graph Neural Networks under Limited Future Knowledge

ICML 2025poster

Temporal graph neural networks (TGNN) have achieved significant momentum in many real-world dynamic graph tasks. While most existing TGNN attack methods assume worst-case scenarios where attackers have complete knowledge of the input graph, the assumption may not always hold in real-world situations…

Cited by 0SourcePDFScholar
2025

Multi-Layered Safety of Redundant Robot Manipulators Via Task-Oriented Planning and Control

ICRA 2025

Ensuring safety is crucial to promote the application of robot manipulators in open workspaces. Factors such as sensor errors or unpredictable collisions make the environment full of uncertainties. In this work, we investigate these potential safety challenges on redundant robot manipulators, and pr

Cited by 1SourcecodeScholar
2025

Synthesizing Performance Constraints for Evaluating and Improving Code Efficiency

NeurIPS 2025poster

Large Language Models (LLMs) have been increasingly used to optimize code efficiency. Evaluating their effectiveness and further suggesting optimization opportunities often rely on high-quality tests to demonstrate the performance bottlenecks presented in the program. However, existing approaches re…

Cited by 0SourcecodeScholar
2025

ToMPC: Task-Oriented Model Predictive Control via ADMM for Safe Robotic Manipulation

RA-L 2025

This paper proposes a task-oriented model predictive control (ToMPC) framework for safe and efficient robotic manipulation in open workspaces. The framework unifies collision-free motion and robot-environment interaction to address diverse scenarios. Additionally, it introduces task-oriented obstacl

Cited by 2SourceScholar
2025

Unsupervised Image-to-Image Style Transfer via Dual-Condition Diffusion Models*

ICASSP 2025accepted

Style transfer is an artistic research topic within a series of generative tasks, and it is quite challenging to stably and reliably guide the generation of images with the expected content and style. Early methods followed a predefined style, defined by a reference image, and applied it to another…

Cited by 0SourceScholar
2024

Active Pose Refinement for Textureless Shiny Objects using the Structured Light Camera

IROS 2024poster

6D pose estimation of textureless shiny objects has become an essential problem in many robotic applications. Many pose estimators require high-quality depth data, often measured by structured light cameras. However, when objects have shiny surfaces (e.g., metal parts), these cameras fail to sense c…

Cited by 3SourceScholar
2024

Augmenting Reasoning Capabilities of LLMs with Graph Structures in Knowledge Base Question Answering

EMNLP 2024finding

Recently, significant progress has been made in employing Large Language Models (LLMs) for semantic parsing to address Knowledge Base Question Answering (KBQA) tasks. Previous work utilize LLMs to generate query statements on Knowledge Bases (KBs) for retrieving answers. However, LLMs often generate…

2024

Efficient Multi-agent Reinforcement Learning by Planning

ICLR 2024poster

Multi-agent reinforcement learning (MARL) algorithms have accomplished remarkable breakthroughs in solving large-scale decision-making tasks. Nonetheless, most existing MARL algorithms are model-free, limiting sample efficiency and hindering their applicability in more challenging scenarios. In cont…

2024

Enhanced Robust Motion Control based on Unknown System Dynamics Estimator for Robot Manipulators

ICRA 2024poster

To achieve high-accuracy manipulation in the presence of unknown disturbances, we propose two novel efficient and robust motion control schemes for high-dimensional robot manipulators. Both controllers incorporate an unknown system dynamics estimator (USDE) to estimate disturbances without requiring…

Cited by 1SourceScholar
2024

Learning Diverse Risk Preferences in Population-Based Self-Play

AAAI 2024technical

Among the remarkable successes of Reinforcement Learning (RL), self-play algorithms have played a crucial role in solving competitive games. However, current self-play RL methods commonly optimize the agent to maximize the expected win-rates against its current or historical copies, resulting in a l…

2024

NeuralPlane: An Efficiently Parallelizable Platform for Fixed-wing Aircraft Control with Reinforcement Learning

NeurIPS 2024poster

Reinforcement learning (RL) demonstrates superior potential over traditional flight control methods for fixed-wing aircraft, particularly under extreme operational conditions. However, the high demand for training samples and the lack of efficient computation in existing simulators hinder its furthe…

2024

S2WAT: Image Style Transfer via Hierarchical Vision Transformer Using Strips Window Attention

AAAI 2024technical

Transformer's recent integration into style transfer leverages its proficiency in establishing long-range dependencies, albeit at the expense of attenuated local modeling. This paper introduces Strips Window Attention Transformer (S2WAT), a novel hierarchical vision transformer designed for style tr…

2024

Single-Trajectory Distributionally Robust Reinforcement Learning

ICML 2024poster

To mitigate the limitation that the classical reinforcement learning (RL) framework heavily relies on identical training and test environments, Distributionally Robust RL (DRRL) has been proposed to enhance performance across a range of environments, possibly including unknown test environments. As…

Cited by 13SourcePDFScholar
2024

Tailoring Vaccine Messaging with Common-Ground Opinions

NAACL 2024findings

One way to personalize chatbot interactions is by establishing common ground with the intended reader. A domain where establishing mutual understanding could be particularly impactful is vaccine concerns and misinformation. Vaccine interventions are forms of messaging which aim to answer concerns ex…

2024

Uncertainty-aware 3D Object-Level Mapping with Deep Shape Priors

ICRA 2024poster

3D object-level mapping is a fundamental problem in robotics, which is especially challenging when object CAD models are unavailable during inference. We propose a framework that can reconstruct high-quality object-level maps for unknown objects. Our approach takes multiple RGB-D images as input and…

Cited by 8SourcecodeScholar
2023

6D Pose Estimation for Textureless Objects on RGB Frames using Multi-View Optimization

ICRA 2023poster

6D pose estimation of textureless objects is a valuable but challenging task for many robotic applications. In this work, we propose a framework to address this challenge using only RGB images acquired from multiple viewpoints. The core idea of our approach is to decouple 6D pose estimation into a s…

Cited by 16SourceScholar
2023

Conservative Offline Policy Adaptation in Multi-Agent Games

NeurIPS 2023poster

Prior research on policy adaptation in multi-agent games has often relied on online interaction with the target agent in training, which can be expensive and impractical in real-world scenarios. Inspired by recent progress in offline reinforcement learn- ing, this paper studies offline policy adapta…

Cited by 2SourcePDFScholar
2023

Enhancing Privacy Preservation in Federated Learning via Learning Rate Perturbation

ICCV 2023poster

Federated learning (FL) is a privacy-enhanced distributed machine learning framework, in which multiple clients collaboratively train a global model by exchanging their model updates without sharing local private data. However, the adversary can use gradient inversion attacks to reveal the clients'…

Cited by 2PDFScholar
2023

Flow to Control: Offline Reinforcement Learning with Lossless Primitive Discovery

AAAI 2023technical

Offline reinforcement learning (RL) enables the agent to effectively learn from logged data, which significantly extends the applicability of RL algorithms in real-world scenarios where exploration can be expensive or unsafe. Previous works have shown that extracting primitive skills from the recurr…

Cited by 18SourcePDFScholar
2023

Understanding and Defending Patched-based Adversarial Attacks for Vision Transformer

ICML 2023poster

Vision Transformer (ViT) is an attention-based model architecture that has demonstrated superior performance on many computer vision tasks. However, its security properties, in particular, the robustness against adversarial attacks, are yet to be thoroughly studied. Recent works have shown that ViT…

Cited by 5SourcePDFScholar
2022

A Multi-Task Learning Method for Weakly Supervised Sound Event Detection

ICASSP 2022accepted

In weakly supervised sound event detection (SED), only coarse-grained labels are available, and thus the supervision information is quite limited. To fully utilize prior knowledge of the time-frequency masks of each sound event, we propose a novel multi-task learning (MTL) method that takes SED as t…

Cited by 0SourceScholar
2022

A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection

ICASSP 2022accepted

Polyphonic sound event localization and detection (SELD) aims at detecting types of sound events with corresponding temporal activities and spatial locations. In this paper, a trackwise ensemble event independent network with a novel data augmentation method is proposed. The proposed model is based…

Cited by 0SourceScholar
2022

APD: Learning Diverse Behaviors for Reinforcement Learning Through Unsupervised Active Pre-Training

RA-L 2022

Unsupervised pre-training in reinforcement learning enables the agent to gain prior environmental knowledge, which is then fine-tuned in the supervised stage to quickly adapt to various downstream tasks. In the absence of task-related rewards, pre-training aims to acquire policies (i.e., behaviors)

Cited by 5SourceScholar
2022

Offline Reinforcement Learning with Value-based Episodic Memory

ICLR 2022poster

Offline reinforcement learning (RL) shows promise of applying RL to real-world problems by effectively utilizing previously collected data. Most existing offline RL algorithms use regularization or constraints to suppress extrapolation error for actions outside the dataset. In this paper, we adopt a…

Cited by 50SourcePDFScholar
2022

Safe Opponent-Exploitation Subgame Refinement

NeurIPS 2022accept

In zero-sum games, an NE strategy tends to be overly conservative confronted with opponents of limited rationality, because it does not actively exploit their weaknesses. From another perspective, best responding to an estimated opponent model is vulnerable to estimation errors and lacks safety guar…

Cited by 9SourcePDFScholar
2022

Where to Attack: A Dynamic Locator Model for Backdoor Attack in Text Classifications

COLING 2022main

Nowadays, deep-learning based NLP models are usually trained with large-scale third-party data which can be easily injected with malicious backdoors. Thus, BackDoor Attack (BDA) study has become a trending research to help promote the robustness of an NLP system. Text-based BDA aims to train a poiso…

2021

Average-Reward Reinforcement Learning with Trust Region Methods

IJCAI 2021poster

Most of reinforcement learning algorithms optimize the discounted criterion which is beneficial to accelerate the convergence and reduce the variance of estimates. Although the discounted criterion is appropriate for certain tasks such as financial related problems, many engineering problems treat f…

Cited by 22SourcePDFScholar
2021

Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning

NeurIPS 2021spotlight

Learning from datasets without interaction with environments (Offline Learning) is an essential step to apply Reinforcement Learning (RL) algorithms in real-world scenarios. However, compared with the single-agent counterpart, offline multi-agent RL introduces more agents with the larger state and a…

2021

Celebrating Diversity in Shared Multi-Agent Reinforcement Learning

NeurIPS 2021poster

Recently, deep multi-agent reinforcement learning (MARL) has shown the promise to solve complex cooperative tasks. Its success is partly because of parameter sharing among agents. However, such sharing may lead agents to behave similarly and limit their coordination capacity. In this paper, we aim t…

Cited by 189SourcePDFScholar
2021

HIGCNN: Hierarchical Interleaved Group Convolutional Neural Networks for Point Clouds Analysis

ICASSP 2021accepted

Although previous works for point clouds analysis have achieved remarkable performance, it is difficult for them to achieve a good trade-off between accuracy and complexity. In this paper, we present an efficient and lightweight neural network for point clouds analysis, named HIGCNN, which can achie…

Cited by 0SourceScholar
2021

Learning to Discover Task-Relevant Features for Interpretable Reinforcement Learning

RA-L 2021

Reinforcement Learning (RL) agents are often fed with large-dimensional observations to achieve the ideal performance in complex environments. Unfortunately, the massive observation space usually contains useless or even adverse features, which leads to low sample efficiency. Existing methods rely o

Cited by 5SourcecodeScholar
2021

ROBI: A Multi-View Dataset for Reflective Objects in Robotic Bin-Picking

IROS 2021poster

In robotic bin-picking applications, the perception of texture-less, highly reflective parts is a valuable but challenging task. The high glossiness can introduce fake edges in RGB images and inaccurate depth measurements, especially in heavily cluttered bin scenarios. In this paper, we present the…

Cited by 55SourceScholar
2019

Fast-rate PAC-Bayes Generalization Bounds via Shifted Rademacher Processes

NeurIPS 2019poster

The developments of Rademacher complexity and PAC-Bayesian theory have been largely independent. One exception is the PAC-Bayes theorem of Kakade, Sridharan, and Tewari (2008), which is established via Rademacher complexity theory by viewing Gibbs classifiers as linear operators. The goal of this pa…

Cited by 38SourcePDFScholar
2017

Deep learning based automatic volume control and limiter system

ICASSP 2017accepted

Automatic speech recognition is now playing an important role in volume control and adjustment of modern smart speakers. According to the recognition results by using the advanced deep neural network technology, this paper proposes an efficient processing system for automatic volume control (AVC) an…

Cited by 0SourceScholar
2016

Estimating ear canal geometry and eardrum reflection coefficient from ear canal input impedance

ICASSP 2016accepted

Based on the signal model of ear canals, a novel method for solving the inverse problem of estimating the unique solution of the ear canal area function and the eardrum reflection coefficient given the acoustic input impedance at the entrance of an ear canal is presented. Up-sampling techniques to i…

Cited by 0SourceScholar