← Search

Jun Wu

62 accepted papers

2026

DualMirage: Hunting Stealthy Multimodal LLM Agents via CAPTCHAs with Contour and Adversarial Illusions

CVPR 2026

The rapid advancement of Multimodal Large Language Models (MLLMs) has given rise to sophisticated autonomous agents capable of performing complex, human-like tasks across the web. However, this also introduces significant security risks, particularly from stealthy MLLM agents that can evade conventi

Cited by 0SourceScholar
2026

Ekka: Automated Diagnosis of Silent Errors in LLM Inference

ICML 2026poster

LLM serving frameworks are quickly evolving with a complex software stack and a vast number of optimizations. The rapid development process can introduce silent errors where output quality silently degrades without any explicit error signals. Diagnosing silent errors is notoriously difficult due to …

Cited by 0SourceScholar
2026

HalluGuard: Demystifying Data-Driven and Reasoning-Driven Hallucinations in LLMs

ICLR 2026poster

The reliability of Large Language Models (LLMs) in high-stakes domains such as healthcare, law, and scientific discovery is often compromised by hallucinations. These failures typically stem from two sources: *data-driven hallucinations* and *reasoning-driven hallucinations*. However, existing detec…

Cited by 0SourcecodeScholar
2026

Look Forward to Walk Backward: Efficient Terrain Memory for Backward Locomotion with Forward Vision

ICRA 2026poster

Legged robots with egocentric forward-facing depth cameras can couple exteroception and proprioception to achieve robust forward agility on complex terrain. When these robots walk backward, the forward-only field of view provides no preview. Purely proprioceptive controllers can remain stable on mod…

2026

Model-Agnostic Sentiment Distribution Stability Analysis for Robust LLM-Generated Texts Detection

AAAI 2026technical

The rapid advancement of large language models (LLMs) has resulted in increasingly sophisticated AI-generated content, posing significant challenges in distinguishing LLM-generated text from human-written language. Existing detection methods, primarily based on lexical heuristics or fine-tuned class

Cited by 0SourcePDFScholar
2026

Multi-Modal Style Transfer-based Prompt Tuning for Efficient Federated Domain Generalization

AAAI 2026technical

Federated Domain Generalization (FDG) aims to collaboratively train a global model across distributed clients that can generalize well on unseen domains. However, existing FDG methods typically struggle with cross-client data heterogeneity and incur significant communication and computation overhead

Cited by 0SourcePDFScholar
2026

Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models

ICLR 2026poster

Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks, but their inference remains computationally inefficient. We observe a common failure mode in many prevalent LLMs, overthinking, where models generate verbose and tangential reasoning traces even for simple quer…

Cited by 0SourceScholar
2026

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization

ICML 2026poster

Despite rapid progress in auto-regressive video diffusion, we identify an emerging system–algorithm bottleneck that limits both deployability and generation quality: KV-cache memory. In auto-regressive video generation models, the KV-cache grows with generation history and quickly dominates GPU memo…

Cited by 0SourceScholar
2026

START: Traversing Sparse Footholds With Terrain Reconstruction

RA-L 2026

Traversing terrains with sparse footholds like legged animals presents a promising yet challenging task for quadruped robots, as it requires precise environmental perception and agile control to secure safe foot placement while maintaining dynamic stability. Model-based hierarchical controllers exce

Cited by 3SourceScholar
2026

START: Traversing Sparse Footholds with Terrain Reconstruction

ICRA 2026poster

Traversing terrains with sparse footholds like legged animals presents a promising yet challenging task for quadruped robots, as it requires precise environmental perception and agile control to secure safe foot placement while maintaining dynamic stability. Model-based hierarchical controllers exce…

2025

Addressing Cold-Start Problem in Click-Through Rate Prediction via Supervised Diffusion Modeling

AAAI 2025technical

Predicting Click-Through Rates is a crucial function within recommendation and advertising platforms, as the output of CTR prediction determines the order of items shown to users. The Embedding and MLP paradigm has become a standard approach for industrial recommendation systems and has been widely…

2025

DORec: Decomposed Object Reconstruction and Segmentation Utilizing 2D Self-Supervised Features

RA-L 2025

Recovering 3D geometry and textures of individual objects is crucial for many robotics applications, such as manipulation, pose estimation, and autonomous driving. However, decomposing a target object from a complex background is challenging. Most existing approaches rely on costly manual labels to

Cited by 1SourceScholar
2025

Invariant Link Selector for Spatial-Temporal Out-of-Distribution Problem

AISTATS 2025poster

In the era of foundation models, Out-of-Distribution (OOD) problems, i.e., the data discrepancy between the training environments and testing environments, hinder AI generalization. Further, relational data like graphs disobeying the Independent and Identically Distributed (IID) condition makes the…

Cited by 0SourcecodeScholar
2025

KGCRR: An Effective Metric-Driven Knowledge Graph Completion Framework by Designing a Novel Upper Bound Function with Adaptive Approximation to Reciprocal Rank

AAAI 2025technical

Knowledge Graph Embedding (KGE) methods have achieved great success in predicting missing links in knowledge graphs, a task also known as Knowledge Graph Completion (KGC). Under this task, the Reciprocal Rank (RR) of ground-truth items serve as a key indicator for evaluating the method’s performance…

2025

LensLLM: Unveiling Fine-Tuning Dynamics for LLM Selection

ICML 2025poster

The proliferation of open-sourced Large Language Models (LLMs) and diverse downstream tasks necessitates efficient model selection, given the impracticality of fine-tuning all candidates due to computational constraints. Despite the recent advances in LLM selection, a fundamental research question l…

2025

MOVE: Multi-Skill Omnidirectional Legged Locomotion With Limited View in 3D Environments

ICRA 2025

Legged robots possess inherent advantages in traversing complex 3D terrains. However, previous work on lowcost quadruped robots with egocentric vision systems has been limited by a narrow front-facing view and exteroceptive noise, restricting omnidirectional mobility in such environments. While buil

Cited by 7SourceScholar
2025

SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment

ACL 2025finding

Direct Preference Optimization (DPO) is broadly utilized for aligning Large Language Models (LLMs) with human values because of its flexibility. Despite its effectiveness, it has been observed that the capability of DPO to generate human-preferred response is limited and the results of DPO are far f…

2024

A Lightweight U-like Network Utilizing Neural Memory Ordinary Differential Equations for Slimming the Decoder

IJCAI 2024poster

In recent years, advanced U-like networks have demonstrated remarkable performance in medical image segmentation tasks. However, their drawbacks, including excessive parameters, high computational complexity, and slow inference speed, pose challenges for practical implementation in scenarios with li…

2024

Decentralized Communication-Maintained Coordination for Multi-Robot Exploration: Achieving Connectivity and Adaptability

IROS 2024poster

The realm of multi-robot autonomous exploration tasks underscores the critical role of communication in coordinating group activities. This paper introduces an innovative decentralized multi-robot exploration algorithm, meticulously crafted to ensure unbroken communication within robotic groups, a c…

Cited by 1SourceScholar
2024

Efficient Global Trajectory Planning for Multi-robot System with Affinely Deformable Formation

IROS 2024

Global trajectory planning is crucial for long-range formation navigation tasks of multi-robot systems in efficiency improvement and energy saving, whose main challenges are the joint space constraints of the whole team and the long-range deployment. To overcome the above difficulties, we reformulat

Cited by 1SourceScholar
2024

Optimal Ber Minimum Precoder Design for OTFS-Based ISAC Systems

ICASSP 2024accepted

This paper investigates the bit error rate (BER) minimum precoder design for an orthogonal time frequency space (OTFS)-based integrated sensing and communications (ISAC) system, which is considered as a promising technique for enabling future wireless networks. In particular, the BER minimum problem…

Cited by 0SourceScholar
2024

PIE: Parkour With Implicit-Explicit Learning Framework for Legged Robots

RA-L 2024

Parkour presents a highly challenging task for legged robots, requiring them to traverse various terrains with agile and smooth locomotion. This necessitates comprehensive understanding of both the robot's own state and the surrounding terrain, despite the inherent unreliability of robot perception

Cited by 49SourceScholar
2024

PJSCC: A Puncturing-Based Joint Source Channel Coding Scheme with Hierarchical Down-Sampling Layer

ICASSP 2024accepted

In this paper, we propose a puncturing-based joint source channel coding scheme with a hierarchical down-sampling layer (PJSCC). The proposed hierarchical down-sampling layer fully exploits both frequency and spatial priors. Moreover, to achieve adaptive compression ratio control, PJSCC utilizes a s…

Cited by 0SourceScholar
2024

PSC: Extending Context Window of Large Language Models via Phase Shift Calibration

EMNLP 2024main

Rotary Position Embedding (RoPE) is an efficient position encoding approach and is widely utilized in numerous large language models (LLMs). Recently, a lot of methods have been put forward to further expand the context window based on RoPE. The core concept of those methods is to predefine or searc…

2024

Surface-Constrained Progressive Feature Preserving Point Cloud Compression

ICASSP 2024accepted

Current point cloud compression methods based on deep learning cannot guarantee that the reconstructed points are constrained to the surface, resulting in low reconstruction quality at low bitrates. Hence, this paper proposes an efficient deep learning-based point cloud geometry compression algorith…

Cited by 0SourceScholar
2024

Toward Understanding Key Estimation in Learning Robust Humanoid Locomotion

IROS 2024poster

Accurate state estimation plays a critical role in ensuring the robust control of humanoid robots, particularly in the context of learning-based control policies for legged robots. However, there is a notable gap in analytical research concerning estimations. Therefore, we endeavor to further unders…

Cited by 6SourceScholar
2023

Exploiting Point-Wise Attention in 6D Object Pose Estimation Based on Bidirectional Prediction

RA-L 2023

Traditional geometric registration based estimation methods only exploit the CAD model implicitly, which leads to their dependence on observation quality and deficiency to occlusion.To address the problem,the letter proposes a bidirectional correspondence prediction network with a point-wise attenti

Cited by 1SourceScholar
2023

Graph-Structured Gaussian Processes for Transferable Graph Learning

NeurIPS 2023poster

Transferable graph learning involves knowledge transferability from a source graph to a relevant target graph. The major challenge of transferable graph learning is the distribution shift between source and target graphs induced by individual node attributes and complex graph structures. To solve th…

2023

Optimizing the Collaboration Structure in Cross-Silo Federated Learning

ICML 2023poster

In federated learning (FL), multiple clients collaborate to train machine learning models together while keeping their data decentralized. Through utilizing more training data, FL suffers from the potential negative transfer problem: the global FL model may even perform worse than the models trained…

2023

PADCLIP: Pseudo-labeling with Adaptive Debiasing in CLIP for Unsupervised Domain Adaptation

ICCV 2023poster

Traditional Unsupervised Domain Adaptation (UDA) leverages the labeled source domain to tackle the learning tasks on the unlabeled target domain. It can be more challenging when a large domain gap exists between the source and the target domain. A more practical setting is to utilize a large-scale p…

Cited by 71PDFScholar
2022

A Right Invariant Extended Kalman Filter for Object Based SLAM

RA-L 2022

With the recent advance of deep learning based object recognition and estimation, it is possible to consider object level SLAM where the pose of each object is estimated in the SLAM process. In this letter, based on a novel Lie group structure, a right invariant extended Kalman filter (RI-EKF) for o

Cited by 37SourceScholar
2022

Discrete Listwise Personalized Ranking for Fast Top-N Recommendation with Implicit Feedback

IJCAI 2022poster

We address the efficiency problem of personalized ranking from implicit feedback by hashing users and items with binary codes, so that top-N recommendation can be fast executed in a Hamming space by bit operations. However, current hashing methods for top-N recommendation fail to align their learnin…

2022

Distribution-Informed Neural Networks for Domain Adaptation Regression

NeurIPS 2022accept

In this paper, we study the problem of domain adaptation regression, which learns a regressor for a target domain by leveraging the knowledge from a relevant source domain. We start by proposing a distribution-informed neural network, which aims to build distribution-aware relationship of inputs and…

Cited by 18SourcePDFScholar
2022

HEA-D: A Hybrid Evolutionary Algorithm for Diversified Top-k Weight Clique Search Problem

IJCAI 2022poster

The diversified top-k weight clique (DTKWC) search problem is an important generalization of the diversified top-k clique (DTKC) search problem with extensive applications, which extends the DTKC search problem by taking into account the weight of vertices. In this paper, we formulate DTKWC search p…

2022

Towards Two-view 6D Object Pose Estimation: A Comparative Study on Fusion Strategy

IROS 2022poster

Current RGB-based 6D object pose estimation methods have achieved noticeable performance on datasets and real world applications. However, predicting 6D pose from single 2D image features is susceptible to disturbance from changing of environment and textureless or resemblant object surfaces. Hence,…

Cited by 3SourceScholar
2022

Vision-Assisted Localization and Terrain Reconstruction with Quadruped Robots

IROS 2022poster

Legged robots, specifically quadruped robots, have good locomotion performance in complex and rugged terrain and are becoming widely used in field exploration and rescue missions. To achieve full autonomy in such scenarios, robots need not only accurate localization but also an accurate understandin…

Cited by 8SourceScholar
2021

DeHiB: Deep Hidden Backdoor Attack on Semi-supervised Learning via Adversarial Perturbation

AAAI 2021technical

The threat of data-poisoning backdoor attacks on learning algorithms typically comes from the labeled data. However, in deep semi-supervised learning (SSL), unknown threats mainly stem from the unlabeled data. In this paper, we propose a novel deep hidden backdoor (DeHiB) attack scheme for SSL-based…

Cited by 51SourcePDFScholar
2021

Fitting the Search Space of Weight-sharing NAS with Graph Convolutional Networks

AAAI 2021technical

Neural architecture search has attracted wide attentions in both academia and industry. To accelerate it, researchers proposed weight-sharing methods which first train a super-network to reuse computation among different operators, from which exponentially many sub-networks can be sampled and effici…

Cited by 20SourcePDFScholar
2021

REDE: End-to-End Object 6D Pose Robust Estimation Using Differentiable Outliers Elimination

RA-L 2021

Object 6D pose estimation is a fundamental task in many applications. Conventional methods solve the task by detecting and matching the keypoints, then estimating the pose. Recent efforts bringing deep learning into the problem mainly overcome the vulnerability of conventional methods to environment

Cited by 41SourcecodeScholar
2020

Learning-based Optimization Algorithms Combining Force Control Strategies for Peg-in-Hole Assembly

IROS 2020poster

In this paper, an approach for automatic peg-in-hole assembly is proposed. The task is divided into two main steps: searching phase and inserting phase. First, a multilayer perceptron network is designed to address the hole search problem and a hybrid force position controller is introduced to ensur…

Cited by 34SourceScholar
2020

Structure Learning for Cyclic Linear Causal Models

UAI 2020poster

We consider the problem of structure learning for linear causal models based on observational data. We treat models given by possibly cyclic mixed graphs, which allow for feedback loops and effects of latent confounders. Generalizing related work on bow-free acyclic graphs, we assume that the unde…

Cited by 22SourcePDFScholar
2019

Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and Evaluation

ICCV 2019oral

Recently, differentiable search methods have made major progress in reducing the computational costs of neural architecture search. However, these approaches often report lower accuracy in evaluating the searched architecture or transferring it to another dataset. This is arguably due to the large g…

Cited by 847PDFcodeScholar
2018

Color-Based Sensing of Bending Deformation on Soft Robots

ICRA 2018poster

This paper introduces a novel approach for sensing the bending deformation on soft robots by leveraging multicolor 3D printing. The measurement of deformation enables to complete the feedback loop of deformation control on soft actuators. The working principle of our approach is based on using compa…

Cited by 15SourceScholar