← Search

Shuai Li

149 accepted papers

2026

AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

ICML 2026poster

Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause of loss spikes by investigating individual factors, we observe that, in practice, such spikes are typically triggered by the confluence of heterogeneou…

Cited by 0SourceScholar
2026

Conditionally Whitened Generative Models for Probabilistic Time Series Forecasting

ICLR 2026poster

Probabilistic forecasting of multivariate time series is challenging due to non-stationarity, inter-variable dependencies, and distribution shifts. While recent diffusion and flow matching models have shown promise, they often ignore informative priors such as conditional means and covariances. In t…

Cited by 0SourcecodeScholar
2026

DECON: Reconstruction of Clothed-Geometric Multiple Humans from a Single Image via Geometry-Guided Decoupling

AAAI 2026technical

3D multi-human reconstruction from single images holds significant potential for advancing AR/VR applications. While remarkable progress has been made in single-human reconstruction, existing methods face challenges when reconstructing multiple humans. These challenges include: (1) severe inter-occl

Cited by 0SourcePDFScholar
2026

DiscoX: Benchmarking Discourse-Level Translation in Expert Domains

ICLR 2026poster

The evaluation of discourse-level translation in expert domains remains inadequate, despite its centrality to knowledge dissemination and cross-lingual scholarly communication. While these translations demand discourse-level coherence and strict terminological precision, current evaluation methods p…

Cited by 0SourcecodeScholar
2026

Evolutionary Multi-View Classification with Label Noise via Gradient and Feature Dual-Perception

ICML 2026spotlight

This paper studies a fundamental yet often overlooked premise in evolutionary multi-view classification (EMVC): the impact of label noise on EMVC, such as distorting fitness landscapes shaped by individual fitness values (e.g., test accuracy). Traditional EMVC assumes training labels are noise-free,…

Cited by 0SourceScholar
2026

FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models

AAAI 2026technical

The powerful generalization of Vision-Language-Action (VLA) models is bottlenecked by their heavy reliance on massive, redundant, and unevenly valued datasets, hindering their widespread application. Existing model-centric optimization paths, such as model compression (which often leads to performan

Cited by 0SourcePDFScholar
2026

GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution

CVPR 2026

Recently, reinforcement learning (RL) has been employed for improving generative image super-resolution (ISR) performance. However, the current efforts are focused on multi-step generative ISR, while one-step generative ISR remains underexplored due to its limited stochasticity. In addition, RL meth

Cited by 0SourcecodeScholar
2026

Hierarchical Frequency-Guided Alignment Transformer for Compressed Video Quality Enhancement

AAAI 2026technical

During the video encoding process, the original spatial domain signal is first transformed into the frequency domain, followed by quantization and compression. As a result, the quality degradation in compressed videos primarily stems from distortions in the frequency domain information. However, exi

Cited by 0SourcePDFScholar
2026

IntentMotion: Learning Intent-Aware Human Motion from Language in 3D Scenes

AAAI 2026technical

Generating human motion in complex 3D scenes from text is a challenging task with broad applications. However, existing methods often overlook realistic physical contact, resulting in visually plausible but physically unrealistic motion, e.g., penetration. To alleviate this, we propose IntentMotion,

Cited by 0SourcePDFScholar
2026

Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

CVPR 2026

Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain focus on short video clips within several seconds, with salient objects visible

Cited by 0SourcecodeScholar
2026

MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants

ICML 2026spotlight

With the rapid advancement of Large Language Models (LLMs) in code generation, human-AI interaction is evolving from static text responses to dynamic, interactive HTML-based applications, which we term **MiniApps**. These applications require models to not only render visual interfaces but also cons…

Cited by 0SourceScholar
2026

MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation

CVPR 2026

Diffusion-based motion generation has advanced rapidly, but current methods still struggle with long-horizon consistency, style control, and multi-condition guidance. A major reason is the fused-conditioning design, where semantic, stylistic, and temporal signals share a single pathway, causing inte

Cited by 0SourceScholar
2026

MotionHiFlow: Text-to-Motion via Hierarchical Flow Matching

CVPR 2026

Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural movements, they usually operate at only one temporal scale, which

Cited by 2SourcecodeScholar
2026

Multi-Subspace Multi-Modal Modeling for Diffusion Models: Estimation, Convergence and Mixture of Experts

ICLR 2026poster

Recently, diffusion models have achieved a great performance with a small dataset of size $n$ and a fast optimization process. Despite the impressive performance, the estimation error suffers from the curse of dimensionality $n^{-1/D}$, where $D$ is the data dimension. Since images are usually a un…

Cited by 0SourceScholar
2026

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization

IJCAI 2026

Offline-to-online reinforcement learning harnesses the stability of offline pretraining and the flexibility of online fine-tuning. A key challenge lies in the non-stationary distribution shift between offline datasets and the evolving online policy. Common approaches often rely on static mixing rati

Cited by 0Scholar
2026

TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

CVPR 2026

Understanding complex surgical scenes requires recognizing multiple interdependent entities--such as instruments, actions, and targets--and maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra-frame label dependencies and

Cited by 0SourceScholar
2026

Task-Adaptive Admittance Control for Human-Quadrotor Cooperative Load Transportation With Dynamic Cable-Length Regulation

RA-L 2026

The collaboration between humans and robots is critical in many robotic applications, especially in those requiring physical human-robot interaction (pHRI). Previous research in pHRI has largely focused on robotic manipulators, employing impedance or admittance control to maintain operational safety

Cited by 0SourceScholar
2026

The Accumulation of Score Estimation Error in Diffusion Models

ICML 2026poster

Diffusion models are widely used for high-quality generation, but their performance is sensitive to the accuracy of the estimated score. We first develop our main bounds in a Gaussian-mixture setting, where the score admits a closed-form structure and the score Hessian can be controlled explicitly, …

Cited by 0SourceScholar
2025

Adaptive Merchant-Centric Risk Control via Unbiased Decision-Making and Dynamic Optimization in E-Commerce

AAAI 2025technical

In the domain of merchant-oriented risk control decisions within e-commerce, balancing the effectiveness of risk management with merchant satisfaction remains a critical challenge. Strict risk control strategies, while effectively mitigating risks, often lead to increased merchant dissatisfaction. C…

Cited by 0SourcePDFScholar
2025

Adaptive Noise Rejection Strategy for Cooperative Motion Control of Dual-Arm Robots

RA-L 2025

Dual-arm robots possess exceptional collaborative capabilities and versatility, demonstrating broad application prospects across various fields. As a significant research area for dual-arm robots, the requirements for coordinated motion control are gradually increasing. In practical applications, ro

Cited by 3SourceScholar
2025

Bandit Learning in Matching Markets with Indifference

ICLR 2025poster

A rich line of recent works studies how participants in matching markets learn their unknown preferences through iterative interactions with each other. The two sides of participants in the market can be respectively formulated as players and arms in the bandit problem. To ensure market stability, t…

Cited by 0SourcePDFScholar
2025

Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs

COLING 2025main

Chain-of-Thought (CoT) has been a widely adopted prompting method, eliciting impressive reasoning abilities of Large Language Models (LLMs). Inspired by the sequential thought structure of CoT, a number of Chain-of-X (CoX) methods have been developed to address challenges across diverse domains and…

Cited by 22SourcePDFScholar
2025

Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control

ICLR 2025poster

Speech-driven 3D talking face method should offer both accurate lip synchronization and controllable expressions. Previous methods solely adopt discrete emotion labels to globally control expressions throughout sequences while limiting flexible fine-grained facial control within the spatiotemporal d…

Cited by 0SourcePDFScholar
2025

Conditional Diffusion Models Based Conditional Independence Testing

AAAI 2025technical

Conditional independence (CI) testing is a fundamental task in modern statistics and machine learning. The conditional randomization test (CRT) was recently introduced to test whether two random variables, X and Y, are conditionally independent given a potentially high-dimensional set of random vari…

2025

Convex Potential Mirror Langevin Algorithm for Efficient Sampling of Energy-Based Models

NeurIPS 2025poster

This paper introduces the Convex Potential Mirror Langevin Algorithm (CPMLA), a novel method to improve sampling efficiency for Energy-Based Models (EBMs). CPMLA uses mirror Langevin dynamics with a convex potential flow as a dynamic mirror map for EBM sampling. This dynamic mirror map enables targe…

Cited by 0SourceScholar
2025

CtrlAvatar: Controllable Avatars Generation via Disentangled Invertible Networks

AAAI 2025technical

As virtual experiences grow in popularity, the demand for realistic, personalized, and animatable human avatars increases. Traditional methods, relying on fixed templates, often produce costly avatars that lack expressiveness and realism. To overcome these challenges, we introduce Controllable Avata…

2025

DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing

NeurIPS 2025spotlight

Leveraging the powerful generation capability of large-scale pretrained text-to-image models, training-free methods have demonstrated impressive image editing results. Conventional diffusion-based methods, as well as recent rectified flow (RF)-based methods, typically reverse synthesis trajectories…

Cited by 0SourceScholar
2025

DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution

NeurIPS 2025poster

Benefiting from pre-trained text-to-image (T2I) diffusion models, real-world image super-resolution (Real-ISR) methods can synthesize rich and realistic details. However, due to the inherent stochasticity of T2I models, different noise inputs often lead to outputs with varying perceptual quality. Al…

Cited by 0SourceScholar
2025

Demonstrating Multi-Suction Item Picking at Scale via Multi-Modal Learning of Pick Success

RSS 2025poster

This work demonstrates how autonomously learning aspects of robotic operation from sparsely-labeled, real-world data of deployed, engineered solutions at industrial scale can provide with solutions that achieve improved performance. Specifically, it focuses on multi-suction robot picking and perfor…

Cited by 0PDFScholar
2025

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

ICCV 2025poster

The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets…

2025

Fast Second-Order Online Kernel Learning Through Incremental Matrix Sketching and Decomposition

IJCAI 2025

Second-order Online Kernel Learning (OKL) has attracted considerable research interest due to its promising predictive performance in streaming environments. However, existing second-order OKL approaches suffer from at least quadratic time complexity with respect to the pre-set budget, rendering the

Cited by 0SourcePDFScholar
2025

Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training

ICCV 2025poster

Impressive results on real-world image super-resolution (Real-ISR) have been achieved by employing pre-trained stable diffusion (SD) models. However, one critical issue of such methods lies in their poor reconstruction of image fine structures, such as small characters and textures, due to the aggre…

2025

Global-Aware Monocular Semantic Scene Completion with State Space Models

ICCV 2025poster

Monocular Semantic Scene Completion (MonoSSC) reconstructs and interprets 3D environments from a single image, enabling diverse real-world applications. However, existing methods are often constrained by the local receptive field of Convolutional Neural Networks (CNNs), making it challenging to hand…

Cited by 0SourcePDFScholar
2025

GraphOTTER: Evolving LLM-based Graph Reasoning for Complex Table Question Answering

COLING 2025main

Complex Table Question Answering involves providing accurate answers to specific questions based on intricate tables that exhibit complex layouts and flexible header locations. Despite considerable progress having been made in the LLM era, the reasoning processes of existing methods are often implic…

2025

Improved Discretization Complexity Analysis of Consistency Models: Variance Exploding Forward Process and Decay Discretization Scheme

ICML 2025poster

Consistency models, a new class of one-step generative models, have shown competitive performance with multi-step diffusion models. The most challenging part of consistency models is the training process, which discretizes the continuous diffusion process into $K$ steps and trains a one-step mapping…

Cited by 0SourcePDFScholar
2025

Large-Scale Mixed-Traffic and Intersection Control using Multi-agent Reinforcement Learning

IROS 2025

Traffic congestion remains a significant challenge in modern urban networks. Autonomous driving technologies have emerged as a potential solution. Among traffic control methods, reinforcement learning has shown superior performance over traffic signals in various scenarios. However, prior research h

Cited by 5SourcecodeScholar
2025

Learning Imperfect Information Extensive-form Games with Last-iterate Convergence under Bandit Feedback

ICML 2025poster

We investigate learning approximate Nash equilibrium (NE) policy profiles in two-player zero-sum imperfect information extensive-form games (IIEFGs) with last-iterate convergence guarantees. Existing algorithms either rely on full-information feedback or provide only asymptotic convergence rates. In…

Cited by 0SourcePDFScholar
2025

Learning Preferences without Interaction for Cooperative AI: A Hybrid Offline-Online Approach

NeurIPS 2025poster

Reinforcement learning (RL) for collaborative agents capable of cooperating with humans to accomplish tasks has long been a central goal in the RL community. While prior approaches have made progress in adapting collaborative agents to diverse human partners, they often focus solely on optimizing ta…

Cited by 0SourceScholar
2025

Learning-Based Slip Detection and Fine Control Using the Tactile Sensor for Robot Stable Grasping

RA-L 2025

Slip detection and control is critical to achieving stable grasping in robotics. However, accurate and robust slip detection and control remains a challenging task. This letter proposes a learning framework with contrastive learning and feature alignment to improve the accuracy of end-to-end slip de

Cited by 3SourceScholar
2025

Logarithmic Regret for Linear Markov Decision Processes with Adversarial Corruptions

AAAI 2025technical

In this work, we study the logarithmic regret for reinforcement learning (RL) with linear function approximation and adversarial corruptions, in the formulation of linear Markov decision processes (MDPs). Specifically, we consider the case where there exist adversarial corruptions over the reward fu…

Cited by 0SourcePDFScholar
2025

MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural Representation

CVPR 2025highlight

This paper presents MetricGrids, a novel grid-based neural representation that combines elementary metric grids in various metric spaces to approximate complex nonlinear signals. While grid-based representations are widely adopted for their efficiency and scalability, the existing feature grids with…

2025

Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion Generation

ICML 2025poster

Offline-to-online Reinforcement Learning (O2O RL) aims to perform online fine-tuning on an offline pre-trained policy to minimize costly online interactions. Existing work used offline datasets to generate data that conform to the online data distribution for data augmentation. However, generated da…

Cited by 0SourcePDFScholar
2025

Point Cloud Registration Based on Adaptively Fused Multimodal Features

RA-L 2025

Point cloud registration is a fundamental task in 3D vision, which plays an important role in various fields but faces challenges in geometrically weak or repetitive scenes. Traditional geometric-based methods struggle in these cases, while recent multimodal approaches improve robustness in weak sce

Cited by 4SourceScholar
2025

Rethinking the Temperature for Federated Heterogeneous Distillation

ICML 2025poster

Federated Distillation (FedKD) relies on lightweight knowledge carriers like logits for efficient client-server communication. Although logit-based methods have demonstrated promise in addressing statistical and architectural heterogeneity in federated learning (FL), current approaches remain const…

Cited by 0SourcePDFScholar
2025

Scaling Laws for Floating–Point Quantization Training

ICML 2025poster

Low-precision training is considered an effective strategy for reducing both training and downstream inference costs. Previous scaling laws for precision mainly focus on integer quantization, which pay less attention to the constituents in floating-point (FP) quantization, and thus cannot well fit t…

Cited by 1SourcePDFScholar
2025

SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction

CVPR 2025poster

Predicting future video frames is essential for decision-making systems, yet RGB frames alone often lack the information needed to fully capture the underlying complexities of the real world. To address this limitation, we propose a multi-modal framework for Synchronous Video Prediction (SyncVP) tha…

2025

The Polynomial Iteration Complexity for Variance Exploding Diffusion Models: Elucidating SDE and ODE Samplers

AISTATS 2025poster

Recently, variance exploding (VE) diffusion models have achieved state-of-the-art (SOTA) performance in two implementations: (1) the SDE-based implementation and (2) the probability flow ODE (PFODE) implementation. However, only a few works analyze the iteration complexity of VE-based models, and mo…

Cited by 0SourceScholar
2025

Towards Provably Efficient Learning of Imperfect Information Extensive-Form Games with Linear Function Approximation

UAI 2025

Despite significant advances in learning imperfect information extensive-form games (IIEFGs), most existing theoretical guarantees are limited to IIEFGs in the tabular case. To permit efficient learning of large-scale IIEFGs, we take the first step in studying two-player zero-sum IIEFGs with linear

2024

Aligning as Debiasing: Causality-Aware Alignment via Reinforcement Learning with Interventional Feedback

NAACL 2024long

Large language models (LLMs) often generate biased outputs containing offensive, toxic, or stereotypical text. Existing LLM alignment methods such as reinforcement learning from human feedback (RLHF) alleviate biases primarily based on reward signals from current model outputs without considering th…

Cited by 6SourcePDFScholar
2024

Arbitrary Motion Style Transfer with Multi-condition Motion Latent Diffusion Model

CVPR 2024poster

Computer animation's quest to bridge content and style has historically been a challenging venture with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality ensuring a harmonious fusion where the core narrative of t…

2024

Calibrating Reasoning in Language Models with Internal Consistency

NeurIPS 2024poster

Large language models (LLMs) have demonstrated impressive capabilities in various reasoning tasks, aided by techniques like chain-of-thought prompting that elicits verbalized reasoning. However, LLMs often generate text with obvious mistakes and contradictions, raising doubts about their ability to…

2024

Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond

ICML 2024poster

We introduce a novel framework of combinatorial multi-armed bandits (CMAB) with multivariant and probabilistically triggering arms (CMAB-MT), where the outcome of each arm is a $d$-dimensional multivariant random variable and the feedback follows a general arm triggering process. Compared with exist…

Cited by 4SourcePDFScholar
2024

Decoupling Degradations with Recurrent Network for Video Restoration in Under-Display Camera

AAAI 2024technical

Under-display camera (UDC) systems are the foundation of full-screen display devices in which the lens mounts under the display. The pixel array of light-emitting diodes used for display diffracts and attenuates incident light, causing various degradations as the light intensity changes. Unlike gene…

2024

Exploring Soft Prompt Initialization Strategy for Few-Shot Continual Text Classification

ICASSP 2024accepted

Few-shot continual learning (FSCL) is a challenging setting as it requires models to learn new knowledge with a few examples over time, and fast adapt to new tasks without forgetting previous knowledge. Prompt-tuning, as an efficient learning approach for language models, has shown competitive perfo…

Cited by 0SourceScholar
2024

Few-Shot Diffusion Models Escape the Curse of Dimensionality

NeurIPS 2024poster

While diffusion models have demonstrated impressive performance, there is a growing need for generating samples tailored to specific user-defined concepts. The customized requirements promote the development of few-shot diffusion models, which use limited $n_{ta}$ target samples to fine-tune a pre-t…

Cited by 1SourcePDFScholar
2024

HOIAnimator: Generating Text-prompt Human-object Animations using Novel Perceptive Diffusion Models

CVPR 2024poster

To date the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehensive physics-centric…

Cited by 11SourcePDFScholar
2024

Hallucination Diversity-Aware Active Learning for Text Summarization

NAACL 2024long

Large Language Models (LLMs) have shown propensity to generate hallucinated outputs, i.e., texts that are factually incorrect or unsupported. Existing methods for alleviating hallucinations typically require costly human annotations to identify and correct hallucinations in LLM outputs. Moreover, mo…

Cited by 6SourcePDFScholar
2024

In-Flight Cable Length Control for Improved Quadrotor-Based Suspended Load Transportation

RA-L 2024

Load transportation (LT) using quadrotor unmanned aerial vehicles (UAVs) has drawn increasing interest from the robotics community in recent years. Two main alternative load carrying modalities have been investigated. Namely, the load is either rigidly attached to the body of the quadrotor (LT-RA),

Cited by 15SourceScholar
2024

Learning Versatile Skills with Curriculum Masking

NeurIPS 2024poster

Masked prediction has emerged as a promising pretraining paradigm in offline reinforcement learning (RL) due to its versatile masking schemes, enabling flexible inference across various downstream tasks with a unified model. Despite the versatility of masked prediction, it remains unclear how to bal…

2024

Leveraging Drift to Improve Sample Complexity of Variance Exploding Diffusion Models

NeurIPS 2024poster

Variance exploding (VE) based diffusion models, an important class of diffusion models, have shown state-of-the-art (SOTA) performance. However, only a few theoretical works analyze VE-based models, and those works suffer from a worse forward convergence rate $1/\text{poly}(T)$ than the $\exp{(-T)}$…

Cited by 0SourcePDFScholar
2024

Motion-adaptive Separable Collaborative Filters for Blind Motion Deblurring

CVPR 2024poster

Eliminating image blur produced by various kinds of motion has been a challenging problem. Dominant approaches rely heavily on model capacity to remove blurring by reconstructing residual from blurry observation in feature space. These practices not only prevent the capture of spatially variable mot…

2024

Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation

ECCV 2024poster

"The scarcity of large-scale 3D-text paired data poses a great challenge on open vocabulary 3D scene understanding, and hence it is popular to leverage internet-scale 2D data and transfer their open vocabulary capabilities to 3D models through knowledge distillation. However, the existing distillati…

2024

SeeSR: Towards Semantics-Aware Real-World Image Super-Resolution

CVPR 2024poster

Owe to the powerful generative priors the pre-trained text-to-image (T2I) diffusion models have become increasingly popular in solving the real-world image super-resolution problem. However as a consequence of the heavy quality degradation of input low-resolution (LR) images the destruction of local…

2024

The Closeness of In-Context Learning and Weight Shifting for Softmax Regression

NeurIPS 2024poster

Large language models (LLMs) are known for their exceptional performance in natural language processing, making them highly effective in many human life-related tasks. The attention mechanism in the Transformer architecture is a critical component of LLMs, as it allows the model to selectively focus…

Cited by 43SourcePDFScholar
2024

Towards a Novel Soft Magnetic Laparoscope for Single Incision Laparoscopic Surgery

ICRA 2024poster

In single-incision laparoscopic surgery (SILS), magnetic anchoring and guidance system (MAGS) is a promising technique to prevent clutter in the surgical workspace and provide a larger vision field. Existing camera designs mainly rely on rigid structure design, resulting in risks of losing magnetic…

Cited by 0SourceScholar
2024

UniVS: Unified and Universal Video Segmentation with Prompts as Queries

CVPR 2024poster

Despite the recent advances in unified image segmentation (IS) developing a unified video segmentation (VS) model remains a challenge. This is mainly because generic category-specified VS tasks need to detect all objects and track them across consecutive frames while prompt-guided VS tasks require r…

2023

Adversarial Attacks on Online Learning to Rank with Click Feedback

NeurIPS 2023poster

Online learning to rank (OLTR) is a sequential decision-making problem where a learning agent selects an ordered list of items and receives feedback through user clicks. Although potential attacks against OLTR algorithms may cause serious losses in real-world applications, there is limited knowledge…

Cited by 6SourcePDFScholar
2023

DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning

IJCAI 2023poster

Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating with others, yet such privacy concern has not been considered in existing works in MARL. We propose the differentially…

2023

Demonstrating Large-Scale Package Manipulation via Learned Metrics of Pick Success

RSS 2023poster

Automating warehouse operations can reduce logistics overhead costs, ultimately driving down the final price for consumers, increasing the speed of delivery, and enhancing the resiliency to workforce fluctuations. The past few years have seen increased interest in automating such repeated tasks but…

Cited by 5SourcePDFScholar
2023

DynaMask: Dynamic Mask Selection for Instance Segmentation

CVPR 2023poster

The representative instance segmentation methods mostly segment different object instances with a mask of the fixed resolution, e.g., 28x 28 grid. However, a low-resolution mask loses rich details, while a high-resolution mask incurs quadratic computation overhead. It is a challenging task to predic…

2023

Efficient Explorative Key-Term Selection Strategies for Conversational Contextual Bandits

AAAI 2023technical

Conversational contextual bandits elicit user preferences by occasionally querying for explicit feedback on key-terms to accelerate learning. However, there are aspects of existing approaches which limit their performance. First, information gained from key-term-level conversations and arm-level re…

2023

FPR: False Positive Rectification for Weakly Supervised Semantic Segmentation

ICCV 2023poster

Many weakly supervised semantic segmentation (WSSS) methods employ the class activation map (CAM) to generate the initial segmentation results. However, CAM often fails to distinguish the foreground from its co-occurred background (e.g., train and railroad), resulting in inaccurate activation from t…

Cited by 46PDFcodeScholar
2023

FSI: Frequency and Spatial Interactive Learning for Image Restoration in Under-Display Cameras

ICCV 2023poster

Under-display camera (UDC) systems remove the screen notch for bezel-free displays and provide a better interactive experience. The main challenge is that the pixel array of light-emitting diodes used for display diffracts and attenuates the incident light, leading to complex degradation. Existing m…

Cited by 22PDFScholar
2023

Few-Shot Composition Learning for Image Retrieval with Prompt Tuning

AAAI 2023technical

We study the problem of composition learning for image retrieval, for which we learn to retrieve target images with search queries in the form of a composition of a reference image and a modification text that describes desired modifications of the image. Existing models of composition learning for…

Cited by 10SourcePDFScholar
2023

Future-conditioned Unsupervised Pretraining for Decision Transformer

ICML 2023poster

Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promising, return conditioning is limited to training data labeled with rewards and therefore faces challenges in learning fr…

2023

Generating Aligned Pseudo-Supervision From Non-Aligned Data for Image Restoration in Under-Display Camera

CVPR 2023poster

Due to the difficulty in collecting large-scale and perfectly aligned paired training data for Under-Display Camera (UDC) image restoration, previous methods resort to monitor-based image systems or simulation-based methods, sacrificing the realness of the data and introducing domain gaps. In this w…

2023

InfoPrompt: Information-Theoretic Soft Prompt Tuning for Natural Language Understanding

NeurIPS 2023poster

Soft prompt tuning achieves superior performances across a wide range of few-shot tasks. However, the performances of prompt tuning can be highly sensitive to the initialization of the prompts. We have also empirically observed that conventional prompt tuning methods cannot encode and learn sufficie…

Cited by 31SourcePDFScholar
2023

K-Nearest-Neighbor Local Sampling Based Conditional Independence Testing

NeurIPS 2023poster

Conditional independence (CI) testing is a fundamental task in statistics and machine learning, but its effectiveness is hindered by the challenges posed by high-dimensional conditioning variables and limited data samples. This article introduces a novel testing approach to address these challenges…

Cited by 2SourcePDFScholar
2023

Learning Adversarial Linear Mixture Markov Decision Processes with Bandit Feedback and Unknown Transition

ICLR 2023poster

We study reinforcement learning (RL) with linear function approximation, unknown transition, and adversarial losses in the bandit feedback setting. Specifically, the unknown transition probability function is a linear mixture model \citep{AyoubJSWY20,ZhouGS21,HeZG22} with a given feature mapping, an…

Cited by 13SourcePDFScholar
2023

Learning Adversarial Low-rank Markov Decision Processes with Unknown Transition and Full-information Feedback

NeurIPS 2023poster

In this work, we study the low-rank MDPs with adversarially changed losses in the full-information feedback setting. In particular, the unknown transition probability kernel admits a low-rank matrix decomposition \citep{REPUCB22}, and the loss functions may change adversarially but are revealed to t…

Cited by 5SourcePDFScholar
2023

MDQE: Mining Discriminative Query Embeddings To Segment Occluded Instances on Challenging Videos

CVPR 2023poster

While impressive progress has been achieved, video instance segmentation (VIS) methods with per-clip input often fail on challenging videos with occluded objects and crowded scenes. This is mainly because instance queries in these methods cannot encode well the discriminative embeddings of instances…

2023

MSF: Motion-Guided Sequential Fusion for Efficient 3D Object Detection From Point Cloud Sequences

CVPR 2023poster

Point cloud sequences are commonly used to accurately detect 3D objects in applications such as autonomous driving. Current top-performing multi-frame detectors mostly follow a Detect-and-Fuse framework, which extracts features from each frame of the sequence and fuses them to detect the objects in…

2023

Nearest-Neighbor Sampling Based Conditional Independence Testing

AAAI 2023technical

The conditional randomization test (CRT) was recently proposed to test whether two random variables X and Y are conditionally independent given random variables Z. The CRT assumes that the conditional distribution of X given Z is known under the null hypothesis and then it is compared to the distrib…

2023

One-to-Few Label Assignment for End-to-End Dense Detection

CVPR 2023poster

One-to-one (o2o) label assignment plays a key role for transformer based end-to-end detection, and it has been recently introduced in fully convolutional detectors for lightweight end-to-end dense detection. However, o2o can largely degrade the feature learning performance due to the limited number…

2023

Online Clustering of Bandits with Misspecified User Models

NeurIPS 2023poster

The contextual linear bandit is an important online learning problem where given arm features, a learning agent selects an arm at each round to maximize the cumulative rewards in the long run. A line of works, called the clustering of bandits (CB), utilize the collaborative effect over user preferen…

Cited by 13SourcePDFScholar
2023

Online Corrupted User Detection and Regret Minimization

NeurIPS 2023poster

In real-world online web systems, multiple users usually arrive sequentially into the system. For applications like click fraud and fake reviews, some users can maliciously perform corrupted (disrupted) behaviors to trick the system. Therefore, it is crucial to design efficient online learning algor…

Cited by 9SourcePDFScholar
2023

SIM: Semantic-Aware Instance Mask Generation for Box-Supervised Instance Segmentation

CVPR 2023poster

Weakly supervised instance segmentation using only bounding box annotations has recently attracted much research attention. Most of the current efforts leverage low-level image features as extra supervision without explicitly exploiting the high-level semantic information of the objects, which will…

2023

Sequential Texts Driven Cohesive Motions Synthesis with Natural Transitions

ICCV 2023poster

The intelligent synthesis/generation of daily-life motion sequences is fundamental and urgently needed for many VR/metaverse-related applications. However, existing approaches commonly focus on monotonic motion generation (e.g., walking, jumping, etc.) based on single instruction-like text, which is…

Cited by 15PDFcodeScholar
2023

Understanding Representation Learnability of Nonlinear Self-Supervised Learning

AAAI 2023technical

Self-supervised learning (SSL) has empirically shown its data representation learnability in many downstream tasks. There are only a few theoretical works on data representation learnability, and many of those focus on final data representation, treating the nonlinear neural network as a ``black box…

2022

A Multiscale Gradient-Backpropagation Optimization Framework for Deformable Convolution Based Compressed Video Enhancement

ICASSP 2022accepted

Deep learning based compressed video quality enhancement has raised lots of interest recently. To explore the information over multiple frames, deformable convolution has been used for temporal alignment. However, in the existing methods, the deformable convolution is used in a relatively naïve way,…

Cited by 1SourceScholar
2022

Class-Balanced Pixel-Level Self-Labeling for Domain Adaptive Semantic Segmentation

CVPR 2022poster

Domain adaptive semantic segmentation aims to learn a model with the supervision of source domain data, and produce satisfactory dense predictions on unlabeled target domain. One popular solution to this challenging task is self-training, which selects high-scoring predictions on target samples as p…

Cited by 111PDFcodeScholar
2022

Context-aware Information-theoretic Causal De-biasing for Interactive Sequence Labeling

EMNLP 2022finding

Supervised training of existing deep learning models for sequence labeling relies on large scale labeled datasets. Such datasets are generally created with crowd-source labeling. However, crowd-source labeling for tasks of sequence labeling can be expensive and time-consuming. Further, crowd-source…

Cited by 7SourcePDFScholar
2022

Discovering Low-rank Subspaces for Language-agnostic Multilingual Representations

EMNLP 2022main

Large pretrained multilingual language models (ML-LMs) have shown remarkable capabilities of zero-shot cross-lingual transfer, without direct cross-lingual supervision. While these results are promising, follow-up works found that, within the multilingual embedding spaces, there exists strong langua…

2022

DisenCite: Graph-Based Disentangled Representation Learning for Context-Specific Citation Generation

AAAI 2022technical

Citing and describing related literature are crucial to scientific writing. Many existing approaches show encouraging performance in citation recommendation, but are unable to accomplish the more challenging and onerous task of citation text generation. In this paper, we propose a novel disentangled…

2022

KUNet: Imaging Knowledge-Inspired Single HDR Image Reconstruction

IJCAI 2022poster

Recently, with the rise of high dynamic range (HDR) display devices, there is a great demand to transfer traditional low dynamic range (LDR) images into HDR versions. The key to success is how to solve the many-to-many mapping problem. However, the existing approaches either do not consider constrai…

2022

Learning to Plan Variable Length Sequences of Actions with a Cascading Bandit Click Model of User Feedback

AISTATS 2022poster

Motivated by problems of ranking with partial information, we introduce a variant of the cascading bandit model that considers flexible length sequences with varying rewards and losses. We formulate two generative models for this problem within the generalized linear setting, and design and analyze…

Cited by 4SourcePDFScholar
2022

PU-Refiner: A Geometry Refiner with Adversarial Learning for Point Cloud Upsampling

ICASSP 2022accepted

We present PU-Refiner, a generative adversarial network for point cloud upsampling. The generator of our network includes a coarse feature expansion module to create coarse upsampled features, a geometry generation module to regress a coarse point cloud from the coarse upsampled features, and a prog…

Cited by 0SourceScholar
2022

Reinforcement Learning-Based Adaptive Biofeedback Engine for Overground Walking Speed Training

RA-L 2022

Wearable biofeedback systems (WBS) have been proposed to aid physical rehabilitation of individuals with motor impairments. Due to significant inter- and intra-individual differences, the effectiveness of a given biofeedback strategy may vary for different users and across therapeutic sessions, as a

Cited by 9SourceScholar
2022

Simultaneously Learning Stochastic and Adversarial Bandits under the Position-Based Model

AAAI 2022technical

Online learning to rank (OLTR) interactively learns to choose lists of items from a large collection based on certain click models that describe users' click behaviors. Most recent works for this problem focus on the stochastic environment where the item attractiveness is assumed to be invariant dur…

Cited by 6SourcePDFScholar
2022

Simultaneously Learning Stochastic and Adversarial Bandits with General Graph Feedback

ICML 2022spotlight

The problem of online learning with graph feedback has been extensively studied in the literature due to its generality and potential to model various learning tasks. Existing works mainly study the adversarial and stochastic feedback separately. If the prior knowledge of the feedback mechanism is u…

Cited by 13SourcePDFScholar
2022

Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection From Point Clouds

CVPR 2022poster

Transformer has demonstrated promising performance in many 2D vision tasks. However, it is cumbersome to apply the self-attention underlying transformer on large-scale point cloud data because point cloud is a long sequence and unevenly distributed in 3D space. To solve this issue, existing methods…

Cited by 211PDFcodeScholar
2021

Category Dictionary Guided Unsupervised Domain Adaptation for Object Detection

AAAI 2021technical

Unsupervised domain adaption (UDA) is a promising solution to enhance the generalization ability of a model from a source domain to a target domain without manually annotating labels for target data. Recent works in cross-domain object detection mostly resort to adversarial feature adaptation to mat…

Cited by 49SourcePDFScholar
2021

Decentralized Circle Formation Control for Fish-like Robots in the Real-world via Reinforcement Learning

ICRA 2021poster

In this paper, the circle formation control problem is addressed for a group of cooperative underactuated fish-like robots involving unknown nonlinear dynamics and disturbances. Based on the reinforcement learning and cognitive consistency theory, we propose a decentralized controller without the kn…

Cited by 27SourceScholar
2021

Removing Diffraction Image Artifacts in Under-Display Camera via Dynamic Skip Connection Network

CVPR 2021poster

Recent development of Under-Display Camera (UDC) systems provides a true bezel-less and notch-free viewing experience on smartphones (and TV, laptops, tablets), while allowing images to be captured from the selfie camera embedded underneath. In a typical UDC system, the microstructure of the semi-tr…

Cited by 74PDFcodeScholar
2021

Self-Guided Community Detection on Networks with Missing Edges

IJCAI 2021poster

The vast majority of community detection algorithms assume that the networks are totally observed. However, in reality many networks cannot be fully observed. On such network is edges-missing network, where some relationships (edges) between two entities are missing. Recently, several works have bee…

Cited by 8SourcePDFScholar
2021

Spatial Feature Calibration and Temporal Fusion for Effective One-Stage Video Instance Segmentation

CVPR 2021poster

Modern one-stage video instance segmentation networks suffer from two limitations. First, convolutional features are neither aligned with anchor boxes nor with ground-truth bounding boxes, reducing the mask sensitivity to spatial location. Second, a video is directly divided into individual frames f…

Cited by 73PDFcodeScholar
2021

The Hardness Analysis of Thompson Sampling for Combinatorial Semi-bandits with Greedy Oracle

NeurIPS 2021poster

Thompson sampling (TS) has attracted a lot of interest in the bandit area. It was introduced in the 1930s but has not been theoretically proven until recent years. All of its analysis in the combinatorial multi-armed bandit (CMAB) setting requires an exact oracle to provide optimal solutions with an…

Cited by 6SourcePDFScholar
2020

LST-Net: Learning a Convolutional Neural Network with a Learnable Sparse Transform

ECCV 2020poster

The 2D convolutional (Conv2d) layer is the fundamental element to a deep convolutional neural network (CNN). Despite the great success of CNN, the conventional Conv2d is still limited in effectively reducing the spatial and channel-wise redundancy of features. In this paper, we propose to mitigate t…

2020

Online Influence Maximization under Linear Threshold Model

NeurIPS 2020poster

Online influence maximization (OIM) is a popular problem in social networks to learn influence propagation model parameters and maximize the influence spread at the same time. Most previous studies focus on the independent cascade (IC) model under the edge-level feedback. In this paper, we address O…

Cited by 52SourcePDFScholar
2020

Towards Understanding the Regularization of Adversarial Robustness on Neural Networks

ICML 2020poster

The problem of adversarial examples has shown that modern Neural Network (NN) models could be rather fragile. Among the more established techniques to solve the problem, one is to require the model to be \emph{$\epsilon$-adversarially robust} (AR); that is, to require the model not to change predict…

Cited by 25SourcePDFScholar
2019

Adaptive Assist-as-needed Control Based on Actor-Critic Reinforcement Learning

IROS 2019poster

In robot-assisted rehabilitation, assist-as-needed (AAN) controllers have been proposed to promote subjects’ active participation, which is thought to lead to better training outcomes. Most of these AAN controllers require a patient-specific manual tuning of the parameters defining the underlying fo…

Cited by 37SourceScholar
2019

Dynamic Anchor Feature Selection for Single-Shot Object Detection

ICCV 2019poster

The design of anchors is critical to the performance of one-stage detectors. Recently, the anchor refinement module (ARM) has been proposed to adjust the initialization of default anchors, providing the detector a better anchor reference. However, this module brings another problem: all pixels at a…

Cited by 53PDFScholar
2019

Manipulability Optimization Control of a Serial Redundant Robot for Robot-assisted Minimally Invasive Surgery

ICRA 2019poster

This paper proposes a manipulability optimization control of a 7-DoF robot manipulator for Robot-Assisted Minimally Invasive Surgery (RAMIS), which at the same time guarantees a Remote Center of Motion (RCM). The first degree of redundancy of the manipulator is used to achieve an RCM constraint, the…

Cited by 62SourceScholar
2019

Tracking Control of Fully-Constrained Cable-Driven Parallel Robots using Adaptive Dynamic Programming

IROS 2019poster

In this paper, a new adaptive tracking controller with learning ability is proposed for fully-constrained cable-driven parallel robots (CDPRs). For these systems, the necessity of maintaining positive and bounded tensions in all cables while coping with disturbances represents a critical control req…

Cited by 5SourceScholar
2018

A Novel Recurrent Neural Network for Improving Redundant Manipulator Motion Planning Completeness

ICRA 2018poster

Recurrent Neural Networks (RNNs) demonstrated advantages on control precision, system robustness and computational efficiency, and have been widely applied to redundant manipulator control optimization. Existing RNN control schemes locally optimize trajectories and are efficient and reliable on obst…

Cited by 33SourceScholar
2018

Independently Recurrent Neural Network (IndRNN): Building a Longer and Deeper RNN

CVPR 2018poster

Recurrent neural networks (RNNs) have been widely used for processing sequential data. However, RNNs are commonly difficult to train due to the well-known gradient vanishing and exploding problems and hard to learn long-term patterns. Long short-term memory (LSTM) and gated recurrent unit (GRU) were…

2018

TopRank: A practical algorithm for online stochastic ranking

NeurIPS 2018poster

Online learning to rank is a sequential decision-making problem where in each round the learning agent chooses a list of items and receives feedback in the form of clicks from the user. Many sample-efficient algorithms have been proposed for this problem that assume a specific click model connecting…

Cited by 86SourcePDFScholar
2017

Improving control precision and motion adaptiveness for surgical robot with recurrent neural network

IROS 2017poster

Surgical robot research is driven by the desire of improving surgical outcomes. This paper proposed a Recurrent Neural Network based controller to address two problems: 1) improving control precision, 2) increasing adaptiveness for robot motion (explained in Section I). RNN was adopted in this work…

Cited by 34SourceScholar
2017

On Context-Dependent Clustering of Bandits

ICML 2017poster

We investigate a novel cluster-of-bandit algorithm CAB for collaborative recommendation tasks that implements the underlying feedback sharing mechanism by estimating user neighborhoods in a context-dependent manner. CAB makes sharp departures from the state of the art by incorporating collaborative…

Cited by 168SourcePDFScholar
2016

A perception system for detecting brake levers in outdoor rail yard environments

IROS 2016poster

A rail yard is a dangerous environment for humans to work in, primarily because of the possibility of serious injuries associated with moving rail cars, locomotives, and uneven terrain. For robots to act autonomously in such environments, there exists a need for a perception system that can act reli…

Cited by 2SourceScholar
2015

A comparative study of contact models for contact-aware state estimation

IROS 2015poster

We study the contact-aware state estimation (CASE) problem, i.e., the problem of estimating the state of an object while it is being actively manipulated by a robot. Several researchers have developed particle filters for this problem. They estimate the state (pose and velocity) of manipulated objec…

Cited by 11SourceScholar