← Search

Jian Li

161 accepted papers

2026

Asynchronous Matching with Dynamic Sampling for Multimodal Dataset Distillation

ICLR 2026poster

Multimodal Dataset Distillation (MDD) has emerged as a vital paradigm for enabling efficient training of vision-language models (VLMs) in the era of multimodal data proliferation. Unlike traditional dataset distillation methods that focus on single-modal tasks, MDD presents distinct challenges: (i)…

Cited by 0SourceScholar
2026

Beyond Frame-Wise Tracking: A Trajectory-Based Paradigm for Efficient Point Cloud Tracking

ICRA 2026poster

LiDAR-based 3D single object tracking (3D SOT) is a critical task in robotics and autonomous systems. Existing methods typically follow frame-wise motion estimation or a sequence-based paradigm. However, the two-frame methods are efficient but lack long-term temporal context, making them vulnerable …

2026

COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learning

CVPR 2026

Recent advances in text-to-image (T2I) generation can produce highly resembling images of copyrighted characters, often indistinguishable from official depictions, raising serious concerns about intellectual property infringement. Consequently, robust detection of copyright character infringement is

Cited by 0SourceScholar
2026

CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions

ICLR 2026poster

Deep learning has demonstrated remarkable capabilities in simulating complex dynamic systems. However, existing methods require known physical properties as supervision or inputs, limiting their applicability under unknown conditions. To explore this challenge, we introduce Cloth Dynamics Grounding…

Cited by 0SourcecodeScholar
2026

Deep Incomplete Multi-View Clustering via Hierarchical Imputation and Alignment

AAAI 2026technical

Incomplete multi-view clustering (IMVC) aims to discover shared cluster structures from multi-view data with partial observations. The core challenges lie in accurately imputing missing views without introducing bias, while maintaining semantic consistency across views and compactness within cluster

Cited by 0SourcePDFScholar
2026

FAVE: A Structured Benchmark for Fine-Grained Audio-Visual Temporal Evaluation in Multimodal LLMs

CVPR 2026

Audio-visual large language models (AVLLMs) have made significant strides in understanding visual and auditory content. However, their ability to capture fine-grained temporal relationships between audio and visual streams remains insufficiently evaluated. To address this, we introduce FAVE (Fine-gr

Cited by 0SourceScholar
2026

Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic Skills

AAAI 2026technical

Autonomous execution of long-horizon, contact-rich manipulation tasks traditionally requires extensive real-world data and expert engineering, posing significant cost and scalability challenges. This paper proposes a novel framework integrating hierarchical semantic decomposition, reinforcement lear

Cited by 0SourcePDFScholar
2026

Kronos: A Foundation Model for the Language of Financial Markets

AAAI 2026technical

The success of large-scale pre-training paradigm, exemplified by Large Language Models (LLMs), has inspired the development of Time Series Foundation Models (TSFMs). However, their application to financial candlestick (K-line) data remains limited, often underperforming non-pre-trained architectures

Cited by 0SourcePDFScholar
2026

L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention

AAAI 2026technical

Recently, Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs), but Vision–Language Models (VLMs) still struggle with multi-step reasoning tasks due to limited multimodal reasoning data. To bridge this gap, researchers have explored methods to

Cited by 0SourcePDFScholar
2026

LLM-Oriented Token-Adaptive Knowledge Distillation

AAAI 2026technical

Knowledge Distillation (KD) is a key technique for compressing Large-scale Language Models (LLMs), but prevailing logit-based methods employ static strategies misaligned with the student’s dynamic learning process. By treating all tokens indiscriminately with a fixed temperature, these methods resul

Cited by 0SourcePDFScholar
2026

Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models

ICLR 2026poster

Large Language Models (LLMs) achieve impressive performance across many tasks but remain prone to hallucination, especially in long-form generation where redundant retrieved contexts and lengthy reasoning chains amplify factual errors. Recent studies highlight a critical phenomenon: the closer key i…

Cited by 0SourceScholar
2026

Navigating the Alpha Jungle: An LLM-Powered MCTS Framework for Formulaic Alpha Factor Mining

AAAI 2026technical

Alpha factor mining is pivotal in quantitative investment for identifying predictive signals from complex financial data. While traditional formulaic alpha mining relies on human expertise, contemporary automated methods, such as those based on genetic programming or reinforcement learning, often st

Cited by 0SourcePDFScholar
2026

Population-Free Pareto Tracking for Sample-Efficient Multi-Policy MORL

ICML 2026poster

Multi-objective reinforcement learning (MORL) is a fundamental framework for real-world decision-making problems involving multiple conflicting criteria. Existing multi-policy (MP) methods typically rely on online evolutionary frameworks that maintain large policy populations, leading to high sample…

Cited by 0SourceScholar
2026

Predicting Video Slot Attention Queries from Random Slot-Feature Pairs

AAAI 2026technical

Unsupervised video Object-Centric Learning (OCL) is promising as it enables object-level scene representation and understanding as we humans do. Mainstream video OCL methods adopt a recurrent architecture: An aggregator aggregates current video frame into object features, termed slots, under some qu

Cited by 0SourcePDFScholar
2026

RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents

ICLR 2026poster

Large language models (LLMs) excel at logical and algorithmic reasoning, yet their emotional intelligence (EQ) still lags far behind their cognitive prowess. While reinforcement learning from verifiable rewards (RLVR) has advanced in other domains, its application to dialogue—especially for emotion…

Cited by 0SourcecodeScholar
2026

RadarMP: Motion Perception for 4D mmWave Radar in Autonomous Driving

AAAI 2026technical

Accurate 3D scene motion perception significantly enhances the safety and reliability of an autonomous driving system. Benefiting from its all-weather operational capability and unique perceptual properties, 4D mmWave radar has emerged as an essential component in advanced autonomous driving. Howeve

Cited by 0SourcePDFScholar
2026

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models

ICML 2026poster

1-bit LLM quantization offers significant advantages in reducing storage and computational costs. However, existing methods typically train 1-bit LLMs from scratch, failing to fully leverage pre-trained models. This results in high training costs and notable accuracy degradation. We identify that th…

Cited by 0SourceScholar
2026

SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM Inference

CVPR 2026

Video Large Language Models (VideoLLMs) face a fundamental performance bottleneck: the token explosion intrinsic to video inputs. The resulting O(N^2) prefill cost makes conventional Transformer inference prohibitively expensive at scale. Existing attempts fall into a hard accuracy-latency dilemma:

Cited by 0SourcecodeScholar
2026

Server-Proximal Aggregation for Federated Domain-Incremental Learning under Partial Participation: Task-Uniform Convergence and Backward Transfer

ICML 2026poster

Real-world federated systems seldom operate on static data: input distributions drift while privacy rules forbid raw data sharing. We study Federated Domain-Incremental Learning (FDIL), where (i) clients are heterogeneous, (ii) tasks arrive sequentially with shifting domains, and (iii) the label spa…

Cited by 0SourceScholar
2026

Sharp-Wave Ripples Learning: A Bio-Inspired Incremental Learning Method

IJCAI 2026

The primary challenge in incremental learning for deep neural networks (DNNs) lies in balancing memory stability with learning plasticity as they incrementally acquire new knowledge. Existing incremental learning methods typically require storing portions of old datasets or relying on specialized ne

Cited by 0Scholar
2026

Stealing Split Learning Bottom Models by Recovering Embedding Geometry

CVPR 2026

Vertical federated learning (VFL) trains models by splitting computation across clients and a server that only exchange intermediate embeddings. Recent work shows that a server even if honest-but-curious can steal a client's bottom model by querying the system and regressing on the returned embeddin

Cited by 0SourceScholar
2026

TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating

ICML 2026poster

Reinforcement learning from verifiable rewards (RLVR) has become an important paradigm for enhancing the reasoning capabilities of large language models, while it also involves a persistent tradeoff between optimization stability and learning efficiency. Token-level importance weighting supports fin…

Cited by 0SourceScholar
2026

Textual Stochastic Gradient Descent: Discrete Optimization of External Memory for Reasoning Language Agents

ICML 2026poster

While Large Language Models (LLMs) possess strong reasoning capabilities, enabling them to learn continuously from experience without parametric retraining remains an open challenge. Existing Retrieval-Augmented Generation (RAG) approaches typically treat memory as a static or append-only corpus, le…

Cited by 0SourceScholar
2026

UniMo: Unified Motion Generation and Understanding with Chain of Thought

AAAI 2026technical

Existing 3D human motion generation and understanding methods often exhibit limited interpretability, restricting effective mutual enhancement between these inherently related tasks. While current unified frameworks based on large language models (LLMs) leverage linguistic priors, they frequently en

Cited by 0SourcePDFScholar
2025

AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning

EMNLP 2025

Continual learning (CL) is essential for deploying large language models (LLMs) in dynamic real-world environments without the need for costly retraining. Recent model merging-based methods have attracted significant attention, but they still struggle to effectively manage the trade-off between lear

2025

Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs

AAAI 2025technical

Despite the superior performance of Large language models on many NLP tasks, they still face significant limitations in memorizing extensive world knowledge. Recent studies have demonstrated that leveraging the Retrieval-Augmented Generation (RAG) framework, combined with Knowledge Graphs that enca…

2025

Advancing High-Resolution and Efficient Automotive Radar Imaging through Domain-Informed 1D Deep Learning

ICASSP 2025accepted

Millimeter-wave (mmWave) radars are critical for autonomous vehicles’ perception tasks, offering reliable performance in adverse weather conditions. However, their application is often hindered by insufficient spatial resolution for detailed semantic scene interpretation. Traditional super-resolutio…

Cited by 0SourceScholar
2025

Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization

EMNLP 2025

Direct Preference Optimization (DPO) is a widely used reinforcement learning from human feedback (RLHF) method across various domains. The study of token importance has attracted widespread attention in DPO. Researchers have found that token importance is crucial for improving the effectiveness of D

2025

CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision

ACL 2025long

Tool invocation significantly enhances the capabilities of Large Language Models (LLMs), yet challenges persist, particularly in complex task scenarios. Current methods, such as instruction-enhanced reasoning and supervised fine-tuning, often result in unnecessarily long reasoning paths and face dif…

Cited by 0SourcePDFScholar
2025

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards

EMNLP 2025

Role-Playing Language Agents (RPLAs) have emerged as a significant application direction for Large Language Models (LLMs). Existing approaches typically rely on prompt engineering or supervised fine-tuning to enable models to imitate character behaviors in specific scenarios, but often neglect the u

Cited by 0SourcePDFScholar
2025

CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models

ACL 2025long

Faithfulness hallucinations are claims generated by a Large Language Model (LLM) not supported by contexts provided to the LLM. Lacking assessment standards, existing benchmarks focus on “factual statements” that rephrase source materials while overlooking “cognitive statements” that involve making…

2025

DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback

ICLR 2025poster

Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward. However, the success of RMAB crucially relies on the availa…

Cited by 2SourcePDFScholar
2025

Decentralized Federated Learning with Model Caching on Mobile Agents

AAAI 2025technical

Federated Learning (FL) trains a shared model using data and computation power on distributed agents coordinated by a central server. Decentralized FL (DFL) utilizes local model exchange and aggregation between agents to reduce the communication and computation overheads on the central server. Howev…

2025

Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement

ICASSP 2025accepted

Deep learning-based speech enhancement (SE) models have recently outperformed traditional techniques, yet their deployment on resource-constrained devices remains challenging due to high computational and memory demands. This paper introduces a novel dynamic frequency-adaptive knowledge distillation…

Cited by 0SourceScholar
2025

Edge-Guided Lighting Adaptation: Real-Time Detection of Transparent Objects for Cell Culture Robot

IROS 2025

In robot-assisted cell culture tasks, fluctuations in lighting conditions can result in blurred boundaries, intensified reflections, and pronounced refractions of transparent objects. These optical phenomena collectively escalate the complexity of image processing and target recognition. To address

Cited by 0SourceScholar
2025

FactorGCL: A Hypergraph-Based Factor Model with Temporal Residual Contrastive Learning for Stock Returns Prediction

AAAI 2025technical

As a fundamental method in economics and finance, the factor model has been extensively utilized in quantitative investment. In recent years, there has been a paradigm shift from traditional linear models with expert-designed factors to more flexible nonlinear machine learning-based models with data…

Cited by 0SourcePDFScholar
2025

Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks

ICLR 2025poster

In this work, we investigate a particular implicit bias in gradient descent training, which we term “Feature Averaging,” and argue that it is one of the principal factors contributing to the non-robustness of deep neural networks. We show that, even when multiple discriminative features are present…

Cited by 1SourcePDFScholar
2025

From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspective

CVPR 2025poster

Ultra-high-definition (UHD) image restoration faces significant challenges due to its high resolution, complex content, and intricate details. To cope with these challenges, we analyze the restoration process in depth through a progressive spectral perspective, and deconstruct the complex UHD restor…

2025

Game-Based Social Learning Particle Swarm Optimizer for Inverse Kinematics of Robotic Arms

RA-L 2025

Robot inverse kinematics is the foundation of robotics and an indispensable part of robot development and applications. The inverse kinematics of a robotic arm involves its specific configuration, the non-convex coupling relationships between joints, and the presence of multiple solutions, among oth

Cited by 2SourceScholar
2025

MBA-RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question Complexity

COLING 2025main

Retrieval Augmented Generation (RAG) has proven to be highly effective in boosting the generative performance of language model in knowledge-intensive tasks. However, existing RAG framework either indiscriminately perform retrieval or rely on rigid single-label classifiers to select retrieval method…

2025

MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving

IROS 2025

This paper introduces MCTrack, a new 3D multi-object tracking method that achieves performance across KITTI, nuScenes, and Waymo datasets. Addressing the gap in existing tracking paradigms, which often perform well on specific datasets but lack generalizability, MCTrack offers a unified solution. Ad

Cited by 27SourcecodeScholar
2025

MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection

ICLR 2025poster

In the field of industrial inspection, Multimodal Large Language Models (MLLMs) have a high potential to renew the paradigms in practical applications due to their robust language capabilities and generalization abilities. However, despite their impressive problem-solving skills in many domains, MLL…

2025

Micro-UAV with Ant-Inspired Bistable Gripper for Adaptive Perching and Wildlife Detection

IROS 2025

With the global ecological environment facing continuous deterioration, effective monitoring of arboreal birds in complex canopy environments remains challenging due to limitations of conventional drones in endurance, size, and habitat disturbance. To address these challenges, this paper presents an

Cited by 0SourceScholar
2025

Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints

AAAI 2025technical

Multi-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominantly focus on transformer-based multimodal fusion, aiming to extract robust multimodal representations. However, ambiguit…

2025

On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared Representations

ICLR 2025poster

Federated reinforcement learning (FedRL) enables multiple agents to collaboratively learn a policy without needing to share the local trajectories collected during agent-environment interactions. However, in practice, the environments faced by different agents are often heterogeneous, but since exis…

Cited by 0SourcePDFScholar
2025

Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior

ICLR 2025poster

Recent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating…

2025

The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement

NeurIPS 2025poster

Large language models (LLMs) have recently transformed from text-based assistants to autonomous agents capable of planning, reasoning, and iteratively improving their actions. While numerical reward signals and verifiers can effectively rank candidate actions, they often provide limited contextual g…

Cited by 0SourceScholar
2025

Towards Universal Dataset Distillation via Task-Driven Diffusion

CVPR 2025poster

Dataset distillation (DD) condenses key information from large-scale datasets into smaller synthetic datasets, reducing storage and computational costs for training networks. However, recent research has primarily focused on image classification tasks, with limited expansion to detection and segment…

Cited by 0SourcePDFScholar
2025

Understanding Constraint Inference in Safety-Critical Inverse Reinforcement Learning

ICLR 2025poster

In practical applications, the underlying constraint knowledge is often unknown and difficult to specify. To address this issue, recent advances in Inverse Constrained Reinforcement Learning (ICRL) have focused on inferring these constraints from expert demonstrations. However, the ICRL approach typ…

Cited by 1SourcePDFScholar
2025

Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws

NeurIPS 2025spotlight

Large Language Models (LLMs) have demonstrated remarkable capabilities across numerous tasks, yet principled explanations for their underlying mechanisms and several phenomena, such as scaling laws, hallucinations, and related behaviors, remain elusive. In this work, we revisit the classical relatio…

Cited by 0SourceScholar
2025

Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly

CVPR 2025poster

**M**ultimodal **L**arge **L**anguage **M**odels (MLLMs) have displayed remarkable performance in multimodal tasks, particularly in visual comprehension. However, we reveal that MLLMs often generate incorrect answers even when they understand the visual content. To this end, we manually construct a…

2025

VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model

NeurIPS 2025poster

With the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models still experience high latency when generating the first audio token during stre…

Cited by 0SourceScholar
2025

pFedGPT: Hierarchically Optimizing LoRA Aggregation Weights for Personalized Federated GPT Models

EMNLP 2025

Federated finetuning of Large Language Models (LLMs) using Low-Rank Adaptation (LoRA) offers computational efficiency and preserves data privacy. However, applying LoRA in federated settings faces significant challenges: standard approaches struggle with data heterogeneity, and existing personalizat

Cited by 0SourcePDFScholar
2024

6-DoF Grasp Detection in Clutter with Enhanced Receptive Field and Graspable Balance Sampling

IROS 2024poster

6-DoF grasp detection of small-scale grasps is crucial for robots to perform specific tasks. This paper focuses on enhancing the recognition capability of small-scale grasping, aiming to improve the overall accuracy of grasping prediction results and the generalization ability of the network. We pro…

Cited by 1SourceScholar
2024

Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game

ACL 2024findings

Human preference alignment is essential to improve the interaction quality of large language models (LLMs). Existing alignment methods depend on manually annotated preference data to guide the LLM optimization directions. However, continuously updating LLMs for alignment raises a distribution gap be…

2024

All Neural Kronecker Product Beamforming for Speech Extraction with Large-Scale Microphone Arrays

ICASSP 2024accepted

Existing frame-wise neural beamformers for speech extraction can obtain promising performance in relatively high signal-to-noise ratio (SNR) scenarios using small microphone arrays, while they still suffer from performance degradation in relatively low SNR environments, e.g., SNR<-5 dB. As an attemp…

Cited by 0SourceScholar
2024

Backdoor Federated Learning by Poisoning Backdoor-Critical Layers

ICLR 2024poster

Federated learning (FL) has been widely deployed to enable machine learning training on sensitive data across distributed devices. However, the decentralized learning paradigm and heterogeneity of FL further extend the attack surface for backdoor attacks. Existing FL attack and defense methodologies…

Cited by 16SourcePDFScholar
2024

Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models

EMNLP 2024main

Retrieval-augmented language model (RALM) represents a significant advancement in mitigating factual hallucination by leveraging external knowledge sources. However, the reliability of the retrieved information is not always guaranteed, and the retrieval of irrelevant data can mislead the response g…

Cited by 106SourcePDFScholar
2024

Cheaper and Faster: Distributed Deep Reinforcement Learning with Serverless Computing

AAAI 2024technical

Deep reinforcement learning (DRL) has gained immense success in many applications, including gaming AI, robotics, and system scheduling. Distributed algorithms and architectures have been vastly proposed (e.g., actor-learner architecture) to accelerate DRL training with large-scale server-based clus…

Cited by 7SourcePDFScholar
2024

Convolutional Spectral Kernel Learning with Generalization Guarantees (Abstract Reprint)

AAAI 2024technical

Kernel methods are powerful tools to capture nonlinear patterns behind given data but often lead to poor performance on complicated tasks compared to convolutional neural networks. The reason is that kernel methods are still shallow and fully connected models, failing to reveal hierarchical features…

Cited by 0SourcePDFScholar
2024

DePRL: Achieving Linear Convergence Speedup in Personalized Decentralized Learning with Shared Representations

AAAI 2024technical

Decentralized learning has emerged as an alternative method to the popular parameter-server framework which suffers from high communication burden, single-point failure and scalability issues due to the need of a central server. However, most existing works focus on a single shared model for all wo…

Cited by 7SourcePDFScholar
2024

Efficient-PIP: Large-scale Pixel-level Aligned Image Pair Generation for Cross-time Infrared-RGB Translation

IROS 2024poster

Generative models are gaining momentum in both academic and industrial applications driven by the availability of large-scale datasets, especially in tasks involving Image-to-Image Translation. Meanwhile, poor human perception of nighttime environment has led to a demand for translation from night-v…

Cited by 0SourcecodeScholar
2024

EmoPrompt-ECPE: Emotion Knowledge-aware Prompt-tuning for Emotion-Cause Pair Extraction

COLING 2024main

Emotion-cause pair extraction (ECPE) main focus is on extracting all potential emotion clauses and corresponding cause clauses from unannotated documents. Existing methods achieve promising results with the help of fine-tuning and prompt paradigms, but they present three downsides. First, most appro…

2024

Fetch and Forge: Efficient Dataset Condensation for Object Detection

NeurIPS 2024poster

Dataset condensation (DC) is an emerging technique capable of creating compact synthetic datasets from large originals while maintaining considerable performance. It is crucial for accelerating network training and reducing data storage requirements. However, current research on DC mainly focuses o…

Cited by 1SourcePDFScholar
2024

High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to Algorithm

AAAI 2024technical

Overparameterization often leads to benign overfitting, where deep neural networks can be trained to overfit the training data but still generalize well on unseen data. However, it lacks a generalized asymptotic framework for nonlinear regressions and connections to conventional complexity notions.…

2024

IMM: An Imitative Reinforcement Learning Approach with Predictive Representation Learning for Automatic Market Making

IJCAI 2024poster

Market making (MM) via Reinforcement Learning (RL) has attracted significant attention in financial trading. Most existing RL-based MM methods focus on optimizing single-price level strategies which fail at frequent order cancellations and loss of queue priority. By comparison, strategies involving…

Cited by 2SourcePDFScholar
2024

Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback

ICML 2024poster

Restless multi-armed bandits (RMAB) play a central role in modeling sequential decision making problems under an instantaneous activation constraint that at most $B$ arms can be activated at any decision epoch. Each restless arm is endowed with a state that evolves independently according to a Marko…

Cited by 3SourcePDFScholar
2024

Rethinking the Uniformity Metric in Self-Supervised Learning

ICLR 2024poster

Uniformity plays an important role in evaluating learned representations, providing insights into self-supervised learning. In our quest for effective uniformity metrics, we pinpoint four principled properties that such metrics should possess. Namely, an effective uniformity metric should remain inv…

2024

Self-supervised Preference Optimization: Enhance Your Language Model with Preference Degree Awareness

EMNLP 2024finding

Recently, there has been significant interest in replacing the reward model in Reinforcement Learning with Human Feedback (RLHF) methods for Large Language Models (LLMs), such as Direct Preference Optimization (DPO) and its variants. These approaches commonly use a binary cross-entropy mechanism on…

2024

Towards Reliable Advertising Image Generation Using Human Feedback

ECCV 2024poster

"In the e-commerce realm, compelling advertising images are pivotal for attracting customer attention. While generative models automate image generation, they often produce substandard images that may mislead customers and require significant labor costs to inspect. This paper delves into increasing…

2024

Trade When Opportunity Comes: Price Movement Forecasting via Locality-Aware Attention and Iterative Refinement Labeling

IJCAI 2024poster

Price movement forecasting, aimed at predicting financial asset trends based on current market information, has achieved promising advancements through machine learning (ML) methods. Most existing ML methods, however, struggle with the extremely low signal-to-noise ratio and stochastic nature of fin…

Cited by 4SourcePDFScholar
2024

USM-Lite: Quantization and Sparsity Aware Fine-Tuning for Speech Recognition with Universal Speech Models

ICASSP 2024accepted

End-to-end automatic speech recognition (ASR) models have seen revolutionary quality gains with the recent development of large-scale universal speech models (USM). However, deploying these massive USMs is extremely expensive due to the enormous memory usage and computational cost. Therefore, model…

Cited by 0SourceScholar
2024

Visual Loop Closure Detection with Thorough Temporal and Spatial Context Exploitation

IROS 2024poster

Despite advancements in visual Simultaneous Localization and Mapping (SLAM), prevailing visual Loop Closure Detection (LCD) methods primarily rely on computationally intensive image similarity comparisons, neglecting temporal-spatial context during long-term exploration. To address this issue, we pr…

Cited by 0SourceScholar
2023

AEC-GAN: Adversarial Error Correction GANs for Auto-Regressive Long Time-Series Generation

AAAI 2023technical

Large-scale high-quality data is critical for training modern deep neural networks. However, data acquisition can be costly or time-consuming for many time-series applications, thus researchers turn to generative models for generating synthetic time-series data. In particular, recent generative adve…

Cited by 11SourcePDFScholar
2023

An Autonomous Surgical Instrument Tracking Framework With a Binocular Camera for a Robotic Flexible Laparoscope

RA-L 2023

In minimally invasive surgery (MIS), the field of view (FOV) plays a vital role. To enhance the stability of FOV and lighten the burden on surgeons, robot-assisted laparoscope systems have been developed and introduced into surgery. However, most of the existing automatic surgical tool tracking sche

Cited by 13SourceScholar
2023

Applied Online Algorithms with Heterogeneous Predictors

ICML 2023poster

For many application domains, the integration of machine learning (ML) models into decision making is hindered by the poor explainability and theoretical guarantees of black box models. Although the emerging area of algorithms with predictions offers a way to leverage ML while enjoying worst-case gu…

Cited by 7SourcePDFScholar
2023

DeFL: Defending against Model Poisoning Attacks in Federated Learning via Critical Learning Periods Awareness

AAAI 2023technical

Federated learning (FL) is known to be susceptible to model poisoning attacks in which malicious clients hamper the accuracy of the global model by sending manipulated model updates to the central server during the FL training process. Existing defenses mainly focus on Byzantine-robust FL aggregati…

Cited by 23SourcePDFScholar
2023

FCC: Feature Clusters Compression for Long-Tailed Visual Recognition

CVPR 2023poster

Deep Neural Networks (DNNs) are rather restrictive in long-tailed data, since they commonly exhibit an under-representation for minority classes. Various remedies have been proposed to tackle this problem from different perspectives, but they ignore the impact of the density of Backbone Features (BF…

2023

Finite-Time Analysis of Whittle Index based Q-Learning for Restless Multi-Armed Bandits with Neural Network Function Approximation

NeurIPS 2023poster

Whittle index policy is a heuristic to the intractable restless multi-armed bandits (RMAB) problem. Although it is provably asymptotically optimal, finding Whittle indices remains difficult. In this paper, we present Neural-Q-Whittle, a Whittle index based Q-learning algorithm for RMAB with neural…

Cited by 19SourcePDFScholar
2023

ImGCL: Revisiting Graph Contrastive Learning on Imbalanced Node Classification

AAAI 2023technical

Graph contrastive learning (GCL) has attracted a surge of attention due to its superior performance for learning node/graph representations without labels. However, in practice, the underlying class distribution of unlabeled nodes for the given graph is usually imbalanced. This highly imbalanced cla…

Cited by 28SourcePDFScholar
2023

Integrating the Sensing and Radio Communications Channel Modelling From Radar Mutual Interference

ICASSP 2023accepted

The current growing interest in integrated sensing and communications (ISAC) for the next generation of radio access networks towards 6G is opening new challenges on the channel estimation and modelling. New frequency bands and novel techniques for joining the sensing and the transmission of informa…

Cited by 0SourceScholar
2023

Learning From Noisy Labels With Decoupled Meta Label Purifier

CVPR 2023poster

Training deep neural networks (DNN) with noisy labels is challenging since DNN can easily memorize inaccurate labels, leading to poor generalization ability. Recently, the meta-learning based label correction strategy is widely adopted to tackle this problem via identifying and correcting potential…

2023

Not All Tasks Are Born Equal: Understanding Zero-Shot Generalization

ICLR 2023top-25%

Recent work has achieved remarkable zero-shot performance with multi-task prompted pretraining, but little has been understood. For the first time, we show that training on a small number of key tasks beats using all the training tasks, while removing these key tasks substantially hurts performance.…

Cited by 14SourcePDFScholar
2023

Online Restless Bandits with Unobserved States

ICML 2023poster

We study the online restless bandit problem, where each arm evolves according to a Markov chain independently, and the reward of pulling an arm depends on both the current state of the corresponding Markov chain and the pulled arm. The agent (decision maker) does not know the transition functions an…

Cited by 7SourcePDFScholar
2023

OpenFE: Automated Feature Generation with Expert-level Performance

ICML 2023poster

The goal of automated feature generation is to liberate machine learning experts from the laborious task of manual feature generation, which is crucial for improving the learning performance of tabular data. The major challenge in automated feature generation is to efficiently and accurately identif…

Cited by 32SourcePDFScholar
2023

Rethinking Mobile Block for Efficient Attention-based Models

ICCV 2023poster

This paper focuses on developing modern, efficient, lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Inverted Residual Block (IRB) serves as the infrastructure for lightweight CNNs, but no counterpart has been recognized by attention-based studies. This…

Cited by 180PDFcodeScholar
2023

Robotic Kinematic Calibration with Only Position Data and Consideration of Non-Geometric Errors Using POE-Based Model and Gaussian Mixture Models

IROS 2023poster

Kinematic calibration is crucial to improve the positioning accuracy of serial robots. This paper proposes a novel algorithm for robotic kinematic calibration based on an augmented product of exponentials (POE)-based kinematic model using Gaussian mixture models (GMMs) with only position data. In th…

Cited by 2SourceScholar
2023

Symphony in the Latent Space: Provably Integrating High-Dimensional Techniques with Non-linear Machine Learning Models

AAAI 2023technical

This paper revisits building machine learning algorithms that involve interactions between entities, such as those between financial assets in an actively managed portfolio, or interactions between users in a social network. Our goal is to forecast the future evolution of ensembles of multivariate…

Cited by 5SourcePDFScholar
2023

Towards Generalizable Reinforcement Learning for Trade Execution

IJCAI 2023poster

Optimized trade execution is to sell (or buy) a given amount of assets in a given time with the lowest possible trading cost. Recently, reinforcement learning (RL) has been applied to optimized trade execution to learn smarter policies from market data. However, we find that many existing RL methods…

2023

Unbiased Gradient Boosting Decision Tree with Unbiased Feature Importance

IJCAI 2023poster

Gradient Boosting Decision Tree (GBDT) has achieved remarkable success in a wide variety of applications. The split finding algorithm, which determines the tree construction process, is one of the most crucial components of GBDT. However, the split finding algorithm has long been criticized for its…

2022

A Novel Single-Arm Stapling Robot for Oral and Maxillofacial Surgery - Design and Verification

RA-L 2022

Because of the limited space, the suture of the oral and maxillofacial surgery is a challenging task, which requires oral surgeons to master excellent techniques. This letter presents a single-arm stapling robot for oral and maxillofacial surgery using magnesium alloy staples, as well the stapling s

Cited by 7SourceScholar
2022

A Surgeon Preference-Guided Autonomous Instrument Tracking Method With a Robotic Flexible Endoscope Based on dVRK Platform

RA-L 2022

In minimally invasive surgery, endoscopes serve as the eyes of surgeon. To avoid fatigue in manual endoscope steering, robotic endoscope holders have been developed. Unfortunately, existing robotic endoscope holders are not widely adopted due to the poor surgeon-robot cooperation. In this work, we d

Cited by 30SourceScholar
2022

A4LidarTag: Depth-Based Fiducial Marker for Extrinsic Calibration of Solid-State Lidar and Camera

RA-L 2022

Visual-based simultaneous localization and mapping (SLAM) systems perform weakly in object tracking and map reconstruction due to the unreliable depth measurement originating from image-only data. Light Detection and Ranging (LiDAR) can be coupled to overcome the drawback of uncertain depth estimati

Cited by 22SourceScholar
2022

Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of Stability

NeurIPS 2022accept

Recent findings demonstrate that modern neural networks trained by full-batch gradient descent typically enter a regime called Edge of Stability (EOS). In this regime, the sharpness, i.e., the maximum Hessian eigenvalue, first increases to the value 2/(step size) (the progressive sharpening phase) a…

Cited by 26SourcePDFScholar
2022

Analyzing and Mitigating Interference in Neural Architecture Search

ICML 2022spotlight

Weight sharing is a popular approach to reduce the training cost of neural architecture search (NAS) by reusing the weights of shared operators from previously trained child models. However, the rank correlation between the estimated accuracy and ground truth accuracy of those child models is low du…

Cited by 0SourcePDFScholar
2022

Controlled Text Generation Using Dictionary Prior in Variational Autoencoders

ACL 2022findings

While variational autoencoders (VAEs) have been widely applied in text generation tasks, they are troubled by two challenges: insufficient representation capacity and poor controllability. The former results from the posterior collapse and restrictive assumption, which impede better representation l…

Cited by 12SourcePDFScholar
2022

Cross-modal Fusion-based Prior Correction for Road Detection in Off-road Environments

IROS 2022poster

Road detection plays a fundamental role in the visual navigation system of autonomous vehicles. However, it's still challenging to achieve robust road detection in off-road scenarios due to their complicated road appearances and ambiguous road structures. Therefore, existing image-based road detecti…

Cited by 2SourceScholar
2022

DIGAT: Modeling News Recommendation with Dual-Graph Interaction

EMNLP 2022finding

News recommendation (NR) is essential for online news services. Existing NR methods typically adopt a news-user representation learning framework, facing two potential limitations. First, in news encoder, single candidate news encoding suffers from an insufficient semantic information problem. Secon…

2022

Design and Control of a Highly Redundant Rigid-flexible Coupling Robot to Assist the COVID-19 Oropharyngeal-Swab Sampling

RA-L 2022

The outbreak of novel coronavirus pneumonia (COVID-19) has caused mortality and morbidity worldwide. Oropharyngeal-swab (OP-swab) sampling is widely used for the diagnosis of COVID-19 in the world. To avoid the clinical staff from being affected by the virus, we developed a 9-degree-of-freedom (DOF)

Cited by 52SourceScholar
2022

FactorVAE: A Probabilistic Dynamic Factor Model Based on Variational Autoencoder for Predicting Cross-Sectional Stock Returns

AAAI 2022technical

As an asset pricing model in economics and finance, factor model has been widely used in quantitative investment. Towards building more effective factor models, recent years have witnessed the paradigm shift from linear models to more flexible nonlinear data-driven machine learning models. However,…

2022

Learning Infinite-Horizon Average-Reward Restless Multi-Action Bandits via Index Awareness

NeurIPS 2022accept

We consider the online restless bandits with average-reward and multiple actions, where the state of each arm evolves according to a Markov decision process (MDP), and the reward of pulling an arm depends on both the current state of the corresponding MDP and the action taken. Since finding the opt…

Cited by 17SourcePDFScholar
2022

Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation

NeurIPS 2022accept

While large-scale neural language models, such as GPT2 and BART, have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms (\textit{e.g.}, greedy search). This phenomenon is counter-intuit…

2022

MINER: Multi-Interest Matching Network for News Recommendation

ACL 2022findings

Personalized news recommendation is an essential technique to help users find interested news. Accurately matching user’s interests and candidate news is the key to news recommendation. Most existing methods learn a single user embedding from user’s historical behaviors to represent the reading inte…

Cited by 87SourcePDFScholar
2022

MTRec: Multi-Task Learning over BERT for News Recommendation

ACL 2022findings

Existing news recommendation methods usually learn news representations solely based on news titles. To sufficiently utilize other fields of news information such as category and entities, some methods treat each field as an additional feature and combine different feature vectors with attentive poo…

Cited by 38SourcePDFScholar
2022

Reinforcement Learning Augmented Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits

AAAI 2022technical

We study a finite-horizon restless multi-armed bandit problem with multiple actions, dubbed as R(MA)^2B. The state of each arm evolves according to a controlled Markov decision process (MDP), and the reward of pulling an arm depends on both the current state and action of the corresponding MDP. Sinc…

Cited by 23SourcePDFScholar
2022

SCSNet: An Efficient Paradigm for Learning Simultaneously Image Colorization and Super-resolution

AAAI 2022technical

In the practical application of restoring low-resolution gray-scale images, we generally need to run three separate processes of image colorization, super-resolution, and dows-sampling operation for the target device. However, this pipeline is redundant and inefficient for the independent processes,…

Cited by 15SourcePDFScholar
2021

Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model

NeurIPS 2021poster

Inspired by biological evolution, we explain the rationality of Vision Transformer by analogy with the proven practical Evolutionary Algorithm (EA) and derive that both of them have consistent mathematical representation. Analogous to the dynamic local population in EA, we improve the existing trans…

Cited by 21SourcePDFScholar
2021

Exploration by Maximizing Renyi Entropy for Reward-Free RL Framework

AAAI 2021technical

Exploration is essential for reinforcement learning (RL). To face the challenges of exploration, we consider a reward-free RL framework that completely separates exploration from exploitation and brings new challenges for exploration algorithms. In the exploration phase, the agent learns an explorat…

2021

Model Adaptation through Hypothesis Transfer with Gradual Knowledge Distillation

IROS 2021poster

The ability to adapt their perception to changing environments is a core characterization of intelligent robots. At present, Unsupervised Domain Adaptation (UDA) methods are used to address this problem where the adaptation task is formulated as a transfer problem from a well-described scenario (sou…

Cited by 21SourceScholar
2021

Multifunctional Robotic Glove with Active-Passive Training Modes for Hand Rehabilitation and Assistance

IROS 2021poster

Soft robotic gloves have shown great advantages in assisting individuals with hand pathologies to perform continuous exercises to restore their hand functions, which could considerably accelerate the rehabilitation process and reduce the costs. However, single rehabilitation mode, difficulty in achi…

Cited by 5SourceScholar
2021

Natural Language Processing Meets Quantum Physics: A Survey and Categorization

EMNLP 2021main

Recent research has investigated quantum NLP, designing algorithms that process natural language in quantum computers, and also quantum-inspired algorithms that improve NLP performance on classical computers. In this survey, we review representative methods at the intersection of NLP and quantum phy…

Cited by 22SourcePDFScholar
2020

Kalman Filtering Attention for User Behavior Modeling in CTR Prediction

NeurIPS 2020spotlight

Click-through rate (CTR) prediction is one of the fundamental tasks for e-commerce search engines. As search becomes more personalized, it is necessary to capture the user interest from rich behavior data. Existing user behavior modeling algorithms develop different attention mechanisms to emphasize…

2020

Online Algorithms for Multi-shop Ski Rental with Machine Learned Advice

NeurIPS 2020poster

We study the problem of augmenting online algorithms with machine learned (ML) advice. In particular, we consider the \emph{multi-shop ski rental} (MSSR) problem, which is a generalization of the classical ski rental problem. In MSSR, each shop has different prices for buying and renting a pair of…

2019

Distributed Quickest Detection of Significant Events in Networks

ICASSP 2019accepted

The problem of quickest detection of significant events in networks is studied. A distributed setting is investigated, where there is no fusion center, and each node only communicates with its neighbors. After an event occurs in the network, a number of nodes are affected, which changes the statisti…

Cited by 0SourceScholar
2019

Tracking by Animation: Unsupervised Learning of Multi-Object Attentive Trackers

CVPR 2019poster

Online Multi-Object Tracking (MOT) from videos is a challenging computer vision task which has been extensively studied for decades. Most of the existing MOT algorithms are based on the Tracking-by-Detection (TBD) paradigm combined with popular machine learning approaches which largely reduce the hu…

Cited by 57PDFcodeScholar
2019

Transform Domain Based Medical Image Super-resolution via Deep Multi-scale Network

ICASSP 2019accepted

This paper proposes a new medical image super-resolution (SR) network, namely deep multi-scale network (DMSN), in the uniform discrete curvelet transform (UDCT) domain. DMSN is made up of a set of cascaded multi-scale fushion (MSF) blocks. In each MSF block, we use convolution kernels of different s…

Cited by 0SourceScholar
2018

BRITS: Bidirectional Recurrent Imputation for Time Series

NeurIPS 2018poster

Time series are widely used as signals in many classification/regression tasks. It is ubiquitous that time series contains many missing values. Given multiple correlated time series data, how to fill in missing values and to predict their class labels? Existing imputation methods often impose strong…

2018

Multi-Class Learning: From Theory to Algorithm

NeurIPS 2018poster

In this paper, we study the generalization performance of multi-class classification and obtain a shaper data-dependent generalization error bound with fast convergence rate, substantially improving the state-of-art bounds in the existing data-dependent generalization analysis. The theoretical analy…

Cited by 58SourcePDFScholar
2017

Compressive pulse-Doppler radar sensing via 1-bit sampling with time-varying threshold

ICASSP 2017accepted

This paper proposes a compressive pulse-Doppler radar that works through one-bit quantization of the received noisy signal. The one-bit quantization is performed by comparing the signal with a time-varying reference level. Considering the sparsity of the targets in the range-Doppler domain, the prob…

Cited by 0SourceScholar
2017

Learning Gradient Descent: Better Generalization and Longer Horizons

ICML 2017poster

Training deep neural networks is a highly nontrivial task, involving carefully selecting appropriate training algorithms, scheduling step sizes and tuning other hyperparameters. Trying different combinations can be quite labor-intensive and time consuming. Recently, researchers have tried to use dee…

2016

Combinatorial Multi-Armed Bandit with General Reward Functions

NeurIPS 2016poster

In this paper, we study the stochastic combinatorial multi-armed bandit (CMAB) framework that allows a general nonlinear reward function, whose expected value may not depend only on the means of the input random variables but possibly on the entire distributions of these variables. Our framework ena…

Cited by 175SourcePDFScholar
2015

Automatic target recognition using discrimination based on optimal transport

ICASSP 2015accepted

The use of distances based on optimal transportation has recently shown promise for discrimination of power spectra. In particular, spectral estimation methods based on ℓ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> regularization as well as…

Cited by 0SourceScholar