← Search

Yi Xu

140 accepted papers

2026

ActiveUMI: Robotic Manipulation with Active Perception from Robot‑Free Human Demonstrations

ICRA 2026poster

We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples a portable VR teleoperation kit with sensorized controllers that mirror the robot's end-effectors, bridging human-robot…

2026

Cost-Sensitive Conformal Training with Provably Controllable Learning Bounds

AAAI 2026technical

Conformal prediction (CP) is a general framework to quantify the predictive uncertainty of machine learning models that uses a set prediction to include the true label with a valid probability. To align the uncertainty measured by CP, conformal training methods minimize the size of the prediction se

Cited by 0SourcePDFScholar
2026

Den-TP: A Density-Balanced Data Curation and Evaluation Framework for Trajectory Prediction

CVPR 2026

Trajectory prediction in autonomous driving has traditionally been studied from a model-centric perspective. However, existing datasets exhibit a strong long-tail distribution in scenario density, where common low-density cases dominate and safety-critical high-density cases are severely underrepres

Cited by 4SourcecodeScholar
2026

Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed Memory

AAAI 2026technical

Sparse neural systems are gaining traction for efficient continual learning due to their modularity and low interference. Architectures like Sparse Distributed Memory Multi-Layer Perceptrons (SDMLP) construct task-specific subnetworks via Top-K activation and have shown resilience against catastroph

Cited by 0SourcePDFScholar
2026

Feed-Forward Taylor-Gaussians-Flow: Towards Non-uniform Motion for Novel View Synthesis from Monocular Video

ICML 2026poster

Long-term non-uniform motion poses a significant challenge for feed-forward Novel View Synthesis (\textbf{NVS}), as it requires modeling higher-order motion, such as acceleration. Existing methods primarily rely on deformation fields or scene flow, which are limited to first-order approximations. Du…

Cited by 0SourceScholar
2026

HumanoidExo: Scalable Whole-Body Humanoid Manipulation Via Wearable Exoskeleton

ICRA 2026poster

A significant bottleneck in humanoid policy learning is the acquisition of large-scale, diverse datasets, as collecting reliable real-world data remains both difficult and cost-prohibitive. To address this limitation, we introduce HumanoidExo, a novel system that transfers human motion to whole-body…

2026

IQ-LUT: INTERPOLATED AND QUANTIZED LUT FOR EFFICIENT IMAGE SUPER-RESOLUTION

ICASSP 2026poster

Lookup table (LUT) methods demonstrate considerable potential in accelerating image super-resolution inference. However, pursuing higher image quality through larger receptive fields and bit-depth triggers exponential growth in the LUT's index space, creating a storage bottleneck that limits deploym…

Cited by 0SourcePDFScholar
2026

Null-Space Filtering for Data-free Continual Model Merging: Preserving Transparency, Promoting Fidelity

ICLR 2026poster

Data-free continual model merging (DFCMM) aims to fuse independently fine-tuned models into a single backbone that evolves with incoming tasks without accessing task data. This paper formulate two fundamental desiderata for DFCMM: transparency, avoiding interference with earlier tasks, and fidelity,…

Cited by 0SourceScholar
2026

Open-World Object Manipulation with Vision-Language-Action Models Via Synthetic Multi-Modal Data

ICRA 2026poster

Imitation learning has proven to be highly effective in teaching robots dexterous manipulation skills. However, it typically relies on large amounts of robot data, which limits its scalability and applicability in dynamic, real-world environments. One key challenge in this context is object generali…

Cited by 0Scholar
2026

SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense

ICLR 2026poster

Large Vision-Language Models (LVLMs) excel in diverse cross-modal tasks. However, object hallucination, where models produce plausible but inaccurate object descriptions, remains a significant challenge. In contrast to previous work focusing on LLM components, this paper is the first to trace LVLM h…

Cited by 0SourcecodeScholar
2026

Score-Based Model for Low-Rank Tensor Recovery

AAAI 2026technical

Low-rank tensor decompositions (TDs) provide an effective framework for multiway data analysis. Traditional TD methods rely on predefined structural assumptions, such as CP or Tucker decompositions. From a probabilistic perspective, these methods effectively model the relationships between latent fa

Cited by 0SourcePDFScholar
2026

Semantically Consistent Language Gaussian Splatting for 3D Point-Level Open-Vocabulary Querying

ICRA 2026poster

Open-vocabulary 3D scene understanding is crucial for robotics applications, such as natural language-driven manipulation, human-robot interaction, and autonomous navigation. Existing methods for querying 3D Gaussian Splatting often struggle with inconsistent 2D mask supervision and lack a robust 3D…

Cited by 0Scholar
2026

Stochastic Gradient Methods under Heavy-Tailed Noises in Weakly Convex Optimization

ICML 2026poster

Recently, many empirical work has shown that, in machine learning, the noise distribution of stochastic gradients often exhibits heavy tails when stochastic optimization methods are employed. Most existing theoretical analyses of heavy-tailed stochastic methods rely on various convexity and smoothne…

Cited by 0SourceScholar
2026

Visual Planning: Let's Think Only with Images

ICLR 2026oral

Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across diverse tasks. However, these models predominantly rely on pure text as the medium for both expressing and structuring reasoning, even when visual informat…

Cited by 0SourcecodeScholar
2025

A Skill-Based Hierarchical Framework with Dangerous Action Masking for Autonomous Navigation of Jumping Robots

IROS 2025

Achieving autonomous navigation for biologically inspired jumping robots remains a long-standing challenge, due to the inherent instability of jumping motions and the limitations in onboard sensor capabilities. This paper proposes a skill-based hierarchical framework with dangerous action masking (S

Cited by 0SourceScholar
2025

ActiveGAMER: Active GAussian Mapping through Efficient Rendering

CVPR 2025poster

We introduce ActiveGAMER, an active mapping system that utilizes 3D Gaussian Splatting (3DGS) to achieve high-quality, real-time scene mapping and exploration. Unlike traditional NeRF-based methods, which are computationally demanding and restrict active mapping performance, our approach leverages t…

Cited by 2SourcePDFScholar
2025

AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs

ACL 2025long

When aligning large language models (LLMs), their performance across various tasks (such as being helpful, harmless, and honest) is heavily influenced by the composition of the training data. However, it is difficult to determine what mixture of data should be used to produce a model with strong per…

2025

BIG-FUSION: Brain-Inspired Global-Local Context Fusion Framework for Multimodal Emotion Recognition in Conversations

AAAI 2025technical

Considering the importance of capturing both global conversational topics and local speaker dependencies for multimodal emotion recognition in conversations, current approaches first utilize sequence models like Transformer to extract global context information, then apply Graph Neural Networks to m…

Cited by 0SourcePDFScholar
2025

ChatVLA-2: Vision-Language-Action Model with Open-World Reasoning

NeurIPS 2025poster

Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks.…

Cited by 0SourceScholar
2025

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

EMNLP 2025

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can’t large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in vision-language-action models (VLA), we identify two key ch

2025

ConformalSAM: Unlocking the Potential of Foundational Segmentation Models in Semi-Supervised Semantic Segmentation with Conformal Prediction

ICCV 2025poster

Pixel-level vision tasks, such as semantic segmentation, require extensive and high-quality annotated data, which is costly to obtain. Semi-supervised semantic segmentation (SSSS) has emerged as a solution to alleviate the labeling burden by leveraging both labeled and unlabeled data through self-tr…

Cited by 0SourcePDFScholar
2025

Efficient ANN-SNN Conversion with Error Compensation Learning

ICML 2025poster

Artificial neural networks (ANNs) have demonstrated outstanding performance in numerous tasks, but deployment in resource-constrained environments remains a challenge due to their high computational and memory requirements. Spiking neural networks (SNNs) operate through discrete spike events and off…

Cited by 0SourcePDFScholar
2025

FaStFact: Faster, Stronger Long-Form Factuality Evaluations in LLMs

EMNLP 2025

Evaluating the factuality of long-form generations from Large Language Models (LLMs) remains challenging due to accuracy issues and costly human assessment. Prior evaluation pipelines attempt this by decomposing text into claims, searching for evidence, and verifying claims, but suffer from critical

2025

MAPS: Motivation-Aware Personalized Search via LLM-Driven Consultation Alignment

ACL 2025long

Personalized product search aims to retrieve and rank items that match users’ preferences and search intent. Despite their effectiveness, existing approaches typically assume that users’ query fully captures their real motivation. However, our analysis of a real-world e-commerce platform reveals tha…

2025

MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging

NeurIPS 2025poster

Continual model merging integrates independently fine-tuned models sequentially without access to the original training data, offering a scalable and efficient solution for continual learning. However, existing methods face two critical challenges: parameter interference among tasks, which leads to…

Cited by 0SourcecodeScholar
2025

PlanarNeRF: Online Learning of Planar Primitives with Neural Radiance Fields

ICRA 2025

Identifying spatially complete planar primitives from visual data is a crucial task in computer vision. Prior methods are largely restricted to either 2D segment recovery or simplifying 3D structures, even with extensive plane annotations. We present PlanarNeRF, a novel framework capable of detectin

Cited by 8SourceScholar
2025

Predicting Spectral Information for Self-Supervised Signal Classification

IJCAI 2025

Deep learning methods have demonstrated remarkable performance across various communication signal processing tasks. However, most signal classification methods require a substantial amount of labeled samples for training, posing significant challenges in the field of communication signals, as label

Cited by 0SourcePDFScholar
2025

Representation Potentials of Foundation Models for Multimodal Alignment: A Survey

EMNLP 2025

Foundation models learn highly transferable representations through large-scale pretraining on diverse data. An increasing body of research indicates that these representations exhibit a remarkable degree of similarity across architectures and modalities. In this survey, we investigate the represent

2025

Similarity = Value? Consultation Value-Assessment and Alignment for Personalized Search

EMNLP 2025

Personalized search systems in e-commerce platforms increasingly involve user interactions with AI assistants, where users consult about products, usage scenarios, and more. Leveraging consultation to personalize search services is trending. Existing methods typically rely on semantic similarity to

2025

SpikingYOLOX: Improved YOLOX Object Detection with Fast Fourier Convolution and Spiking Neural Networks

AAAI 2025technical

In recent years, with the advancements in brain science, spiking neural networks (SNNs) have garnered significant attention. SNNs can generate spikes that mimic the function of neurons transmission in humans brain, thereby significantly reducing computational costs by the event-driven nature during…

Cited by 0SourcePDFScholar
2025

Uncertainty-Guided Enhancement on Driving Perception System Via Foundation Models

ICRA 2025

Multimodal foundation models offer promising advancements for enhancing driving perception systems, but their high computational and financial costs pose challenges. We develop a method that leverages foundation models to refine predictions from existing driving perception modelssuch as enhancing ob

Cited by 4SourceScholar
2025

Understanding while Exploring: Semantics-driven Active Mapping

NeurIPS 2025poster

Effective robotic autonomy in unknown environments demands proactive exploration and precise understanding of both geometry and semantics. In this paper, we propose ActiveSGM, an active semantic mapping framework designed to predict the informativeness of potential observations before execution. Bui…

Cited by 0SourceScholar
2025

VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision

CoRL 2025poster

Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized to mimic driving patterns observed in data, without capturing the underlying reasoning processes. This limitation constr…

Cited by 0SourceScholar
2024

Better Representations via Adversarial Training in Pre-Training: A Theoretical Perspective

AISTATS 2024poster

Pre-training is known to generate universal representations for downstream tasks in large-scale deep learning such as large language models. Existing literature, e.g., Kim et al. (2020), empirically observe that the downstream tasks can inherit the adversarial robustness of the pre-trained model. We…

2024

Diffusion Models for Multi-Task Generative Modeling

ICLR 2024poster

Diffusion-based generative modeling has been achieving state-of-the-art results on various generation tasks. Most diffusion models, however, are limited to a single-generation modeling. Can we generalize diffusion models with the ability of multi-modal generative training for more generalizable mode…

Cited by 7SourcePDFScholar
2024

Diverse and Stable 2D Diffusion Guided Text to 3D Generation with Noise Recalibration

AAAI 2024technical

In recent years, following the success of text guided image generation, text guided 3D generation has gained increasing attention among researchers. Dreamfusion is a notable approach that enhances generation quality by utilizing 2D text guided diffusion models and introducing SDS loss, a technique f…

2024

Dual-Consistency Model Inversion for Non-Exemplar Class Incremental Learning

CVPR 2024poster

Non-exemplar class incremental learning (NECIL) aims to continuously assimilate new knowledge without forgetting previously acquired ones when historical data are unavailable. One of the generative NECIL methods is to invert the images of old classes for joint training. However these synthetic image…

Cited by 4SourcePDFScholar
2024

Evolutionary Contrastive Distillation for Language Model Alignment

EMNLP 2024finding

The ability of large language models (LLMs) to execute complex instructions is essential for their real-world applications. However, several recent studies indicate that LLMs struggle with challenging instructions. In this paper, we propose Evolutionary Contrastive Distillation (ECD), a novel method…

2024

FBLG: A Local Graph Based Approach for Handling Dual Skewed Non-IID Data in Federated Learning

IJCAI 2024poster

In real-world situations, federated learning often needs to process non-IID (non-independent and identically distributed) data with multiple skews, causing inadequate model performance. Existing federated learning methods mainly focus on addressing the problem with a single skew of non-IID, and henc…

2024

Facilitating Multimodal Classification via Dynamically Learning Modality Gap

NeurIPS 2024poster

Multimodal learning falls into the trap of the optimization dilemma due to the modality imbalance phenomenon, leading to unsatisfactory performance in real applications. A core reason for modality imbalance is that the models of each modality converge at different rates. Many attempts naturally focu…

2024

HuRef: HUman-REadable Fingerprint for Large Language Models

NeurIPS 2024poster

Protecting the copyright of large language models (LLMs) has become crucial due to their resource-intensive training and accompanying carefully designed licenses. However, identifying the original base model of an LLM is challenging due to potential parameter alterations. In this study, we introduce…

2024

Is Reference Necessary in the Evaluation of NLG Systems? When and Where?

NAACL 2024long

The majority of automatic metrics for evaluating NLG systems are reference-based. However, the challenge of collecting human annotation results in a lack of reliable references in numerous application scenarios. Despite recent advancements in reference-free metrics, it has not been well understood w…

2024

Learn to Optimize Denoising Scores: A Unified and Improved Diffusion Prior for 3D Generation

ECCV 2024poster

"In this paper, we propose a unified framework aimed at enhancing the diffusion priors for 3D generation tasks. Despite the critical importance of these tasks, existing methodologies often struggle to generate high-caliber results. We begin by examining the inherent limitations in previous diffusion…

2024

NARUTO: Neural Active Reconstruction from Uncertain Target Observations

CVPR 2024poster

We present NARUTO a neural active reconstruction system that combines a hybrid neural representation with uncertainty learning enabling high-fidelity surface reconstruction. Our approach leverages a multi-resolution hash-grid as the mapping backbone chosen for its exceptional convergence speed and c…

2024

OOSTraj: Out-of-Sight Trajectory Prediction With Vision-Positioning Denoising

CVPR 2024poster

Trajectory prediction is fundamental in computer vision and autonomous driving particularly for understanding pedestrian behavior and enabling proactive decision-making. Existing approaches in this field often assume precise and complete observational data neglecting the challenges associated with o…

2024

PanoFree: Tuning-Free Holistic Multi-view Image Generation with Cross-view Self-Guidance

ECCV 2024poster

"Immersive scene generation, notably panorama creation, benefits significantly from the adaptation of large pre-trained text-to-image (T2I) models for multi-view image generation. Due to the high cost of acquiring multi-view images, tuning-free generation is preferred. However, existing methods are…

2024

ParCo: Part-Coordinating Text-to-Motion Synthesis

ECCV 2024poster

"We study a challenging task: text-to-motion synthesis, aiming to generate motions that align with textual descriptions and exhibit coordinated movements. Currently, the part-based methods introduce part partition into the motion synthesis process to achieve finer-grained generation. However, these…

2024

RepEval: Effective Text Evaluation with LLM Representation

EMNLP 2024main

The era of Large Language Models (LLMs) raises new demands for automatic evaluation metrics, which should be adaptable to various application scenarios while maintaining low cost and effectiveness. Traditional metrics for automatic text evaluation are often tailored to specific scenarios, while LLM-…

2024

Robust Multi-Task Learning with Excess Risks

ICML 2024poster

Multi-task learning (MTL) considers learning a joint model for multiple tasks by optimizing a convex combination of all task losses. To solve the optimization problem, existing methods use an adaptive weight updating scheme, where task weights are dynamically adjusted based on their respective losse…

2024

Shopping MMLU: A Massive Multi-Task Online Shopping Benchmark for Large Language Models

NeurIPS 2024poster

Online shopping is a complex multi-task, few-shot learning problem with a wide and evolving range of entities, relations, and tasks. However, existing models and benchmarks are commonly tailored to specific tasks, falling short of capturing the full complexity of online shopping. Large Language Mode…

2024

Skeleton-Based Human Action Recognition with Noisy Labels

IROS 2024poster

Understanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are o…

Cited by 5SourcecodeScholar
2024

Spacetime Gaussian Feature Splatting for Real-Time Dynamic View Synthesis

CVPR 2024poster

Novel view synthesis of dynamic scenes has been an intriguing yet challenging problem. Despite recent advancements simultaneously achieving high-resolution photorealistic results real-time rendering and compact storage remains a formidable task. To address these challenges we propose Spacetime Gauss…

2024

Spectrum AUC Difference (SAUCD): Human-aligned 3D Shape Evaluation

CVPR 2024poster

Existing 3D mesh shape evaluation metrics mainly focus on the overall shape but are usually less sensitive to local details. This makes them inconsistent with human evaluation as human perception cares about both overall and detailed shape. In this paper we propose an analytic metric named Spectrum…

Cited by 7SourcePDFScholar
2024

Stereo-NEC: Enhancing Stereo Visual-Inertial SLAM Initialization with Normal Epipolar Constraints

ICRA 2024poster

We propose an accurate and robust initialization approach for stereo visual-inertial SLAM systems. Unlike the current state-of-the-art method, which heavily relies on the accuracy of a pure visual SLAM system to estimate inertial variables without updating camera poses, potentially compromising accu…

Cited by 10SourcecodeScholar
2024

SynPrompt: Syntax-aware Enhanced Prompt Engineering for Aspect-based Sentiment Analysis

COLING 2024main

Although there have been some works using prompt learning for the Aspect-based Sentiment Analysis(ABSA) tasks, their methods of prompt-tuning are simple and crude. Compared with vanilla fine-tuning methods, prompt learning intuitively bridges the objective form gap between pre-training and fine-tuni…

2024

The Pitfalls and Promise of Conformal Inference Under Adversarial Attacks

ICML 2024poster

In safety-critical applications such as medical imaging and autonomous driving, where decisions have profound implications for patient health and road safety, it is imperative to maintain both high adversarial robustness to protect against potential adversarial attacks and reliable uncertainty quant…

2023

3D-Aware Facial Landmark Detection via Multi-View Consistent Training on Synthetic Data

CVPR 2023poster

Accurate facial landmark detection on wild images plays an essential role in human-computer interaction, entertainment, and medical applications. Existing approaches have limitations in enforcing 3D consistency while detecting 3D/2D facial landmarks due to the lack of multi-view in-the-wild training…

2023

AdamsFormer for Spatial Action Localization in the Future

CVPR 2023poster

Predicting future action locations is vital for applications like human-robot collaboration. While some computer vision tasks have made progress in predicting human actions, accurately localizing these actions in future frames remains an area with room for improvement. We introduce a new task called…

2023

Design and Optimization of a Miniature Locust-Inspired Stable Jumping Robot

RA-L 2023

Jumping is a key locomotion for miniature robots, but it is difficult for a robot to jump a long distance without flipping. To solve this problem, we develop a miniature locust-inspired jumping robot, which has a body length of 10 cm and weight of 60 g. On the basis of the extracted skeletal muscle

Cited by 15SourceScholar
2023

Exploring and Verbalizing Academic Ideas by Concept Co-occurrence

ACL 2023long

Researchers usually come up with new ideas only after thoroughly comprehending vast quantities of literature. The difficulty of this procedure is exacerbated by the fact that the number of academic publications is growing exponentially. In this study, we devise a framework based on concept co-occurr…

2023

High Fidelity 3D Hand Shape Reconstruction via Scalable Graph Frequency Decomposition

CVPR 2023poster

Despite the impressive performance obtained by recent single-image hand modeling techniques, they lack the capability to capture sufficient details of the 3D hand mesh. This deficiency greatly limits their applications when high fidelity hand modeling is required, e.g., personalized hand modeling. T…

2023

In-sample Actor Critic for Offline Reinforcement Learning

ICLR 2023poster

Offline reinforcement learning suffers from out-of-distribution issue and extrapolation error. Most methods penalize the out-of-distribution state-action pairs or regularize the trained policy towards the behavior policy but cannot guarantee to get rid of extrapolation error. We propose In-sample…

Cited by 13SourcePDFScholar
2023

Interventional Bag Multi-Instance Learning on Whole-Slide Pathological Images

CVPR 2023highlight

Multi-instance learning (MIL) is an effective paradigm for whole-slide pathological images (WSIs) classification to handle the gigapixel resolution and slide-level label. Prevailing MIL methods primarily focus on improving the feature extractor and aggregator. However, one deficiency of these method…

2023

Mining and Applying Composition Knowledge of Dance Moves for Style-Concentrated Dance Generation

AAAI 2023technical

Choreography refers to creation of dance motions according to both music and dance knowledge, where the created dances should be style-specific and consistent. However, most of the existing methods generate dances using the given music as the only reference, lacking the stylized dancing knowledge, n…

2023

NeuRBF: A Neural Fields Representation with Adaptive Radial Basis Functions

ICCV 2023oral

We present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for storing local neural features and N-dimensional linear kernels for interpolating features at continuous query points. The…

Cited by 82PDFcodeScholar
2023

Not All Out-of-Distribution Data Are Harmful to Open-Set Active Learning

NeurIPS 2023poster

Active learning (AL) methods have been proven to be an effective way to reduce the labeling effort by intelligently selecting valuable instances for annotation. Despite their great success with in-distribution (ID) scenarios, AL methods suffer from performance degradation in many real-world applicat…

2023

OpenIllumination: A Multi-Illumination Dataset for Inverse Rendering Evaluation on Real Objects

NeurIPS 2023poster

We introduce OpenIllumination, a real-world dataset containing over 108K images of 64 objects with diverse materials, captured under 72 camera views and a large number of different illuminations. For each image in the dataset, we provide accurate camera parameters, illumination ground truth, and for…

2023

RIAV-MVS: Recurrent-Indexing an Asymmetric Volume for Multi-View Stereo

CVPR 2023poster

This paper presents a learning-based method for multi-view depth estimation from posed images. Our core idea is a "learning-to-optimize" paradigm that iteratively indexes a plane-sweeping cost volume and regresses the depth map via a convolutional Gated Recurrent Unit (GRU). Since the cost volume pl…

2023

ReAugKD: Retrieval-Augmented Knowledge Distillation For Pre-trained Language Models

ACL 2023short

Knowledge Distillation (KD) is one of the most effective approaches to deploying large-scale pre-trained language models in low-latency environments by transferring the knowledge contained in the large-scale models to smaller student models. Prior KD approaches use the soft labels and intermediate a…

Cited by 23SourcePDFScholar
2023

Supported Trust Region Optimization for Offline Reinforcement Learning

ICML 2023poster

Offline reinforcement learning suffers from the out-of-distribution issue and extrapolation error. Most policy constraint methods regularize the density of the trained policy towards the behavior policy, which is too restrictive in most cases. We propose Supported Trust Region optimization (STR) whi…

Cited by 16SourcePDFScholar
2023

Supported Value Regularization for Offline Reinforcement Learning

NeurIPS 2023poster

Offline reinforcement learning suffers from the extrapolation error and value overestimation caused by out-of-distribution (OOD) actions. To mitigate this issue, value regularization approaches aim to penalize the learned value functions to assign lower values to OOD actions. However, existing value…

2023

TWINS: A Fine-Tuning Framework for Improved Transferability of Adversarial Robustness and Generalization

CVPR 2023poster

Recent years have seen the ever-increasing importance of pre-trained models and their downstream training in deep learning research and applications. At the same time, the defense for adversarial examples has been mainly investigated in the context of training from random initialization on simple cl…

2023

Temporal Knowledge Graph Reasoning with Historical Contrastive Learning

AAAI 2023technical

Temporal knowledge graph, serving as an effective way to store and model dynamic relations, shows promising prospects in event forecasting. However, most temporal knowledge graph reasoning methods are highly dependent on the recurrence or periodicity of events, which brings challenges to inferring f…

2023

Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting

ICCV 2023poster

Hand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space which is inadequate for 3D real-world applications. In this paper, we set up an…

Cited by 19PDFcodeScholar
2023

Uncovering the Missing Pattern: Unified Framework Towards Trajectory Imputation and Prediction

CVPR 2023poster

Trajectory prediction is a crucial undertaking in understanding entity movement or human behavior from observed sequences. However, current methods often assume that the observed sequences are complete while ignoring the potential for missing values caused by object occlusion, scope limitation, sens…

2023

Understanding and Constructing Latent Modality Structures in Multi-Modal Representation Learning

CVPR 2023poster

Contrastive loss has been increasingly used in learning representations from multiple modalities. In the limit, the nature of the contrastive loss encourages modalities to exactly match each other in the latent space. Yet it remains an open question how the modality alignment affects the downstream…

Cited by 54SourcePDFScholar
2023

Unsupervised Graph-Text Mutual Conversion with a Unified Pretrained Language Model

ACL 2023long

Graph-to-text (G2T) generation and text-to-graph (T2G) triple extraction are two essential tasks for knowledge graphs. Existing unsupervised approaches become suitable candidates for jointly learning the two tasks due to their avoidance of using graph-text parallel data. However, they adopt multiple…

Cited by 3SourcePDFScholar
2022

AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-Time Image Enhancement

CVPR 2022poster

The 3D Lookup Table (3D LUT) is a highly-efficient tool for real-time image enhancement tasks, which models a non-linear 3D color transform by sparsely sampling it into a discretized 3D lattice. Previous works have made efforts to learn image-adaptive output color values of LUTs for flexible enhance…

Cited by 82PDFcodeScholar
2022

Asynchronous Convergence in Multi-Task Learning via Knowledge Distillation from Converged Tasks

NAACL 2022industry

Multi-task learning (MTL) aims to solve multiple tasks jointly by sharing a base representation among them. This can lead to more efficient learning and better generalization, as compared to learning each task individually. However, one issue that often arises in MTL is the convergence speed between…

Cited by 4SourcePDFScholar
2022

Deformable VisTR: Spatio Temporal Deformable Attention for Video Instance Segmentation

ICASSP 2022accepted

Video instance segmentation (VIS) task requires classifying, segmenting, and tracking object instances over all frames in a video clip. Recently, VisTR [1] has been proposed as end-to-end transformer-based VIS framework, while demonstrating state-of-the-art performance. However, VisTR is slow to con…

Cited by 0SourceScholar
2022

DynaMaR: Dynamic Prompt with Mask Token Representation

EMNLP 2022industry

Recent research has shown that large language models pretrained using unsupervised approaches can achieve significant performance improvement on many downstream tasks. Typically when adapting these language models to downstream tasks, like a classification or regression task, we employ a fine-tuning…

Cited by 1SourcePDFScholar
2022

EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text Classification

NAACL 2022long

Recent works have empirically shown the effectiveness of data augmentation (DA) in NLP tasks, especially for those suffering from data scarcity. Intuitively, given the size of generated data, their diversity and quality are crucial to the performance of targeted tasks. However, to the best of our kn…

2022

Effective Model Sparsification by Scheduled Grow-and-Prune Methods

ICLR 2022poster

Deep neural networks (DNNs) are effective in solving many real-world problems. Larger DNN models usually exhibit better quality (e.g., accuracy) but their excessive computation results in long inference time. Model sparsification can reduce the computation and memory cost while maintaining model qua…

2022

GeoRefine: Self-Supervised Online Depth Refinement for Accurate Dense Mapping

ECCV 2022poster

"We present a robust and accurate depth refinement system, named GeoRefine, for geometrically-consistent dense mapping from monocular sequences. GeoRefine consists of three modules: a hybrid SLAM module using learning-based priors, an online depth refinement module leveraging self-supervision, and a…

Cited by 11SourcePDFScholar
2022

HCSC: Hierarchical Contrastive Selective Coding

CVPR 2022poster

Hierarchical semantic structures naturally exist in an image dataset, in which several semantically relevant image clusters can be further integrated into a larger cluster with coarser-grained semantics. Capturing such structures with image representations can greatly benefit the semantic understand…

Cited by 101PDFcodeScholar
2022

Improved Fine-Tuning by Better Leveraging Pre-Training Data

NeurIPS 2022accept

As a dominant paradigm, fine-tuning a pre-trained model on the target data is widely used in many deep learning applications, especially for small data sets. However, recent studies have empirically shown that training from scratch has the final performance that is no worse than this pre-training st…

2022

Interventional Multi-Instance Learning with Deconfounded Instance-Level Prediction

AAAI 2022technical

When applying multi-instance learning (MIL) to make predictions for bags of instances, the prediction accuracy of an instance often depends on not only the instance itself but also its context in the corresponding bag. From the viewpoint of causal inference, such bag contextual prior works as a conf…

Cited by 27SourcePDFScholar
2022

Learning From Untrimmed Videos: Self-Supervised Video Representation Learning With Hierarchical Consistency

CVPR 2022poster

Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos, leading to limited diversity in visual patterns and limited performance gain. In this work, we aim to learn representat…

Cited by 20PDFScholar
2022

MemREIN: Rein the Domain Shift for Cross-Domain Few-Shot Learning

IJCAI 2022poster

Few-shot learning aims to enable models generalize to new categories (query instances) with only limited labeled samples (support instances) from each category. Metric-based mechanism is a promising direction which compares feature embeddings via different metrics. However, it always fail to general…

Cited by 12SourcePDFScholar
2022

PlaneMVS: 3D Plane Reconstruction From Multi-View Stereo

CVPR 2022poster

We present a novel framework named PlaneMVS for 3D plane reconstruction from multiple input views with known camera poses. Most previous learning-based plane reconstruction methods reconstruct 3D planes from single images, which highly rely on single-view regression and suffer from depth scale ambig…

Cited by 49PDFcodeScholar
2022

Posterior Refinement on Metric Matrix Improves Generalization Bound in Metric Learning

ECCV 2022poster

"Deep metric learning (DML) attempts to learn a representation model as well as a metric function with a limited generalization gap, so that the model trained on finite known data can achieve similitude performance on infinite unseen data. While considerable efforts have been made to bound the gener…

Cited by 0SourcePDFScholar
2022

SepLUT: Separable Image-Adaptive Lookup Tables for Real-Time Image Enhancement

ECCV 2022poster

"Image-adaptive lookup tables (LUTs) have achieved great success in real-time image enhancement tasks due to their high efficiency for modeling color transforms. However, they embed the complete transform, including the color component-independent and the component-correlated parts, into only a sing…

2022

Vision-Language Pre-Training With Triple Contrastive Learning

CVPR 2022poster

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information (MI) between an image and its matched text. However, simply…

Cited by 351PDFcodeScholar
2022

Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias Perspective

NeurIPS 2022accept

Contrastive learning (CL) has been the de facto technique for self-supervised representation learning (SSL), with impressive empirical success such as multi-modal representation learning. However, traditional CL loss only considers negative samples from a minibatch, which could cause biased gradient…

Cited by 40SourcePDFScholar
2021

An Online Method for A Class of Distributionally Robust Optimization with Non-convex Objectives

NeurIPS 2021poster

In this paper, we propose a practical online method for solving a class of distributional robust optimization (DRO) with non-convex objectives, which has important applications in machine learning for improving the robustness of neural networks. In the literature, most methods for solving DRO are ba…

2021

DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled Samples

NeurIPS 2021poster

The scarcity of labeled data is a critical obstacle to deep learning. Semi-supervised learning (SSL) provides a promising way to leverage unlabeled data by pseudo labels. However, when the size of labeled data is very small (say a few labeled samples per class), SSL performs poorly and unstably, pos…

Cited by 36SourcePDFScholar
2021

Dash: Semi-Supervised Learning with Dynamic Thresholding

ICML 2021oral

While semi-supervised learning (SSL) has received tremendous attentions in many machine learning tasks due to its successful use of unlabeled data, existing SSL algorithms use either all unlabeled examples or the unlabeled examples with a fixed high-confidence prediction during the training progress…

2021

Federated Deep AUC Maximization for Hetergeneous Data with a Constant Communication Complexity

ICML 2021spotlight

Deep AUC (area under the ROC curve) Maximization (DAM) has attracted much attention recently due to its great potential for imbalanced data classification. However, the research on Federated Deep AUC Maximization (FDAM) is still limited. Compared with standard federated learning (FL) approaches that…

2021

GIF Thumbnails: Attract More Clicks to Your Videos

AAAI 2021technical

With the rapid increase of mobile devices and online media, more and more people prefer posting/viewing videos online. Generally, these videos are presented on video streaming sites with image thumbnails and text titles. While facing huge amounts of videos, a viewer clicks through a certain video wi…

2021

MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor Environments

ICCV 2021poster

Self-supervised depth estimation for indoor environments is more challenging than its outdoor counterpart in at least the following two aspects: (i) the depth range of indoor sequences varies a lot across different frames, making it difficult for the depth network to induce consistent depth cues, wh…

Cited by 96PDFScholar
2021

Tra2Tra: Trajectory-to-Trajectory Prediction With a Global Social Spatial-Temporal Attentive Neural Network

RA-L 2021

Accurate trajectory prediction plays a key role in robot navigation. It is beneficial for planning a collision-free and appropriate path for the autonomous robots, especially in crowded scenes. However, it is a particularly challenging task because there are complex and subtle interactions among ped

Cited by 42SourceScholar
2021

Weakly-supervised Text Classification Based on Keyword Graph

EMNLP 2021main

Weakly-supervised text classification has received much attention in recent years for it can alleviate the heavy burden of annotating massive data. Among them, keyword-driven methods are the mainstream where user-provided keywords are exploited to generate pseudo-labels for unlabeled texts. However,…

2020

Adaptive Fractional Dilated Convolution Network for Image Aesthetics Assessment

CVPR 2020poster

To leverage deep learning for image aesthetics assessment, one critical but unsolved issue is how to seamlessly incorporate the information of image aspect ratios to learn more robust models. In this paper, an adaptive fractional dilated convolution (AFDC), which is aspect-ratio-embedded, compositio…

Cited by 113PDFScholar
2020

Optimal Epoch Stochastic Gradient Descent Ascent Methods for Min-Max Optimization

NeurIPS 2020poster

Epoch gradient descent method (a.k.a. Epoch-GD) proposed by (Hazan and Kale, 2011) was deemeda breakthrough for stochastic strongly convex minimization, which achieves theoptimal convergence rate of O(1/T) with T iterative updates for the objective gap. However, its extension to solving stochastic m…

Cited by 72SourcePDFScholar
2020

Stochastic Optimization for Non-convex Inf-Projection Problems

ICML 2020poster

In this paper, we study a family of non-convex and possibly non-smooth inf-projection minimization problems, where the target objective function is equal to minimization of a joint function over another variable. This problem include difference of convex (DC) functions and a family of bi-convex func…

Cited by 6SourcePDFScholar
2020

Talking-head Generation with Rhythmic Head Motion

ECCV 2020poster

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys linguistic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing works either generate still talking-face videos or rely on l…

2019

Katalyst: Boosting Convex Katayusha for Non-Convex Problems with a Large Condition Number

ICML 2019oral

An important class of non-convex objectives that has wide applications in machine learning consists of a sum of $n$ smooth functions and a non-smooth convex function. Tremendous studies have been devoted to conquering these problems by leveraging one of the two types of variance reduction techniques…

Cited by 4SourcePDFScholar
2019

Non-asymptotic Analysis of Stochastic Methods for Non-Smooth Non-Convex Regularized Problems

NeurIPS 2019poster

Stochastic Proximal Gradient (SPG) methods have been widely used for solving optimization problems with a simple (possibly non-smooth) regularizer in machine learning and statistics. However, to the best of our knowledge no non-asymptotic convergence analysis of SPG exists for non-convex optimizati…

Cited by 28SourcePDFScholar
2019

Stochastic Optimization for DC Functions and Non-smooth Non-convex Regularizers with Non-asymptotic Convergence

ICML 2019oral

Difference of convex (DC) functions cover a broad family of non-convex and possibly non-smooth and non-differentiable functions, and have wide applications in machine learning and statistics. Although deterministic algorithms for DC functions have been extensively studied, stochastic optimization th…

Cited by 50SourcePDFScholar
2018

Crowd Counting via Adversarial Cross-Scale Consistency Pursuit

CVPR 2018poster

Crowd counting or density estimation is a challenging task in computer vision due to large scale variations, perspective distortions and serious occlusions, etc. Existing methods generally suffers from two issues: 1) the model averaging effects in multi-scale CNNs induced by the widely adopted L2 re…

Cited by 414SourcePDFScholar
2018

First-order Stochastic Algorithms for Escaping From Saddle Points in Almost Linear Time

NeurIPS 2018poster

(This is a theory paper) In this paper, we consider first-order methods for solving stochastic non-convex optimization problems. The key building block of the proposed algorithms is first-order procedures to extract negative curvature from the Hessian matrix through a principled sequence starting fr…

Cited by 145SourcePDFScholar
2018

Geometric Constrained Joint Lane Segmentation and Lane Boundary Detection

ECCV 2018poster

Lane detection is playing an indispensable role in advanced driver assistance systems. The existing approaches for lane detection can be categorized as lane area segmentation and lane boundary detection. Most of these methods abandon a great quantity of complementary information, such as geometric p…

2017

ADMM without a Fixed Penalty Parameter: Faster Convergence with New Adaptive Penalization

NeurIPS 2017poster

Alternating direction method of multipliers (ADMM) has received tremendous interest for solving numerous problems in machine learning, statistics and signal processing. However, it is known that the performance of ADMM and many of its variants is very sensitive to the penalty parameter of a quadrat…

Cited by 68SourcePDFScholar
2017

Adaptive SVRG Methods under Error Bound Conditions with Unknown Growth Parameter

NeurIPS 2017poster

Error bound, an inherent property of an optimization problem, has recently revived in the development of algorithms with improved global convergence without strong convexity. The most studied error bound is the quadratic error bound, which generalizes strong convexity and is satisfied by a large fa…

Cited by 24SourcePDFScholar
2017

Grasp quality evaluation and planning for objects with negative curvature

ICRA 2017poster

We consider the problem of grasping concave objects, i.e., objects whose surface includes regions with negative curvature. When a multifingered hand is used to restrain these objects, these areas can be advantageously used to determine grasps capable of more robustly resisting to external disturbanc…

Cited by 3SourceScholar
2017

Stochastic Convex Optimization: Faster Local Growth Implies Faster Global Convergence

ICML 2017poster

In this paper, a new theory is developed for first-order stochastic convex optimization, showing that the global convergence rate is sufficiently quantified by a local growth rate of the objective function in a neighborhood of the optimal solutions. In particular, if the objective function $F(\mathb…

Cited by 52SourcePDFScholar
2016

Homotopy Smoothing for Non-Smooth Problems with Lower Complexity than $O(1/\epsilon)$

NeurIPS 2016poster

In this paper, we develop a novel {\bf ho}moto{\bf p}y {\bf s}moothing (HOPS) algorithm for solving a family of non-smooth problems that is composed of a non-smooth term with an explicit max-structure and a smooth term or a simple non-smooth term whose proximal mapping is easy to compute. The bes…

Cited by 27SourcePDFScholar