← Search

HUI LI

98 accepted papers

2026

A Data-Observation Hybrid Compensation Method for Precise Force Control of Cable-Driven Wrist Exoskeletons in Teleoperation

RA-L 2026

High-precision force control of wearable exoskeletons enables highly transparent force interaction operations, effectively improving the feasibility of teleoperation tasks. We adopted a cable-driven spherical parallel wrist exoskeleton (SPWE) which enables it to reduce the volume and enhance operati

Cited by 0SourceScholar
2026

ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions

CVPR 2026

Existing hand-object interactions (HOI) methods are largely limited to rigid objects, while 4D reconstruction methods of articulated objects generally require pre-scanning the object or even multi-view videos. It remains an unexplored but significant challenge to reconstruct 4D human-articulated-obj

Cited by 0SourcecodeScholar
2026

Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusion

CVPR 2026

Infrared and visible image fusion (IVIF) aims to synthesise complementary information from the two source modalities while preserving natural textures and salient thermal signatures simultaneously. Existing solutions predominantly rely on extensive sets of rigidly aligned image pairs for training. H

Cited by 0SourcecodeScholar
2026

Bi-Bridge: Bidirectional Diffusion Bridges for Low-Light Image Enhancement

CVPR 2026

Low-Light Image Enhancement (LLIE) is a challenging task, as severe information loss means a single input can correspond to multiple plausible restorations. This inherent ambiguity causes conventional regression-based models to produce overly-smooth results that lack detail. While recent generative

Cited by 0SourceScholar
2026

Bias-Spectrum Neural Processes for Parametric PDEs: Architecture Priors Meet PDE Constraints

ICML 2026poster

Parametric partial differential equations (PDEs) serve as fundamental models across science and engineering, yet constructing fast and accurate surrogate models from sparse, irregularly sampled observations with reliable uncertainty quantification remains challenging. Existing approaches struggle to…

Cited by 0SourceScholar
2026

CRAFT: Long-Horizon Cable Routing Algorithm and Low-Friction Caging Gripper

ICRA 2026poster

Cable routing is a common manipulation task in assembly and manufacturing, yet it remains challenging due to the deformable nature of cables and the constraints of cluttered routing environments. In this paper, we present CRAFT: Cable Routing Around Fixtures using Two grippers, a novel hardware plus…

Cited by 0Scholar
2026

ChartE$^{3}$: A Comprehensive Benchmark for End-to-End Chart Editing

ICML 2026poster

Charts are a fundamental visualization format for structured data analysis. Enabling end-to-end chart editing according to user intent is of great practical value, yet remains challenging due to the need for both fine-grained control and global structural consistency. Most existing approaches adopt …

Cited by 0SourceScholar
2026

FedCARE: Federated Unlearning with Conflict-Aware Projection and Relearning-Resistant Recovery

IJCAI 2026

Federated learning (FL) enables collaborative model training without centralizing raw data, but privacy regulations such as the right to be forgotten require FL systems to remove the influence of previously used training data upon request. Retraining a federated model from scratch is prohibitively e

Cited by 0Scholar
2026

FedCDWA: Decoupled Federated Prototype Distillation with Hierarchical Wasserstein Aggregation

ICML 2026poster

Federated learning enables decentralized clients to collaboratively train models without sharing local data. However, heterogeneous client distributions often induce client drift and hinder convergence. This paper proposes FedCDWA, a decoupled hierarchical federated distillation framework. FedCDWA d…

Cited by 0SourceScholar
2026

FusionRegister: Every Infrared and Visible Image Fusion Deserves Registration

CVPR 2026

Spatial registration across different visual modalities is a critical but formidable step in multi-modality image fusion for real-world perception. Although several methods are proposed to address this issue, the existing registration-based fusion methods typically require extensive pre-registration

Cited by 0SourcecodeScholar
2026

Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AI

ICLR 2026poster

While Large Language Models (LLMs) show immense promise as planners for embodied AI, their stochastic nature and lack of formal reasoning capabilities prevent the strict safety guarantees required for physical deployment. Current approaches fall short: they either rely on other unreliable LLMs for s…

Cited by 0SourceScholar
2026

HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models

AAAI 2026technical

3D understanding has drawn significant attention recently, leveraging Vision-Language Models (VLMs) to enable multi-modal reasoning between point cloud and text data. Current 3D-VLMs directly embed the 3D point clouds into 3D tokens, following large 2D-VLMs with powerful reasoning capabilities. Howe

Cited by 0SourcePDFScholar
2026

Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation

CVPR 2026

Transformers rely on explicit positional encoding to model structure in data. WhileRotary Position Embedding (RoPE) excels in 1D domains, its application to image generation reveals significant limitations such as fine-grained spatial relationmodeling, color cues, and object counting. This paper ide

Cited by 0SourcecodeScholar
2026

Learning Task-Invariant Properties Via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots

ICRA 2026poster

Achieving quadruped robot locomotion across diverse and dynamic terrains presents significant challenges, primarily due to the discrepancies between simulation environments and real-world conditions. Traditional sim-to-real transfer methods often rely on manual feature design or costly real-world fi…

2026

Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving

AAAI 2026technical

Generative reasoning with large language models (LLMs) often involves long decoding sequences, leading to substantial memory and latency overheads from accumulating key-value (KV) caches. While existing KV compression methods primarily focus on reducing prefill memory from long input sequences, they

Cited by 0SourcePDFScholar
2026

MetaGameBO: Hierarchical Game-Theoretic Driven Robust Meta-Learning for Bayesian Optimization

AAAI 2026technical

Meta-learning for Bayesian optimization accelerates optimization by leveraging knowledge from previous tasks, but existing methods optimize for average performance and fail on challenging outlier tasks critical in practice. These limitations become particularly severe when target tasks exhibit distr

Cited by 0SourcePDFScholar
2026

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture

CVPR 2026

This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at the training timestep is the corresponding ground-truth noisy data that is an interpolation of the noise and the data, and duri

Cited by 0SourcecodeScholar
2026

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models

ICML 2026poster

Modern Vision-Language Models (VLMs) pose significant individual-level privacy risks by linking fragmented multimodal data to identifiable individuals through hierarchical chain-of-thought reasoning. However, existing privacy benchmarks remain structurally insufficient for this threat, as they prima…

Cited by 0SourceScholar
2026

PMCE: Probabilistic Multi-Granularity Semantics with Caption-Guided Enhancement for Few-Shot Learning

IJCAI 2026

Few-shot learning aims to recognize novel categories from limited labeled samples, where prototypes estimated from 1--5 supports per class are often unreliable. Semantic-based approaches alleviate this by introducing class-level priors, but they often ignore instance-level cues and rarely optimize q

Cited by 0Scholar
2026

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

ICML 2026poster

Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visual latents throughout denoising. In this setting, we observe a prompt forgetting phenomenon: the semantics of the prompt …

Cited by 0SourceScholar
2026

RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry

ICML 2026poster

Recent image generators produce photo-realistic content that undermines the reliability of downstream recognition systems. As visual appearance cues become less pronounced, appearance-driven detectors that rely on forensic cues or high-level representations lose stability. This motivates a shift fro…

Cited by 0SourceScholar
2026

RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression

ICML 2026poster

Vector quantization is a fundamental tool for compressing high-dimensional embeddings, yet existing multi-codebook methods rely on static codebooks that limit expressiveness under heterogeneous data geometry. While recent dynamic quantizers like QINCo adapt codebooks to individual inputs and improve…

Cited by 0SourceScholar
2026

RSA-CR: Resisting Shilling Attacks in Citation Recommendation via Dumbbell Inductive Learning

AAAI 2026technical

Citation recommendation aims to provide researchers with the most relevant references for their manuscripts, helping them swiftly discover pertinent studies and bolster the reliability of their arguments. However, some individuals manipulate these recommendation systems by injecting false informatio

Cited by 0SourcePDFScholar
2026

RcAE: Recursive Reconstruction Framework for Unsupervised Industrial Anomaly Detection

AAAI 2026technical

Unsupervised industrial anomaly detection requires accurately identifying defects without labeled data. Traditional autoencoder-based methods often struggle with incomplete anomaly suppression and loss of fine details, as their single-pass decoding fails to effectively handle anomalies with varying

Cited by 0SourcePDFScholar
2026

Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space

ICML 2026poster

Infrared and visible image fusion aims to integrate complementary information from both modalities. However, most existing methods rely on Euclidean representations, which inherently impose geometric constraints that hinder effective semantic modelling. Specifically, Euclidean geometry imposes rigid…

Cited by 0SourceScholar
2026

The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVA

CVPR 2026

Multimodal Large Language Models (MLLMs) like LLaVA have demonstrated remarkable capabilities in multi-modal understanding and generation. This success motivates us to investigate whether the inherent prior knowledge embedded within such MLLMs contains sufficient spatial awareness for dense predicti

Cited by 0SourcecodeScholar
2026

Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators

AAAI 2026technical

Quadrupedal robots with manipulators offer strong mobility and adaptability for grasping in unstructured, dynamic environments through coordinated whole-body control. However, existing research has predominantly focused on static-object grasping, neglecting the challenges posed by dynamic targets an

Cited by 0SourcePDFScholar
2025

A Multi-Agent Framework with Automated Decision Rule Optimization for Cross-Domain Misinformation Detection

EMNLP 2025

Misinformation spans various domains, but detection methods trained on specific domains often perform poorly when applied to others. With the rapid development of Large Language Models (LLMs), researchers have begun to utilize LLMs for cross-domain misinformation detection. However, existing LLM-bas

Cited by 0SourcePDFScholar
2025

ARCH: Hierarchical Hybrid Learning for Long-Horizon Contact-Rich Robotic Assembly

CoRL 2025poster

Generalizable long-horizon robotic assembly requires reasoning at multiple levels of abstraction. While end-to-end imitation learning (IL) is a promising approach, it typically requires large amounts of expert demonstration data and often struggles to achieve the high precision demanded by assembly…

Cited by 0SourceScholar
2025

Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific Loss

ACL 2025long

Recent studies have explored Continual Instruction Tuning (CIT) in Multimodal Large Language Models (MLLMs), with a primary focus on Task-incremental CIT, where MLLMs are required to continuously acquire new tasks. However, the more practical and challenging Domain-incremental CIT, focused on the co…

Cited by 0SourcePDFScholar
2025

CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completion

EMNLP 2025

Repository-level code completion automatically predicts the unfinished code based on the broader information from the repository. Recent strides in Code Large Language Models (code LLMs) have spurred the development of repository-level code completion methods, yielding promising results. Nevertheles

2025

DADM: Dual Alignment of Domain and Modality for Face Anti-spoofing

ICCV 2025poster

With the availability of diverse sensor modalities (i.e., RGB, Depth, Infrared) and the success of multi-modal learning, multi-modal face anti-spoofing (FAS) has emerged as a prominent research focus. The intuition behind it is that leveraging multiple modalities can uncover more intrinsic spoofing…

2025

Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution

ICCV 2025poster

Discrete Wavelet Transform (DWT) has been widely explored to enhance the performance of image super-resolution (SR). Despite some DWT-based methods improving SR by capturing fine-grained frequency signals, most existing approaches neglect the interrelations among multi-scale frequency sub-bands, res…

Cited by 0SourcePDFScholar
2025

DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors

AAAI 2025technical

Dynamic 3D interaction has been attracting a lot of attention recently. However, creating such 4D content remains challenging. One solution is to animate 3D scenes with physics-based simulation, which requires manually assigning precise physical properties to the object or the simulated results woul…

2025

Expected Hypervolume Improvement Is a Particular Hypervolume Improvement

AAAI 2025technical

Multi-objective Bayesian optimization (MOBO) aims to optimize multiple competing objective functions in the expensive-to-evaluate scenario. The Expected Hypervolume Improvement (EHVI) is a commonly used acquisition function for MOBO and shows a good performance. However, the computation of EHVI beco…

2025

FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching

ICCV 2025poster

Facial attractiveness prediction (FAP) has long been an important computer vision task, which could be widely applied in live videos with facial retouching. However, previous FAP datasets are either small or closed-source. Moreover, the corresponding FAP models exhibit limited generalization and ada…

2025

Fabrica: Dual-Arm Assembly of General Multi-Part Objects via Integrated Planning and Learning

CoRL 2025oral

Multi-part assembly poses significant challenges for robotic systems to execute long-horizon, contact-rich manipulation with generalization across complex geometries. We present a dual-arm robotic system capable of end-to-end planning and control for autonomous assembly of general multi-part objects…

Cited by 0SourceScholar
2025

Flow-based Domain Randomization for Learning and Sequencing Robotic Skills

ICML 2025poster

Domain randomization in reinforcement learning is an established technique for increasing the robustness of control policies learned in simulation. By randomizing properties of the environment during training, the learned policy can be robust to uncertainty along the randomized dimensions. While the…

2025

FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors

ICCV 2025poster

Interactive image editing allows users to modify images through visual interaction operations such as drawing, clicking, and dragging. Existing methods construct such supervision signals from videos, as they capture how objects change with various physical interactions. However, these models are usu…

2025

Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

ICLR 2025poster

Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities.First, we extend…

2025

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

CVPR 2025poster

Existing methodologies for animating portrait images face significant challenges, particularly in handling non-frontal perspectives, rendering dynamic objects around the portrait, and generating immersive, realistic backgrounds. In this paper, we introduce the first application of a pretrained trans…

2025

ICRT: In-Context Imitation Learning via Next-Token Prediction

ICRA 2025

In-context imitation learning is the capability to perform novel tasks when prompted with task demonstration examples. In-Context Robot Transformer (ICRT) is a causal transformer that performs autoregressive prediction on sensorimotor trajectories, which include images, proprioceptive states, and ac

Cited by 53SourceScholar
2025

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners

ICLR 2025poster

Contrastive language-image pretraining (CLIP) has significantly advanced image-based vision learning. A pressing topic subsequently arises: how can we effectively adapt CLIP to the video domain? Recent studies have focused on adjusting either the textual or visual branch of CLIP for action recogniti…

2025

LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph Completion

NeurIPS 2025poster

Multi-modal Knowledge Graph Completion (MMKGC) aims to predict missing entities, relations, or attributes in knowledge graphs by collaboratively modeling the triple structure and multimodal information (e.g., text, images, videos) associated with entities. This approach facilitates the automatic dis…

Cited by 0SourcecodeScholar
2025

Learning Robust Neural Processes with Risk-Averse Stochastic Optimization

ICML 2025poster

Neural processes (NPs) are a promising paradigm to enable skill transfer learning across tasks with the aid of the distribution of functions. The previous NPs employ the empirical risk minimization principle in optimization. However, the fast adaption ability to different tasks can vary widely, and…

Cited by 0SourcePDFScholar
2025

Learning Transition Patterns by Large Language Models for Sequential Recommendation

COLING 2025main

Large Language Models (LLMs) have demonstrated powerful performance in sequential recommendation due to their robust language modeling and comprehension capabilities. In such paradigms, the item texts of interaction sequences are formulated as sentences and LLMs are utilized to learn language repres…

Cited by 0SourcePDFScholar
2025

Learning to Generalize: An Information Perspective on Neural Processes

NeurIPS 2025poster

Neural Processes (NPs) combine the adaptability of neural networks with the efficiency of meta-learning, offering a powerful framework for modeling stochastic processes. However, existing methods focus on empirical performance while lacking a rigorous theoretical understanding of generalization. To…

Cited by 0SourceScholar
2025

Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System

ACL 2025long

The rapid advancement of scientific progress requires innovative tools that can accelerate knowledge discovery. Although recent AI methods, particularly large language models (LLMs), have shown promise in tasks such as hypothesis generation and experimental design, they fall short of replicating the…

2025

MetaCon: Revitalizing Internet Congestion Control with Meta-Reinforcement Learning

ICASSP 2025accepted

Effective congestion control algorithms (CCAs) are crucial for the smooth operation of Internet communication infrastructure. CCAs adjust transmission rates based on congestion signals, optimizing resource utilization and user experience. However, existing studies, both rule-based and learning-based…

Cited by 0SourceScholar
2025

Monte Carlo Tree Search Based Prompt Autogeneration for Jailbreak Attacks against LLMs

COLING 2025main

Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails. With the recent blossom of large language models (LLMs), there’s a growing focus on jailbreak att…

2025

Multimodal Quantitative Language for Generative Recommendation

ICLR 2025poster

Generative recommendation has emerged as a promising paradigm aiming at directly generating the identifiers of the target candidates. Most existing methods attempt to leverage prior knowledge embedded in Pre-trained Language Models (PLMs) to improve the recommendation performance. However, they ofte…

Cited by 0SourcePDFScholar
2025

Non-Autoregressive Image Captioning with Multi-Label Classification and Self-Critical Sequence Training

ICASSP 2025accepted

Most current image captioning models rely on the autoregressive approach, which unfortunately results in significant inference delays that hinder their practical use. In contrast, non-autoregressive methods show promising potential for increasing inference speeds. However, there is often a performan…

Cited by 0SourceScholar
2025

One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion

CVPR 2025poster

Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction thro…

2025

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

CVPR 2025highlight

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality, human-centric video datasets presents a challenge to progress in…

Cited by 2SourcePDFScholar
2025

Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws

NeurIPS 2025spotlight

Existing infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain limited. Moreover, the lack of interpretability in modal informat…

Cited by 0SourceScholar
2025

SiQA: A Large Multi-Modal Question Answering Model for Structured Images Based on RAG

ICASSP 2025accepted

Existing Large Multimodal Models (LMMs) demonstrate excellent performance in handling visual tasks in everyday scenarios. However, they still face challenges in understanding structured images, such as flowcharts and organizational charts, which are characterized by text-rich and complex hierarchica…

Cited by 0SourceScholar
2025

Towards Neurorobotic Interface for Finger Joint Angle Estimation: A Multi-Stage CNN-LSTM Network with Transfer Learning

ICRA 2025

To maximize the autonomy of individuals with upper limb amputations in daily activities, leveraging forearm muscle information to infer movement intent is a promising research direction. While current prosthetic hand technologies can utilize forearm muscle data to achieve basic movements such as gra

Cited by 3SourceScholar
2024

A Multi-Scale Bimodal Fusion Network for Robust and Accurate Online Handwriting Recognition

ICASSP 2024accepted

Online handwriting recognition based on sensor trajectory information faces several unresolved challenges: 1) sensor signals lack sufficient global spatial context; 2) different recognition tasks have inconsistent requirements for feature receptive fields. This is due to the inconsistent scales of t…

Cited by 0SourceScholar
2024

ASAP: Automated Sequence Planning for Complex Robotic Assembly with Physical Feasibility

ICRA 2024poster

The automated assembly of complex products requires a system that can automatically plan a physically feasible sequence of actions for assembling many parts together. In this paper, we present ASAP, a physics-based planning approach for automatically generating such a sequence for general-shaped ass…

Cited by 23SourceScholar
2024

Bridging the Sim-to-Real Gap with Dynamic Compliance Tuning for Industrial Insertion

ICRA 2024poster

Contact-rich manipulation tasks often exhibit a large sim-to-real gap. For instance, industrial assembly tasks frequently involve tight insertions where the clearance is less than 0.1 mm and can even be negative when dealing with a deformable receptacle. This narrow clearance leads to complex contac…

Cited by 10SourcecodeScholar
2024

Code Membership Inference for Detecting Unauthorized Data Use in Code Pre-trained Language Models

EMNLP 2024finding

Code pre-trained language models (CPLMs) have received great attention since they can benefit various tasks that facilitate software development and maintenance. However, CPLMs are trained on massive open-source code, raising concerns about potential data infringement. This paper launches the study…

2024

MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization

COLING 2024main

Given the long textual product information and the product image, Multi-modal Product Summarization (MPS) aims to increase customers’ desire to purchase by highlighting product characteristics with a short textual summary. Existing MPS methods can produce promising results. Nevertheless, they still…

2024

MSG-BART: Multi-Granularity Scene Graph-Enhanced Encoder-Decoder Language Model for Video-Grounded Dialogue Generation

ICASSP 2024accepted

Generating dialogue grounded in videos requires a high level of understanding and reasoning about the visual scenes in the videos. However, existing large visual-language models are not effective due to their latent features and decoder-only structure, especially with respect to spatio-temporal rela…

Cited by 0SourceScholar
2024

Nukplex: An Efficient Local Search Algorithm for Maximum K-Plex Problem

IJCAI 2024poster

The maximum k-plex problem (MKPP) is an significant relaxation version of the maximum clique problem with extensive applications. Recently, lots of researchers have proposed many heuristic algorithms based on various methods to solve the MKPP. In this work, to further improve the performance of solv…

2024

PaReNeRF: Toward Fast Large-scale Dynamic NeRF with Patch-based Reference

CVPR 2024poster

With photo-realistic image generation Neural Radiance Field (NeRF) is widely used for large-scale dynamic scene reconstruction as autonomous driving simulator. However large-scale scene reconstruction still suffers from extremely long training time and rendering time. Low-resolution (LR) rendering c…

Cited by 1SourcePDFScholar
2024

Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory Matching

ICLR 2024poster

The ultimate goal of Dataset Distillation is to synthesize a small synthetic dataset such that a model trained on this synthetic set will perform equally well as a model trained on the full, real dataset. Until now, no method of Dataset Distillation has reached this completely lossless goal, in part…

2024

VF-Detector: Making Multi-Granularity Code Changes on Vulnerability Fix Detector Robust to Mislabeled Changes

IJCAI 2024poster

As software development projects increasingly rely on open-source software, users face the risk of security vulnerabilities from third-party libraries. To address label and character noise in code changes, we present VF-Detector to automatically identifying bug-fix commits in actual noise developmen…

2023

CORE: Co-planarity Regularized Monocular Geometry Estimation with Weak Supervision

ICCV 2023poster

The ill-posed nature of monocular 3D geometry (depth map and surface normals) estimation makes it rely mostly on data-driven approaches such as Deep Neural Networks (DNN). However, data acquisition of surface normals, especially the reliable normals, is acknowledged difficult. Commonly, reconstructi…

Cited by 0PDFScholar
2023

Contrastive Learning at the Relation and Event Level for Rumor Detection

ICASSP 2023accepted

Existing studies for rumor detection rely heavily on a large number of labeled data to operate in a fully-supervised manner. However, manual data annotation in realistic cases is very expensive and time-consuming. In this paper, we propose a novel self-supervised Relation-Event based Contrastive Lea…

Cited by 0SourceScholar
2023

Dimensional Optimization and Anti-Disturbance Analysis of an Upgraded Feed Mechanism in FAST

ICRA 2023poster

Five-hundred-meter aperture spherical radio telescope (FAST) is a very famous large-scale scientific facility with excellent performance for astronomical observation in the world, but it currently fails to observe the center of the Milky Way Galaxy due to the limited observation angle that is affect…

Cited by 2SourceScholar
2023

Injecting Multimodal Information into Rigid Protein Docking via Bi-level Optimization

NeurIPS 2023poster

The structure of protein-protein complexes is critical for understanding binding dynamics, biological mechanisms, and intervention strategies. Rigid protein docking, a fundamental problem in this field, aims to predict the 3D structure of complexes from their unbound states without conformational ch…

Cited by 6SourcePDFScholar
2023

PDF: Point Diffusion Implicit Function for Large-scale Scene Neural Representation

NeurIPS 2023poster

Recent advances in implicit neural representations have achieved impressive results by sampling and fusing individual points along sampling rays in the sampling space. However, due to the explosively growing sampling space, finely representing and synthesizing detailed textures remains a challenge f…

Cited by 5SourcePDFScholar
2023

Practical Cross-System Shilling Attacks with Limited Access to Data

AAAI 2023technical

In shilling attacks, an adversarial party injects a few fake user profiles into a Recommender System (RS) so that the target item can be promoted or demoted. Although much effort has been devoted to developing shilling attack methods, we find that existing approaches are still far from practical. In…

2023

RECESS Vaccine for Federated Learning: Proactive Defense Against Model Poisoning Attacks

NeurIPS 2023poster

Model poisoning attacks greatly jeopardize the application of federated learning (FL). The effectiveness of existing defenses is susceptible to the latest model poisoning attacks, leading to a decrease in prediction accuracy. Besides, these defenses are intractable to distinguish benign outliers fro…

Cited by 13SourcePDFScholar
2023

Rethinking Feature-Based Knowledge Distillation for Face Recognition

CVPR 2023poster

With the continual expansion of face datasets, feature-based distillation prevails for large-scale face recognition. In this work, we attempt to remove identity supervision in student training, to spare the GPU memory from saving massive class centers. However, this naive removal leads to inferior d…

Cited by 38SourcePDFScholar
2023

SMUG: Towards Robust Mri Reconstruction by Smoothed Unrolling

ICASSP 2023accepted

Although deep learning (DL) has gained much popularity for accelerated magnetic resonance imaging (MRI), recent studies have shown that DL-based MRI reconstruction models could be over-sensitive to tiny input perturbations (that are called ‘adversarial perturbations’), which cause unstable, low-qual…

Cited by 0SourceScholar
2023

Safe Self-Supervised Learning in Real of Visuo-Tactile Feedback Policies for Industrial Insertion

ICRA 2023poster

Industrial insertion tasks are often performed repetitively with parts that are subject to tight tolerances and prone to breakage. Learning an industrial insertion policy in real is challenging as the collision between the parts and the environment can cause slippage or breakage of the part. In this…

Cited by 22SourceScholar
2022

ARCANE: An Efficient Architecture for Exact Machine Unlearning

IJCAI 2022poster

Recently users’ right-to-be-forgotten is stipulated by many laws and regulations. However, only removing the data from the dataset is not enough, as machine learning models would memorize the training data once the data is involved in model training, increasing the risk of exposing users’ privacy. T…

Cited by 117SourcePDFScholar
2022

BIT-DMR: A Humanoid Dual-Arm Mobile Robot for Complex Rescue Operations

RA-L 2022

Using robots to assist or even replace rescuers for searching and rescuing has always been a research hotspot. Robots that can carry out dexterous operations at the scene are of great significance to reduce the life threat of rescue workers. Aiming at the characteristics of narrow space and frequent

Cited by 43SourceScholar
2022

Pushing the Performance Limit of Scene Text Recognizer Without Human Annotation

CVPR 2022poster

Scene text recognition (STR) attracts much attention over the years because of its wide application. Most methods train STR model in a fully supervised manner which requires large amounts of labeled data. Although synthetic data contributes a lot to STR, it suffers from the real-to-synthetic domain…

Cited by 24PDFScholar
2022

Towards Accurate Facial Landmark Detection via Cascaded Transformers

CVPR 2022poster

Accurate facial landmarks are essential prerequisites for many tasks related to human faces. In this paper, an accurate facial landmark detector is proposed based on cascaded transformers. We formulate facial landmark detection as a coordinate regression task such that the model can be trained end-t…

Cited by 49PDFScholar
2021

MatchVIE: Exploiting Match Relevancy between Entities for Visual Information Extraction

IJCAI 2021poster

Visual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or classification problem, which requires models to carefully identify eac…

Cited by 33SourcePDFScholar
2019

Dynamic Experience Replay

CoRL 2019

We present a novel technique called Dynamic Experience Replay (DER) that allows Reinforcement Learning (RL) algorithms to use experience replay samples not only from human demonstrations but also successful transitions generated by RL agents during training and therefore improve training efficiency.

Cited by 0SourcePDFScholar
2019

Generative Adversarial User Model for Reinforcement Learning Based Recommendation System

ICML 2019oral

There are great interests as well as many challenges in applying reinforcement learning (RL) to recommendation systems. In this setting, an online user is the environment; neither the reward function nor the environment dynamics are clearly defined, making the application of RL challenging. In this…