← Search

di wu

125 accepted papers

2026

A Survey on Soft Robot Adaptability: Implementations, Applications, and Prospects

ICRA 2026poster

Soft robots,compared to rigid robots,possess inherent advantages,including higher degrees of freedom, compliance,and enhanced safety,which have contributed to their increasing application across various fields. Among these benefits, adaptability is particularly noteworthy. In this paper, adaptabilit…

2026

Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLMs) extend foundation models to real-world applications by integrating inputs such as text and vision. However, their broad knowledge capacity raises growing concerns about privacy leakage, toxicity mitigation, and intellectual property violations. Machine Unlear

Cited by 0SourcePDFScholar
2026

DREAM: Document Recognition with Explicit Adaptive Memory

CVPR 2026

Large multimodal models (LMMs) have shown promising performance for various document recognition tasks. However, LMMs adopt implicit modeling, and the parameters lack interpretability. Inspired by recent advances in human memory and learning research, we propose an explicit multiscale prototype memo

Cited by 0SourcecodeScholar
2026

Density-Aware Point Cloud Upsampling via Relational Graph Flow Matching

RA-L 2026

Real-world point clouds exhibit non-uniform density distributions, varying across distance and scale. Conventional upsampling methods typically treat points homogeneously, which over-smooths sparse regions while over-processing dense regions. We propose PURF, a density-aware point cloud upsampling f

Cited by 0SourceScholar
2026

EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations

ICRA 2026poster

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate head and hand movements, continuously reposition their viewpoin…

2026

Escaping Policy Contraction: Contraction-Aware PPO (CaPPO) for Stable Language Model Fine-Tuning

ICLR 2026poster

Reinforcement learning from human feedback (RLHF) with proximal policy optimization (PPO) is widely used but often yields less diverse outputs than supervised fine-tuning, suggesting an effect in which the policy’s support contracts during on-policy optimization. We formalize this “policy contractio…

Cited by 0SourceScholar
2026

FUSE: Full‑spectrum Unlearnable Examples via Spectral Equalization

ICML 2026poster

Unlearnable examples (UEs) protect training data by injecting imperceptible perturbations so that models fail to extract exploitable representations. In this paper, we reveal that existing UEs exhibit a critical failure once low-pass filtering is applied, indicating that the effective perturbation s…

Cited by 0SourceScholar
2026

FedHPro: Federated Hyper-Prototype Learning via Gradient Matching

ICML 2026poster

Federated Learning (FL) enables collaborative training of distributed clients while protecting privacy. To enhance generalization capability in FL, prototype-based FL is in the spotlight, since shared global prototypes offer semantic anchors for aligning client-specific local prototypes. However, ex…

Cited by 0SourceScholar
2026

GIFSplat: Generative Prior-Guided Iterative Feed-Forward 3D Gaussian Splatting from Sparse Views

CVPR 2026

Feed-forward 3D reconstruction offers substantial runtime advantages over per-scene optimization, which remains slow at inference and often fragile under sparse views. However, existing feed-forward methods still have potential for further performance gains, especially for out-of-domain data, and st

Cited by 3SourcecodeScholar
2026

GUIDE: Gated Uncertainty-Informed Disentangled Experts for Long-tailed Recognition

ICLR 2026poster

Long-Tailed Recognition (LTR) remains a significant challenge in deep learning. While multi-expert architectures are a prominent paradigm, we argue that their efficacy is fundamentally limited by a series of deeply entangled problems at the levels of representation, policy, and optimization. These e…

Cited by 0SourceScholar
2026

Generalizable and Actionable Parts Pose Estimation with Symmetry Annotation-Free Learning Strategy

ICML 2026poster

Urgently needed generalizable robot object interaction and manipulation requires high-quality Cross-Category object perception. As a pioneer of this area, Generalizable and Actionable Parts (GAParts) understanding has attracted increasing attention from relevant researchers. However, most recent wor…

Cited by 0SourceScholar
2026

Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding

AAAI 2026technical

Spoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating contex

Cited by 0SourcePDFScholar
2026

Learning Dynamics as Feedback: An Adaptive Entropy Flow Dynamics Framework for Long-tailed Human Action Recognition

AAAI 2026technical

Deep human action recognition models trained on real-world data are often challenged by long-tailed distributions, where performance on rare classes is severely degraded. Current solutions typically apply static or heuristic interventions that are disconnected from the model

Cited by 0SourcePDFScholar
2026

Long-tailed Test-Time Adaptation for Vision-Language Models

ICLR 2026poster

Test-Time Adaptation (TTA) aims to further adapt models to unlabeled test sets arriving in a sequential datastream, thereby progressively strengthening the model's generalization ability. While existing TTA methods for Vision-Language Models (VLMs) are primarily designed and evaluated on (nearly) ba…

Cited by 0SourcecodeScholar
2026

Meta-FC: Meta-Learning with Feature Consistency for Robust and Generalizable Watermarking

CVPR 2026

Deep learning-based watermarking has made remarkable progress in recent years. To achieve robustness against various distortions, current methods commonly adopt a training strategy where a \underline s ingle \underline r andom \underline d istortion (SRD) is chosen as the noise layer in each trainin

Cited by 0SourcecodeScholar
2026

PAGPL: Privacy-Aware Graph Prompt Learning Scheme via Adaptive Perturbation-Estimated Topology Recovery

AAAI 2026technical

Graph prompt learning (GPL) serves as a crucial framework for mitigating the knowledge transfer by reconciling the substantial mismatch between pre-training models and downstream tasks. However, prevalent GPL paradigm fail to accommodate graph data affected by privacy-induced noise. Specifically, 1)

Cited by 0SourcePDFScholar
2026

QuantWear: Quantum-scale Wear Particle Detection for Jet Engine Diagnosis

ICML 2026poster

The quantity and 3-D shape of wear particles are essential indicators for assessing the health of jet engines, enabling early detection of potential damage and preventing accidents caused by catastrophic failures. However, capturing wear particles is difficult due to their minute sizes and ultra hig…

Cited by 0SourceScholar
2026

REArtGS++: Generalizable Articulation Reconstruction with Temporal Geometry Constraint via Planar Gaussian Splatting

CVPR 2026

Articulated objects are pervasive in daily environments, such as drawers and refrigerators. Towards their part-level surface reconstruction and joint parameter estimation, REArtGS [??] introduces a category-agnostic approach using multi-view RGB images at two different states. However, we observe th

Cited by 0SourceScholar
2026

Rethinking Crystal Symmetry Prediction: A Decoupled Perspective

AAAI 2026technical

Efficiently and accurately determining the symmetry is a crucial step in the structural analysis of crystalline materials. Existing methods usually mindlessly apply deep learning models while ignoring the underlying chemical rules. More importantly, experiments show that they face a serious sub-prop

Cited by 0SourcePDFScholar
2026

Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective

AAAI 2026technical

Point cloud completion aims to reconstruct complete 3D shapes from partial observations, which is a challenging problem due to severe occlusions and missing geometry. Despite recent advances in multimodal techniques that leverage complementary RGB images to compensate for missing geometry, most meth

Cited by 0SourcePDFScholar
2026

SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object Interaction

AAAI 2026technical

Articulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models

Cited by 0SourcePDFScholar
2026

Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery

CVPR 2026

Vision-language foundation models (VLFMs) promise zero-shot and retrieval understanding for Earth observation. While operational satellite systems often lack full multi-spectral coverage, making RGB-only inference highly desirable for scalable deployment, the adoption of VLFMs for satellite imagery

Cited by 0SourcecodeScholar
2026

When Priors Backfire: On the Vulnerability of Unlearnable Examples to Pretraining

ICLR 2026poster

Unlearnable Examples (UEs) are introduced as a data protection strategy that generates imperceptible perturbations to mislead models into learning spurious correlations rather than real semantics. In this paper, we reveal a fundamental vulnerability of UEs that emerges when learning starts from a pr…

Cited by 0SourcecodeScholar
2026

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

ICML 2026oral

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (\textit{i}) producing precise low-level actions from high-dimensional observations, (\t…

Cited by 0SourcecodeScholar
2025

A Cost-effective Solution for Remote Sensing Image Segmentation via Train/Test-Time Adaptation

ICASSP 2025accepted

Remote Sensing Image (RSI) segmentation has made significant strides, emerging as a leading solution for interpreting remote sensing data. However, due to the substantial domain gap between different remote sensors and limited computational resources, existing RSI segmentation methods often suffer f…

Cited by 0SourceScholar
2025

A Label Co-occurrence Transformation Network for Joint Empathy Detection and Empathy Intent Classification

ICASSP 2025accepted

Empathy detection (ED) aims to understand the user’s empathy direction, while empathy intent classification (EIC) focuses on identifying the empathy intent behind the user’s utterance. Both tasks have garnered significant attention. Recent studies have shown that jointly training these tasks can imp…

Cited by 0SourceScholar
2025

A Novel Sparse Active Online Learning Framework for Fast and Accurate Streaming Anomaly Detection Over Data Streams

IJCAI 2025

Online Anomaly Detection (OAD) is critical for identifying rare yet important data points in large, dynamic, and complex data streams. A key challenge lies in achieving accurate and consistent detection of anomalies while maintaining computational and memory efficiency. Conventional OAD approaches,

Cited by 0SourcePDFScholar
2025

ARNet: Self-Supervised FG-SBIR with Unified Sample Feature Alignment and Multi-Scale Token Recycling

AAAI 2025technical

Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) aims to minimize the distance between sketches and corresponding images in the embedding space. However, scalability is hindered by the growing complexity of solutions, mainly due to the abstract nature of fine-grained sketches. In this paper, we p…

2025

Accurate 3D Facial Paralysis Analysis Using Multi-View Infrared Structured Light System

ICASSP 2025accepted

Facial paralysis is a prevalent disorder affecting the facial nerve. In clinical settings, the severity of facial paralysis is typically assessed by physicians based on their subjective experience, evaluating the range of facial muscle movements and facial symmetry. To address these limitations, thi…

Cited by 0SourceScholar
2025

BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression

NAACL 2025findings

Retrieval-augmented generation (RAG) can supplement large language models (LLMs) by integrating external knowledge. However, as the number of retrieved documents increases, the input length to LLMs grows linearly, causing a dramatic increase in latency and a degradation in long-context understanding…

2025

Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?

ICLR 2025spotlight

Extremely low-resource (XLR) languages lack substantial corpora for training NLP models, motivating the use of all available resources such as dictionaries and grammar books. Machine Translation from One Book (Tanzer et al., 2024) suggests that prompting long-context LLMs with one grammar book enabl…

Cited by 5SourcePDFScholar
2025

Disentangled Representation Learning for Chinese Handwriting Recognition

ICASSP 2025accepted

Deep learning-based sequence modeling methods have improved the performance in Chinese handwriting recognition tasks. However, the implicit representations learned in current deep neural network models usually lack explainability and generalization ability for practical handwriting samples with dive…

Cited by 0SourceScholar
2025

Fully-Scalable Massively Parallel Algorithm for k-center with Outliers

AAAI 2025technical

In this paper, we consider the k-center problem with outliers (the (k, z)-center problem) in the context of Massively Parallel Computation (MPC). Existing MPC algorithms for the (k, z)-center problem typically require Ω(k) local space per machine. While this may be feasible when k is small, these al…

Cited by 0SourcePDFScholar
2025

GraphDAE-PU: Graph Denosing Auto-Encoder for Arbitrary-Scale Point Cloud Upsampling

ICASSP 2025accepted

Existing learning-based arbitrary-scale point cloud upsampling methods are usually challenged with limited point cloud feature representation and noise-sensitive refinement of coarse point cloud. In this paper, we introduce GraphDAE-PU, a novel framework for point cloud upsampling that addresses the…

Cited by 0SourceScholar
2025

How to Learn in a Noisy World? Self-Correcting the Real-World Data Noise in Machine Translation

NAACL 2025findings

The massive amounts of web-mined parallel data often contain large amounts of noise. Semantic misalignment, as the primary source of the noise, poses a challenge for training machine translation systems. In this paper, we first introduce a process for simulating misalignment controlled by semantic s…

2025

INT: Establishing Information Transfer for Multilingual Intent Detection and Slot Filling

ACL 2025finding

Multilingual spoken language understanding (SLU) involves intent detection (ID) and slot filling (SF) across multiple languages. The inherent linguistic diversity presents significant challenges in achieving performance comparable to traditional SLU. Recent studies have attempted to improve multilin…

Cited by 0SourcePDFScholar
2025

LKSNeXt: An Efficient Medical Image Segmentation Network with Large Kernels and Lightweight Structure

ICASSP 2025accepted

U-shaped architectures play a critical role in medical image segmentation. Traditional fully convolutional U-shaped networks, however, encounter numerous challenges in processing medical images, particularly in capturing long-range dependencies and global contextual information.Recently, hybrid arch…

Cited by 0SourceScholar
2025

LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

ICLR 2025poster

Recent large language model (LLM)-driven chat assistant systems have integrated memory components to track user-assistant chat histories, enabling more accurate and personalized responses. However, their long-term memory capabilities in sustained interactions remain underexplored. We introduce LongM…

2025

Many-Objective Motion Generation Method for Redundant Manipulators by Solving Pathwise Inverse Kinematics

IROS 2025

Modern robots are required to operate in complex environments and perform diverse tasks, resulting in redundant degrees of freedom (DoF) for flexibility. However, managing redundancy is challenging due to the high-dimensional and non-convex nature of robotic kinematics. When executing complex tracki

Cited by 0SourceScholar
2025

MonoSG: Monocular 3D Object Detection With Stereo Guidance

RA-L 2025

In the context of autonomous driving, monocular 3D detection is regarded as a fundamental and essential task due to its convenience, speed, and low cost. However, the lack of depth information in monocular images presents significant challenges for predicting object 3D information. Although existing

Cited by 5SourceScholar
2025

Please Translate Again: Two Simple Experiments on Whether Human-Like Reasoning Helps Translation

EMNLP 2025

Large Language Models (LLMs) demonstrate strong reasoning capabilities for many tasks, often by explicitly decomposing the task via Chain-of-Thought (CoT) reasoning. Recent work on LLM-based translation designs hand-crafted prompts to decompose translation, or trains models to incorporate intermedia

Cited by 0SourcePDFScholar
2025

REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints

NeurIPS 2025poster

Articulated objects, as prevalent entities in human life, their 3D representations play crucial roles across various applications. However, achieving both high-fidelity textured surface reconstruction and dynamic generation for articulated objects remains challenging for existing methods. In this pa…

Cited by 0SourceScholar
2025

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

RSS 2025poster

Developing robust and general-purpose manipulation policies is a key goal in robotics. To achieve effective generalization, it is essential to construct comprehensive datasets that encompass a large number of demonstration trajectories and diverse tasks. Unlike vision or language data, which can be…

Cited by 20PDFScholar
2025

Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models

ACL 2025finding

Although large language models (LLMs) achieve effective safety alignment at the time of release, they still face various safety challenges. A key issue is that fine-tuning often compromises the safety alignment of LLMs. To address this issue, we propose a method named IRR (Identify, Remove, and Reca…

2025

Towards Effective Federated Graph Foundation Model via Mitigating Knowledge Entanglement

NeurIPS 2025poster

Recent advances in graph machine learning have shifted to data-centric paradigms, driven by two emerging research fields: (1) Federated graph learning (FGL) facilitates multi-client collaboration but struggles with data and task heterogeneity, resulting in limited practicality; (2) Graph fo…

Cited by 0SourceScholar
2025

Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings

ICLR 2025poster

Recent advancements in brain-computer interfaces (BCIs) and deep learning have made decoding lexical tones from intracranial recordings possible, providing the potential to restore the communication ability of speech-impaired tonal language speakers. However, data heterogeneity induced by both physi…

Cited by 0SourcePDFScholar
2025

Ultrasound-Guided Robotic Blood Drawing and In Vivo Studies on Submillimetre Vessels of Rats

ICRA 2025

Billions of vascular access procedures are performed annually worldwide, serving as a crucial first step in various clinical diagnostic and therapeutic procedures. For pediatric or elderly individuals, whose vessels are small in size (typically 2 to 3 mm in diameter for adults and <1 mm in children)

Cited by 2SourceScholar
2025

Utterance as A Bridge: Few-shot Joint Learning of Empathy Detection and Empathy Intent Classification

ICASSP 2025accepted

Empathy detection (ED) and empathy intent classification (EIC) aim to identify the empathy direction expressed in user utterances and the underlying empathy intent behind them. Previous studies show that facilitating information transfer between tasks can enhance model performance. However, the inte…

Cited by 0SourceScholar
2024

BADFSS: Backdoor Attacks on Federated Self-Supervised Learning

IJCAI 2024poster

Self-supervised learning (SSL) is capable of learning remarkable representations from centrally available data. Recent works further implement federated learning with SSL to learn from rapidly growing decentralized unlabeled images (e.g., from cameras and phones), often resulting from privacy constr…

Cited by 3SourcePDFScholar
2024

DESectBot: Design and Validation of a Novel Two-Segment Decoupled Continuum Robotic System for Endoscopic Submucosal Dissection

IROS 2024poster

Endoscopic Submucosal Dissection (ESD) is a minimally invasive procedure designed to remove precancerous and cancerous lesions from the gastrointestinal (GI) tract. Given the GI tract’s tortuous and narrow shape, along with the need for varied movements during dissection, this requires highly flexib…

Cited by 0SourceScholar
2024

Domain-Slot Aware Contrastive Learning for Improved Dialogue State Tracking

ICASSP 2024accepted

Large-scale pre-trained neural language model has facilitated to achieve the state-of-the-art performance on Dialogue State Tracking (DST) tasks. One of the existing works models the semantic correlation between the dialogue context and (domain, slot) pair encoded by BERT and make the prediction. De…

Cited by 0SourceScholar
2024

Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language Understanding

ICASSP 2024accepted

Multi-intent spoken language understanding consists of two typical subtasks: multi-intent detection and slot filling. Existing approach suffers from two limitations: (1) It fails to explicitly model the information transfer between slots associated within the same intent clause; (2) Using a co-occur…

Cited by 0SourceScholar
2024

Energy Sharing Mechanism for Freeform Robots Utilizing Conductive Spherical Sliding Surfaces

IROS 2024poster

Energy sharing among modular robots enables sustainable operation of the system by maintaining energy balance among the modules. In this paper, we propose a novel energy sharing mechanism for FreeSN, a modular self-reconfigurable robot consisting of node and strut modules. Utilizing the feature that…

Cited by 0SourceScholar
2024

Exploring Effective Stimulus Encoding via Vision System Modeling for Visual Prostheses

ICLR 2024poster

Visual prostheses are potential devices to restore vision for blind people, which highly depends on the quality of stimulation patterns of the implanted electrode array. However, existing processing frameworks prioritize the generation of stimulation while disregarding the potential impact of restor…

Cited by 0SourcePDFScholar
2024

Fact-Aware Summarization with Contrastive Learning for Few-Shot Dialogue State Tracking

ICASSP 2024accepted

Dialogue state tracking (DST) is a crucial component of task-oriented dialogue systems, as it aims to accurately track the user’s goals throughout the dialogue history. However, DST models struggle with new domains due to limited annotated data, leading to poor performance. To solve this key challen…

Cited by 0SourceScholar
2024

FedInverse: Evaluating Privacy Leakage in Federated Learning

ICLR 2024poster

Federated Learning (FL) is a distributed machine learning technique where multiple devices (such as smartphones or IoT devices) train a shared global model by using their local data. FL claims that the data privacy of local participants is preserved well because local data will not be shared with ei…

2024

FedLMT: Tackling System Heterogeneity of Federated Learning via Low-Rank Model Training with Theoretical Guarantees

ICML 2024poster

Federated learning (FL) is an emerging machine learning paradigm for preserving data privacy. However, diverse client hardware often has varying computation resources. Such system heterogeneity limits the participation of resource-constrained clients in FL, and hence degrades the global model accura…

Cited by 2SourcePDFScholar
2024

FedTAD: Topology-aware Data-free Knowledge Distillation for Subgraph Federated Learning

IJCAI 2024poster

Subgraph federated learning (subgraph-FL) is a new distributed paradigm that facilitates the collaborative training of graph neural networks (GNNs) by multi-client subgraphs. Unfortunately, a significant challenge of subgraph-FL arises from subgraph heterogeneity, which stems from node and topology…

Cited by 14SourcePDFScholar
2024

FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning

ICML 2024poster

In this work, we investigate how to leverage pre-trained visual-language models (VLM) for online Reinforcement Learning (RL). In particular, we focus on sparse reward tasks with pre-defined textual task descriptions. We first identify the problem of reward misalignment when applying VLM as a reward…

2024

How Far can 100 Samples Go? Unlocking Zero-Shot Translation with Tiny Multi-Parallel Data

ACL 2024findings

Zero-shot translation aims to translate between language pairs not seen during training in Multilingual Machine Translation (MMT) and is widely considered an open problem. A common, albeit resource-consuming, solution is to add as many related translation directions as possible to the training corpu…

2024

KMatrix: A Flexible Heterogeneous Knowledge Enhancement Toolkit for Large Language Model

EMNLP 2024system demonstrations

Knowledge-Enhanced Large Language Models (K-LLMs) system enhances Large Language Models (LLMs) abilities using external knowledge. Existing K-LLMs toolkits mainly focus on free-textual knowledge, lacking support for heterogeneous knowledge like tables and knowledge graphs, and fall short in comprehe…

2024

MKG-FENN: A Multimodal Knowledge Graph Fused End-to-End Neural Network for Accurate Drug–Drug Interaction Prediction

AAAI 2024technical

Taking incompatible multiple drugs together may cause adverse interactions and side effects on the body. Accurate prediction of drug-drug interaction (DDI) events is essential for avoiding this issue. Recently, various artificial intelligence-based approaches have been proposed for predicting DDI ev…

2024

MMGNN: A Molecular Merged Graph Neural Network for Explainable Solvation Free Energy Prediction

IJCAI 2024poster

In this paper, we address the challenge of accurately modeling and predicting Gibbs free energy in solute-solvent interactions, a pivotal yet complex aspect in the field of chemical modeling. Traditional approaches, primarily relying on deep learning models, face limitations in capturing the intrica…

Cited by 5SourcePDFScholar
2024

Meta-Task Prompting Elicits Embeddings from Large Language Models

ACL 2024long

We introduce a new unsupervised text embedding method, Meta-Task Prompting with Explicit One-Word Limitation (MetaEOL), for generating high-quality sentence embeddings from Large Language Models (LLMs) without the need for model fine-tuning. Leveraging meta-task prompting, MetaEOL guides LLMs to pro…

2024

MogaNet: Multi-order Gated Aggregation Network

ICLR 2024poster

By contextualizing the kernel as global as possible, Modern ConvNets have shown great potential in computer vision tasks. However, recent progress on \textit{multi-order game-theoretic interaction} within deep neural networks (DNNs) reveals the representation bottleneck of modern ConvNets, where the…

2024

Neural Network Approximation for Pessimistic Offline Reinforcement Learning

AAAI 2024technical

Deep reinforcement learning (RL) has shown remarkable success in specific offline decision-making scenarios, yet its theoretical guarantees are still under development. Existing works on offline RL theory primarily emphasize a few trivial settings, such as linear MDP or general function approximatio…

Cited by 3SourcePDFScholar
2024

Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine Translation

EMNLP 2024main

Training a unified multilingual model promotes knowledge transfer but inevitably introduces negative interference. Language-specific modeling methods show promise in reducing interference. However, they often rely on heuristics to distribute capacity and struggle to foster cross-lingual transfer via…

2024

On Leveraging Encoder-only Pre-trained Language Models for Effective Keyphrase Generation

COLING 2024main

This study addresses the application of encoder-only Pre-trained Language Models (PLMs) in keyphrase generation (KPG) amidst the broader availability of domain-tailored encoder-only models compared to encoder-decoder models. We investigate three core inquiries: (1) the efficacy of encoder-only PLMs…

2024

Repoformer: Selective Retrieval for Repository-Level Code Completion

ICML 2024oral

Recent advances in retrieval-augmented generation (RAG) have initiated a new era in repository-level code completion. However, the invariable use of retrieval in existing methods exposes issues in both efficiency and robustness, with a large proportion of the retrieved contexts proving unhelpful or…

Cited by 30SourcePDFScholar
2024

Representational Isomorphism and Alignment of Multilingual Large Language Models

EMNLP 2024finding

In this paper, we investigate the capability of Large Language Models (LLMs) to represent texts in multilingual contexts. Our findings show that sentence representations derived from LLMs exhibit a high degree of isomorphism across languages.This existing isomorphism can facilitate representational…

Cited by 1SourcePDFScholar
2024

Robot Policy Learning with Temporal Optimal Transport Reward

NeurIPS 2024poster

Reward specification is one of the most tricky problems in Reinforcement Learning, which usually requires tedious hand engineering in practice. One promising approach to tackle this challenge is to adopt existing expert video demonstrations for policy learning. Some recent work investigates how to l…

2024

Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation

EMNLP 2024main

Retrieval-augmented language models (RALMs) have shown strong performance and wide applicability in knowledge-intensive tasks. However, there are significant trustworthiness concerns as RALMs are prone to generating unfaithful outputs, including baseless information or contradictions with the retrie…

2024

The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention

EMNLP 2024main

Prompt-based “diversity interventions” are commonly adopted to improve the diversity of Text-to-Image (T2I) models depicting individuals with various racial or gender traits. However, will this strategy result in nonfactual demographic distribution, especially when generating real historical figures…

2024

Unsupervised Anomaly Detection via Masked Diffusion Posterior Sampling

IJCAI 2024poster

Reconstruction-based methods have been commonly used for unsupervised anomaly detection, in which a normal image is reconstructed and compared with the given test image to detect and locate anomalies. Recently, diffusion models have shown promising applications for anomaly detection due to their pow…

Cited by 3SourcePDFScholar
2024

VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling

ICML 2024poster

Similar to natural language models, pre-trained genome language models are proposed to capture the underlying intricacies within genomes with unsupervised sequence modeling. They have become essential tools for researchers and practitioners in biology. However, the hand-crafted tokenization policies…

Cited by 9SourcePDFScholar
2023

Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks

EMNLP 2023long main

Instruction tuning (IT) achieves impressive zero-shot generalization results by training large language models (LLMs) on a massive amount of diverse tasks with instructions. However, how to select new tasks to improve the performance and generalizability of IT models remains an open question. Traini…

Cited by 0SourcecodeScholar
2023

Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN

ICML 2023poster

Masked image modeling, an emerging self-supervised pre-training method, has shown impressive success across numerous downstream vision tasks with Vision transformers. Its underlying idea is simple: a portion of the input image is masked out and then reconstructed via a pre-text task. However, the wo…

Cited by 47SourcePDFScholar
2023

BARA: Efficient Incentive Mechanism with Online Reward Budget Allocation in Cross-Silo Federated Learning

IJCAI 2023poster

Federated learning (FL) is a prospective distributed machine learning framework that can preserve data privacy. In particular, cross-silo FL can complete model training by making isolated data islands of different organizations collaborate with a parameter server (PS) via exchanging model parameter…

Cited by 7SourcePDFScholar
2023

Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine Translation

EMNLP 2023long main

Using a shared vocabulary is common practice in Multilingual Neural Machine Translation (MNMT). In addition to its simple design, shared tokens play an important role in positive knowledge transfer, which manifests naturally when the shared tokens refer to similar meanings across languages. However,…

Cited by 0SourcecodeScholar
2023

DPAUC: Differentially Private AUC Computation in Federated Learning

AAAI 2023technical

Federated learning (FL) has gained significant attention recently as a privacy-enhancing tool to jointly train a machine learning model by multiple participants. The prior work on FL has mostly studied how to protect label privacy during model training. However, model evaluation in FL might also le…

2023

Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames

ICASSP 2023accepted

Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy and latency. In this paper, we present fast-U2++, an enhanced version of U2++ to further reduce partial latency. The cor…

Cited by 0SourceScholar
2023

FedDWA: Personalized Federated Learning with Dynamic Weight Adjustment

IJCAI 2023poster

Different from conventional federated learning, personalized federated learning (PFL) is able to train a customized model for each individual client according to its unique requirement. The mainstream approach is to adopt a kind of weighted aggregation method to generate personalized models, in whic…

2023

Mitigating Domain Dependency for Improved Speech Enhancement Via SNR Loss Boosting

ICASSP 2023accepted

Current supervised speech enhancement methods based on deep learning typically utilize amplitude-based loss functions for optimization, such as Mean Absolute Error (MAE) or Mean Square Error (MSE) loss, which measures the difference between the amplitudes of the estimated and clean speech signals. H…

Cited by 0SourceScholar
2023

Online Semi-supervised Learning with Mix-Typed Streaming Features

AAAI 2023technical

Online learning with feature spaces that are not fixed but can vary over time renders a seemingly flexible learning paradigm thus has drawn much attention. Unfortunately, two restrictions prohibit a ubiquitous application of this learning paradigm in practice. First, whereas prior studies mainly ass…

2023

Optimal Parameterized Joints Selection to Improve Motion Planning Performance of Redundant Manipulators

ICRA 2023poster

The redundant manipulators' analytical solutions can be obtained by the parameterization method. Multiple parameterized joints and their corresponding parametric representations exist for a redundant manipulator. However, how to select the optimal parameterized joints has yet to be well-addressed. T…

Cited by 2SourceScholar
2023

Relightable Neural Human Assets From Multi-View Gradient Illuminations

CVPR 2023poster

Human modeling and relighting are two fundamental problems in computer vision and graphics, where high-quality datasets can largely facilitate related research. However, most existing human datasets only provide multi-view human images captured under the same illumination. Although valuable for mode…

2023

Rethinking Model Selection and Decoding for Keyphrase Generation with Pre-trained Sequence-to-Sequence Models

EMNLP 2023long main

Keyphrase Generation (KPG) is a longstanding task in NLP with widespread applications. The advent of sequence-to-sequence (seq2seq) pre-trained language models (PLMs) has ushered in a transformative era for KPG, yielding promising performance improvements. However, many design decisions remain unexp…

Cited by 0SourcecodeScholar
2023

Sensor Fusion for Shape Reconstruction Using Electromagnetic Tracking Sensors and Multi-Core Optical Fiber

RA-L 2023

Optical fiber-based shape sensing is gaining popularity in cardiac catheterization lately. Typically, these procedures are taking place under the guidance of fluoroscopy. However, fluoroscopy has several disadvantages. Thanks to fiber optic shape sensing and Electromagnetic Tracking (EMT), the 3D ca

Cited by 21SourceScholar
2023

Spatial Self-Distillation for Object Detection with Inaccurate Bounding Boxes

ICCV 2023poster

Object detection via inaccurate bounding box supervision has boosted a broad interest due to the expensive high-quality annotation data or the occasional inevitability of low annotation quality (e.g. tiny objects). The previous works usually utilize multiple instance learning (MIL), which highly dep…

Cited by 18PDFcodeScholar
2023

ToThePoint: Efficient Contrastive Learning of 3D Point Clouds via Recycling

CVPR 2023poster

Recent years have witnessed significant developments in point cloud processing, including classification and segmentation. However, supervised learning approaches need a lot of well-labeled data for training, and annotation is labor- and time-intensive. Self-supervised learning, on the other hand, u…

2023

TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty

ICASSP 2023accepted

In this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which do…

Cited by 0SourceScholar
2022

AutoMix: Unveiling the Power of Mixup for Stronger Classifiers

ECCV 2022poster

"Data mixing augmentation have proved to be effective for improving the generalization ability of deep neural networks. While early methods mix samples by hand-crafted policies (\textit{e.g.}, linear interpolation), recent methods utilize saliency information to match the mixed samples and labels vi…

2022

Constrained Adaptive Projection with Pretrained Features for Anomaly Detection

IJCAI 2022poster

Anomaly detection aims to separate anomalies from normal samples, and the pretrained network is promising for anomaly detection. However, adapting the pretrained features would be confronted with the risk of pattern collapse when finetuning on one-class training data. In this paper, we propose an an…

2022

Contact Localization of Continuum and Flexible Robot Using Data-Driven Approach

RA-L 2022

Continuum robots such as robotic catheters are increasingly being used in minimally invasive surgery. Compliance contributes to enhanced safety during e.g. catheter insertion, however, estimation of contact force and location may help clinicians avoiding exerting excessive force. Ultimately this cou

Cited by 21SourceScholar
2022

DLME: Deep Local-Flatness Manifold Embedding

ECCV 2022poster

"Manifold learning (ML) aims to seek low-dimensional embedding from high-dimensional data. The problem is challenging on real-world datasets, especially with under-sampling data, and we find that previous methods perform poorly in this case. Generally, ML methods first transform input data into a lo…

2022

Deep-Learning-Based Compliant Motion Control of a Pneumatically-Driven Robotic Catheter

RA-L 2022

In cardiovascular interventions, when steering catheters and especially robotic catheters, great care should be paid to prevent applying too large forces on the vessel walls as this could dislodge calcifications, induce scars or even cause perforation. To address this challenge, this paper presents

Cited by 39SourceScholar
2022

Object Localization Under Single Coarse Point Supervision

CVPR 2022poster

Point-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance for the inconsistency of annotated points. Existing POL methods heavily r…

Cited by 33PDFcodeScholar
2022

Online Adaptive Identification and Switching of Soft Contact Model Based on ART-II Method

ICRA 2022poster

In order to obtain a high-precision contact model that can properly describe the target soft tissue, this paper proposes a hybrid soft contact model based on a clustering algorithm ART-II, which selects the most suitable soft contact model according to the surgical environment. The least-square meth…

Cited by 3SourceScholar
2022

Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting

AAAI 2022technical

Time series data appears in many real-world fields such as energy, transportation, communication systems. Accurate modelling and forecasting of time series data can be of significant importance to improve the efficiency of these systems. Extensive research efforts have been taken for time series pro…

2022

Representation Learning for Resource-Constrained Keyphrase Generation

EMNLP 2022finding

State-of-the-art keyphrase generation methods generally depend on large annotated datasets, limiting their performance in domains with limited annotated data. To overcome this challenge, we design a data-oriented approach that first identifies salient information using retrieval-based corpus-level s…

2022

WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition

ICASSP 2022accepted

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours in total. We collect the data from YouTube and Podcast, which covers a variety of…

Cited by 0SourceScholar
2021

Hysteresis Modeling of Robotic Catheters Based on Long Short-Term Memory Network for Improved Environment Reconstruction

RA-L 2021

Catheters are increasingly being used to tackle problems in the cardiovascular system. However, positioning precision of the catheter tip is negatively affected by hysteresis. To ensure tissue damage due to imprecise positioning is avoided, hysteresis is to be understood and compensated for. This wo

Cited by 52SourceScholar
2021

Optimizing Cellular Networks via Continuously Moving Base Stations on Road Networks

ICRA 2021poster

Although existing cellular network base stations are typically immobile, the recent development of small form factor base stations and self driving cars has enabled the possibility of deploying a team of continuously moving base stations that can reorganize the network infrastructure to adapt to cha…

Cited by 1SourceScholar
2021

T-IK: An Efficient Multi-Objective Evolutionary Algorithm for Analytical Inverse Kinematics of Redundant Manipulator

RA-L 2021

This letter proposes a new method combined by the parameterization method and T-IK to solve the inverse kinematics problem of redundant manipulators in the position domain. T-IK is an improved multi-objective optimization algorithm based on NSGA-II. By adding population migration strategy and adapti

Cited by 29SourceScholar
2020

A Fully Actuated Body-Mounted Robotic Assistant for MRI-Guided Low Back Pain Injection

ICRA 2020poster

This paper reports the development of a fully actuated body-mounted robotic assistant for MRI-guided low back pain injection. The robot is designed with a 4-DOF needle alignment module and a 2-DOF remotely actuated needle driver module. The 6-DOF fully actuated robot can operate inside the scanner b…

Cited by 23SourceScholar
2020

Context-Aware Cross-Attention for Non-Autoregressive Translation

COLING 2020main

Non-autoregressive translation (NAT) significantly accelerates the inference process by predicting the entire target sequence. However, due to the lack of target dependency modelling in the decoder, the conditional generation process heavily depends on the cross-attention. In this paper, we reveal a…

Cited by 48SourcePDFScholar
2020

Interval Search Genetic Algorithm Based on Trajectory to Solve Inverse Kinematics of Redundant Manipulators and Its Application

ICRA 2020poster

In this paper, a new method is proposed to solve the inverse kinematics problem of redundant manipulators. This method demonstrates superior performance on continuous motion by combining interval search genetic algorithm based on trajectory which we propose with parametric joint angle method. In thi…

Cited by 13SourceScholar
2016

Convergence-optimized variable node structure for stochastic LDPC decoder

ICASSP 2016accepted

By using stochastic computation, a fully-parallel low-density parity-check (LDPC) decoder can be implemented using a lower wire complexity. In order to enhance the decoder performance, probability tracers, such as up/down counters, are added at each edge between variable nodes and check nodes, as de…

Cited by 3SourceScholar