← Search

Di Wang

168 accepted papers

2026

A Hybrid Space Model for Misaligned Multi-modality Image Fusion

AAAI 2026technical

Infrared and visible image fusion aims to integrate complementary information, such as thermal saliency from infrared imagery and fine-grained texture details from visible imagery. However, real-world multi-modal misalignment and geometric deformation often introduce severe artifacts. Most existing

Cited by 1SourcePDFScholar
2026

Algorithmic Recourse of In-Context Learning for Tabular Data

ICML 2026poster

As predictive models are increasingly deployed in high-stakes settings such as credit approval, there is a growing need for post-hoc methods that provide recourse to affected individuals. Many such models operate on tabular data, where features correspond to real-world attributes. Recently, in-conte…

Cited by 0SourceScholar
2026

Anatomical Region-Guided Contrastive Decoding: A Plug-and-Play Strategy for Mitigating Hallucinations in Medical VLMs

AAAI 2026technical

Medical Vision-Language Models (MedVLMs) show immense promise in clinical applicability. However, their reliability is hindered by hallucinations, where models often fail to derive answers from visual evidence, instead relying on learned textual priors. Existing mitigation strategies for MedVLMs hav

Cited by 0SourcePDFScholar
2026

Any2Any: Unified Arbitrary Modality Translation for Remote Sensing

ICML 2026poster

Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-modal translation methods treat each modality pair as an independent task, resulting in quadratic complexity and limited ge…

Cited by 0SourceScholar
2026

Benign Overfitting in Adversarial Training for Vision Transformers

ICML 2026poster

Despite the remarkable success of Vision Transformers (ViTs) across a wide range of vision tasks, recent studies have revealed that they remain vulnerable to adversarial examples, much like Convolutional Neural Networks (CNNs). A common empirical defense strategy is adversarial training, yet the the…

Cited by 1SourceScholar
2026

Beyond Fixed Biases: Decoding the Role of Reasoning Uncertainty in MLLM Modality Conflicts

ICML 2026poster

Multimodal Large Language Models (MLLMs) must resolve conflicts when modalities provide contradictory information, a process we term "modality following". We propose a framework that deconstructs this behavior into case-specific relative reasoning uncertainty and a model's stable inherent preference…

Cited by 0SourceScholar
2026

Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability

ICML 2026poster

Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning. We introduce TRACED, a framework that assesses reasoning quality through theoretically grounded geometric kinematics. By decomposing reasoning traces into Progress (displacement) and Stab…

Cited by 0SourceScholar
2026

BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point Clouds

CVPR 2026

We introduce BuildAnyPoint, a novel generative framework for structured 3D building reconstruction from point clouds with diverse distributions, such as those captured by airborne LiDAR and Structure-from-Motion.To recover artist-created building abstraction in this highly underconstrained setting,

Cited by 0SourceScholar
2026

ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Senisng

CVPR 2026

Spatiotemporal image generation is a highly meaningful task, which can generate future scenes conditioned on given observations. However, existing change generation methods can only handle event-driven changes (e.g., new buildings) and fail to model cross-temporal variations (e.g., seasonal shifts).

Cited by 0SourcecodeScholar
2026

CompetitorFormer: Mitigating Query Conflicts for 3D Instance Segmentation via Competitive Strategy

CVPR 2026

Transformer-based approaches have recently become the dominant paradigm for 3D instance segmentation. These methods typically employ a multi-layer decoder that iteratively refines a set of learnable queries into instance mask predictions. However, we observe that multiple queries often target the sa

Cited by 0SourcecodeScholar
2026

Degradation-Aware Metric Prompting for Hyperspectral Image Restoration

ICML 2026poster

Unified hyperspectral image (HSI) restoration aims to recover diverse degradations within a single model. However, current methods often rely on impractical explicit priors or opaque black-box representations that overfit to training distributions, hampering generalization to unseen scenarios. To br…

Cited by 0SourceScholar
2026

Discrete Survival Knowledge Distillation for Competing Risks Analysis

ICML 2026poster

Accurate prediction in survival analysis with competing risks is challenged by rare event rates and limited effective sample sizes. Knowledge distillation offers a promising way to transfer information from an external teacher to improve a local student, but existing methods are overwhelmingly devel…

Cited by 0SourceScholar
2026

Dissecting Representation Misalignment in Contrastive Learning via Influence Function

ICLR 2026poster

Contrastive learning, commonly applied in large-scale multimodal models, often relies on data from diverse and often unreliable sources, which can include misaligned or mislabeled text-image pairs. This frequently leads to robustness issues and hallucinations, ultimately causing performance degradat…

Cited by 0SourceScholar
2026

Dual-Kernel Adapter: Expanding Spatial Horizons for Data-Constrained Medical Image Analysis

ICLR 2026poster

Adapters have become a widely adopted strategy for efficient fine-tuning of foundation models, particularly in resource-constrained settings. However, their performance under extreme data scarcity—common in medical imaging due to high annotation costs, privacy regulations, and fragmented datasets—re…

Cited by 0SourceScholar
2026

Efficient Few-Step Solution Generation via Discrete Flow Matching for Combinatorial Optimization

AAAI 2026technical

Combinatorial optimization problems (COPs) are fundamental to many real-world applications where efficiently producing high-quality solutions is critical. Recent advances in diffusion-based non-autoregressive models have reformulated solving COPs as a generative process, achieving promising results.

Cited by 0SourcePDFScholar
2026

Finding Differentially Private Second Order Stationary Points in Stochastic Minimax Optimization

ICML 2026poster

We provide the first study of the problem of finding differentially private (DP) second-order stationary points (SOSP) in stochastic (non-convex) minimax optimization. Existing literature either focuses only on first-order stationary points for minimax problems or on SOSP for classical stochastic mi…

Cited by 0SourceScholar
2026

GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization

CVPR 2026

Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date satellite imagery is unavailable. It further underexploits compl

Cited by 0SourcecodeScholar
2026

HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph Completion

AAAI 2026technical

Multi-modal knowledge graph completion (MMKGC) aims to infer missing entities of triples by leveraging heterogeneous information in knowledge graph (KG). However, existing approaches often struggle with inconsistent modality alignment, limited reasoning depth, and insufficient negative sample qualit

Cited by 0SourcePDFScholar
2026

RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images

AAAI 2026technical

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing images based on free-form natural language expressions. Existing approaches are typically constrained to closed-set vocabularies, limiting their applicability in open-world scenarios. While recent attempts to leverage

Cited by 0SourcePDFScholar
2026

Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models

ICML 2026poster

Large Reasoning Models (LRMs) suffer from sycophantic behavior, where models tend to agree with users' incorrect beliefs and follow misinformation rather than maintain independent reasoning. This behavior undermines model reliability and poses societal risks. Mitigating LRM sycophancy requires monit…

Cited by 0SourceScholar
2026

Residual Diffusion Bridge Model for Image Restoration

CVPR 2026

Diffusion bridge models establish probabilistic paths between arbitrary paired distributions and exhibit great potential for universal image restoration. Most existing methods merely treat them as simple variants of stochastic interpolants, lacking a unified analytical perspective. Besides, they ind

Cited by 0SourcecodeScholar
2026

SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias

AAAI 2026technical

Large vision-language models such as CLIP have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models often develop multimodal spurious biases, the undesirable tendency to rely on spurious features. For example, CLIP may infer

Cited by 0SourcePDFScholar
2026

SAMCL: Empowering SAM to Continually Learn from Dynamic Domains with Extreme Storage Efficiency

AAAI 2026technical

Segment Anything Model (SAM) struggles in open-world scenarios with diverse domains. In such settings, naive fine-tuning with a well-designed learning module is inadequate and often causes catastrophic forgetting issue when learning incrementally. To address this issue, we propose a novel continual

Cited by 0SourcePDFScholar
2026

SARMAE: Masked Autoencoder for SAR Representation Learning

CVPR 2026

Synthetic Aperture Radar (SAR) imagery plays a critical role in all-weather, day-and-night remote sensing applications. However, existing SAR-oriented deep learning is constrained by data scarcity, while the physically grounded speckle noise in SAR imagery further hampers fine-grained semantic repre

Cited by 0SourcecodeScholar
2026

Scaling Inference-Time Computation via Opponent Simulation: Enabling Online Strategic Adaptation in Repeated Negotiation

ICML 2026poster

While large language models (LLMs) have emerged as powerful decision-makers across a wide range of single-agent and stationary environments, fewer efforts have been devoted to settings where LLMs must engage in \emph{repeated} and \emph{strategic} interactions with unknown or dynamic opponents. In s…

Cited by 0SourceScholar
2026

Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding

ICML 2026poster

Multimodal reasoning for ultra-high-resolution (UHR) remote sensing (RS) is usually bottlenecked by visual evidence acquisition: the model necessities localizing tiny task-relevant regions in massive pixel spaces. While Agentic Reinforcement Learning with Verifiable Rewards (RLVR) using zoom-in tool…

Cited by 0SourceScholar
2026

Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models

CVPR 2026

Machine unlearning aims to erase requested data from trained models without full retraining. For Reasoning Multimodal Large Language Models (RMLLMs), this is uniquely challenging: intermediate chain-of-thought steps can still leak sensitive information even when final answers are forgotten, and over

Cited by 0SourceScholar
2026

Trajectory-Aware Certified Decentralized Unlearning via SGD Stability

ICML 2026poster

Decentralized Unlearning (DU) aims to remove the influence of specific clients from a collaboratively trained global model. However, existing methods suffer from strong reliance on static, problem-specific hyperparameters or restrictive convexity assumptions, limiting their general applicability. To…

Cited by 0SourceScholar
2026

TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model

AAAI 2026technical

Transformers are the cornerstone of modern large language models, but their quadratic computational complexity limits efficiency in long-sequence processing. Recent advancements in Mamba, a state space model (SSM) with linear complexity, offer promising efficiency gains but suffer from unstable cont

Cited by 0SourcePDFScholar
2026

Understanding Private Learning From Feature Perspective

ICML 2026poster

Differentially private Stochastic Gradient Descent (DP-SGD) has become integral to privacy-preserving machine learning, ensuring robust privacy guarantees in sensitive domains. Despite notable empirical advances leveraging features from non-private, pre-trained models to enhance DP-SGD training, a t…

Cited by 0SourceScholar
2026

Understanding and Improving Continuous LLM Adversarial Training via In-context Learning Theory

ICLR 2026poster

Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To improve the efficiency of AT for LLMs, recent studies propose continuous AT (CAT) that searches for adversarial inputs within the continuous embedding…

Cited by 0SourcecodeScholar
2026

UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes

CVPR 2026

Instruction-driven segmentation in remote sensing generates masks from guidance, offering great potential for accessible and generalizable applications. However, existing methods suffer from fragmented task formulations and limited instruction data, hindering effective understanding and generalizati

Cited by 0SourcecodeScholar
2026

When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models

AAAI 2026technical

Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented this tendency, the internal mechanisms that enable such behavior remain poorly understood. In this paper, we provide a mec

Cited by 0SourcePDFScholar
2025

A Benchmark for Semantic Sensitive Information in LLMs Outputs

ICLR 2025poster

Large language models (LLMs) can output sensitive information, which has emerged as a novel safety concern. Previous works focus on structured sensitive information (e.g. personal identifiable information). However, we notice that sensitive information can also be at semantic level, i.e. semantic s…

2025

ABNet: Mitigating Sample Imbalance in Anomaly Detection Within Dynamic Graphs

IJCAI 2025

In dynamic graphs, detecting anomalous nodes faces challenges due to sample imbalance, stemming from the scarcity of anomalous samples and feature representation bias. Existing methods often use unsupervised or semi-supervised learning to extract anomalous samples from unlabeled data, but struggle t

Cited by 0SourcePDFScholar
2025

CODEMENV: Benchmarking Large Language Models on Code Migration

ACL 2025finding

Large language models (LLMs) have demonstrated remarkable proficiency in handling a wide range of tasks within the software engineering domain, but their ability to perform code migration—adapting code to different environments—remains underexplored. In this work, we propose a novel benchmark, : Cod…

2025

COMPKE: Complex Question Answering under Knowledge Editing

ACL 2025finding

Knowledge Editing-Efficiently modifying the knowledge in large language models has gathered great attention. Current benchmarks primarily use multi-hop question answering to assess and analyze newly injected or updated knowledge. However, we argue that these benchmarks fail to effectively evaluate h…

Cited by 0SourcePDFScholar
2025

Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation

EMNLP 2025

Suicide remains a major global mental health challenge, and early intervention hinges on recognizing signs of suicidal ideation. In private conversations, such ideation is often expressed in subtle or conflicted ways, making detection especially difficult. Existing data sets are mainly based on publ

Cited by 0SourcePDFScholar
2025

CognitionCapturer: Decoding Visual Stimuli from Human EEG Signal with Multimodal Information

AAAI 2025technical

Electroencephalogram (EEG) signals have attracted significant attention from researchers due to their non-invasive nature and high temporal sensitivity in decoding visual stimuli. However, most recent studies have focused solely on the relationship between EEG and image data pairs, neglecting the va…

2025

DGL: Dynamic Global-Local Information Aggregation for Scalable VRP Generalization with Self-Improvement Learning

IJCAI 2025

The Vehicle Routing Problem (VRP) is a critical combinatorial optimization problem with wide-reaching real-world applications, particularly in logistics, transportation. While neural network-based VRP solvers have shown impressive results on test instances similar to training data, their performance

2025

DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image Restoration

NeurIPS 2025poster

Diffusion models have achieved remarkable progress in universal image restoration. However, existing methods perform naive inference in the reverse process, which leads to cumulative errors under limited sampling steps and large step intervals. Moreover, they struggle to balance the commonality of d…

Cited by 0SourcecodeScholar
2025

EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification

NeurIPS 2025poster

Understanding the internal mechanisms of transformer-based language models remains challenging. Mechanistic interpretability based on circuit discovery aims to reverse engineer neural networks by analyzing their internal processes at the level of computational subgraphs. In this paper, we revisit ex…

Cited by 0SourceScholar
2025

ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge

EMNLP 2025

We introduce ESGenius , a comprehensive benchmark for evaluating and enhancing the proficiency of Large Language Models (LLMs) in Environmental, Social, and Governance (ESG) and sustainability-focused question answering. ESGenius comprises two key components: (i) ESGenius-QA , a collection of 1,136

2025

Editable Concept Bottleneck Models

ICML 2025poster

Concept Bottleneck Models (CBMs) have garnered much attention for their ability to elucidate the prediction process through a human-understandable concept layer. However, most previous studies focused on cases where the data, including concepts, are clean. In many scenarios, we always need to remove…

Cited by 10SourcePDFScholar
2025

FD2-Net: Frequency-Driven Feature Decomposition Network for Infrared-Visible Object Detection

AAAI 2025technical

Infrared-visible object detection (IVOD) seeks to harness the complementary information in infrared and visible images, thereby enhancing the performance of detectors in complex environments. However, existing methods often neglect the frequency characteristics of complementary information, such as…

Cited by 2SourcePDFScholar
2025

Fair Text-to-Image Diffusion via Fair Mapping

AAAI 2025technical

In this paper, we address the limitations of existing text-to-image diffusion models in generating demographically fair results when given human-related descriptions. These models often struggle to disentangle the target language context from sociocultural biases, resulting in biased image generatio…

Cited by 14SourcePDFScholar
2025

Fraud-R1 : A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducements

ACL 2025finding

With the increasing integration of large language models (LLMs) into real-world applications such as finance, e-commerce, and recommendation systems, their susceptibility to misinformation and adversarial manipulation poses significant risks. Existing fraud detection benchmarks primarily focus on si…

2025

GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution

NeurIPS 2025spotlight

Ultra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To addre…

Cited by 0SourcecodeScholar
2025

HMoE: Heterogeneous Mixture of Experts for Language Modeling

EMNLP 2025

Mixture of Experts (MoE) offers remarkable performance and computational efficiency by selectively activating subsets of model parameters. Traditionally, MoE models use homogeneous experts, each with identical capacity. However, varying complexity in input data necessitates experts with diverse capa

2025

Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling

ICCV 2025poster

Masked Image Modeling (MIM) has become an essential method for building foundational visual models in remote sensing (RS). However, the limitations in size and diversity of existing RS datasets restrict the ability of MIM methods to learn generalizable representations. Additionally, conventional MIM…

2025

Heterogeneous Sensor Fusion and Active Perception for Transparent Object Reconstruction with a PDM2 Sensor and a Camera

ICRA 2025

Transparent household objects present a challenge for domestic service robots, since neither regular cameras nor RGB-D cameras can provide accurate points for shape reconstruction. The new type of pretouch dual-modality distance and material sensor (PDM<sup xmlns:mml="http://www.w3.org/1998/Math/Mat

Cited by 2SourceScholar
2025

HumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Dataset

CVPR 2025highlight

With the rapid evolution of 3D generation algorithms, the cost of producing 3D humanoid character models has plummeted, yet the field is impeded by the lack of a comprehensive dataset for automatic rigging--a pivotal step in character animation. Addressing this gap, we present HumanRig, the first la…

2025

Improved Rates of Differentially Private Nonconvex-Strongly-Concave Minimax Optimization

AAAI 2025technical

In this paper, we study the problem of (finite sum) minimax optimization in the Differential Privacy (DP) model. Unlike most of the previous studies on the (strongly) convex-concave settings or loss functions satisfying the Polyak-Lojasiewicz condition, here we mainly focus on the nonconvex-strongly…

Cited by 0SourcePDFScholar
2025

Locate-then-edit for Multi-hop Factual Recall under Knowledge Editing

ICML 2025poster

The locate-then-edit paradigm has shown significant promise for knowledge editing (KE) in Large Language Models (LLMs). While previous methods perform well on single-hop fact recall tasks, they consistently struggle with multi-hop factual recall tasks involving newly edited knowledge. In this paper,…

Cited by 4SourcePDFScholar
2025

MQA-KEAL: Multi-hop Question Answering under Knowledge Editing for Arabic Language

COLING 2025main

Large Language Models (LLMs) have demonstrated significant capabilities across numerous application domains. A key challenge is to keep these models updated with latest available information, which limits the true potential of these models for the end-applications. Although, there have been numerous…

2025

Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning

NAACL 2025findings

Transformer-based language models have achieved significant success; however, their internal mechanisms remain largely opaque due to the complexity of non-linear interactions and high-dimensional operations. While previous studies have demonstrated that these models implicitly embed reasoning trees,…

Cited by 0SourcePDFScholar
2025

Multiple Rotation Averaging with Constrained Reweighting Deep Matrix Factorization

ICRA 2025

Multiple rotation averaging plays a crucial role in computer vision and robotics domains. The conventional optimization-based methods optimize a nonlinear cost function based on certain noise assumptions, while most previous learning-based methods require ground truth labels in the supervised traini

Cited by 0SourceScholar
2025

Nearly Optimal Differentially Private ReLU Regression

UAI 2025

In this paper, we investigate one of the most fundamental non-convex learning problems-ReLU regression-in the Differential Privacy (DP) model. Previous studies on private ReLU regression heavily rely on stringent assumptions, such as constant-bounded norms for feature vectors and labels. We relax th

Cited by 0SourcePDFScholar
2025

Object-Level Backdoor Attacks in RGB-T Semantic Segmentation with Cross-Modality Trigger Optimization

IJCAI 2025

The escalating threat of backdoor risks in deep vision models is a pressing concern. Existing research on backdoor attacks is often confined to a single modality, neglecting the challenges posed by multi-modality scene perception. This work is a pioneer of backdoor attacks in RGB-Thermal (RGB-T) sem

Cited by 0SourcePDFScholar
2025

Predicting Depression in Screening Interviews from Interactive Multi-Theme Collaboration

ACL 2025finding

Automatic depression detection provides cues for early clinical intervention by clinicians. Clinical interviews for depression detection involve dialogues centered around multiple themes. Existing studies primarily design end-to-end neural network models to capture the hierarchical structure of clin…

Cited by 0SourcePDFScholar
2025

Privacy-Preserving Low-Rank Adaptation Against Membership Inference Attacks for Latent Diffusion Models

AAAI 2025technical

Low-rank adaptation (LoRA) is an efficient strategy for adapting latent diffusion models (LDMs) on a private dataset to generate specific images by minimizing the adaptation loss. However, the LoRA-adapted LDMs are vulnerable to membership inference (MI) attacks that can judge whether a particular d…

2025

Private Training Large-scale Models with Efficient DP-SGD

NeurIPS 2025poster

As large language models (LLMs) increasingly underpin technological advancements, the privacy of their training data emerges as a critical concern. Differential Privacy (DP) serves as a rigorous mechanism to protect this data, yet its integration via Differentially Private Stochastic Gradient Descen…

Cited by 0SourcecodeScholar
2025

ReMask-Animate: Refined Character Image Animation Using Mask-Guided Adapters

AAAI 2025technical

Pose-controlled human video generation is of significant interest and finds extensive applications in areas such as automated advertising and content creation on social media platforms. While existing methods employing pose sequences and reference images for human image animation have exhibited nota…

Cited by 0SourcePDFScholar
2025

RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing

NeurIPS 2025poster

Recent advances in self-supervised learning for Vision Transformers (ViTs) have fueled breakthroughs in remote sensing (RS) foundation models. However, the quadratic complexity of self-attention poses a significant barrier to scalability, particularly for large models and high-resolution images. Whi…

Cited by 0SourcecodeScholar
2025

Robust and Efficient 3D Gaussian Splatting for Urban Scene Reconstruction

ICCV 2025poster

We present a framework that enables fast reconstruction and real-time rendering of urban-scale scenes while maintaining robustness against appearance variations across multi-view captures. Our approach begins with scene partitioning for parallel training, employing a visibility-based image selection…

2025

Scaling Laws for Floating–Point Quantization Training

ICML 2025poster

Low-precision training is considered an effective strategy for reducing both training and downstream inference costs. Previous scaling laws for precision mainly focus on integer quantization, which pay less attention to the constituents in floating-point (FP) quantization, and thus cannot well fit t…

Cited by 1SourcePDFScholar
2025

Second-Order Convergence in Private Stochastic Non-Convex Optimization

NeurIPS 2025poster

We investigate the problem of finding second-order stationary points (SOSP) in differentially private (DP) stochastic non-convex optimization. Existing methods suffer from two key limitations: \textbf{(i)} inaccurate convergence error rate due to overlooking gradient variance in the saddle point esc…

Cited by 0SourceScholar
2025

Semi-supervised Concept Bottleneck Models

ICCV 2025poster

Concept Bottleneck Models (CBMs) have garnered increasing attention due to their ability to provide concept-based explanations for black-box deep learning models while achieving high final prediction accuracy using human-like concepts. However, the training of current CBMs is heavily dependent on th…

Cited by 0SourcePDFScholar
2025

Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence

NeurIPS 2025poster

Jailbreak attacks against large language models (LLMs) aim to induce harmful behaviors in LLMs through carefully crafted adversarial prompts. To mitigate attacks, one way is to perform adversarial training (AT)-based alignment, i.e., training LLMs on some of the most adversarial prompts to help them…

Cited by 0SourcecodeScholar
2025

Speculating LLMs’ Chinese Training Data Pollution from Their Tokens

EMNLP 2025

Tokens are basic elements in the datasets for LLM training. It is well-known that many tokens representing Chinese phrases in the vocabulary of GPT (4o/4o-mini/o1/o3/4.5/4.1/o4-mini) are indicating contents like pornography or online gambling. Based on this observation, our goal is to locate Pollute

2025

TEMPO: Temporal Multi-scale Autoregressive Generation of Protein Conformational Ensembles

NeurIPS 2025poster

Understanding the dynamic behavior of proteins is critical to elucidating their functional mechanisms, yet generating realistic, temporally coherent trajectories of protein ensembles remains a significant challenge. In this work, we introduce a novel hierarchical autoregressive framework for modelin…

Cited by 0SourceScholar
2025

TextMEF: Text-guided Prompt Learning for Multi-exposure Image Fusion

IJCAI 2025

Multi-exposure image fusion~(MEF) aims to integrate a set of low dynamic range images, producing a single image with a higher dynamic range than either one. Despite significant advancements, current MEF approaches still struggle to handle extremely over- or under-exposed conditions, resulting in uns

Cited by 0SourcePDFScholar
2025

The Security Threat of Compressed Projectors in Large Vision-Language Models

EMNLP 2025

The choice of a suitable visual language projector (VLP) is critical to the successful training of large visual language models (LVLMs). Mainstream VLPs can be broadly categorized into compressed and uncompressed projectors, and each offers distinct advantages in performance and computational effici

2025

Topology-Aware Learning of Tubular Manifolds via SE(3)-Equivariant Network on Ball B-Spline Curve

NeurIPS 2025poster

Tubular-like system shape analysis is quite difficult in geometry and topology, while it is widely used in plants and organs analysis in practice. However, traditional discrete representations such as voxels and point clouds often require substantial storage and may lead to the loss of fine-grained…

Cited by 0SourceScholar
2025

TraffiDent: A Dataset for Understanding the Interplay Between Traffic Dynamics and Incidents

NeurIPS 2025poster

Long-separated research has been conducted on two highly correlated tracks: traffic and incidents. Traffic track witnesses complicating deep learning models, e.g., to push the prediction a few percent more accurate, and the incident track only studies the incidents alone, e.g., to infer the incident…

Cited by 0SourcecodeScholar
2025

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models

ICCV 2025poster

The rapid advancements in Vision Language Models (VLMs) have prompted the development of multi-modal medical assistant systems. Despite this progress, current models still have inherent probabilistic uncertainties, often producing erroneous or unverified responses--an issue with serious implications…

2025

Understanding How Value Neurons Shape the Generation of Specified Values in LLMs

EMNLP 2025

Rapid integration of large language models (LLMs) into societal applications has intensified concerns about their alignment with universal ethical principles, as their internal value representations remain opaque despite behavioral alignment advancements. Current approaches struggle to systematicall

Cited by 0SourcePDFScholar
2025

Understanding the Dark Side of LLMs’ Intrinsic Self-Correction

ACL 2025long

Intrinsic self-correction was initially proposed to improve LLMs’ responses via feedback solely based on their inherent capability. However, recent works show that LLMs’ intrinsic self-correction fails without oracle labels as feedback. In this paper, our research goal is to *interpret LLMs’ intrins…

Cited by 0SourcePDFScholar
2025

Understanding the Repeat Curse in Large Language Models from a Feature Perspective

ACL 2025finding

Large language models (LLMs) have made remarkable progress in various domains, yet they often suffer from repetitive text generation, a phenomenon we refer to as the ”Repeat Curse”. While previous studies have proposed decoding strategies to mitigate repetition, the underlying mechanism behind this…

2025

XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?

CVPR 2025highlight

The astonishing breakthrough of multimodal large language models (MLLMs) has necessitated new benchmarks to quantitatively assess their capabilities, reveal their limitations, and indicate future research directions. However, this is challenging in the context of remote sensing (RS), since the image…

2024

A Comprehensive Framework for Occluded Human Pose Estimation

ICASSP 2024accepted

Occlusion presents a significant challenge in human pose estimation. The challenges posed by occlusion can be attributed to the following factors: 1) Data: The collection and annotation of occluded human pose samples are relatively challenging. 2) Feature: Occlusion can cause feature confusion due t…

Cited by 0SourceScholar
2024

A Field Guide for Pacing Budget and ROS Constraints

ICML 2024poster

Budget pacing is a popular service that has been offered by major internet advertising platforms since their inception. In the past few years, autobidding products that provide real-time bidding as a service to advertisers have seen a prominent rise in adoption. A popular autobidding stategy is valu…

Cited by 4SourcePDFScholar
2024

An LLM can Fool Itself: A Prompt-Based Adversarial Attack

ICLR 2024poster

The wide-ranging applications of large language models (LLMs), especially in safety-critical domains, necessitate the proper evaluation of the LLM’s adversarial robustness. This paper proposes an efficient tool to audit the LLM’s adversarial robustness via a prompt-based adversarial attack (PromptAt…

2024

Anchoring Path for Inductive Relation Prediction in Knowledge Graphs

AAAI 2024technical

Aiming to accurately predict missing edges representing relations between entities, which are pervasive in real-world Knowledge Graphs (KGs), relation prediction plays a critical role in enhancing the comprehensiveness and utility of KGs. Recent research focuses on path-based methods due to their in…

2024

Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality

ACL 2024findings

Autonomous artificial intelligence (AI) agents have emerged as promising protocols for automatically understanding the language-based environment, particularly with the exponential development of large language models (LLMs). However, a fine-grained, comprehensive understanding of multimodal environ…

2024

Calibration System and Algorithm Design for a Soft Hinged Micro Scanning Mirror With a Triaxial Hall Effect Sensor

RA-L 2024

Micro scanning mirrors (MSMs) extend the range and field of view of LiDARs, medical imaging devices, and laser projectors. However, a new class of soft-hinged MSMs contains out-of-plane translation in addition to the 2 <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http:

Cited by 1SourceScholar
2024

Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization

ICML 2024spotlight

The current state-of-the-art theoretical analysis of Actor-Critic (AC) algorithms significantly lags in addressing the practical aspects of AC implementations. This crucial gap needs bridging to bring the analysis in line with practical implementations of AC. To address this, we advocate for conside…

Cited by 2SourcePDFScholar
2024

Coupled Active Perception and Manipulation Planning for a Mobile Manipulator in Precision Agriculture Applications

ICRA 2024poster

A mobile manipulator often finds itself in an application where it needs to take a close-up view before performing a manipulation task. Named this as a coupled active perception and manipulation (CAPM) problem, we model the uncertainty in the perception process and devise a key state/task planning a…

Cited by 3SourceScholar
2024

Dissecting Fine-Tuning Unlearning in Large Language Models

EMNLP 2024main

Fine-tuning-based unlearning methods prevail for erasing targeted harmful, sensitive, or copyrighted information within large language models while preserving overall capabilities. However, the true effectiveness of the methods is unclear. In this paper, we delve into the limitations of fine-tuning-…

2024

Distilling Autoregressive Models to Obtain High-Performance Non-autoregressive Solvers for Vehicle Routing Problems with Faster Inference Speed

AAAI 2024technical

Neural construction models have shown promising performance for Vehicle Routing Problems (VRPs) by adopting either the Autoregressive (AR) or Non-Autoregressive (NAR) learning approach. While AR models produce high-quality solutions, they generally have a high inference latency due to their sequenti…

2024

Faithful Vision-Language Interpretation via Concept Bottleneck Models

ICLR 2024poster

The demand for transparency in healthcare and finance has led to interpretable machine learning (IML) models, notably the concept bottleneck models (CBMs), valued for their potential in performance and insights into deep neural networks. However, CBM's reliance on manually annotated data poses chall…

Cited by 35SourcePDFScholar
2024

Image Content Generation with Causal Reasoning

AAAI 2024technical

The emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed the reasoning potential reflected in the generated textual content. However, this current ability for causal reasoning…

2024

Improved Analysis of Sparse Linear Regression in Local Differential Privacy Model

ICLR 2024poster

In this paper, we revisit the problem of sparse linear regression in the local differential privacy (LDP) model. Existing research in the non-interactive and sequentially local models has focused on obtaining the lower bounds for the case where the underlying parameter is $1$-sparse, and extending…

Cited by 4SourcePDFScholar
2024

Improving Interpretation Faithfulness for Vision Transformers

ICML 2024spotlight

Vision Transformers (ViTs) have achieved state-of-the-art performance for various vision tasks. One reason behind the success lies in their ability to provide plausible innate explanations for the behavior of neural architectures. However, ViTs suffer from issues with explanation faithfulness, as th…

Cited by 4SourcePDFScholar
2024

LAFA: Multimodal Knowledge Graph Completion with Link Aware Fusion and Aggregation

AAAI 2024technical

Recently, an enormous amount of research has emerged on multimodal knowledge graph completion (MKGC), which seeks to extract knowledge from multimodal data and predict the most plausible missing facts to complete a given multimodal knowledge graph (MKG). However, existing MKGC approaches largely ign…

Cited by 14SourcePDFScholar
2024

LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation

IJCAI 2024poster

Due to spatial redundancy in remote sensing images, sparse tokens containing rich information are usually involved in self-attention (SA) to reduce the overall token numbers within the calculation, avoiding the high computational cost issue in Vision Transformers. However, such methods usually obtai…

2024

Matching Distance and Geometric Distribution Aided Learning Multiview Point Cloud Registration

RA-L 2024

Multiview point cloud registration plays a crucial role in robotics, automation, and computer vision fields. This letter concentrates on pose graph construction and motion synchronization within multiview registration. Previous methods for pose graph construction often pruned fully connected graphs

Cited by 8SourcecodeScholar
2024

Mixed Geometry Message and Trainable Convolutional Attention Network for Knowledge Graph Completion

AAAI 2024technical

Knowledge graph completion (KGC) aims to study the embedding representation to solve the incompleteness of knowledge graphs (KGs). Recently, graph convolutional networks (GCNs) and graph attention networks (GATs) have been widely used in KGC tasks by capturing neighbor information of entities. Howev…

Cited by 10SourcePDFScholar
2024

Perplexity-aware Correction for Robust Alignment with Noisy Preferences

NeurIPS 2024poster

Alignment techniques are critical in ensuring that large language models (LLMs) output helpful and harmless content by enforcing the LLM-generated content to align with human preferences. However, the existence of noisy preferences (NPs), where the responses are mistakenly labelled as chosen or rej…

2024

Private Language Models via Truncated Laplacian Mechanism

EMNLP 2024main

Recently it has been shown that deep learning models for NLP tasks are prone to attacks that can even reconstruct the verbatim training texts. To prevent privacy leakage, researchers have investigated word-level perturbations, relying on the formal guarantees of differential privacy (DP) in the embe…

Cited by 1SourcePDFScholar
2024

Refining Latent Homophilic Structures over Heterophilic Graphs for Robust Graph Convolution Networks

AAAI 2024technical

Graph convolution networks (GCNs) are extensively utilized in various graph tasks to mine knowledge from spatial data. Our study marks the pioneering attempt to quantitatively investigate the GCN robustness over omnipresent heterophilic graphs for node classification. We uncover that the predominant…

Cited by 12SourcePDFScholar
2024

Revisiting Differentially Private ReLU Regression

NeurIPS 2024poster

As one of the most fundamental non-convex learning problems, ReLU regression under differential privacy (DP) constraints, especially in high-dimensional settings, remains a challenging area in privacy-preserving machine learning. Existing results are limited to the assumptions of bounded norm $ \|\m…

Cited by 1SourcePDFScholar
2024

SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation

AAAI 2024technical

High-resolution representation is essential for achieving good performance in human pose estimation models. To obtain such features, existing works utilize high-resolution input images or fine-grained image tokens. However, this dense high-resolution representation brings a significant computational…

2024

Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling

NeurIPS 2024poster

In current deep learning tasks, Adam-style optimizers—such as Adam, Adagrad, RMSprop, Adafactor, and Lion—have been widely used as alternatives to SGD-style optimizers. These optimizers typically update model parameters using the sign of gradients, resulting in more stable convergence curves. The l…

Cited by 6SourcePDFScholar
2024

TMFN: A Target-oriented Multi-grained Fusion Network for End-to-end Aspect-based Multimodal Sentiment Analysis

COLING 2024main

End-to-end multimodal aspect-based sentiment analysis (MABSA) combines multimodal aspect terms extraction (MATE) with multimodal aspect sentiment classification (MASC), aiming to simultaneously extract aspect words and classify the sentiment polarity of each aspect. However, existing MABSA methods h…

Cited by 3SourcePDFScholar
2024

Toward Precise Robotic Weed Flaming Using a Mobile Manipulator with a Blowtorch

IROS 2024poster

Robotic weed flaming is a new and environmentally friendly approach to weed removal in the agricultural field. Using a mobile manipulator equipped with a blowtorch, we design a new system and algorithm to enable effective weed flaming, which requires robotic manipulation with a soft and deformable e…

Cited by 0SourceScholar
2024

Towards Multi-dimensional Explanation Alignment for Medical Classification

NeurIPS 2024poster

The lack of interpretability in the field of medical image analysis has significant ethical and legal implications. Existing interpretable methods in this domain encounter several challenges, including dependency on specific models, difficulties in understanding and visualization, and issues related…

Cited by 1SourcePDFScholar
2024

Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning

AAAI 2024technical

Despite the great success of large language models (LLMs) in various tasks, they suffer from generating hallucinations. We introduce Truth Forest, a method that enhances truthfulness in LLMs by uncovering hidden truth representations using multi-dimensional orthogonal probes. Specifically, it create…

2024

Truthful High Dimensional Sparse Linear Regression

NeurIPS 2024poster

We study the problem of fitting the high dimensional sparse linear regression model, where the data are provided by strategic or self-interested agents (individuals) who prioritize their privacy of data disclosure. In contrast to the classical setting, our focus is on designing mechanisms that can e…

Cited by 1SourcePDFScholar
2024

Unleashing Channel Potential: Space-Frequency Selection Convolution for SAR Object Detection

CVPR 2024poster

Deep Convolutional Neural Networks (DCNNs) have achieved remarkable performance in synthetic aperture radar (SAR) object detection but this comes at the cost of tremendous computational resources partly due to extracting redundant features within a single convolutional layer. Recent works either del…

Cited by 14SourcePDFScholar
2023

A Pretouch Perception Algorithm for Object Material and Structure Mapping to Assist Grasp and Manipulation Using a DMDSM Sensor

IROS 2023poster

We report a new material and structure mapping (MSM) algorithm to assist robotic grasping and manipulation. Building on our new sensor development, the algorithm has four main components: 1) detection of time-of-flight (ToF) durations for the dual modalities of optoacoustic (OA) and pulse-echo ultra…

Cited by 3SourceScholar
2023

DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text

EMNLP 2023long findings

With the rapid progress of Large language models (LLMs) and the huge amount of text they generate, it becomes impractical to manually distinguish whether a text is machine-generated. The growing use of LLMs in social media and education, prompts us to develop methods to detect machine-generated text…

Cited by 0SourcecodeScholar
2023

Differentially Private Episodic Reinforcement Learning with Heavy-tailed Rewards

ICML 2023poster

In this paper we study the problem of (finite horizon tabular) Markov decision processes (MDPs) with heavy-tailed rewards under the constraint of differential privacy (DP). Compared with the previous studies for private reinforcement learning that typically assume rewards are sampled from some bound…

Cited by 1SourcePDFScholar
2023

Differentially Private Stochastic Convex Optimization in (Non)-Euclidean Space Revisited

UAI 2023poster

In this paper, we revisit the problem of Differentially Private Stochastic Convex Optimization (DP-SCO) in Euclidean and general $\ell_p^d$ spaces. Specifically, we focus on three settings that are still far from well understood: (1) DP-SCO over a constrained and bounded (convex) set in Euclidean…

Cited by 6SourcePDFScholar
2023

GRI: Graph-based Relative Isomorphism of Word Embedding Spaces

EMNLP 2023long findings

Automated construction of bi-lingual dictionaries using monolingual embedding spaces is a core challenge in machine translation. The end performance of these dictionaries relies on the geometric similarity of individual spaces, i.e., their degree of isomorphism. Existing attempts aimed at controllin…

Cited by 0SourcecodeScholar
2023

Multi-Aspect Explainable Inductive Relation Prediction by Sentence Transformer

AAAI 2023technical

Recent studies on knowledge graphs (KGs) show that path-based methods empowered by pre-trained language models perform well in the provision of inductive and explainable relation predictions. In this paper, we introduce the concepts of relation path coverage and relation path confidence to filter ou…

2023

Robust Budget Pacing with a Single Sample

ICML 2023oral

Major Internet advertising platforms offer budget pacing tools as a standard service for advertisers to manage their ad campaigns. Given the inherent non-stationarity in an advertiser's value and also competing advertisers' values over time, a commonly used approach is to learn a target expenditure…

Cited by 4SourcePDFScholar
2023

SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model

NeurIPS 2023poster

The success of the Segment Anything Model (SAM) demonstrates the significance of data-centric machine learning. However, due to the difficulties and high costs associated with annotating Remote Sensing (RS) images, a large amount of valuable RS data remains unlabeled, particularly at the pixel level…

2023

The Third Generation (G3) Dual-Modal and Dual Sensing Mechanisms (DMDSM) Pretouch Sensor for Robotic Grasping

ICRA 2023poster

Fingertip-mounted pretouch sensors are very useful for robotic grasping. In this paper, we report a new (G3) dual-modal and dual sensing mechanisms (DMDSM) pretouch sensor for near-distance ranging and material sensing, which is based on pulse-echo ultrasound (US) and optoacoustics (OA). Different f…

Cited by 4SourceScholar
2022

IM2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation

EMNLP 2022main

Evaluation metrics shine the light on the best models and thus strongly influence the research directions, such as the recently developed dialogue metrics USR, FED, and GRADE. However, most current metrics evaluate the dialogue data as isolated and static because they only focus on a single quality…

2022

On Facility Location Problem in the Local Differential Privacy Model

AISTATS 2022poster

We study the facility location problem under the constraints imposed by local differential privacy (LDP). Recently, Gupta et al. (2010) and Esencayi et al. (2019) proposed lower and upper bounds for the problem on the central differential privacy (DP) model where a trusted curator first collects all…

Cited by 3SourcePDFScholar
2022

Optimal Rates of (Locally) Differentially Private Heavy-tailed Multi-Armed Bandits

AISTATS 2022poster

In this paper we investigate the problem of stochastic multi-armed bandits (MAB) in the (local) differential privacy (DP/LDP) model. Unlike previous results that assume bounded/sub-Gaussian reward distributions, we focus on the setting where each arm’s reward distribution only has $(1+v)$-th moment…

Cited by 38SourcePDFScholar
2022

Private Stochastic Convex Optimization and Sparse Learning with Heavy-tailed Data Revisited

IJCAI 2022poster

In this paper, we revisit the problem of Differentially Private Stochastic Convex Optimization (DP-SCO) with heavy-tailed data, where the gradient of the loss function has bounded moments. Instead of the case where the loss function is Lipschitz or each coordinate of the gradient has bounded second…

Cited by 13SourcePDFScholar
2022

Semantic-aware Texture-Structure Feature Collaboration for Underwater Image Enhancement

ICRA 2022poster

Underwater image enhancement has become an attractive topic as a significant technology in marine engi-neering and aquatic robotics. However, the limited number of datasets and imperfect hand-crafted ground truth weaken its robustness to unseen scenarios, and hamper the application to high-level vis…

Cited by 34SourcecodeScholar
2022

The Second Generation (G2) Fingertip Sensor for Near-Distance Ranging and Material Sensing in Robotic Grasping

ICRA 2022poster

To continuously improve robotic grasping, we are interested in developing a contactless fingertip-mounted sensor for near-distance ranging and material sensing. Previously, we demonstrated a dual-modal and dual sensing mechanisms (DMDSM) pretouch sensor prototype based on pulse-echo ultrasound and o…

Cited by 9SourceScholar
2022

Unsupervised Misaligned Infrared and Visible Image Fusion via Cross-Modality Image Generation and Registration

IJCAI 2022poster

Recent learning-based image fusion methods have marked numerous progress in pre-registered multi-modality data, but suffered serious ghosts dealing with misaligned multi-modality data, due to the spatial deformation and the difficulty narrowing cross-modality discrepancy. To overcome the obstacles,…

2021

CARE: Commonsense-Aware Emotional Response Generation with Latent Concepts

AAAI 2021technical

Rationality and emotion are two fundamental elements of humans. Endowing agents with rationality and emotion has been one of the major milestones in AI. However, in the field of conversational AI, most existing models only specialize in one aspect and neglect the other, which often leads to dull or…

2021

Concept-Based Label Embedding via Dynamic Routing for Hierarchical Text Classification

ACL 2021long

Hierarchical Text Classification (HTC) is a challenging task that categorizes a textual description within a taxonomic hierarchy. Most of the existing methods focus on modeling the text. Recently, researchers attempt to model the class representations with some resources (e.g., external dictionaries…

2021

Device Design and System Integration of a Two-Axis Water-immersible Micro Scanning Mirror (WIMSM) to Enable Dual-modal Optical and Acoustic Communication and Ranging for Underwater Vehicles

ICRA 2021poster

To address the communication and ranging challenges caused by underwater environment, we design dual modal devices for autonomous underwater vehicles (AUVs). The dual-modal design builds upon a co-axial ultrasonic and green laser beams which leverage different signal diverging patterns and different…

Cited by 6SourceScholar
2021

Fingertip Pulse-Echo Ultrasound and Optoacoustic Dual-Modal and Dual Sensing Mechanisms Near-Distance Sensor for Ranging and Material Sensing in Robotic Grasping

ICRA 2021poster

To improve robotic grasping, we are interested in developing a new non-contact fingertip-mounted sensor for near-distance ranging and material sensing. Here we report new progress in combining direct pulse-echo ultrasound and optoacoustic effects in sensor design to deal with optically and/or acoust…

Cited by 10SourceScholar
2021

PLOME: Pre-training with Misspelled Knowledge for Chinese Spelling Correction

ACL 2021long

Chinese spelling correction (CSC) is a task to detect and correct spelling errors in texts. CSC is essentially a linguistic problem, thus the ability of language understanding is crucial to this task. In this paper, we propose a Pre-trained masked Language model with Misspelled knowledgE (PLOME) for…

2020

CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional Correspondence

IROS 2020poster

In this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective functi…

Cited by 15SourcecodeScholar
2020

Fingertip Non-Contact Optoacoustic Sensor for Near-Distance Ranging and Thickness Differentiation for Robotic Grasping

IROS 2020poster

We report the feasibility study of a new optoacoustic sensor for both near-distance ranging and material thickness classification for robotic grasping. It is based on the optoacoustic effect where focused laser pulses are used to generate wideband ultrasound signals in the target. With a much smalle…

Cited by 11SourceScholar
2020

On Differentially Private Stochastic Convex Optimization with Heavy-tailed Data

ICML 2020poster

In this paper, we consider the problem of designing Differentially Private (DP) algorithms for Stochastic Convex Optimization (SCO) on heavy-tailed data. The irregularity of such data violates some key assumptions used in almost all existing DP-SCO and DP-ERM methods, resulting in failure to provide…

Cited by 71SourcePDFScholar
2020

Robust Pedestrian Tracking in Crowd Scenarios Using an Adaptive GMM-based Framework

IROS 2020poster

In this paper, we address the issue of pedestrian tracking in crowd scenarios. People in close social relationships tend to act as a group which is a great challenge to individually discriminate and track pedestrians on a LiDAR system. In this paper, we integrally model groups of people and track th…

Cited by 4SourceScholar
2019

Differentially Private Empirical Risk Minimization with Non-convex Loss Functions

ICML 2019oral

We study the problem of Empirical Risk Minimization (ERM) with (smooth) non-convex loss functions under the differential-privacy (DP) model. Existing approaches for this problem mainly adopt gradient norms to measure the error, which in general cannot guarantee the quality of the solution. To addres…

Cited by 105SourcePDFScholar
2019

Facility Location Problem in Differential Privacy Model Revisited

NeurIPS 2019poster

In this paper we study the facility location problem in the model of differential privacy (DP) with uniform facility cost. Specifically, we first show that under the hierarchically well-separated tree (HST) metrics and the super-set output setting that was introduced in Gupta et. al., there is an $\…

Cited by 12SourcePDFScholar
2019

On the Tunable Sparse Graph Solver for Pose Graph Optimization in Visual SLAM Problems

IROS 2019poster

We report a tunable sparse optimization solver that can trade a slight decrease in accuracy for significant speed improvement in pose graph optimization in visual simultaneous localization and mapping (vSLAM). The solver is designed for devices with significant computation and power constraints such…

Cited by 9SourceScholar
2019

Precise Correntropy-based 3D Object Modelling With Geometrical Traffic Prior

IROS 2019poster

Robust 3D perception using LiDAR is of prime importance for robotics, and its fundamental core lies in precise object modelling resisting to noise and outliers. In this paper, a precise 3D object modelling algorithm is designed especially for the intelligent vehicles. The proposed algorithm is advan…

Cited by 1SourceScholar
2019

Toward Fingertip Non-Contact Material Recognition and Near-Distance Ranging for Robotic Grasping

ICRA 2019poster

We report the feasibility study of a new acoustic and optical bi-modal distance & material sensor for robotic grasping. The new sensor is designed to be mounted on the robot fingertip to provide last-moment perception before contact happens. It is based on both pulse-echo ultrasound and optoacoustic…

Cited by 19SourceScholar
2019

Virtual Lane Boundary Generation for Human-Compatible Autonomous Driving: A Tight Coupling between Perception and Planning

IROS 2019poster

Existing autonomous vehicle (AV) navigation algorithms treat lane recognition, obstacle avoidance, local path planning, and lane following as separate functional modules which result in driving behavior that is incompatible with human drivers. It is imperative to design human-compatible navigation a…

Cited by 10SourceScholar
2018

Empirical Risk Minimization in Non-interactive Local Differential Privacy Revisited

NeurIPS 2018poster

In this paper, we revisit the Empirical Risk Minimization problem in the non-interactive local model of differential privacy. In the case of constant or low dimensions ($p\ll n$), we first show that if the loss function is $(\infty, T)$-smooth, we can avoid a dependence of the sample complexity,…

Cited by 76SourcePDFScholar
2018

Encoder-Camera-Ground Penetrating Radar Tri-Sensor Mapping for Surface and Subsurface Transportation Infrastructure Inspection

ICRA 2018poster

We report system and algorithmic development for a sensing suite comprising multiple sensors for both surface and subsurface transportation infrastructure inspection focusing on multi-modal mapping for inspection. The sensing suite contains a camera, a ground penetrating radar (GPR), and a wheel enc…

Cited by 19SourceScholar
2018

Robotic Subsurface Pipeline Mapping with a Ground-penetrating Radar and a Camera

IROS 2018poster

We propose a novel subsurface pipeline mapping method by fusing Ground Penetrating Radar (GPR) scans and camera images. To facilitate the simultaneous detection of multiple pipelines, we model the GPR sensing process and prove hyperbola response for general scanning with non-perpendicular angles. Fu…

Cited by 9SourceScholar
2017

Capacity Releasing Diffusion for Speed and Locality

ICML 2017poster

Diffusions and related random walk procedures are of central importance in many areas of machine learning, data analysis, and applied mathematics. Because they spread mass agnostically at each step in an iterative manner, they can sometimes spread mass “too aggressively,” thereby failing to find the…

Cited by 49SourcePDFScholar
2017

Differentially Private Empirical Risk Minimization Revisited: Faster and More General

NeurIPS 2017poster

In this paper we study differentially private Empirical Risk Minimization(ERM) in different settings. For smooth (strongly) convex loss function with or without (non)-smooth regularization, we give algorithms which achieve either optimal or near optimal utility bound with less gradient complexity co…

Cited by 343SourcePDFScholar