← Search

Chao Chen

143 accepted papers

2026

Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning

CVPR 2026

Computational pathology has advanced rapidly in recent years, driven by domain-specific image encoders and growing interest in using vision-language models to answer natural-language questions about diseases. Yet, the core problem behind pathology question-answering remains unsolved, considering tha

Cited by 0SourcecodeScholar
2026

Approximation Algorithm for Constrained k-Center Clustering: A Local Search Approach

AAAI 2026technical

Clustering is a long-standing research problem and a fundamental tool in AI and data analysis. The traditional k-center problem, known as a fundamental theoretical challenge in clustering, has a best possible approximation ratio of 2, and any improvement to a ratio of 2 - ε would imply P = NP. In th

Cited by 0SourcePDFScholar
2026

CoCo-MILP: Inter-Variable Contrastive and Intra-Constraint Competitive MILP Solution Prediction

AAAI 2026technical

Mixed-Integer Linear Programming (MILP) is a cornerstone of combinatorial optimization, yet solving large-scale instances remains a significant computational challenge. Recently, Graph Neural Networks (GNNs) have shown promise in accelerating MILP solvers by predicting high-quality solutions. Howev

Cited by 0SourcePDFScholar
2026

Disturbance-Aware Hybrid Learning for Robust and Adaptive UAV Flight in Extreme Winds

IJCAI 2026

Safe and precise maneuvering of quadrotor unmanned aerial vehicles (UAVs) in high-speed wind environments remains a critical challenge. Wind disturbances are nonlinear, time-varying, and difficult to model, causing traditional controllers to struggle with perception and compensation, especially unde

Cited by 0Scholar
2026

FinMathBench: A Formula-Driven Benchmark for Evaluating LLMs’ Math Reasoning Capabilities in Finance

AAAI 2026technical

Many existing financial math reasoning benchmarks suffer from data contamination and high manual construction costs. To address this, we propose a novel formula-driven approach to dynamically construct math reasoning benchmarks in finance. Our two-stage approach: (1) generates single-formula questio

Cited by 0SourcePDFScholar
2026

Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition

ICRA 2026poster

Sequential Visual Place Recognition (Seq-VPR) leverages transformers to capture spatio-temporal features effectively. In practice, a transformer-based Seq-VPR model should be flexible to the number of frames per sequence (sequence length), deliver fast inference, and use little memory to meet real-t…

2026

From Attribution to Action: Jointly ALIGNing Predictions and Explanations

AAAI 2026technical

Explanation-guided learning (EGL) has shown promise in aligning model predictions with interpretable reasoning, particularly in computer vision tasks. However, most approaches rely on external annotations or heuristic-based segmentation to supervise model explanations, which can be noisy, imprecise

Cited by 0SourcePDFScholar
2026

HE-VPR: Height Estimation Enabled Aerial Visual Place Recognition against Scale Variance

ICRA 2026poster

In this work, we propose HE-VPR, a visual place recognition (VPR) framework that incorporates height estimation. Our system decouples height inference from place recognition, allowing both modules to share a frozen DINOv2 backbone. Two lightweight bypass adapter branches are integrated into our syst…

2026

Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval

ICLR 2026poster

Retrieval-Augmented Generation (RAG) has significantly enhanced LLMs by incorporating external information. However, prevailing agentic RAG approaches are constrained by a critical limitation: they treat the retrieval process as a black-box querying operation. This confines agents' actions to query…

Cited by 0SourceScholar
2026

MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation

AAAI 2026technical

Semantic segmentation is a fundamental task in computer vision with wide-ranging applications, including autonomous driving and robotics. While RGB-based methods have achieved strong performance with CNNs and Transformers, their effectiveness degrades under fast motion, low-light, or high dynamic ra

Cited by 0SourcePDFScholar
2026

OPIC: Enhancing Language Model Merging via Optimizing In-Context Capability

ICML 2026poster

Task-vector–based model merging enables low-cost, training-free multi-task learning for large language models, but suffers from severe performance degradation due to task conflict. Prior mitigation strategies largely rely on validation data for costly hyperparameter tuning, limiting both interpretab…

Cited by 0SourceScholar
2026

Optimized Algorithms for Text Clustering with LLM-Generated Constraints

AAAI 2026technical

Clustering is a fundamental tool that has garnered significant interest across a wide range of applications including text analysis. To improve clustering accuracy, many researchers have proposed incorporating background knowledge, typically in the form of must‑link and cannot‑link constraints, to

Cited by 0SourcePDFScholar
2026

SA-MPPI: Sensitivity-Aware Model Predictive Path Integral Control for Robust and Agile Quadrotor Flight

ICRA 2026poster

Reliable quadrotor control in dynamic environments remains challenging due to external disturbances and internal uncertainties. While Model Predictive Path Integral (MPPI) control enables agile maneuvers through samplingbased optimization, its performance often degrades under such unmodeled uncertai…

Cited by 0Scholar
2026

The Convergent Representation of Vision-Language Contrastive Learning: Geometry, Modality Gap and Shared Space Alignment

ICML 2026poster

Multimodal contrastive learning (MCL) aims to embed data from two modalities in a shared embedding space. However, in practice, representations of images and text occupy completely separate regions of embedding space, a phenomenon called the modality gap. Moreover, experimental findings on how the s…

Cited by 0SourceScholar
2026

UltraVPR: Unsupervised Lightweight Rotation-Invariant Aerial Visual Place Recognition

ICRA 2026poster

Aerial Visual Place Recognition (VPR) is critical for Unmanned Aerial Vehicles (UAVs) localization, especially in environments with unstable or unavailable GPS signals. While neural network-based VPR methods have become mainstream, they face significant challenges on UAV platforms. Traditional CNN-b…

2026

Unleashing Humanoid Reaching Potential Via Real-World-Ready Skill Space

ICRA 2026poster

Humans possess a large reachable space in the 3D world, enabling interactions with objects at varying heights and distances. However, realizing such large-space reaching on humanoids is a complex whole-body control (WBC) problem. Learning from scratch often leads to optimization difficulty and poor …

2026

Unleashing Humanoid Reaching Potential via Real-World-Ready Skill Space

RA-L 2026

Humans possess a large reachable space in the 3D world, enabling interactions with objects at varying heights and distances. However, realizing such large-space reaching on humanoids is a complex whole-body control (WBC) problem. Learning from scratch often leads to optimization difficulty and poor

Cited by 23SourcecodeScholar
2025

Backdooring Vision-Language Models with Out-Of-Distribution Data

ICLR 2025poster

The emergence of Vision-Language Models (VLMs) represents a significant advancement in integrating computer vision with Large Language Models (LLMs) to generate detailed text descriptions from visual inputs. Despite their growing importance, the security of VLMs, particularly against backdoor attack…

Cited by 3SourcePDFScholar
2025

Controlling Thinking Speed in Reasoning Models

NeurIPS 2025spotlight

Human cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking. While current Large Reasoning Models (LRMs) excel at System 2 thinking, their inability to perform fast thinking leads to high computational overhead and latency. In this w…

Cited by 0SourceScholar
2025

Focus-Then-Reuse: Fast Adaptation in Visual Perturbation Environments

NeurIPS 2025poster

Visual reinforcement learning has shown promise in various real-world applications. However, deploying policies in complex real-world environments with visual perturbations remains a significant challenge. We notice that humans tend to filter information at the object level prior to decision-making,…

Cited by 0SourcecodeScholar
2025

Geometry of Long-Tailed Representation Learning: Rebalancing Features for Skewed Distributions

ICLR 2025poster

Deep learning has achieved significant success by training on balanced datasets. However, real-world data often exhibit long-tailed distributions. Empirical studies have revealed that long-tailed data skew data representations, where head classes dominate the feature space. Many methods have been pr…

Cited by 0SourcePDFScholar
2025

Heteroscedastic Bayesian Optimization-Based Dynamic PID Tuning for Accurate and Robust UAV Trajectory Tracking

IROS 2025

Unmanned Aerial Vehicles (UAVs) play an important role in various applications, where precise trajectory tracking is crucial. However, conventional control algorithms for trajectory tracking often exhibit limited performance due to the underactuated, nonlinear, and highly coupled dynamics of quadrot

Cited by 1SourceScholar
2025

Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification

CVPR 2025poster

Despite remarkable advancements in text-to-image person re-identification (TIReID) facilitated by the breakthrough of cross-modal embedding models, existing methods often struggle to distinguish challenging candidate images due to intrinsic limitations, such as network architecture and data quality.…

2025

Learning Bijective Surface Parameterization for Inferring Signed Distance Functions from Sparse Point Clouds with Grid Deformation

CVPR 2025poster

Inferring signed distance functions (SDFs) from sparse point clouds remains a challenge in surface reconstruction. The key lies in the lack of detailed geometric information in sparse point clouds, which is essential for learning a continuous field. To resolve this issue, we present a novel approach…

Cited by 3SourcePDFScholar
2025

MATCH: Multi-faceted Adaptive Topo-Consistency for Semi-Supervised Histopathology Segmentation

NeurIPS 2025poster

In semi-supervised segmentation, capturing meaningful semantic structures from unlabeled data is essential. This is particularly challenging in histopathology image analysis, where objects are densely distributed. To address this issue, we propose a semi-supervised segmentation framework designed to…

Cited by 0SourcecodeScholar
2025

MERGE: Multi-faceted Hierarchical Graph-based GNN for Gene Expression Prediction from Whole Slide Histopathology Images

CVPR 2025poster

Recent advances in Spatial Transcriptomics (ST) pair histology images with spatially resolved gene expression profiles, enabling predictions of gene expression across different tissue locations based on image patches. This opens up new possibilities for enhancing whole slide image (WSI) prediction t…

2025

MSR: A Multifaceted Self-Retrieval Framework for Microscopic Cascade Prediction

AAAI 2025technical

The microscopic cascade prediction task has wide applications in downstream areas like ''rumor detection''. Its goal is to forecast the diffusion routines of information cascade within networks. Existing works typically formulate it as a classification task, which fails to well align with the Social…

Cited by 0SourcePDFScholar
2025

NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban Environments

ICRA 2025

Visual place recognition (VPR) enables autonomous robots to identify previously visited locations, which contributes to tasks like simultaneous localization and mapping (SLAM). VPR faces challenges such as accurate image neighbor retrieval and appearance change in scenery. Event cameras, also known

Cited by 8SourcecodeScholar
2025

OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part Detection

ICCV 2025poster

In recent years, affordance detection has become essential for robotic manipulation in real-world scenes, where robots must autonomously interpret commands and perform actions. Current methods often focus on individual point cloud objects or simple semantic queries, limiting their effectiveness in d…

Cited by 0SourcePDFScholar
2025

ROPO: Robust Preference Optimization for Large Language Models

ICML 2025poster

The prevalent noise in the preference data unavoidably poses significant challenges to the preference alignment of large language models (LLMs). Existing efforts for this problem either marginally alleviate the impact of noise without noise reduction, or rely on external LLMs that incur substantial…

Cited by 2SourcePDFScholar
2025

ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQL

ICLR 2025poster

Despite the significant advancements in Text-to-SQL (Text2SQL) facilitated by large language models (LLMs), the latest state-of-the-art techniques are still trapped in the in-context learning of closed-source LLMs (e.g., GPT-4), which limits their applicability in open scenarios. To address this ch…

2025

RoME: Domain-Robust Mixture-of-Experts for MILP Solution Prediction across Domains

NeurIPS 2025poster

Mixed-Integer Linear Programming (MILP) is a fundamental and powerful framework for modeling complex optimization problems across diverse domains. Recently, learning-based methods have shown great promise in accelerating MILP solvers by predicting high-quality solutions. However, most existing appro…

Cited by 0SourceScholar
2025

SKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMs

CVPR 2025poster

Generating layouts from textual descriptions by large language models (LLMs) plays a crucial role in precise spatial reasoning-induced domains such as robotic object rearrangement and text-to-image generation. However, current methods face challenges in limited real-world examples, handling diverse…

Cited by 0SourcePDFScholar
2025

SLTNet: Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based Networks

IROS 2025

Event-based semantic segmentation has great potential in autonomous driving and robotics due to the advantages of event cameras, such as high dynamic range, low latency, and low power cost. Unfortunately, current artificial neural network (ANN)-based segmentation methods suffer from high computation

Cited by 1SourcecodeScholar
2025

Self-Supervised Place Recognition by Refining Temporal and Featural Pseudo Labels From Panoramic Data

RA-L 2025

Visual place recognition (VPR) using deep networks has achieved state-of-the-art performance. However, most of them require a training set with ground truth sensor poses to obtain positive and negative samples of each observation's spatial neighborhood for supervised learning. When such information

Cited by 6SourcecodeScholar
2025

Sharpening Neural Implicit Functions with Frequency Consolidation Priors

AAAI 2025technical

Signed Distance Functions (SDFs) are vital implicit representations to represent high fidelity 3D surfaces. Current methods mainly leverage a neural network to learn an SDF from various supervisions including signed distances, 3D point clouds, or multi-view images. However, due to various reasons in…

2025

Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series

ICASSP 2025accepted

Time series analysis is vital for numerous applications, and transformers have become increasingly prominent in this domain. Leading methods customize the transformer architecture from NLP and CV, utilizing a patching technique to convert continuous signals into segments. Yet, time series data is un…

Cited by 0SourceScholar
2025

TopoCellGen: Generating Histopathology Cell Topology with a Diffusion Model

CVPR 2025poster

Accurately modeling multi-class cell topology is crucial in digital pathology, as it provides critical insights into tissue structure and pathology. The synthetic generation of cell topology enables realistic simulations of complex tissue environments, enhances downstream tasks by augmenting trainin…

2025

UltraVPR: Unsupervised Lightweight Rotation- Invariant Aerial Visual Place Recognition

RA-L 2025

Aerial Visual Place Recognition (VPR) is critical for Unmanned Aerial Vehicles (UAVs) localization, especially in environments with unstable or unavailable GPS signals. While neural network-based VPR methods have become mainstream, they face significant challenges on UAV platforms. Traditional CNN-b

Cited by 0SourcecodeScholar
2025

UniTR: A Unified Framework for Joint Representation Learning of Trajectories and Road Networks

AAAI 2025technical

Representation learning of urban spatial-temporal data is fundamental and critical, serving a wide range of intelligent applications. Given that road networks and trajectories are inherently interrelated, their joint representation learning can significantly enhance the accuracy and utility of these…

2025

Wasserstein-Regularized Conformal Prediction under General Distribution Shift

ICLR 2025poster

Conformal prediction yields a prediction set with guaranteed $1-\alpha$ coverage of the true target under the i.i.d. assumption, which can fail and lead to a gap between $1-\alpha$ and the actual coverage. Prior studies bound the gap using total variation distance, which cannot identify the gap cha…

Cited by 0SourcePDFScholar
2024

A Novel Wide-Area Multiobject Detection System with High-Probability Region Searching

ICRA 2024poster

In recent years, wide-area visual surveillance systems have been widely applied in various industrial and transportation scenarios. These systems, however, face significant challenges when implementing multi-object detection due to conflicts arising from the need for high-resolution imaging, efficie…

Cited by 5SourceScholar
2024

AerialVL: A Dataset, Baseline and Algorithm Framework for Aerial-Based Visual Localization With Reference Map

RA-L 2024

Visual localization plays an essential role in the autonomous flight of Unmanned Aerial Vehicles (UAVs) especially for the Global Navigation Satellite System (GNSS) denied environments. Existing aerial-based visual localization methods mainly focus on eliminating image variance between database map

Cited by 15SourceScholar
2024

An Iterative Associative Memory Model for Empathetic Response Generation

ACL 2024long

Empathetic response generation aims to comprehend the cognitive and emotional states in dialogue utterances and generate proper responses. Psychological theories posit that comprehending emotional and cognitive states necessitates iteratively capturing and understanding associated words across dialo…

2024

CTSM: Combining Trait and State Emotions for Empathetic Response Model

COLING 2024main

Empathetic response generation endeavors to empower dialogue systems to perceive speakers’ emotions and generate empathetic responses accordingly. Psychological research demonstrates that emotion, as an essential factor in empathy, encompasses trait emotions, which are static and context-independent…

2024

Detector Collapse: Backdooring Object Detection to Catastrophic Overload or Blindness in the Physical World

IJCAI 2024poster

Object detection tasks, crucial in safety-critical systems like autonomous driving, focus on pinpointing object locations. These detectors are known to be susceptible to backdoor attacks. However, existing backdoor techniques have primarily been adapted from classification tasks, overlooking deeper…

Cited by 13SourcePDFScholar
2024

Enhancing LLM’s Cognition via Structurization

NeurIPS 2024poster

When reading long-form text, human cognition is complex and structurized. While large language models (LLMs) process input contexts through a causal and sequential perspective, this approach can potentially limit their ability to handle intricate and complex inputs effectively. To enhance LLM’s cogn…

2024

Focus-Then-Decide: Segmentation-Assisted Reinforcement Learning

AAAI 2024technical

Visual Reinforcement Learning (RL) is a promising approach to achieve human-like intelligence. However, it currently faces challenges in learning efficiently within noisy environments. In contrast, humans can quickly identify task-relevant objects in distraction-filled surroundings by applying previ…

2024

GeoCluster: Enhancing Visual Place Recognition in Spatial Domain on Aerial Vehicle Platforms

RA-L 2024

Visual Place Recognition (VPR) is a critical technology for achieving robust long-term visual geo-localization. During the past few years, VPR research mainly focused on ground-based platforms in the street-level captured scenes with deep learning methods (e.g. NetVLAD, GeM), but little attention wa

Cited by 5SourceScholar
2024

INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection

ICLR 2024poster

Knowledge hallucination have raised widespread concerns for the security and reliability of deployed LLMs. Previous efforts in detecting hallucinations have been employed at logit-level uncertainty estimation or language-level self-consistency evaluation, where the semantic information is inevitably…

2024

Inferring Neural Signed Distance Functions by Overfitting on Single Noisy Point Clouds through Finetuning Data-Driven based Priors

NeurIPS 2024poster

It is important to estimate an accurate signed distance function (SDF) from a point cloud in many computer vision applications. The latest methods learn neural SDFs using either a data-driven based or an overfitting-based strategy. However, these two kinds of methods are with either poor generalizat…

2024

Learning Local Pattern Modularization for Point Cloud Reconstruction from Unseen Classes

ECCV 2024poster

"It is challenging to reconstruct 3D point clouds in unseen classes from single 2D images. Instead of object-centered coordinate system, current methods generalized global priors learned in seen classes to reconstruct 3D shapes from unseen classes in viewer-centered coordinate system. However, the r…

2024

Linear Uncertainty Quantification of Graphical Model Inference

NeurIPS 2024poster

Uncertainty Quantification (UQ) is vital for decision makers as it offers insights into the potential reliability of data and model, enabling more informed and risk-aware decision-making. Graphical models, capable of representing data with complex dependencies, are widely used across domains. Exist…

Cited by 0SourcePDFScholar
2024

MultiPull: Detailing Signed Distance Functions by Pulling Multi-Level Queries at Multi-Step

NeurIPS 2024poster

Reconstructing a continuous surface from a raw 3D point cloud is a challenging task. Latest methods employ supervised learning or pretrained priors to learn a signed distance function (SDF). However, neural networks tend to smooth local details due to the lack of ground truth signed distnaces or nor…

Cited by 8SourcePDFScholar
2024

OWL: A Large Language Model for IT Operations

ICLR 2024poster

With the rapid advancement of IT operations, managing and analyzing large data volumes efficiently for practical applications has become increasingly critical. Natural Language Processing (NLP) techniques have demonstrated remarkable capabilities in various tasks, including named entity recognition,…

2024

Rethinking Out-of-Distribution Detection on Imbalanced Data Distribution

NeurIPS 2024poster

Detecting and rejecting unknown out-of-distribution (OOD) samples is critical for deployed neural networks to void unreliable predictions. In real-world scenarios, however, the efficacy of existing OOD detection methods is often impeded by the inherent imbalance of in-distribution (ID) data, which c…

2024

SI-MIL: Taming Deep MIL for Self-Interpretability in Gigapixel Histopathology

CVPR 2024poster

Introducing interpretability and reasoning into Multiple Instance Learning (MIL) methods for Whole Slide Image (WSI) analysis is challenging given the complexity of gigapixel slides. Traditionally MIL interpretability is limited to identifying salient regions deemed pertinent for downstream tasks of…

2024

Semi-supervised Segmentation of Histopathology Images with Noise-Aware Topological Consistency

ECCV 2024poster

"In digital pathology, segmenting densely distributed objects like glands and nuclei is crucial for downstream analysis. Since detailed pixel-wise annotations are very time-consuming, we need semi-supervised segmentation methods that can learn from unlabeled images. Existing semi-supervised methods…

2024

Sketch and Refine: Towards Fast and Accurate Lane Detection

AAAI 2024technical

Lane detection is to determine the precise location and shape of lanes on the road. Despite efforts made by current methods, it remains a challenging task due to the complexity of real-world scenarios. Existing approaches, whether proposal-based or keypoint-based, suffer from depicting lanes effecti…

2024

Task-Agnostic Detector for Insertion-Based Backdoor Attacks

NAACL 2024findings

Textual backdoor attacks pose significant security threats. Current detection approaches, typically relying on intermediate feature representation or reconstructing potential triggers, are task-specific and less effective beyond sentence classification, struggling with tasks like question answering…

2024

Training for Stable Explanation for Free

NeurIPS 2024poster

To foster trust in machine learning models, explanations must be faithful and stable for consistent insights. Existing relevant works rely on the $\ell_p$ distance for stability assessment, which diverges from human perception. Besides, existing adversarial training (AT) associated with intensive co…

2024

TrojVLM: Backdoor Attack Against Vision Language Models

ECCV 2024poster

"The emergence of Vision Language Models (VLMs) is a significant advancement in integrating computer vision with Large Language Models (LLMs) to produce detailed text descriptions based on visual inputs, yet it introduces new security vulnerabilities. Unlike prior work that centered on single modali…

Cited by 13SourcePDFScholar
2023

Attention-Enhancing Backdoor Attacks Against BERT-based Models

EMNLP 2023long findings

Recent studies have revealed that Backdoor Attacks can threaten the safety of natural language processing (NLP) models. Investigating the strategies of backdoor attacks will help to understand the model's vulnerability. Most existing textual backdoor attacks focus on generating stealthy triggers or…

Cited by 0SourceScholar
2023

Category-Extensible Out-of-Distribution Detection via Hierarchical Context Descriptions

NeurIPS 2023poster

The key to OOD detection has two aspects: generalized feature representation and precise category description. Recently, vision-language models such as CLIP provide significant advances in both two issues, but constructing precise category descriptions is still in its infancy due to the absence of u…

2023

Coco-LIC: Continuous-Time Tightly-Coupled LiDAR-Inertial-Camera Odometry Using Non-Uniform B-Spline

RA-L 2023

In this letter, we propose an efficient continuous-time LiDAR-Inertial-Camera Odometry, utilizing non-uniform B-splines to tightly couple measurements from the LiDAR, IMU, and camera. In contrast to uniform B-spline-based continuous-time methods, our non-uniform B-spline approach offers significant

Cited by 39SourcecodeScholar
2023

Deep Anomaly Detection and Search via Reinforcement Learning (Student Abstract)

AAAI 2023technical

Semi-supervised anomaly detection is a data mining task which aims at learning features from partially-labeled datasets. We propose Deep Anomaly Detection and Search (DADS) with reinforcement learning. During the training process, the agent searches for possible anomalies in unlabeled dataset to enh…

Cited by 1SourcePDFScholar
2023

DeepMapping2: Self-Supervised Large-Scale LiDAR Map Optimization

CVPR 2023poster

LiDAR mapping is important yet challenging in self-driving and mobile robotics. To tackle such a global point cloud registration problem, DeepMapping converts the complex map estimation into a self-supervised training of simple deep networks. Despite its broad convergence range on small datasets, De…

2023

Denial-of-Service or Fine-Grained Control: Towards Flexible Model Poisoning Attacks on Federated Learning

IJCAI 2023poster

Federated learning (FL) is vulnerable to poisoning attacks, where adversaries corrupt the global aggregation results and cause denial-of-service (DoS). Unlike recent model poisoning attacks that optimize the amplitude of malicious perturbations along certain prescribed directions to cause DoS, we pr…

Cited by 17SourcePDFScholar
2023

End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge Distillation

AAAI 2023technical

Most existing Human-Object Interaction (HOI) Detection methods rely heavily on full annotations with predefined HOI categories, which is limited in diversity and costly to scale further. We aim at advancing zero-shot HOI detection to detect both seen and unseen HOIs simultaneously. The fundamental c…

2023

Enhancing Modality-Agnostic Representations via Meta-Learning for Brain Tumor Segmentation

ICCV 2023poster

In medical vision, different imaging modalities provide complementary information. However, in practice, not all modalities may be available during inference or even training. Previous approaches, e.g., knowledge distillation or image synthesis, often assume the availability of full modalities for a…

Cited by 20PDFScholar
2023

Enhancing Neural Topic Model with Multi-Level Supervisions from Seed Words

ACL 2023findings

Efforts have been made to apply topic seed words to improve the topic interpretability of topic models. However, due to the semantic diversity of natural language, supervisions from seed words could be ambiguous, making it hard to be incorporated into the current neural topic models. In this paper,…

Cited by 11SourcePDFScholar
2023

FoPro: Few-Shot Guided Robust Webly-Supervised Prototypical Learning

AAAI 2023technical

Recently, webly supervised learning (WSL) has been studied to leverage numerous and accessible data from the Internet. Most existing methods focus on learning noise-robust models from web images while neglecting the performance drop caused by the differences between web domain and real-world domain.…

2023

From Coarse to Fine: Hierarchical Pixel Integration for Lightweight Image Super-resolution

AAAI 2023technical

Image super-resolution (SR) serves as a fundamental tool for the processing and transmission of multimedia data. Recently, Transformer-based models have achieved competitive performances in image SR. They divide images into fixed-size patches and apply self-attention on these patches to model long-r…

2023

Geo-Localization With Transformer-Based 2D-3D Match Network

RA-L 2023

This letter presents a novel method for geographical localization by registering satellite maps with LiDAR point clouds. This method includes a Transformer-based 2D-3D matching network called D-GLSNet that directly matches the LiDAR point clouds and satellite images through end-to-end learning. With

Cited by 18SourceScholar
2023

Graph Signal Sampling for Inductive One-Bit Matrix Completion: a Closed-form Solution

ICLR 2023poster

Inductive one-bit matrix completion is motivated by modern applications such as recommender systems, where new users would appear at test stage with the ratings consisting of only ones and no zeros. We propose a unified graph signal sampling framework which enjoys the benefits of graph signal analys…

2023

GridPull: Towards Scalability in Learning Implicit Representations from 3D Point Clouds

ICCV 2023poster

Learning implicit representations has been a widely used solution for surface reconstruction from 3D point clouds. The latest methods infer a distance or occupancy field by overfitting a neural network on a single point cloud. However, these methods suffer from a slow inference due to the slow conve…

Cited by 26PDFcodeScholar
2023

Learning Probabilistic Topological Representations Using Discrete Morse Theory

ICLR 2023top-25%

Accurate delineation of fine-scale structures is a very important yet challenging problem. Existing methods use topological information as an additional training loss, but are ultimately making pixel-wise predictions. In this paper, we propose a novel deep learning based method to learn topological/…

Cited by 19SourcePDFScholar
2023

Learning to Segment from Noisy Annotations: A Spatial Correction Approach

ICLR 2023poster

Noisy labels can significantly affect the performance of deep neural networks (DNNs). In medical image segmentation tasks, annotations are error-prone due to the high demand in annotation time and in the annotators' expertise. Existing methods mostly tackle label noise in classification tasks. Their…

2023

Modelling of Tendon-Driven Continuum Robot Based on Constraint Analysis and Pseudo-Rigid Body Model

RA-L 2023

Quasi-static models of tendon-driven continuum robots (TDCR) require consideration of both the kinematic and static conditions simultaneously. While the Pseudo-Rigid Body (PRB-3R) model has been demonstrated to be efficient, existing works ignore the mechanical effect of the tendons such as elongati

Cited by 13SourceScholar
2023

Optimal Parameter and Neuron Pruning for Out-of-Distribution Detection

NeurIPS 2023poster

For a machine learning model deployed in real world scenarios, the ability of detecting out-of-distribution (OOD) samples is indispensable and challenging. Most existing OOD detection methods focused on exploring advanced training skills or training-free tricks to prevent the model from yielding ove…

Cited by 5SourcePDFScholar
2023

Topology-Aware Uncertainty for Image Segmentation

NeurIPS 2023poster

Segmentation of curvilinear structures such as vasculature and road networks is challenging due to relatively weak signals and complex geometry/topology. To facilitate and accelerate large scale annotation, one has to adopt semi-automatic approaches such as proofreading by experts. In this work, we…

2023

Topology-Guided Multi-Class Cell Context Generation for Digital Pathology

CVPR 2023poster

In digital pathology, the spatial context of cells is important for cell classification, cancer diagnosis and prognosis. To model such complex cell context, however, is challenging. Cells form different mixtures, lineages, clusters and holes. To model such structural patterns in a learnable fashion,…

Cited by 15SourcePDFScholar
2023

UniFusion: Unified Multi-View Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-View

ICCV 2023poster

Bird's eye view (BEV) representation is a new perception formulation for autonomous driving, which is based on spatial fusion. Further, temporal fusion is also introduced in BEV representation and gains great success. In this work, we propose a new method that unifies both spatial and temporal fusio…

Cited by 53PDFScholar
2023

Unsupervised Inference of Signed Distance Functions From Single Sparse Point Clouds Without Learning Priors

CVPR 2023poster

It is vital to infer signed distance functions (SDFs) from 3D point clouds. The latest methods rely on generalizing the priors learned from large scale supervision. However, the learned priors do not generalize well to various geometric variations that are unseen during training, especially for extr…

2022

A Manifold View of Adversarial Risk

AISTATS 2022poster

The adversarial risk of a machine learning model has been widely studied. Most previous works assume that the data lies in the whole ambient space. We propose to take a new angle and take the manifold assumption into consideration. Assuming data lies in a manifold, we investigate two new types of ad…

Cited by 4SourcePDFScholar
2022

Cycle Representation Learning for Inductive Relation Prediction

ICML 2022spotlight

In recent years, algebraic topology and its modern development, the theory of persistent homology, has shown great potential in graph representation learning. In this paper, based on the mathematics of algebraic topology, we propose a novel solution for inductive relation prediction, an important le…

2022

DIFNet: Boosting Visual Information Flow for Image Captioning

CVPR 2022poster

Current Image captioning (IC) methods predict textual words sequentially based on the input visual information from the visual feature extractor and the partially generated sentence information. However, for most cases, the partially generated sentence may dominate the target word prediction due to…

Cited by 62PDFScholar
2022

ECO-TR: Efficient Correspondences Finding via Coarse-to-Fine Refinement

ECCV 2022poster

"Abstract. Modeling sparse and dense image matching within a unified functional model has recently attracted increasing research interest. However, existing efforts mainly focus on improving matching accuracy while ignoring its efficiency, which is crucial for real-world applications. In this paper,…

2022

Learning Topological Interactions for Multi-Class Medical Image Segmentation

ECCV 2022poster

"Deep learning methods have achieved impressive performance for multi-class medical image segmentation. However, they are limited in their ability to encode topological interactions among different classes (e.g., containment and exclusion). These constraints naturally arise in biomedical images and…

2022

MDOE: A Spatiotemporal Event Representation Considering the Magnitude and Density of Events

RA-L 2022

Event-based sensors (e.g., DVS cameras) are capable of higher dynamic range, higher temporal resolution, lower time latency, and better power efficiency compared to conventional devices (e.g., RGB cameras). However, learning from these sensors remains challenging; event-based sensors output a stream

Cited by 4SourceScholar
2022

Neural Approximation of Graph Topological Features

NeurIPS 2022accept

Topological features based on persistent homology capture high-order structural information so as to augment graph neural network methods. However, computing extended persistent homology summaries remains slow for large and dense graphs and can be a serious bottleneck for the learning pipeline. Insp…

2022

Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 Lesions

ICASSP 2022accepted

Automatic segmentation of COVID-19 lesions is essential for computer-aided diagnosis. However, this task remains challenging because widely-used supervised based methods require large-scale annotated data that is difficult to obtain. Although an unsupervised method based on anomaly detection has sho…

Cited by 0SourceScholar
2022

PixelFolder: An Efficient Progressive Pixel Synthesis Network for Image Generation

ECCV 2022poster

"Pixel synthesis is a promising research paradigm for image generation, which can well exploit pixel-wise prior knowledge for generation. However, existing methods still suffer from excessive memory footprint and computation overhead. In this paper, we propose a progressive pixel synthesis network t…

2022

Resistance Training Using Prior Bias: Toward Unbiased Scene Graph Generation

AAAI 2022technical

Scene Graph Generation (SGG) aims to build a structured representation of a scene using objects and pairwise relationships, which benefits downstream tasks. However, current SGG methods usually suffer from sub-optimal scene graph generation because of the long-tailed distribution of training data. T…

2022

SeqTR: A Simple Yet Universal Network for Visual Grounding

ECCV 2022poster

"In this paper, we propose a simple yet universal network termed SeqTR for visual grounding tasks, e.g., phrase localization, referring expression comprehension (REC) and segmentation (RES). The canonical paradigms for visual grounding often require substantial expertise in designing network archite…

2022

Stability of SGD: Tightness analysis and improved bounds

UAI 2022poster

Stochastic Gradient Descent (SGD) based methods have been widely used for training large-scale machine learning models that also generalize well in practice. Several explanations have been offered for this generalization performance, a prominent one being algorithmic stability Hardt et al [2016]. Ho…

Cited by 41SourcePDFScholar
2022

Temporal Context Matters: Enhancing Single Image Prediction With Disease Progression Representations

CVPR 2022oral

Clinical outcome or severity prediction from medical images has largely focused on learning representations from single-timepoint or snapshot scans. It has been shown that disease progression can be better characterized by temporal imaging. We therefore hypothesized that outcome predictions can be i…

Cited by 21PDFScholar
2022

Trigger Hunting with a Topological Prior for Trojan Detection

ICLR 2022poster

Despite their success and popularity, deep neural networks (DNNs) are vulnerable when facing backdoor attacks. This impedes their wider adoption, especially in mission critical applications. This paper tackles the problem of Trojan detection, namely, identifying Trojaned models – models trained with…

2021

Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation

ICML 2021spotlight

User interests are usually dynamic in the real world, which poses both theoretical and practical challenges for learning accurate preferences from rich behavior data. Among existing user behavior modeling solutions, attention networks are widely adopted for its effectiveness and relative simplicity.…

2021

Learning with Feature-Dependent Label Noise: A Progressive Approach

ICLR 2021spotlight

Label noise is frequently observed in real-world large-scale datasets. The noise is introduced due to a variety of reasons; it is heterogeneous and feature-dependent. Most existing approaches to handling noisy labels fall into two categories: they either assume an ideal feature-independent noise, or…

2021

Link Prediction with Persistent Homology: An Interactive View

ICML 2021spotlight

Link prediction is an important learning task for graph-structured data. In this paper, we propose a novel topological approach to characterize interactions between two nodes. Our topological feature, based on the extended persistent homology, encodes rich structural information regarding the multi-…

2021

Localization in the Crowd with Topological Constraints

AAAI 2021technical

We address the problem of crowd localization, i.e., the prediction of dots corresponding to people in a crowded scene. Due to various challenges, a localization method is prone to spatial semantic errors, i.e., predicting multiple dots within a same person or collapsing multiple dots in a cluttered…

2021

Multi-Class Cell Detection Using Spatial Context Representation

ICCV 2021poster

In digital pathology, both detection and classification of cells are important for automatic diagnostic and prognostic tasks. Classifying cells into subtypes, such as tumor cells, lymphocytes or stromal cells is particularly challenging. Existing methods focus on morphological appearance of individu…

Cited by 42PDFcodeScholar
2021

Scalable and Explainable 1-Bit Matrix Completion via Graph Signal Learning

AAAI 2021technical

One-bit matrix completion is an important class of positive-unlabeled (PU) learning problems where the observations consist of only positive examples, e.g., in top-N recommender systems. For the first time, we show that 1-bit matrix completion can be formulated as the problem of recovering clean gra…

2021

Topological Detection of Trojaned Neural Networks

NeurIPS 2021poster

Deep neural networks are known to have security issues. One particular threat is the Trojan attack. It occurs when the attackers stealthily manipulate the model's behavior through Trojaned training samples, which can later be exploited. Guided by basic neuroscientific principles, we discover subtle…

Cited by 58SourcePDFScholar
2021

Topology-Aware Segmentation Using Discrete Morse Theory

ICLR 2021spotlight

In the segmentation of fine-scale structures from natural and biomedical images, per-pixel accuracy is not the only metric of concern. Topological correctness, such as vessel connectivity and membrane closure, is crucial for downstream analysis tasks. In this paper, we propose a new approach to trai…

Cited by 112SourcePDFScholar
2021

Unsupervised Learning of Fine Structure Generation for 3D Point Clouds by 2D Projections Matching

ICCV 2021poster

Learning to generate 3D point clouds without 3D supervision is an important but challenging problem. Current solutions leverage various differentiable renderers to project the generated 3D point clouds onto a 2D image plane, and train deep neural networks using the per-pixel difference with 2D groun…

Cited by 47PDFcodeScholar
2020

A Topological Filter for Learning with Label Noise

NeurIPS 2020poster

Noisy labels can impair the performance of deep neural networks. To tackle this problem, in this paper, we propose a new method for filtering label noise. Unlike most existing methods relying on the posterior probability of a noisy classifier, we focus on the much richer spatial behavior of data in…

2020

DRWR: A Differentiable Renderer without Rendering for Unsupervised 3D Structure Learning from Silhouette Images

ICML 2020poster

Differentiable renderers have been used successfully for unsupervised 3D structure learning from 2D images because they can bridge the gap between 3D and 2D. To optimize 3D shape parameters, current renderers rely on pixel-wise losses between rendered images of 3D reconstructions and ground truth im…

Cited by 61SourcePDFScholar
2020

Error-Bounded Correction of Noisy Labels

ICML 2020poster

To collect large scale annotated data, it is inevitable to introduce label noise, i.e., incorrect class labels. To be robust against label noise, many successful methods rely on the noisy classifiers (i.e., models trained on the noisy training data) to determine whether a label is trustworthy. Howev…

2020

Learn distributed GAN with Temporary Discriminators

ECCV 2020poster

In this work, we propose a method for training distributed GAN with sequential temporary discriminators. Our proposed method tackles the challenge of training GAN in the federated learning manner: How to update the generator with a flow of temporary discriminators? We apply our proposed method to le…

2020

Reliable Weighted Optimal Transport for Unsupervised Domain Adaptation

CVPR 2020poster

Recently, extensive researches have been proposed to address the UDA problem, which aims to learn transferrable models for the unlabeled target domain. Among them, the optimal transport is a promising metric to align the representations of the source and target domains. However, most existing works…

Cited by 181PDFScholar
2020

Selective Transfer With Reinforced Transfer Network for Partial Domain Adaptation

CVPR 2020poster

One crucial aspect of partial domain adaptation (PDA) is how to select the relevant source samples in the shared classes for knowledge transfer. Previous PDA methods tackle this problem by re-weighting the source samples based on their high-level information (deep features). However, since the domai…

Cited by 80PDFScholar
2020

Synthetic Learning: Learn From Distributed Asynchronized Discriminator GAN Without Sharing Medical Image Data

CVPR 2020poster

In this paper, we propose a data privacy-preserving and communication efficient distributed GAN learning framework named Distributed Asynchronized Discriminator GAN (AsynDGAN). Our proposed framework aims to train a central generator learns from distributed discriminator, and use the generated synth…

Cited by 113PDFcodeScholar
2019

A Topological Regularizer for Classifiers via Persistent Homology

AISTATS 2019poster

Regularization plays a crucial role in supervised learning. Most existing methods enforce a global regularization in a structure agnostic manner. In this paper, we initiate a new direction and propose to enforce the structural simplicity of the classification boundary by regularizing over its topolo…

Cited by 156SourcePDFScholar
2019

ClusterNet: Deep Hierarchical Cluster Network With Rigorously Rotation-Invariant Representation for Point Cloud Analysis

CVPR 2019poster

Current neural networks for 3D object recognition are vulnerable to 3D rotation. Existing works mostly rely on massive amounts of rotation-augmented data to alleviate the problem, which lacks solid guarantee of the 3D rotation invariance. In this paper, we address the issue by introducing a novel po…

Cited by 217PDFScholar
2018

Dynamic Simulation of Planetary Rovers with Terrain Property Mapping

ICRA 2018poster

Simulation of planetary rovers moving on complex terrains is critical for Mars exploration. Equivalent stiffness is proposed and used to characterize the pressure-sinkage property of terrain, while friction angle to characterize the shearing property. Terramechanics model for calculating forces betw…

Cited by 5SourceScholar
2017

Composing Tree Graphical Models with Persistent Homology Features for Clustering Mixed-Type Data

ICML 2017poster

Clustering data with both continuous and discrete attributes is a challenging task. Existing methods lack a principled probabilistic formulation. In this paper, we propose a clustering method based on a tree-structured graphical model to describe the generation process of mixed-type data. Our tree-s…

Cited by 25SourcePDFScholar
2017

Mixture-Rank Matrix Approximation for Collaborative Filtering

NeurIPS 2017poster

Low-rank matrix approximation (LRMA) methods have achieved excellent accuracy among today's collaborative filtering (CF) methods. In existing LRMA methods, the rank of user/item feature matrices is typically fixed, i.e., the same rank is adopted to describe all users/items. However, our studies show…

Cited by 42SourcePDFScholar
2016

Low-Rank Matrix Approximation with Stability

ICML 2016poster

Low-rank matrix approximation has been widely adopted in machine learning applications with sparse data, such as recommender systems. However, the sparsity of the data, incomplete and noisy, introduces challenges to the algorithm stability – small changes in the training data may significantly chang…

2015

Kinodynamic motion planning with Space-Time Exploration Guided Heuristic Search for car-like robots in dynamic environments

IROS 2015poster

The Space Exploration Guided Heuristic Search (SEHS) method solves the motion planning problem, especially for car-like robots, in two steps: a circle-based space exploration in the workspace followed by a circle-guided heuristic search in the configuration space. This paper extends this approach fo…

Cited by 28SourceScholar