← Search

Bo Dong

37 accepted papers

2026

Dynamic Weight Adaptation in Spiking Neural Networks Inspired by Biological Homeostasis

AAAI 2026technical

Homeostatic mechanisms play a crucial role in maintaining optimal functionality within the neural circuits of the brain. By regulating physiological and biochemical processes, these mechanisms ensure the stability of an organism’s internal environment, enabling it to better adapt to external changes

Cited by 0SourcePDFScholar
2026

Enhancing Pre-training Data Detection in LLMs Through Discriminative and Symmetric Prefix Selection

AAAI 2026technical

The rapid development of large language models (LLMs) has relied on access to high-quality, large-scale datasets, yet growing concerns around data privacy and security have spurred substantial research into pre-training data detection. While state-of-the-art (SOTA) methods such as RECALL and CON-REC

Cited by 0SourcePDFScholar
2026

Generalist Graph Anomaly Detection via Prototype-Based Distillation

ICML 2026poster

Driven by the pressing demand for graph anomaly detection (GAD) in high-stakes domains, the generalist GAD paradigm, which trains a single detector transferable across new graphs, has recently gained growing attention. However, existing methods often rely on scarce and costly annotations for trainin…

Cited by 0SourceScholar
2026

PolarDepth: Monocular Transparent Object Depth from Polar-Physics Priors

ICML 2026poster

Depth estimation for transparent objects remains a fundamental challenge, as RGB-based cues often fail in regions affected by refraction and light transmission. Polarization provides physically grounded information related to surface orientation and material properties, offering reliable geometric c…

Cited by 0SourceScholar
2026

Scope Delineation Before Localization: A Two-Stage Framework for Enhancing Failure Attribution in Multi-Agent Systems

AAAI 2026technical

Large language models (LLMs) are seeing growing adoption in multi-agent systems. In these systems, efficient failure attribution is critical for ensuring robustness and interpretability. Current LLM-based attribution methods often face challenges with lengthy logs and lacking expert knowledge. Drawi

Cited by 0SourcePDFScholar
2026

View-on-Graph: Zero-Shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs

AAAI 2026technical

3D visual grounding (3DVG) identifies objects in 3D scenes from language descriptions. Existing zero-shot approaches leverage 2D vision–language models (VLMs) by converting 3D spatial information (SI) into forms amenable to VLM processing, typically as composite inputs such as specified-view renderi

Cited by 0SourcePDFScholar
2025

Fully Autonomous Neuromorphic Navigation and Dynamic Obstacle Avoidance

NeurIPS 2025spotlight

Unmanned aerial vehicles could accurately accomplish complex navigation and obstacle avoidance tasks under external control. However, enabling unmanned aerial vehicles (UAVs) to rely solely on onboard computation and sensing for real-time navigation and dynamic obstacle avoidance remains a significa…

Cited by 0SourceScholar
2025

Out-of-Distribution Generalization on Graphs via Progressive Inference

AAAI 2025technical

The development and evaluation of graph neural networks (GNNs) generally follow the independent and identically distributed (i.i.d.) assumption. Yet this assumption is often untenable in practice due to the uncontrollable data generation mechanism. In particular, when the data distribution shows a s…

2025

Revisiting Graph Contrastive Learning on Anomaly Detection: A Structural Imbalance Perspective

AAAI 2025technical

The superiority of graph contrastive learning (GCL) has prompted its application to anomaly detection tasks for more powerful risk warning systems. Unfortunately, existing GCL-based models tend to excessively prioritize overall detection performance while neglecting robustness to structural imbalanc…

2025

Separating the Wheat from the Chaff: Spatio-Temporal Transformer with View-interweaved Attention for Photon-Efficient Depth Sensing

AAAI 2025technical

Time-resolved imaging is an emerging sensing modality that has been shown to enable advanced applications, including remote sensing, fluorescence lifetime imaging, and even non-line-of-sight sensing. Single-photon avalanche diodes (SPADs) outperform relevant time-resolved imaging technologies thanks…

Cited by 0SourcePDFScholar
2025

UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback

NeurIPS 2025poster

Relighting is a crucial task with both practical demand and artistic value, and recent diffusion models have shown strong potential by enabling rich and controllable lighting effects. However, as they are typically optimized in semantic latent space, where proximity does not guarantee physical corre…

Cited by 0SourcecodeScholar
2025

VERO: Verification and Zero-Shot Feedback Acquisition for Few-Shot Multimodal Aspect-Level Sentiment Classification

AAAI 2025technical

Deep learning approaches for multimodal aspect-level sentiment classification (MALSC) often require extensive data, which is costly and time-consuming to obtain. To mitigate this, current methods typically fine-tune small-scale pretrained models like BERT and BART with few-shot examples. While these…

2024

Apprenticeship-Inspired Elegance: Synergistic Knowledge Distillation Empowers Spiking Neural Networks for Efficient Single-Eye Emotion Recognition

IJCAI 2024poster

We introduce a novel multimodality synergistic knowledge distillation scheme tailored for efficient single-eye motion recognition tasks. This method allows a lightweight, unimodal student spiking neural network (SNN) to extract rich knowledge from an event-frame multimodal teacher network. The core…

Cited by 1SourcePDFScholar
2024

Efficient Federated Multi-View Clustering with Integrated Matrix Factorization and K-Means

IJCAI 2024poster

Multi-view clustering is a popular unsupervised multi-view learning method. Real-world multi-view data are often distributed across multiple entities, presenting a challenge for performing multi-view clustering. Federated learning provides a solution by enabling multiple entities to collaboratively…

Cited by 1SourcePDFScholar
2024

Estimating Noisy Class Posterior with Part-level Labels for Noisy Label Learning

CVPR 2024poster

In noisy label learning estimating noisy class posteriors plays a fundamental role for developing consistent classifiers as it forms the basis for estimating clean class posteriors and the transition matrix. Existing methods typically learn noisy class posteriors by training a classification model w…

2024

Federated Multi-View Clustering via Tensor Factorization

IJCAI 2024poster

Multi-view clustering is an effective method to process massive unlabeled multi-view data. Since data of different views may be collected and held by different parties, it becomes impractical to train a multi-view clustering model in a centralized way, for the sake of privacy. However, federated mul…

Cited by 1SourcePDFScholar
2024

Partial Multi-View Clustering via Self-Supervised Network

AAAI 2024technical

Partial multi-view clustering is a challenging and practical research problem for data analysis in real-world applications, due to the potential data missing issue in different views. However, most existing methods have not fully explored the correlation information among various incomplete views. I…

Cited by 7SourcePDFScholar
2024

RR-PU: A Synergistic Two-Stage Positive and Unlabeled Learning Framework for Robust Tax Evasion Detection

AAAI 2024technical

Tax evasion, an unlawful practice in which taxpayers deliberately conceal information to avoid paying tax liabilities, poses significant challenges for tax authorities. Effective tax evasion detection is critical for assisting tax authorities in mitigating tax revenue loss. Recently, machine-learnin…

Cited by 4SourcePDFScholar
2024

The Evidence Contraction Issue in Deep Evidential Regression: Discussion and Solution

AAAI 2024technical

Deep Evidential Regression (DER) places a prior on the original Gaussian likelihood and treats learning as an evidence acquisition process to quantify uncertainty. For the validity of the evidence theory, DER requires specialized activation functions to ensure that the prior parameters remain non-ne…

2023

Compressing Context to Enhance Inference Efficiency of Large Language Models

EMNLP 2023long main

Large language models (LLMs) achieved remarkable performance across various tasks. However, they face challenges in managing long documents and extended conversations, due to significantly increased computational requirements, both in memory and inference time, and potential context truncation when…

Cited by 0SourcecodeScholar
2023

Dichotomous Image Segmentation with Frequency Priors

IJCAI 2023poster

Dichotomous image segmentation (DIS) has a wide range of real-world applications and gained increasing research attention in recent years. In this paper, we propose to tackle DIS with informative frequency priors. Our model, called FP-DIS, stems from the fact that prior knowledge in the frequency do…

2023

Multi-view Spectral Polarization Propagation for Video Glass Segmentation

ICCV 2023poster

In this paper, we present the first polarization-guided video glass segmentation propagation solution (PGVS-Net) that can robustly and coherently propagate glass segmentation in RGB-P video sequences. By leveraging spatiotemporal polarization and color information, our method combines multi-view pol…

Cited by 8PDFScholar
2023

NerCo: A Contrastive Learning Based Two-Stage Chinese NER Method

IJCAI 2023poster

Sequence labeling serves as the most commonly used scheme for Chinese named entity recognition(NER). However, traditional sequence labeling methods classify tokens within an entity into different classes according to their positions. As a result, different tokens in the same entity may be learned wi…

2023

Single Depth-image 3D Reflection Symmetry and Shape Prediction

ICCV 2023poster

In this paper, we present Iterative Symmetry Completion Network (ISCNet), a single depth-image shape completion method that exploits reflective symmetry cues to obtain more detailed shapes. The efficacy of single depth-image shape completion methods is often sensitive to the accuracy of the symmetry…

Cited by 7PDFScholar
2022

All You Need Is RAW: Defending against Adversarial Attacks with Camera Image Pipelines

ECCV 2022poster

"Existing neural networks for computer vision tasks are vulnerable to adversarial attacks: adding imperceptible perturbations to the input images can fool these models to make a false prediction on an image that was correctly predicted without the perturbation. Various defense methods have proposed…

Cited by 11SourcePDFScholar
2022

Biologically Inspired Dynamic Thresholds for Spiking Neural Networks

NeurIPS 2022accept

The dynamic membrane potential threshold, as one of the essential properties of a biological neuron, is a spontaneous regulation mechanism that maintains neuronal homeostasis, i.e., the constant overall spiking firing rate of a neuron. As such, the neuron firing rate is regulated by a dynamic spikin…

Cited by 35SourcePDFScholar
2022

Glass Segmentation Using Intensity and Spectral Polarization Cues

CVPR 2022poster

Transparent and semi-transparent materials pose significant challenges for existing scene understanding and segmentation algorithms due to their lack of RGB texture which impedes the extraction of meaningful features. In this work, we exploit that the light-matter interactions on glass materials pro…

Cited by 93PDFScholar
2022

Regularized Modal Regression on Markov-Dependent Observations: A Theoretical Assessment

AAAI 2022technical

Modal regression, a widely used regression protocol, has been extensively investigated in statistical and machine learning communities due to its robustness to outlier and heavy-tailed noises. Understanding modal regression's theoretical behavior can be fundamental in learning theory. Despite signif…

Cited by 1SourcePDFScholar
2022

Spiking Transformers for Event-Based Single Object Tracking

CVPR 2022poster

Event-based cameras bring a unique capability to tracking, being able to function in challenging real-world conditions as a direct result of their high temporal resolution and high dynamic range. These imagers capture events asynchronously that encode rich temporal and spatial information. However,…

Cited by 190PDFScholar
2022

TNTC: Two-Stream Network with Transformer-Based Complementarity for Gait-Based Emotion Recognition

ICASSP 2022accepted

Recognizing the human emotion automatically from visual characteristics plays a vital role in many intelligent applications. Recently, gait-based emotion recognition, especially gait skeletons-based characteristic, has attracted much attention, while many available methods have been proposed gradual…

Cited by 0SourceScholar
2021

Object Tracking by Jointly Exploiting Frame and Event Domain

ICCV 2021poster

Inspired by the complementarity between conventional frame-based and bio-inspired event-based cameras, we propose a multi-modal based approach to fuse visual cues from the frame- and event-domain to enhance the single object tracking performance, especially in degraded conditions (e.g., scenes with…

Cited by 112PDFScholar