← Search

Baocai Yin

48 accepted papers

2026

AlignTrack: Top-Down Spatiotemporal Resolution Alignment for RGB-Event Visual Tracking

AAAI 2026technical

Most existing RGB-Event trackers rely on strictly aligned datasets, overlooking the asynchronous spatio-temporal resolutions common in real-world scenarios. This methodological limitation impedes effective RGB-Event feature alignment and ultimately degrades tracking performance. To overcome this li

Cited by 0SourcePDFScholar
2026

BiOTPrompt: Bidirectional Optimal Transport Guided Prompting for Disease Evolution-aware Radiology Report Generation

CVPR 2026

Radiology report generation (RRG) aims to automatically describe medical images via free-text reports. In clinical practice, comparing current and prior chest X-rays is essential for assessing disease progression, motivating the development of longitudinal RRG methods. However, most existing approac

Cited by 0SourcecodeScholar
2026

Binary-Gaussian: Compact and Progressive Representation for 3D Gaussian Segmentation

AAAI 2026technical

3D Gaussian Splatting (3D-GS) has emerged as an efficient 3D representation and a promising foundation for semantic tasks like segmentation. However, existing 3D-GS-based segmentation methods typically rely on high-dimensional category features, which introduce substantial memory overhead. Moreover,

Cited by 0SourcePDFScholar
2026

CompEvent: Complex-valued Event-RGB Fusion for Low-light Video Enhancement and Deblurring

AAAI 2026technical

Low-light video deblurring poses significant challenges in applications like nighttime surveillance and autonomous driving due to dim lighting and long exposures. While event cameras offer potential solutions with superior low-light sensitivity and high temporal resolution, existing fusion methods t

Cited by 0SourcePDFScholar
2026

E-MaT:Event-oriented Mamba for Egocentric Point Tracking

AAAI 2026technical

Egocentric point tracking aims to localize points on object surfaces from a first-person perspective and serves as a critical step toward embodied intelligence. Recent methods rely on video input, tracking query points through feature matching across consecutive frames. However, these methods strug

Cited by 0SourcePDFScholar
2026

GeoEvo: Identity-Aware Potential Game with Geometric Evolution for Personalized Multimodal Federated Learning

ICML 2026poster

We reconceptualize Personalized Multimodal Federated Learning (PMFL) by treating missing modalities as intrinsic structural identities that constrain each client to a distinct Riemannian submanifold, rather than deficiencies to be compensated. To resolve the tension between identity preservation and…

Cited by 0SourceScholar
2026

MARE: Multimodal Analogical Reasoning for Disease Evolution-Aware Radiology Report Generation

AAAI 2026technical

Radiology report generation from longitudinal medical data is critical for assessing disease progression and automating diagnostic workflows. While recent methods incorporate longitudinal information, they primarily rely on multimodal feature fusion, with limited capacity for explicit disease evolut

Cited by 0SourcePDFScholar
2026

TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment

CVPR 2026

Tables are pervasive in diverse documents, making table recognition (TR) a fundamental task in document analysis. Existing modular TR pipelines separately model table structure and content, leading to suboptimal integration and complex workflows.End-to-end approaches rely heavily on large-scale TR d

Cited by 0SourcecodeScholar
2026

TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation

ICRA 2026poster

Vision-Language Navigation (VLN) presents a unique challenge for Large Vision-Language Models (VLMs) due to their inherent architectural mismatch: VLMs are primarily pretrained on static, disembodied vision-language tasks, which fundamentally clash with the dynamic, embodied, and spatially-structure…

2025

Col-OLHTR: A Novel Framework for Multimodal Online Handwritten Text Recognition

ICASSP 2025accepted

Online Handwritten Text Recognition (OLHTR) has gained considerable attention for its diverse range of applications. Current approaches usually treat OLHTR as a sequence recognition task, employing either a single trajectory or image encoder, or multi-stream encoders, combined with a CTC or attentio…

Cited by 0SourceScholar
2025

Efficient Fine-tuning Strategies for Enhancing Face Recognition Performance in Challenging Scenarios

ICASSP 2025accepted

Face recognition plays a crucial role in human life, prompting numerous excellent research efforts. However, face recognition in real-world applications presents various scenarios such as occluded, overexposed and near-infrared face recognition. Due to domain discrepancy and a lack of large-scale tr…

Cited by 0SourceScholar
2025

Exploring Historical Information for RGBE Visual Tracking with Mamba

CVPR 2025poster

Combining the advantages of conventional and event cameras for robust visual tracing has drawn extensive interest. However, existing tracking approaches heavily engage in complex cross-modal fusion modules, leading to higher computational complexity and training challenges. Besides, these methods ge…

Cited by 0SourcePDFScholar
2025

HC-LLM: Historical-Constrained Large Language Models for Radiology Report Generation

AAAI 2025technical

Radiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient follow-ups. Traditional methods usually struggle with long sequence dependencies when incorporating historical informatio…

2025

MaskPrompt: Open-Vocabulary Affordance Segmentation with Object Shape Mask Prompts

AAAI 2025technical

Affordance refers to the interactable functional properties of an object, and affordance segmentation aims to pixel-level segment the object functional parts in a given image, which is crucial for various interactive vision tasks. Existing methods address the affordance segmentation problem by utili…

Cited by 0SourcePDFScholar
2025

RFL: Simplifying Chemical Structure Recognition with Ring-Free Language

AAAI 2025technical

The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional structures of molecules, particularly those with rings and multiple branches, present significant challenges for current…

2025

Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning

CVPR 2025poster

Generalized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires consistent visual-semantic alignment. Existing approaches fine-tune the visual backbone by seen-class data to obtain sema…

Cited by 0SourcePDFScholar
2024

1DFormer: A Transformer Architecture Learning 1D Landmark Representations for Facial Landmark Tracking

IJCAI 2024poster

Recently, heatmap regression methods based on 1D landmark representations have shown prominent performance on locating facial landmarks. However, previous methods ignored to make deep explorations on the good potentials of 1D landmark representations for sequential and structural modeling of multi…

Cited by 0SourcePDFScholar
2024

AutoFGNN: A Framework for Extracting All Frequency Information from Large-Scale Graphs

ICASSP 2024accepted

As a powerful model for deep learning on graph-structured data, the scalability limitation of Graph Neural Networks (GNNs) are receiving increasing attention. To tackle this limitation, two categories of scalable GNNs have been proposed: sampling-based and model simplification methods. However, samp…

Cited by 0SourceScholar
2024

Graph Neural Networks with Soft Association between Topology and Attribute

AAAI 2024technical

Graph Neural Networks (GNNs) have shown great performance in learning representations for graph-structured data. However, recent studies have found that the interference between topology and attribute can lead to distorted node representations. Most GNNs are designed based on homophily assumptions,…

2024

Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph Reasoning

NeurIPS 2024poster

Temporal Knowledge Graph Reasoning (TKGR) is the process of utilizing temporal information to capture complex relations within a Temporal Knowledge Graph (TKG) to infer new knowledge. Conventional methods in TKGR typically depend on deep learning algorithms or temporal logical rules. However, deep l…

2024

NAMER: Non-Autoregressive Modeling for Handwritten Mathematical Expression Recognition

ECCV 2024poster

"Recently, Handwritten Mathematical Expression Recognition (HMER) has gained considerable attention in pattern recognition for its diverse applications in document understanding. Current methods typically approach HMER as an image-to-sequence generation task within an autoregressive (AR) encoder-dec…

Cited by 2SourcePDFScholar
2024

TAU: Trajectory Data Augmentation with Uncertainty for Next POI Recommendation

AAAI 2024technical

Next Point-of-Interest (POI) recommendation has been proven effective at utilizing sparse, intricate spatial-temporal trajectory data to recommend subsequent POIs to users. While existing methods commonly alleviate the problem of data sparsity by integrating spatial-temporal context information, POI…

Cited by 12SourcePDFScholar
2023

Center Focusing Network for Real-Time LiDAR Panoptic Segmentation

CVPR 2023poster

LiDAR panoptic segmentation facilitates an autonomous vehicle to comprehensively understand the surrounding objects and scenes and is required to run in real time. The recent proposal-free methods accelerate the algorithm, but their effectiveness and efficiency are still limited owing to the difficu…

2023

Discriminative Active Learning for Robotic Grasping in Cluttered Scene

RA-L 2023

Robotic grasping is a challenging task due to the diversity of object shapes. A sufficiently labeled dataset is essential for the grasp pose detection methods based on deep learning. However, data annotation is a costly procedure. Active learning aims to mitigate the greedy need for massive labeled

Cited by 18SourceScholar
2023

Frame-Event Alignment and Fusion Network for High Frame Rate Tracking

CVPR 2023poster

Most existing RGB-based trackers target low frame rate benchmarks of around 30 frames per second. This setting restricts the tracker's functionality in the real world, especially for fast motion. Event-based cameras as bioinspired sensors provide considerable potential for high frame rate tracking d…

Cited by 44SourcePDFScholar
2023

Learning Interaction Regions and Motion Trajectories Simultaneously From Egocentric Demonstration Videos

RA-L 2023

Learning to interact with objects is significant for robots to integrate into human environments. When the interaction semantic is definite, manually guiding the manipulator is a commonly used method to teach robots how to interact with objects. However, the learning results are robot-dependent beca

Cited by 9SourceScholar
2023

Referring Image Segmentation Using Text Supervision

ICCV 2023poster

Existing Referring Image Segmentation (RIS) methods typically require expensive pixel-level or box-level annotations for supervision. In this paper, we observe that the referring texts used in RIS already provide sufficient information to localize the target object. Hence, we propose a novel weakly-…

Cited by 34PDFcodeScholar
2023

Single Depth-image 3D Reflection Symmetry and Shape Prediction

ICCV 2023poster

In this paper, we present Iterative Symmetry Completion Network (ISCNet), a single depth-image shape completion method that exploits reflective symmetry cues to obtain more detailed shapes. The efficacy of single depth-image shape completion methods is often sensitive to the accuracy of the symmetry…

Cited by 7PDFScholar
2023

The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And Recognition

ICASSP 2023accepted

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two trac…

Cited by 0SourceScholar
2022

A New Perspective on the Effects of Spectrum in Graph Neural Networks

ICML 2022spotlight

Many improvements on GNNs can be deemed as operations on the spectrum of the underlying graph matrix, which motivates us to directly study the characteristics of the spectrum and their effects on GNN performance. By generalizing most existing GNN architectures, we show that the correlation issue cau…

2022

Bi-Directional Object-Context Prioritization Learning for Saliency Ranking

CVPR 2022poster

The saliency ranking task is recently proposed to study the visual behavior that humans would typically shift their attention over different objects of a scene based on their degrees of saliency. Existing approaches focus on learning either object-object or object-scene relations. Such a strategy fo…

Cited by 38PDFcodeScholar
2022

Biologically Inspired Dynamic Thresholds for Spiking Neural Networks

NeurIPS 2022accept

The dynamic membrane potential threshold, as one of the essential properties of a biological neuron, is a spontaneous regulation mechanism that maintains neuronal homeostasis, i.e., the constant overall spiking firing rate of a neuron. As such, the neuron firing rate is regulated by a dynamic spikin…

Cited by 35SourcePDFScholar
2022

Recommending Fine-Grained Tool Consistent With Common Sense Knowledge for Robot

RA-L 2022

When robots carry out task, selecting an appropriate tool is necessary. The current research ignores the fine-grained characteristic of tasks, and mainly focuses on whether the task can be completed. Little consideration is paid for the object being manipulated, which affects the task completion qua

Cited by 2SourceScholar
2022

Spiking Transformers for Event-Based Single Object Tracking

CVPR 2022poster

Event-based cameras bring a unique capability to tracking, being able to function in challenging real-world conditions as a direct result of their high temporal resolution and high dynamic range. These imagers capture events asynchronously that encode rich temporal and spatial information. However,…

Cited by 190PDFScholar
2021

A Vision-based Irregular Obstacle Avoidance Framework via Deep Reinforcement Learning

IROS 2021poster

Deep reinforcement learning has achieved great success in laser-based collision avoidance work because the laser can sense accurate depth information without too much redundant data, which can maintain the robustness of the algorithm when it is migrated from the simulation environment to the real wo…

Cited by 20SourceScholar
2021

Hierarchical Graph Convolution Network for Traffic Forecasting

AAAI 2021technical

Traffic forecasting is attracting considerable interest due to its widespread application in intelligent transportation systems. Given the complex and dynamic traffic data, many methods focus on how to establish a spatial-temporal model to express the non-stationary traffic patterns. Recently, the l…

2021

Object Tracking by Jointly Exploiting Frame and Event Domain

ICCV 2021poster

Inspired by the complementarity between conventional frame-based and bio-inspired event-based cameras, we propose a multi-modal based approach to fuse visual cues from the frame- and event-domain to enhance the single object tracking performance, especially in degraded conditions (e.g., scenes with…

Cited by 112PDFScholar
2019

Double Nuclear Norm Based Low Rank Representation on Grassmann Manifolds for Clustering

CVPR 2019poster

Unsupervised clustering for high-dimension data (such as imageset or video) is a hard issue in data processing and data mining area since these data always lie on a manifold (such as Grassmann manifold). Inspired of Low Rank representation theory, researchers proposed a series of effective clusterin…

Cited by 25PDFScholar
2017

Learning Uncertain Convolutional Features for Accurate Saliency Detection

ICCV 2017poster

Deep convolutional neural networks (CNNs) have delivered superior performance in many computer vision tasks. In this paper, we propose a novel deep fully convolutional network model for accurate salient object detection. The key contribution of this work is to learn deep uncertain convolutional feat…

Cited by 447PDFScholar
2017

Learning to Detect Salient Objects With Image-Level Supervision

CVPR 2017poster

Deep Neural Networks (DNNs) have substantially improved the state-of-the-art in salient object detection. However, training DNNs requires costly pixel-level annotations. In this paper, we leverage the observation that image-level tags provide important cues of foreground salient objects, and develop…

Cited by 1450PDFScholar
2016

Mixture of Bilateral-Projection Two-Dimensional Probabilistic Principal Component Analysis

CVPR 2016poster

The probabilistic principal component analysis (PPCA) is built upon a global linear mapping, with which it is insufficient to model complex data variation. This paper proposes a mixture of bilateral-projection probabilistic principal component analysis model (mixB2DPPCA) on 2D data. With multi-compo…

Cited by 5PDFScholar