← Search

Shaoyi Du

20 accepted papers

2026

Cog-RAG: Cognitive-Inspired Dual-Hypergraph with Theme Alignment Retrieval-Augmented Generation

AAAI 2026technical

Retrieval-Augmented Generation (RAG) enhances the response quality and domain-specific performance of large language models (LLMs) by incorporating external knowledge to combat hallucinations. In recent research, graph structures have been integrated into RAG to enhance the capture of semantic relat

Cited by 0SourcePDFScholar
2026

Cross-Modal Dynamic Hypergraph Computation via Functional-Structural Brain Network for Brain Disorder Diagnosis

IJCAI 2026

Cross-modal brain networks characterize the complex connections between different brain regions from both functional and structural perspectives, which is of significant importance for brain network analysis and the diagnosis of brain diseases. However, existing methods have failed to fully exploit

Cited by 0Scholar
2026

DAPE: Harmonizing Content-Position Encoding for Versatile Dense Visual Prediction

AAAI 2026technical

Dense visual prediction tasks, including object detection and segmentation, inherently require precise and discriminative positional information to delineate object boundaries and pixel regions. Recent DETR-based frameworks advance dense prediction tasks through iterative attention applied to conten

Cited by 0SourcePDFScholar
2026

MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation

AAAI 2026technical

Multi-modal Retrieval-Augmented Generation (MMRAG) enables highly credible generation by integrating external multi-modal knowledge, thus demonstrating impressive performance in complex multi-modal scenarios. However, existing MMRAG methods fail to clarify the reasoning logic behind retrieval and re

Cited by 0SourcePDFScholar
2026

Revisiting Weight Regularization for Low-Rank Continual Learning

ICLR 2026poster

Continual Learning (CL) with large-scale pre-trained models (PTMs) has recently gained wide attention, shifting the focus from training from scratch to continually adapting PTMs. This has given rise to a promising paradigm: parameter-efficient continual learning (PECL), where task interference is ty…

Cited by 0SourcecodeScholar
2026

Role Hypergraph Contrastive Learning for Multivariate Time-Series Analysis

AAAI 2026technical

Multivariate Time-Series (MTS) analysis is crucial across various domains. Considering the spatial and temporal consistency of MTS, existing methods leverage graph structures with temporal augmentation and contrastive learning to achieve robust learning of spatial dependencies and temporal patterns.

Cited by 0SourcePDFScholar
2025

Beyond Graphs: Can Large Language Models Comprehend Hypergraphs?

ICLR 2025poster

Existing benchmarks like NLGraph and GraphQA evaluate LLMs on graphs by focusing mainly on pairwise relationships, overlooking the high-order correlations found in real-world data. Hypergraphs, which can model complex beyond-pairwise relationships, offer a more robust framework but are still underex…

2025

Cross-Template-Based Hypergraph Transformer

ICASSP 2025accepted

Single-template-based brain functional network analysis methods can provide limited functional connectivity information, which constrains the performance of brain disease diagnosis. Previous works have explored multi-template functional network analysis but failed to integrate the high-order correla…

Cited by 0SourceScholar
2025

ERetinex: Event Camera Meets Retinex Theory for Low-Light Image Enhancement

ICRA 2025

Low-light image enhancement aims to restore the under-exposure image captured in dark scenarios. Under such scenarios, traditional frame-based cameras may fail to capture the structure and color information due to the exposure time limitation. Event cameras are bio-inspired vision sensors that respo

Cited by 5SourcecodeScholar
2025

Point-MaDi: Masked Autoencoding with Diffusion for Point Cloud Pre-training

NeurIPS 2025poster

Self-supervised pre-training is essential for 3D point cloud representation learning, as annotating their irregular, topology-free structures is costly and labor-intensive. Masked autoencoders (MAEs) offer a promising framework but rely on explicit positional embeddings, such as patch center coordin…

Cited by 0SourceScholar
2024

ColorPCR: Color Point Cloud Registration with Multi-Stage Geometric-Color Fusion

CVPR 2024poster

Point cloud registration is still a challenging and open problem. For example when the overlap between two point clouds is extremely low geo-only features may be not sufficient. Therefore it is important to further explore how to utilize color data in this task. Under such circumstances we propose C…

Cited by 6SourcePDFScholar
2024

ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition

IROS 2024poster

Place recognition is an important task for robots and autonomous cars to localize themselves and close loops in pre-built maps. While single-modal sensor-based methods have shown satisfactory performance, cross-modal place recognition that retrieving images from a point-cloud database remains a chal…

Cited by 3SourcecodeScholar
2024

PHFormer: Multi-Fragment Assembly Using Proxy-Level Hybrid Transformer

AAAI 2024technical

Fragment assembly involves restoring broken objects to their original geometries, and has many applications, such as archaeological restoration. Existing learning based frameworks have shown potential for solving part assembly problems with semantic decomposition, but cannot handle such geometrical…

2024

Semantic Flow: Learning Semantic Fields of Dynamic Scenes from Monocular Videos

ICLR 2024poster

In this work, we pioneer Semantic Flow, a neural semantic representation of dynamic scenes from monocular videos. In contrast to previous NeRF methods that reconstruct dynamic scenes from the colors and volume densities of individual points, Semantic Flow learns semantics from continuous flows that…

Cited by 5SourcePDFScholar
2023

Decompose More and Aggregate Better: Two Closer Looks at Frequency Representation Learning for Human Motion Prediction

CVPR 2023poster

Encouraged by the effectiveness of encoding temporal dynamics within the frequency domain, recent human motion prediction systems prefer to first convert the motion representation from the original pose space into the frequency space. In this paper, we introduce two closer looks at effective frequen…

Cited by 23SourcePDFScholar
2023

MonoNeRF: Learning a Generalizable Dynamic Radiance Field from Monocular Videos

ICCV 2023poster

In this paper, we target at the problem of learning a generalizable dynamic radiance field from monocular videos. Different from most existing NeRF methods that are based on multiple views, monocular videos only contain one view at each timestamp, thereby suffering from ambiguity along the view dire…

Cited by 43PDFcodeScholar
2022

C-CAM: Causal CAM for Weakly Supervised Semantic Segmentation on Medical Image

CVPR 2022poster

Recently, many excellent weakly supervised semantic segmentation (WSSS) works are proposed based on class activation mapping (CAM). However, there are few works that consider the characteristics of medical images. In this paper, we find that there are mainly two challenges of medical images in WSSS:…

Cited by 118PDFcodeScholar
2020

CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional Correspondence

IROS 2020poster

In this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective functi…

Cited by 15SourcecodeScholar