← Search

Rui Feng

43 accepted papers

2026

Embodied-DETR: End-to-End Temporal 3D Object Detection in Egocentric Views

ICML 2026poster

Embodied 3D object detection is a fundamental perception capability for embodied agents, where observations are partial, heavily occluded, and sequential, requiring modeling of temporal continuity. However, existing benchmarks and methods are primarily designed for fully reconstructed global scenes …

Cited by 0SourceScholar
2026

UniMapping: Unified SLAM Framework for Map-Centric Embodied Perception

ICML 2026spotlight

Simultaneous Localization and Mapping (SLAM) is increasingly expected to provide reusable spatial representations for downstream perception. However, existing approaches often struggle with scale-consistency and producing maps that lack the geometric fidelity required for reliable perception. We pro…

Cited by 0SourceScholar
2025

AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation

NeurIPS 2025poster

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding,…

Cited by 0SourcecodeScholar
2025

AS-Det: Active Sampling for Adaptive 3D Object Detection in Point Clouds

AAAI 2025technical

3D object detection in point clouds is critical in 3D computer vision, autonomous driving, and robotics. Existing point-based detectors, tailored to handle unstructured raw point clouds, often rely on simplistic sampling strategies to select a subset of points for local representation learning and d…

2025

Can Automated Speech Recognition Errors Provide Valuable Clues for Alzheimer's Disease Detection?

ICASSP 2025accepted

Recent advances in automatic speech recognition (ASR) technology have boosted the viability of fully automated Alzheimer’s disease (AD) detection via ASR transcripts. However, there is a lack of understanding of how ASR errors affect the performance of AD detection. This paper addresses that gap. Fi…

Cited by 0SourceScholar
2025

ConTrack3D: Contrastive Learning Contributes Concise 3D Multi-Object Tracking

ICRA 2025

Online object detection and tracking are crucial for embodied intelligence systems, including autonomous vehicles and robotics. Traditional approaches employ a pipeline structure to perform detection and tracking separately, which can not fully leverage information from the detector. Moreover, most

Cited by 0SourceScholar
2025

EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos

ICLR 2025poster

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and…

Cited by 0SourcePDFScholar
2025

EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues

NAACL 2025long

Role-playing agents (RPAs) powered by large language models (LLMs) have been widely utilized in dialogue systems for their capability to deliver personalized interactions. Current evaluations of RPAs mainly focus on personality fidelity, tone imitation, and knowledge consistency, while overlooking e…

Cited by 0SourcePDFScholar
2025

GLiM: Integrating Graph Transformer and LLM for Document-Level Biomedical Relation Extraction with Incomplete Labeling

ACL 2025finding

Document-level relation extraction (DocRE) identifies relations between entities across an entire document. However, as the number and complexity of entities and entity-pair relations grow, the problem space expands quadratically, causing incomplete annotations and frequent false negatives, especial…

2025

Human Simulacra: Benchmarking the Personification of Large Language Models

ICLR 2025poster

Large Language Models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted the attention of the social science community, who see the potential in leveraging LLMs to replace human participants in experiments, thereby reducing research costs…

2025

Minimizing Disparities between Real and Pseudo Queries for Unsupervised Visual Grounding

ICASSP 2025accepted

Visual grounding involves the identification and localization of image regions given textual descriptions. To reduce the manual labeling effort on region-text pairs, unsupervised visual grounding aims to generate pseudo bounding box and query pairs for training grounding models. However, there exist…

Cited by 0SourceScholar
2025

RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial Documents

COLING 2025main

Randomized Controlled Trials (RCTs) are rigorous clinical studies crucial for reliable decision-making, but their credibility can be compromised by bias. The Cochrane Risk of Bias tool (RoB 2) assesses this risk, yet manual assessments are time-consuming and labor-intensive. Previous approaches have…

Cited by 0SourcePDFScholar
2025

Sketch-based Point Cloud Generation with Diffusion Model and Pre-training Enhancement

ICASSP 2025accepted

Diffusion models, known for their success in various generative tasks like image generation and super-resolution, are applied in this study for point cloud generation, a field that has not been extensively explored due to the complexity of point clouds. We propose a novel method using a diffusion mo…

Cited by 0SourceScholar
2025

The USTC System for EEG-Music Emotion Recognition Challenge

ICASSP 2025accepted

This paper presents the Neural Harmony team’s submission to Task 1 (Person Identification) of the ICASSP 2025 EEG-Music Emotion Recognition Challenge, which aims to identify the subject from a given EEG segment. To enhance performance, we propose a novel architecture incorporating the Multiscale Con…

Cited by 0SourceScholar
2025

Uncertainty-Aware Dynamic Fusion for Multimodal Clinical Prediction Tasks

ICASSP 2025accepted

Multimodal fusion offers significant potential for enhancing medical diagnosis, particularly in the Intensive Care Unit (ICU), where integrating diverse data sources is crucial. Traditional static fusion models often fail to account for sample-wise variations in modality importance, which can impact…

Cited by 0SourceScholar
2024

ControlCap: Controllable Captioning via No-Fuss Lexicon

ICASSP 2024accepted

Controllable captioning has received much attention in recent years. Although substantial progress has been made, existing methods still face challenges such as high training costs, intricate control signals and limited control capabilities. To address these issues, we propose a straightforward and…

Cited by 0SourceScholar
2024

Cross-Image Distillation for Semi-Supervised Semantic Segmentation

ICASSP 2024accepted

Semi-supervised semantic segmentation approaches have drawn much more attention in recent years, which aim to exploit a large amount of unlabeled data together with a small number of labeled data. However, existing models usually regarded segmentation as pixel-wise classification, neglecting global…

Cited by 0SourceScholar
2024

DeepPointMap: Advancing LiDAR SLAM with Unified Neural Descriptors

AAAI 2024technical

Point clouds have shown significant potential in various domains, including Simultaneous Localization and Mapping (SLAM). However, existing approaches either rely on dense point clouds to achieve high localization accuracy or use generalized descriptors to reduce map size. Unfortunately, these two a…

2024

Denoising Diffusion-Augmented Hybrid Video Anomaly Detection via Reconstructing Noised Frames

IJCAI 2024poster

Video Anomaly Detection (VAD) is crucial for enhancing security and surveillance systems through automatic identification of irregular events, thereby enabling timely responses and augmenting overall situational awareness. Although existing methods have achieved decent detection performances on benc…

Cited by 3SourcePDFScholar
2024

Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning

ICASSP 2024accepted

Video paragraph captioning task aims to generate a detailed, fluent and relevant paragraph for a given video. Prior studies often focus on isolating visual objects (potential main components in a sentence) from the overall video content. They rarely explore the latent semantic relations between obje…

Cited by 0SourceScholar
2024

Retrieval-Augmented Egocentric Video Captioning

CVPR 2024poster

Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper (1) we develop EgoI…

Cited by 38SourcePDFScholar
2024

Tag2Text: Guiding Vision-Language Model via Image Tagging

ICLR 2024poster

This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a…

Cited by 84SourcePDFScholar
2024

Towards Evidential and Class Separable Open Set Object Detection

AAAI 2024technical

Detecting in open-world scenarios poses a formidable challenge for models intended for real-world deployment. The advanced closed set object detectors achieve impressive performance under the closed set setting, but often produce overconfident misprediction on unknown objects due to the lack of supe…

2023

Boosting Fine-Grained Sketch-Based Image Retrieval with Self-Supervised Learning

ICASSP 2023accepted

Fine-grained sketch-based image retrieval (FG-SBIR) aims at aligning images and sketches at the instance level. It is a challenging task as there are significant differences between sketch and image. Existing methods usually produce less desired performance due to the lack of large-scale fine-graine…

Cited by 0SourceScholar
2023

Diabetic Retinopathy Grading with Weakly-Supervised Lesion Priors

ICASSP 2023accepted

Explicit information of lesions can provide visual instructions for diabetic retinopathy (DR) grading on fundus images. However, pixel-level lesion annotations are extremely difficult and time-consuming to acquire. In this work, we propose a novel weakly-supervised lesion-aware network for DR gradin…

Cited by 0SourceScholar
2023

Large Language Models are Complex Table Parsers

EMNLP 2023long main

With the Generative Pre-trained Transformer 3.5 (GPT-3.5) exhibiting remarkable reasoning and comprehension abilities in Natural Language Processing (NLP), most Question Answering (QA) research has primarily centered around general QA tasks based on GPT, neglecting the specific challenges posed by C…

Cited by 0SourceScholar
2023

Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision

CVPR 2023poster

In this paper, we consider the problem of open-vocabulary semantic segmentation (OVS), which aims to segment objects of arbitrary classes instead of pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor,…

2023

May the Force be with You: Unified Force-Centric Pre-Training for 3D Molecular Conformations

NeurIPS 2023poster

Recent works have shown the promise of learning pre-trained models for 3D molecular representation. However, existing pre-training models focus predominantly on equilibrium data and largely overlook off-equilibrium conformations. It is challenging to extend these methods to off-equilibrium data beca…

Cited by 10SourcePDFScholar
2023

Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge

ICASSP 2023accepted

Video paragraph captioning task aims at generating a fine-grained, coherent and relevant paragraph for a video. Different from the images where objects are static, the temporal states of objects are changing in videos. The dynamic information could be contributed to understanding the whole video con…

Cited by 0SourceScholar
2023

SCSGNet: Spatial-Correlated and Shape-Guided Network for Breast Mass Segmentation

ICASSP 2023accepted

Automatic and accurate breast mass segmentation plays a crucial role in the early diagnosis of breast cancer. However, it has been a challenging task for two main reasons: (1) Breast masses are diverse; and (2) The boundaries of masses are ambiguous. To address these problems, we propose a Spatial-C…

Cited by 0SourceScholar
2023

Sparse Frame Grouping Network with Action Centered for Untrimmed Video Paragraph Captioning

EMNLP 2023long findings

Generating paragraph captions for untrimmed videos without event annotations is challenging, especially when aiming to enhance precision and minimize repetition at the same time. To address this challenge, we propose a module called Sparse Frame Grouping (SFG). It dynamically groups event informatio…

Cited by 0SourceScholar
2023

Video Captioning via Relation-Aware Graph Learning

ICASSP 2023accepted

Recent neural models for video captioning usually employed an encoder-decoder framework. However, most approaches either neglected the spatial and temporal interactions between objects in a video or implicitly modelled the interactions, resulting in less desired performance. In this paper, we propos…

Cited by 0SourceScholar
2022

CERES: Pretraining of Graph-Conditioned Transformer for Semi-Structured Session Data

NAACL 2022long

User sessions empower many search and recommendation tasks on a daily basis. Such session data are semi-structured, which encode heterogeneous relations between queries and products, and each item is described by the unstructured text. Despite recent advances in self-supervised learning for text or…

Cited by 3SourcePDFScholar
2022

CREAM: Weakly Supervised Object Localization via Class RE-Activation Mapping

CVPR 2022poster

Weakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) derived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomple…

Cited by 45PDFcodeScholar
2022

End-to-end Stochastic Optimization with Energy-based Model

NeurIPS 2022accept

Decision-focused learning (DFL) was recently proposed for stochastic optimization problems that involve unknown parameters. By integrating predictive modeling with an implicitly differentiable optimization layer, DFL has shown superior performance to the standard two-stage predict-then-optimize pipe…

2022

TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp Segmentation

IJCAI 2022poster

Automatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully superv…

2021

Improving Multimodal Speech Enhancement by Incorporating Self-Supervised and Curriculum Learning

ICASSP 2021accepted

Speech enhancement in realistic scenarios still remains many challenges, such as complex background signals and data limitations. In this paper, we present a co-attention based framework that incorporates self-supervised and curriculum learning to derive the target speech in noisy environments. Spec…

Cited by 0SourceScholar