← Search

Yuejie Zhang

27 accepted papers

2026

Embodied-DETR: End-to-End Temporal 3D Object Detection in Egocentric Views

ICML 2026poster

Embodied 3D object detection is a fundamental perception capability for embodied agents, where observations are partial, heavily occluded, and sequential, requiring modeling of temporal continuity. However, existing benchmarks and methods are primarily designed for fully reconstructed global scenes …

Cited by 0SourceScholar
2026

UniMapping: Unified SLAM Framework for Map-Centric Embodied Perception

ICML 2026spotlight

Simultaneous Localization and Mapping (SLAM) is increasingly expected to provide reusable spatial representations for downstream perception. However, existing approaches often struggle with scale-consistency and producing maps that lack the geometric fidelity required for reliable perception. We pro…

Cited by 0SourceScholar
2025

AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation

NeurIPS 2025poster

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding,…

Cited by 0SourcecodeScholar
2025

An Empirical Analysis of Uncertainty in Large Language Model Evaluations

ICLR 2025poster

As LLM-as-a-Judge emerges as a new paradigm for assessing large language models (LLMs), concerns have been raised regarding the alignment, bias, and stability of LLM evaluators. While substantial work has focused on alignment and bias, little research has concentrated on the stability of LLM evaluat…

2025

EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos

ICLR 2025poster

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and…

Cited by 0SourcePDFScholar
2025

EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues

NAACL 2025long

Role-playing agents (RPAs) powered by large language models (LLMs) have been widely utilized in dialogue systems for their capability to deliver personalized interactions. Current evaluations of RPAs mainly focus on personality fidelity, tone imitation, and knowledge consistency, while overlooking e…

Cited by 0SourcePDFScholar
2025

GLiM: Integrating Graph Transformer and LLM for Document-Level Biomedical Relation Extraction with Incomplete Labeling

ACL 2025finding

Document-level relation extraction (DocRE) identifies relations between entities across an entire document. However, as the number and complexity of entities and entity-pair relations grow, the problem space expands quadratically, causing incomplete annotations and frequent false negatives, especial…

2025

Human Simulacra: Benchmarking the Personification of Large Language Models

ICLR 2025poster

Large Language Models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted the attention of the social science community, who see the potential in leveraging LLMs to replace human participants in experiments, thereby reducing research costs…

2025

Minimizing Disparities between Real and Pseudo Queries for Unsupervised Visual Grounding

ICASSP 2025accepted

Visual grounding involves the identification and localization of image regions given textual descriptions. To reduce the manual labeling effort on region-text pairs, unsupervised visual grounding aims to generate pseudo bounding box and query pairs for training grounding models. However, there exist…

Cited by 0SourceScholar
2025

RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial Documents

COLING 2025main

Randomized Controlled Trials (RCTs) are rigorous clinical studies crucial for reliable decision-making, but their credibility can be compromised by bias. The Cochrane Risk of Bias tool (RoB 2) assesses this risk, yet manual assessments are time-consuming and labor-intensive. Previous approaches have…

Cited by 0SourcePDFScholar
2025

Sketch-based Point Cloud Generation with Diffusion Model and Pre-training Enhancement

ICASSP 2025accepted

Diffusion models, known for their success in various generative tasks like image generation and super-resolution, are applied in this study for point cloud generation, a field that has not been extensively explored due to the complexity of point clouds. We propose a novel method using a diffusion mo…

Cited by 0SourceScholar
2025

Uncertainty-Aware Dynamic Fusion for Multimodal Clinical Prediction Tasks

ICASSP 2025accepted

Multimodal fusion offers significant potential for enhancing medical diagnosis, particularly in the Intensive Care Unit (ICU), where integrating diverse data sources is crucial. Traditional static fusion models often fail to account for sample-wise variations in modality importance, which can impact…

Cited by 0SourceScholar
2024

ControlCap: Controllable Captioning via No-Fuss Lexicon

ICASSP 2024accepted

Controllable captioning has received much attention in recent years. Although substantial progress has been made, existing methods still face challenges such as high training costs, intricate control signals and limited control capabilities. To address these issues, we propose a straightforward and…

Cited by 0SourceScholar
2024

DeepPointMap: Advancing LiDAR SLAM with Unified Neural Descriptors

AAAI 2024technical

Point clouds have shown significant potential in various domains, including Simultaneous Localization and Mapping (SLAM). However, existing approaches either rely on dense point clouds to achieve high localization accuracy or use generalized descriptors to reduce map size. Unfortunately, these two a…

2024

Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning

ICASSP 2024accepted

Video paragraph captioning task aims to generate a detailed, fluent and relevant paragraph for a given video. Prior studies often focus on isolating visual objects (potential main components in a sentence) from the overall video content. They rarely explore the latent semantic relations between obje…

Cited by 0SourceScholar
2024

Retrieval-Augmented Egocentric Video Captioning

CVPR 2024poster

Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper (1) we develop EgoI…

Cited by 38SourcePDFScholar
2024

Tag2Text: Guiding Vision-Language Model via Image Tagging

ICLR 2024poster

This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a…

Cited by 84SourcePDFScholar
2023

Boosting Fine-Grained Sketch-Based Image Retrieval with Self-Supervised Learning

ICASSP 2023accepted

Fine-grained sketch-based image retrieval (FG-SBIR) aims at aligning images and sketches at the instance level. It is a challenging task as there are significant differences between sketch and image. Existing methods usually produce less desired performance due to the lack of large-scale fine-graine…

Cited by 0SourceScholar
2023

Large Language Models are Complex Table Parsers

EMNLP 2023long main

With the Generative Pre-trained Transformer 3.5 (GPT-3.5) exhibiting remarkable reasoning and comprehension abilities in Natural Language Processing (NLP), most Question Answering (QA) research has primarily centered around general QA tasks based on GPT, neglecting the specific challenges posed by C…

Cited by 0SourceScholar
2023

Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision

CVPR 2023poster

In this paper, we consider the problem of open-vocabulary semantic segmentation (OVS), which aims to segment objects of arbitrary classes instead of pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor,…

2023

Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge

ICASSP 2023accepted

Video paragraph captioning task aims at generating a fine-grained, coherent and relevant paragraph for a video. Different from the images where objects are static, the temporal states of objects are changing in videos. The dynamic information could be contributed to understanding the whole video con…

Cited by 0SourceScholar
2023

SCSGNet: Spatial-Correlated and Shape-Guided Network for Breast Mass Segmentation

ICASSP 2023accepted

Automatic and accurate breast mass segmentation plays a crucial role in the early diagnosis of breast cancer. However, it has been a challenging task for two main reasons: (1) Breast masses are diverse; and (2) The boundaries of masses are ambiguous. To address these problems, we propose a Spatial-C…

Cited by 0SourceScholar
2023

Sparse Frame Grouping Network with Action Centered for Untrimmed Video Paragraph Captioning

EMNLP 2023long findings

Generating paragraph captions for untrimmed videos without event annotations is challenging, especially when aiming to enhance precision and minimize repetition at the same time. To address this challenge, we propose a module called Sparse Frame Grouping (SFG). It dynamically groups event informatio…

Cited by 0SourceScholar
2023

Video Captioning via Relation-Aware Graph Learning

ICASSP 2023accepted

Recent neural models for video captioning usually employed an encoder-decoder framework. However, most approaches either neglected the spatial and temporal interactions between objects in a video or implicitly modelled the interactions, resulting in less desired performance. In this paper, we propos…

Cited by 0SourceScholar
2022

CREAM: Weakly Supervised Object Localization via Class RE-Activation Mapping

CVPR 2022poster

Weakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) derived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomple…

Cited by 45PDFcodeScholar
2022

TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp Segmentation

IJCAI 2022poster

Automatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully superv…