← Search

Yirui Li

5 accepted papers

2026

Action-and-object Aware Alignment for Partially Relevant Video Retrieval

AAAI 2026technical

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing relevant moments for a given text query. This task is extremely challenging, as untrimmed videos often include numerous actions and objects unrelated to the query. However, existing methods usually struggle with f

Cited by 0SourcePDFScholar
2026

FAM: Fine-Grained Alignment Matters in Multimodal Embedding Learning with Large Vision-Language Models

AAAI 2026technical

Learning multimodal representation is a fundamental task that supports a wide range of applications such as visual-text retrieval. While pioneering approaches e.g., CLIP paves the way by learning separated encoders for different modalities, they struggle to model complex interactions between modalit

Cited by 0SourcePDFScholar
2025

Core Context Aware Transformers for Long Context Language Modeling

ICML 2025poster

Transformer-based Large Language Models (LLMs) have exhibited remarkable success in extensive tasks primarily attributed to self-attention mechanism, which requires a token to consider all preceding tokens as its context to compute attention. However, when the context length L becomes very large (e.…

Cited by 12SourcePDFScholar
2024

DDS-SLAM: Dense Semantic Neural SLAM for Deformable Endoscopic Scenes

IROS 2024poster

Estimating camera motion and continuously reconstructing dense scenes in deformable environments presents a complex and open challenge. Many existing approaches tend to rely on assumptions about the scene’s topology or the nature of deformable motion. However, these assumptions do not hold true in m…

Cited by 2SourcecodeScholar
2024

SoftNeRF: A Self-Modeling Soft Robot Plugin for Various Tasks

IROS 2024poster

Building a self-model for robots, enabling them to simulate their physical selves and predict future states without direct interaction with the physical world, is crucial for robot motion planning and control. Existing self-modeling methods primarily focus on rigid robots and typically require signi…

Cited by 0SourcecodeScholar