← Search

Anran Liu

4 accepted papers

2026

ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid Reasoning

AAAI 2026technical

Retrieval-augmented generation (RAG) has greatly improved Large Language Models (LLMs) by adding external knowledge. However, current RAG-based methods face difficulties with long-context video understanding due to two main challenges. First, Current RAG-based methods for long-context video understa

Cited by 0SourcePDFScholar
2026

VideoTrace-R1: Long Video-based Retrieval-Augmented Generation via Temporal Path Graph Understanding

ICML 2026poster

Long-video temporal reasoning remains challenging for Large Video Language Models (LVLMs). Recent reasoning-enhanced models apply reinforcement learning with outcome supervision to improve temporal understanding. However, outcome-only rewards cannot distinguish whether a model arrived at the correct…

Cited by 0SourceScholar
2025

FuzzyMIL: Decoupling Pathological Phenotypes through Deep Fuzzy Clustering for Efficient Whole Slide Image Analysis

ICASSP 2025accepted

In Multiple Instance Learning (MIL) for Whole Slide Image (WSI) analysis, attention mechanisms are often employed to weigh the importance of different instances. However, global attention may lead to feature homogenization and overlook tissue differences. Adding local attention can capture these var…

Cited by 0SourceScholar
2024

Generate Like Experts: Multi-Stage Font Generation by Incorporating Font Transfer Process into Diffusion Models

CVPR 2024poster

Few-shot font generation (FFG) produces stylized font images with a limited number of reference samples which can significantly reduce labor costs in manual font designs. Most existing FFG methods follow the style-content disentanglement paradigm and employ the Generative Adversarial Network (GAN) t…