← Search

Ruifei Ma

3 accepted papers

2026

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

CVPR 2026

The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-level understanding, their ability to capture fine-grained motion details remains limited, primarily due to their focus on

Cited by 0SourceScholar
2025

CTSG: Integrating Context and Way Topology Into Scene Graph for Zero-shot Navigation

IROS 2025

A robust environment representation is critical for enabling robot systems to accomplish embodied navigation tasks. While offering efficient and sparse representations of environments compared to dense semantic maps, traditional 3D Scene Graphs often rely on multi-level semantic hierarchies that ris

Cited by 1SourceScholar
2025

MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation

AAAI 2025technical

Controllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs provide a suitable data representation that facilitates these applications. Howe…