← Search

Dayun Ju

4 accepted papers

2026

ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting

CVPR 2026

Recent advancements in Video Large Language Models (VideoLLMs) have enabled strong performance across diverse multimodal video tasks. To reduce the high computational cost of processing dense video frames, efficiency-oriented methods such as frame selection have been widely adopted. While effective

Cited by 0SourcecodeScholar
2025

Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation

CVPR 2025poster

Open-Vocabulary Semantic Segmentation (OVSS) has advanced with recent vision-language models (VLMs), enabling segmentation beyond predefined categories through various learning schemes. Notably, training-free methods offer scalable, easily deployable solutions for handling unseen data, a key goal of…

Cited by 1SourcePDFScholar
2025

Rare Text Semantics Were Always There in Your Diffusion Transformer

NeurIPS 2025poster

Starting from flow- and diffusion-based transformers, Multi-modal Diffusion Transformers (MM-DiTs) have reshaped text-to-vision generation, gaining acclaim for exceptional visual fidelity. As these models advance, users continually push the boundary with imaginative or rare prompts, which advanced m…

Cited by 0SourceScholar
2024

EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation

CVPR 2024highlight

Semantic segmentation has innately relied on extensive pixel-level annotated data leading to the emergence of unsupervised methodologies. Among them leveraging self-supervised Vision Transformers for unsupervised semantic segmentation (USS) has been making steady progress with expressive deep featur…