← Search

Minjoon Jung

4 accepted papers

2025

Confidence-guided Refinement Reasoning for Zero-shot Question Answering

EMNLP 2025

We propose Confidence-guided Refinement Reasoning (C2R), a novel training-free framework applicable to question-answering (QA) tasks across text, image, and video domains. C2R strategically constructs and refines sub-questions and their answers (sub-QAs), deriving a better confidence score for the t

Cited by 0SourcePDFScholar
2025

On the Consistency of Video Large Language Models in Temporal Comprehension

CVPR 2025poster

Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction consistency -- a key indicator for robustness and trustworthiness of…

2024

PGA: Personalizing Grasping Agents with Single Human-Robot Interaction

IROS 2024poster

Language-Conditioned Robotic Grasping (LCRG) aims to develop robots that comprehend and grasp objects based on natural language instructions. While the ability to understand personal objects like my wallet facilitates more natural interaction with human users, current LCRG systems only allow generic…

Cited by 2SourcecodeScholar
2022

Modal-specific Pseudo Query Generation for Video Corpus Moment Retrieval

EMNLP 2022main

Video corpus moment retrieval (VCMR) is the task to retrieve the most relevant video moment from a large video corpus using a natural language query.For narrative videos, e.g., drama or movies, the holistic understanding of temporal dynamics and multimodal reasoning are crucial.Previous works have s…