← Search

Jaywon Koo

4 accepted papers

2026

ProxyThinker: Test-Time Guidance through Small Visual Reasoners

ICLR 2026poster

Recent advancements in reinforcement learning with verifiable rewards have pushed the boundaries of the visual reasoning capabilities in large vision-language models (LVLMs). However, training LVLMs with reinforcement fine-tuning (RFT) is computationally expensive, posing a significant challenge to…

Cited by 0SourcecodeScholar
2024

Beyond Grounding: Extracting Fine-Grained Event Hierarchies across Modalities

AAAI 2024technical

Events describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if events across textual and visual (video) domains are identical…

2024

Multimodal Multi-loss Fusion Network for Sentiment Analysis

NAACL 2024long

This paper investigates the optimal selection and fusion of feature encoders across multiple modalities and combines these in one neural network to improve sentiment detection. We compare different fusion methods and examine the impact of multi-loss training within the multi-modality fusion network,…

2024

PropTest: Automatic Property Testing for Improved Visual Programming

EMNLP 2024finding

Visual Programming has recently emerged as an alternative to end-to-end black-box visual reasoning models. This type of method leverages Large Language Models (LLMs) to generate the source code for an executable computer program that solves a given problem. This strategy has the advantage of offerin…

Cited by 4SourcePDFScholar