← Search

Minwoo Choi

9 accepted papers

2026

Infinite-Story: A Training-Free Consistent Text-to-Image Generation

AAAI 2026technical

We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style in

Cited by 0SourcePDFScholar
2026

Scale-Invariant and View-Relational Representation Learning for Full Surround Monocular Depth

RA-L 2026

Recent foundation models demonstrate strong generalization capabilities in monocular depth estimation. However, directly applying these models to Full Surround Monocular Depth Estimation (FSMDE) presents two major challenges: (1) high computational cost, which limits real-time performance, and (2) d

Cited by 0SourceScholar
2026

Scale-Invariant and View-Relational Representation Learning for Full Surround Monocular Depth

ICRA 2026poster

Recent foundation models demonstrate strong generalization capabilities in monocular depth estimation. However, directly applying these models to Full Surround Monocular Depth Estimation (FSMDE) presents two major challenges: (1) high computational cost, which limits real-time performance, and (2) d…

2026

TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task Optimization

CVPR 2026

Multi-task learning (MTL) involves the simultaneous optimization of multiple task-specific losses, often leading to gradient conflicts and scale imbalances that result in negative transfer. While existing multi-task optimization methods attempt to mitigate these challenges, they either lack the stoc

Cited by 0SourceScholar
2025

Intrinsic Image Decomposition for Robust Self-supervised Monocular Depth Estimation on Reflective Surfaces

AAAI 2025technical

Self-supervised monocular depth estimation (SSMDE) has gained attention in the field of deep learning as it estimates depth without requiring ground truth depth maps. This approach typically uses a photometric consistency loss between a synthesized image, generated from the estimated depth, and the…

Cited by 0SourcePDFScholar
2025

LOMM: Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

ICCV 2025poster

In this paper, we present Latest Object Memory Management (LOMM) for temporally consistent video instance segmentation that significantly improves long-term instance tracking. At the core of our method is Latest Object Memory (LOM), which robustly tracks and continuously updates the latest states of…

Cited by 0SourcePDFScholar
2025

Self-supervised Monocular Depth Estimation Robust to Reflective Surface Leveraged by Triplet Mining

ICLR 2025poster

Self-supervised monocular depth estimation (SSMDE) aims to predict the dense depth map of a monocular image, by learning depth from RGB image sequences, eliminating the need for ground-truth depth labels. Although this approach simplifies data acquisition compared to supervised methods, it struggles…

Cited by 1SourcePDFScholar
2022

ADAS: A Direct Adaptation Strategy for Multi-Target Domain Adaptive Semantic Segmentation

CVPR 2022poster

In this paper, we present a direct adaptation strategy (ADAS), which aims to directly adapt a single model to multiple target domains in a semantic segmentation task without pretrained domain-specific models. To do so, we design a multi-target domain transfer network (MTDT-Net) that aligns visual at…

Cited by 29PDFcodeScholar