← Search

Dokwan Oh

13 accepted papers

2026

Time Without Time: Pseudo-Temporal Representation for Space-Time Super-Resolution

CVPR 2026

Space-time video super-resolution (STVSR) is a task aimed at simultaneously upsampling a video in both spatial and temporal dimensions. Previous studies on STVSR have primarily focused on task-specific architectures and modeling paradigms, while effective pretraining strategies remain underexplored.

Cited by 0SourceScholar
2025

Diffusion on Demand: Selective Caching and Modulation for Efficient Generation

NeurIPS 2025poster

Diffusion transformers demonstrate significant potential for various generation tasks but are challenged by high computational cost. Recently, feature caching methods have been introduced to improve inference efficiency by storing features at certain timesteps and reusing them at subsequent timestep…

Cited by 0SourceScholar
2025

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

ICML 2025poster

Attention mechanisms are central to the success of large language models (LLMs), enabling them to capture intricate token dependencies and implicitly assign importance to each token. Recent studies have revealed the sink token, which receives disproportionately high attention despite their limited s…

Cited by 0SourcePDFScholar
2024

Diversify, Contextualize, and Adapt: Efficient Entropy Modeling for Neural Image Codec

NeurIPS 2024poster

Designing a fast and effective entropy model is challenging but essential for practical application of neural codecs. Beyond spatial autoregressive entropy models, more efficient backward adaptation-based entropy models have been recently developed. They not only reduce decoding time by using smalle…

Cited by 0SourcePDFScholar
2024

Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity

ICML 2024poster

Recent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have significantly degraded pixel-wise fidelity, limiting their practicality. To fill this gap, we develop a new text-guided image co…

2023

D-3DLD: Depth-Aware Voxel Space Mapping for Monocular 3D Lane Detection with Uncertainty

ICASSP 2023accepted

The estimation of 3D lanes from monocular RGB images is a fundamentally ill-posed problem. Previous studies have assumed that all lanes are on a flat ground plane. However, we argue that the algorithms based on this assumption have difficulty in detecting various lanes in actual driving environments…

Cited by 0SourceScholar
2022

DaDA: Distortion-aware Domain Adaptation for Unsupervised Semantic Segmentation

NeurIPS 2022accept

Distributional shifts in photometry and texture have been extensively studied for unsupervised domain adaptation, but their counterparts in optical distortion have been largely neglected. In this work, we tackle the task of unsupervised domain adaptation for semantic image segmentation where unknown…

2022

SeeThroughNet: Resurrection of Auxiliary Loss by Preserving Class Probability Information

CVPR 2022poster

Auxiliary loss is additional loss besides the main branch loss to help optimize the learning process of neural networks. In order to calculate the auxiliary loss between the feature maps of intermediate layers and the ground truth in the field of semantic segmentation, the size of each feature map m…

Cited by 7PDFScholar
2021

Unsupervised Representation Transfer for Small Networks: I Believe I Can Distill On-the-Fly

NeurIPS 2021poster

A current remarkable improvement of unsupervised visual representation learning is based on heavy networks with large-batch training. While recent methods have greatly reduced the gap between supervised and unsupervised performance of deep models such as ResNet-50, this development has been relative…

Cited by 14SourcePDFScholar
2020

Segmenting 2K-Videos at 36.5 FPS with 24.3 GFLOPs: Accurate and Lightweight Realtime Semantic Segmentation Network

ICRA 2020poster

We propose a fast and lightweight end-to-end convolutional network architecture for real-time segmentation of high resolution videos, NfS-SegNet, that can segement 2K-videos at 36.5 FPS with 24.3 GFLOPS. This speed and computation-efficiency is due to following reasons: 1) The encoder network, NfS-N…

Cited by 11SourceScholar