← Search

Hanseok Ko

12 accepted papers

2025

Less is more: Efficient Scene Graph Generation with reparameterization

ICASSP 2025accepted

Scene Graph Generation (SGG) aims to identify objects and their relationships in visual scenes but faces two key challenges: high computational overhead, particularly for real-time applications, and the long-tailed distribution of predicates, which biases models toward frequent relationships. To add…

Cited by 0SourceScholar
2024

Towards Multi-Domain Face Landmark Detection with Synthetic Data from Diffusion Model

ICASSP 2024accepted

Recently, deep learning-based facial landmark detection for in-the-wild faces has achieved significant improvement. However, there are still challenges in face landmark detection in other domains (e.g. cartoon, caricature, etc). This is due to the scarcity of extensively annotated training data. To…

Cited by 0SourceScholar
2024

ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models

CVPR 2024highlight

Generating novel views of an object from a single image is a challenging task. It requires an understanding of the underlying 3D structure of the object from an image and rendering high-quality spatially consistent new views. While recent methods for view synthesis based on diffusion have shown grea…

Cited by 35SourcePDFScholar
2023

MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation

ICASSP 2023accepted

When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may no…

Cited by 0SourceScholar
2022

Injecting 3D Perception of Controllable NeRF-GAN into StyleGAN for Editable Portrait Image Synthesis

ECCV 2022poster

"Over the years, 2D GANs have achieved great successes in photorealistic portrait generation. However, they lack 3D understanding in the generation process, thus they suffer from multi-view inconsistency problem. To alleviate the issue, many 3D-aware GANs have been proposed and shown notable results…

2021

Memory-based Semantic Segmentation for Off-road Unstructured Natural Environments

IROS 2021poster

With the availability of many datasets tailored for autonomous driving in real-world urban scenes, semantic segmentation for urban driving scenes achieves significant progress. However, semantic segmentation for off-road, unstructured environments is not widely studied. Directly applying existing se…

Cited by 16SourceScholar
2020

CAFE-GAN: Arbitrary Face Attribute Editing with Complementary Attention Feature

ECCV 2020poster

The goal of face attribute editing is altering a facial image according to given target attributes such as hair color, mustache, gender, etc. It belongs to the image-to-image domain transfer problem with a set of attributes considered as a distinctive domain. There have been some works in multi-doma…

Cited by 38SourcePDFScholar
2018

Precise Regression for Bounding Box Correction for Improved Tracking Based on Deep Reinforcement Learning

ICASSP 2018accepted

In this paper, we propose a precise regression approach for correcting imprecise bounding box using deep reinforcement learning. Object tracking task essentially builds trajectory of a moving object based on detection and tracking algorithms and its current state is indicated by having the object en…

Cited by 0SourceScholar
2017

Deep Neural Network based learning and transferring mid-level audio features for acoustic scene classification

ICASSP 2017accepted

Deep Neural Network (DNN) based transfer learning has been shown to be effective in Visual Object Classification (VOC) for complementing the deficit of target domain training samples by adapting classifiers that have been pre-trained for other large-scaled DataBase (DB). Although there exists an abu…

Cited by 0SourceScholar
2017

Subspace projection cepstral coefficients for noise robust acoustic event recognition

ICASSP 2017accepted

In this paper, a novel feature for noise robust sound event recognition is proposed. The proposed feature is obtained by a two-step procedure. First, a subspace bank is established via target event analysis in complex vector space. Then, by projecting observation vectors onto the subspace bank, nois…

Cited by 0SourceScholar