← Search

Jihwan Park

13 accepted papers

2026

GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection

AAAI 2026technical

Knowledge distillation (KD) is a promising compression technique for reducing the computational burden of large language models (LLMs). Depending on access to the teacher model’s internal parameters, KD is typically categorized into white-box and black-box KD. While white-box KD benefits from full a

Cited by 0SourcePDFScholar
2026

RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection

CVPR 2026

Weakly-supervised Human-Object Interaction (HOI) detection is essential for scalable scene understanding, as it learns interactions from only image-level annotations. Due to the lack of localization signals, prior works typically rely on an external object detector to generate candidate pairs and th

Cited by 0SourcecodeScholar
2026

Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization

AAAI 2026technical

Vision-Language Models (VLMs) have been widely used in various visual recognition tasks due to their remarkable generalization capabilities. As these models grow in size and complexity, fine-tuning becomes costly, emphasizing the need to reuse adaptation knowledge from

Cited by 0SourcePDFScholar
2025

Super-Class Guided Transformer for Zero-Shot Attribute Classification

AAAI 2025technical

Attribute classification is crucial for identifying specific characteristics within image regions. Vision-Language Models (VLMs) have been effective in zero-shot tasks by leveraging their general knowledge from large-scale datasets. Recent studies demonstrate that transformer-based models with class…

2025

Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection

NeurIPS 2025poster

Zero-shot Human-Object Interaction detection aims to localize humans and objects in an image and recognize their interaction, even when specific verb-object pairs are unseen during training. Recent works have shown promising results using prompt learning with pretrained vision-language models such a…

Cited by 0SourcecodeScholar
2024

Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-based Visual Relationship Detection

CVPR 2024poster

Visual Relationship Detection (VRD) has seen significant advancements with Transformer-based architectures recently. However we identify two key limitations in a conventional label assignment for training Transformer-based VRD models which is a process of mapping a ground-truth (GT) to a prediction.…

2024

Learning Contextualized Representation on Discrete Space Via Hierarchical Product Quantization

ICASSP 2024accepted

Self-supervised learning has recently demonstrated significant success in various speech processing applications. Recent studies report that pre-training with contextualized continuous targets plays a crucial role in fine-tuning for better speech downstream tasks. However, unlike the continuous targ…

Cited by 0SourceScholar
2023

Joint Unsupervised and Supervised Learning for Context-Aware Language Identification

ICASSP 2023accepted

Language identification (LID) recognizes the language of a spoken utterance automatically. According to recent studies, LID models trained with an automatic speech recognition (ASR) task perform better than those trained with a LID task only. However, we need additional text labels to train the mode…

Cited by 0SourceScholar
2023

Metric Learning for User-Defined Keyword Spotting

ICASSP 2023accepted

The goal of this work is to detect new spoken terms defined by users. While most previous works address Keyword Spotting (KWS) as a closed-set classification problem, this limits their transferability to unseen terms. The ability to define custom keywords has advantages in terms of user experience.I…

Cited by 0SourceScholar
2023

Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models

ICCV 2023poster

Video Question Answering (VideoQA) is a challenging task that entails complex multi-modal reasoning. In contrast to multiple-choice VideoQA which aims to predict the answer given several options, the goal of open-ended VideoQA is to answer questions without restricting candidate answers. However, th…

Cited by 7PDFcodeScholar
2022

Consistency Learning via Decoding Path Augmentation for Transformers in Human Object Interaction Detection

CVPR 2022poster

Human-Object Interaction detection is a holistic visual recognition task that entails object detection as well as interaction classification. Previous works of HOI detection has been addressed by the various compositions of subset predictions, e.g., Image -> HO -> I, Image -> HI -> O. Recently, tran…

Cited by 31PDFcodeScholar
2016

Dual-microphone voice activity detection based on using optimally weighted maximum a posteriori probabilities

ICASSP 2016accepted

In this paper, we propose to improve the dual-microphone voice activity detection (VAD) technique for which a discriminative weight training is applied to achieve optimally weighted spatial features. In our approach, we first derive the maximum a posteriori (MAP) probabilities from the spatial featu…

Cited by 0SourceScholar