← Search

Kyung-Min Kim

7 accepted papers

2023

Direct Preference-based Policy Optimization without Reward Modeling

NeurIPS 2023poster

Preference-based reinforcement learning (PbRL) is an approach that enables RL agents to learn from preference, which is particularly useful when formulating a reward function is challenging. Existing PbRL methods generally involve a two-step procedure: they first learn a reward model based on given…

2023

Scaling Law for Recommendation Models: Towards General-Purpose User Representations

AAAI 2023technical

Recent advancement of large-scale pretrained models such as BERT, GPT-3, CLIP, and Gopher, has shown astonishing achievements across various task domains. Unlike vision recognition and language models, studies on general-purpose user representation at scale still remain underexplored. Here we explor…

Cited by 40SourcePDFScholar
2022

Exploiting Numerical-Contextual Knowledge to Improve Numerical Reasoning in Question Answering

NAACL 2022findings

Numerical reasoning over text is a challenging subtask in question answering (QA) that requires both the understanding of texts and numbers. However, existing language models in these numerical reasoning QA models tend to overly rely on the pre-existing parametric knowledge at inference time, which…

Cited by 8SourcePDFScholar
2022

Know Your Action Set: Learning Action Relations for Reinforcement Learning

ICLR 2022poster

Intelligent agents can solve tasks in various ways depending on their available set of actions. However, conventional reinforcement learning (RL) assumes a fixed action set. This work asserts that tasks with varying action sets require reasoning of the relations between the available actions. For in…

2021

Have You Seen That Number? Investigating Extrapolation in Question Answering Models

EMNLP 2021main

Numerical reasoning in machine reading comprehension (MRC) has shown drastic improvements over the past few years. While the previous models for numerical MRC are able to interpolate the learned numerical reasoning capabilities, it is not clear whether they can perform just as well on numbers unseen…

Cited by 27SourcePDFScholar
2021

Metropolis-Hastings Data Augmentation for Graph Neural Networks

NeurIPS 2021poster

Graph Neural Networks (GNNs) often suffer from weak-generalization due to sparsely labeled data despite their promising results on various graph-based tasks. Data augmentation is a prevalent remedy to improve the generalization ability of models in many domains. However, due to the non-Euclidean nat…

Cited by 62SourcePDFScholar
2018

Multimodal Dual Attention Memory for Video Story Question Answering

ECCV 2018poster

We propose a video story question-answering (QA) architecture, Multimodal Dual Attention Memory (MDAM). The key idea is to use a dual attention mechanism with late fusion. MDAM uses self-attention to learn the latent concepts in scene frames and captions. Given a question, MDAM uses the second atten…

Cited by 97SourcePDFScholar