← Search

Dongwon Kim

18 accepted papers

2026

Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model

CVPR 2026

World models provide a powerful framework for simulating environment dynamics conditioned on actions or instructions, enabling downstream tasks such as action planning or policy learning.Recent approaches leverage world models as learned simulators, but its application to decision-time planning rema

Cited by 0SourcecodeScholar
2026

SyncMos: Scalable Motion Synchronisation for Multi-Agent Scene Interaction

CVPR 2026

Text-guided motion generation in 3D scenes has advanced the synthesis of human-scene interactions, contributing to embodied AI, scene understanding, and virtual agent simulation. While recent studies have begun exploring multi-agent scenarios, achieving temporally synchronised interactions among mul

Cited by 0SourceScholar
2025

Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

ICCV 2025poster

Image tokenizers form the foundation of modern text-toimage generative models but are notoriously difficult to train. Furthermore, most existing text-to-image models rely on large-scale, high-quality private datasets, making them challenging to replicate. In this work, we introduce **T**ext-**A**war…

Cited by 0SourcePDFScholar
2025

Federated Learning for Feature Generalization with Convex Constraints

ICML 2025poster

Federated learning (FL) often struggles with generalization due to heterogeneous client data. Local models are prone to overfitting their local data distributions, and even transferable features can be distorted during aggregation. To address these challenges, we propose FedCONST, an approach that a…

Cited by 0SourcePDFScholar
2024

Extending CLIP’s Image-Text Alignment to Referring Image Segmentation

NAACL 2024long

Referring Image Segmentation (RIS) is a cross-modal task that aims to segment an instance described by a natural language expression. Recent methods leverage large-scale pretrained unimodal models as backbones along with fusion techniques for joint reasoning across modalities. However, the inherent…

Cited by 9SourcePDFScholar
2024

PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery

ECCV 2024poster

"Text-based person search, employing free-form text queries to identify individuals within a vast image collection, presents a unique challenge in aligning visual and textual representations, particularly at the human part level. Existing methods often struggle with part feature extraction and align…

Cited by 4SourcePDFScholar
2023

Shatter and Gather: Learning Referring Image Segmentation with Text Supervision

ICCV 2023poster

Referring image segmentation, the task of segmenting any arbitrary entities described in free-form texts, opens up a variety of vision applications. However, manual labeling of training data for this task is prohibitively costly, leading to lack of labeled data for training. We address this issue b…

Cited by 22PDFcodeScholar
2022

ReSTR: Convolution-Free Referring Image Segmentation Using Transformers

CVPR 2022poster

Referring image segmentation is an advanced semantic segmentation task where target is not a predefined class but is described in natural language. Most of existing methods for this task rely heavily on convolutional neural networks, which however have trouble capturing long-range dependencies betwe…

Cited by 176PDFScholar
2020

Simultaneous Estimations of Joint Angle and Torque in Interactions with Environments using EMG

ICRA 2020poster

We develop a decoding technique that estimates both the position and torque of a joint of the limb in interaction with an environment based on activities of the agonist-antagonist pair of muscles using electromyography in real time. The long short-term memory (LSTM) network is employed as the core p…

Cited by 11SourceScholar
2017

Impedance control with structural compliance and a sensorless strategy for contact tasks

IROS 2017poster

This study proposes an actuator whose design relies on a coordinated approach to control and hardware design: impedance control is supplemented with the introduction of a spring-damper coupler between the actuator and reference (ground). The coupler between the actuator and reference has the effect…

Cited by 2SourceScholar