← Search

Minjun Kim

19 accepted papers

2026

CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop Reasoning

CVPR 2026

Real-world reasoning often requires combining information across modalities, connecting textual context with visual cues in a multi-hop process. Yet, most multimodal benchmarks fail to capture this ability: they typically rely on single images or set of images, where answers can be inferred from a s

Cited by 0SourceScholar
2026

LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision Transformers

AAAI 2026technical

How can we accurately quantize a pre-trained Vision Transformer model? Quantization algorithms compress Vision Transformers (ViTs) into low-bit formats, reducing memory and computation demands with minimal accuracy degradation. However, existing methods rely on uniform precision, ignoring the divers

Cited by 0SourcePDFScholar
2026

Neural Collapse-Informed Initialization with Perturbation Injection in Classification-based Metric Learning

AAAI 2026technical

Recent studies have revealed Neural Collapse (NC) in deep classifiers, where last-layer weights and features align into an equiangular tight frame (ETF), concentrating class information along specific embedding directions. However, conventional fine-tuning typically disregards this structure, initi

Cited by 0SourcePDFScholar
2026

Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model Compression

ICLR 2026poster

What happens when multiple compression methods are combined—does the order in which they are applied matter? Joint model compression has emerged as a powerful strategy to achieve higher efficiency by combining multiple methods such as pruning and quantization. A central but underexplored factor in j…

Cited by 0SourcecodeScholar
2025

Can LLMs Truly Plan? A Comprehensive Evaluation of Planning Capabilities

EMNLP 2025

The existing assessments of planning capabilities of large language models (LLMs) remain largely limited to single-language or specific representation formats. To address this gap, we introduce the Multi-Plan benchmark comprising 204 multilingual and multi-format travel planning scenarios. In experi

Cited by 0SourcePDFScholar
2025

Exact Fractional Order Impedance Rendering for Highly Flexible and Multi-Jointed Robots Using Time-Delay Estimation

RA-L 2025

Fractional-order (FO) control, which employs non-integer orders of differintegration, has been adopted for its advantages in generalized controller design or reduced modeling complexity. For compliance control, FO impedance is employed as a generalization of conventional integer-order (IO) impedance

Cited by 1SourceScholar
2025

Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models

ACL 2025long

How can we quantize large language models while preserving accuracy? Quantization is essential for deploying large language models (LLMs) efficiently. Binary-coding quantization (BCQ) and uniform quantization (UQ) are promising quantization schemes that have strong expressiveness and optimizability,…

2025

VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

COLING 2025main

We propose the VLR-Bench, a visual question answering (VQA) benchmark for evaluating vision language models (VLMs) based on retrieval augmented generation (RAG). Unlike existing evaluation datasets for external knowledge-based VQA, the proposed VLR-Bench includes five input passages. This allows tes…

Cited by 1SourcePDFScholar
2024

BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining

AAAI 2024technical

The current research direction in generative models, such as the recently developed GPT4, aims to find relevant knowledge information for multimodal and multilingual inputs to provide answers. Under these research circumstances, the demand for multilingual evaluation of visual question answering (VQ…

Cited by 5SourcePDFScholar
2024

X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment

NAACL 2024findings

The impressive development of large language models (LLMs) is expanding into the realm of large multimodal models (LMMs), which incorporate multiple types of data beyond text. However, the nature of multimodal models leads to significant expenses in the creation of training data. Furthermore, constr…

2023

Closed-Loop Control of Magnetic Modular Cubes for 2D Self-Assembly

RA-L 2023

Reconfigurable modular robots can dynamically assemble/disassemble to accomplish the desired task better. Magnetic modular cubes are scalable modular subunits with embedded permanent magnets in a 3D-printed cubic body and can be wirelessly controlled by an external, uniform, time-varying magnetic fi

Cited by 8SourceScholar
2021

Enumeration of Polyominoes & Polycubes Composed of Magnetic Cubes

IROS 2021poster

This paper examines a family of designs for magnetic cubes and counts how many configurations are possible for each design as a function of the number of modules. Magnetic modular cubes are cubes with magnets arranged on their faces. The magnets are positioned so that each face has either magnetic s…

Cited by 4SourceScholar
2020

Magnetically Actuated Simple Millirobots for Complex Navigation and Modular Assembly

RA-L 2020

Magnetic millirobots can be controlled remotely by external magnetic fields, making them promising candidates for biomedical and engineering applications. This letter presents a low-cost millirobot that is simple in design, easy to fabricate, highly scalable, and can be used as modular sub-units wit

Cited by 26SourceScholar
2016

Powered upper-limb control using passivity-based nonlinear disturbance observer for unknown payload carrying applications

ICRA 2016

This paper proposes a passivity-based nonlinear disturbance observer (DOB) design for a powered upper-limb robot control. The proposed DOB allows for the nonlinearities of the robot dynamics, whereas the typical DOB designs cannot. Moreover, by virtue of the passivity property, human operator and en

Cited by 5SourceScholar