← Search

Minhae Kwon

7 accepted papers

2026

$\texttt{Multi}^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments

ICML 2026poster

A central goal of large language model (LLM) research is to build agentic systems that can plan, act, and adapt through sustained interaction with dynamic environments. While recent LLM-based agents exhibit impressive contextual reasoning, their long-horizon decision-making remains fragile, often su…

Cited by 0SourceScholar
2026

Fed-ADE: Adaptive Learning Rate for Federated Post-adaptation under Distribution Shift

CVPR 2026

Federated learning (FL) in post-deployment settings must adapt to non-stationary data streams across heterogeneous clients without access to ground-truth labels. A major challenge is learning rate selection under client-specific, time-varying distribution shifts, where fixed learning rates often lea

Cited by 2SourcecodeScholar
2026

Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap Balancing

ICML 2026poster

Model merging has become a practical post-training strategy for building a single multi-task large language model (LLM) by combining multiple task-specialized models, avoiding costly joint training. However, most existing approaches rely on post-hoc merging, in which task-specific models are merged …

Cited by 0SourceScholar
2025

Temporal Distance-aware Transition Augmentation for Offline Model-based Reinforcement Learning

ICML 2025poster

The goal of offline reinforcement learning (RL) is to extract the best possible policy from the previously collected dataset considering the *out-of-distribution* (OOD) sample issue. Offline model-based RL (MBRL) is a captivating solution capable of alleviating such issues through a \textit{state-ac…

Cited by 0SourcePDFScholar
2024

AD4RL: Autonomous Driving Benchmarks for Offline Reinforcement Learning with Value-based Dataset

ICRA 2024poster

Offline reinforcement learning has emerged as a promising technology by enhancing its practicality through the use of pre-collected large datasets. Despite its practical benefits, most algorithm development research in offline reinforcement learning still relies on game tasks with synthetic datasets…

Cited by 12SourceScholar
2020

Inverse Rational Control with Partially Observable Continuous Nonlinear Dynamics

NeurIPS 2020poster

A fundamental question in neuroscience is how the brain creates an internal model of the world to guide actions using sequences of ambiguous sensory information. This is naturally formulated as a reinforcement learning problem under partial observations, where an agent must estimate relevant latent…

Cited by 41SourcePDFScholar