← Search

Seungjun Lee

19 accepted papers

2026

D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation

CVPR 2026

Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D Vision-Language-Planning Model (D3D-VLP). Our model introduces t

Cited by 0SourcecodeScholar
2026

EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding

CVPR 2026

Understanding a 3D scene immediately with its exploration is essential for embodied tasks, where an agent must construct and comprehend the 3D representation in an online and nearly real-time manner. In this study, we propose EmbodiedSplat, an online feed-forward 3DGS for open-vocabulary scene under

Cited by 0SourcecodeScholar
2026

Learning Optimal Strategies for Needle Handover in Surgical Suturing

ICRA 2026poster

Automation of suturing subtasks, such as needle handover, has the potential to reduce surgeons' fatigue and improve surgical efficiency. Needle handover is particularly challenging due to the combinatorial nature of grasping and handover strategies, uncertainties in needle pose estimation, and inacc…

Cited by 0Scholar
2026

Segment Any Events with Language

ICLR 2026poster

Scene understanding with free-form language has been widely explored within diverse modalities such as images, point clouds, and LiDAR. However, related studies on event sensors are scarce or narrowly centered on semantic-level understanding. We introduce **SEAL**, the first Semantic-aware Segment A…

Cited by 0SourceScholar
2025

DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting

CVPR 2025poster

Reconstructing sharp 3D representations from blurry multi-view images are long-standing problem in computer vision. Recent works attempt to enhance high-quality novel view synthesis from the motion blur by leveraging event-based cameras, benefiting from high dynamic range and microsecond temporal re…

Cited by 0SourcePDFScholar
2025

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation

NeurIPS 2025oral

Vision-and-Language Navigation (VLN) is a core task where embodied agents leverage their spatial mobility to navigate in 3D environments toward designated destinations based on natural language instructions. Recently, video-language large models (Video-VLMs) with strong generalization capabilities a…

Cited by 0SourcecodeScholar
2025

ImagePiece: Content-aware Re-tokenization for Efficient Image Recognition

AAAI 2025technical

Vision Transformers (ViTs) have achieved remarkable success in various computer vision tasks. However, ViTs have a huge computational cost due to their inherent reliance on multi-head self-attention (MHSA), prompting efforts to accelerate ViTs for practical applications. To this end, recent works ai…

Cited by 0SourcePDFScholar
2025

PDEfuncta: Spectrally-Aware Neural Representation for PDE Solution Modeling

NeurIPS 2025poster

Scientific machine learning often involves representing complex solution fields that exhibit high-frequency features such as sharp transitions, fine-scale oscillations, and localized structures. While implicit neural representations (INRs) have shown promise for continuous function modeling, capturi…

Cited by 0SourceScholar
2025

SCENT: Robust Spatiotemporal Learning for Continuous Scientific Data via Scalable Conditioned Neural Fields

ICML 2025poster

Spatiotemporal learning is challenging due to the intricate interplay between spatial and temporal dependencies, the high dimensionality of the data, and scalability constraints. These challenges are further amplified in scientific domains, where data is often irregularly distributed (e.g., missing…

Cited by 0SourcePDFScholar
2024

Inducing Point Operator Transformer: A Flexible and Scalable Architecture for Solving PDEs

AAAI 2024technical

Solving partial differential equations (PDEs) by learning the solution operators has emerged as an attractive alternative to traditional numerical methods. However, implementing such architectures presents two main challenges: flexibility in handling irregular and arbitrary input and output formats…

2024

KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models

ACL 2024findings

The evolution of large language models (LLMs) has culminated in a multitask model paradigm where prompts drive the generation of user-specific outputs. However, this advancement has revealed a critical challenge: LLMs frequently produce outputs against socially acceptable commonsense standards in va…

2024

Length-aware Byte Pair Encoding for Mitigating Over-segmentation in Korean Machine Translation

ACL 2024findings

Byte Pair Encoding is an effective approach in machine translation across several languages. However, our analysis indicates that BPE is prone to over-segmentation in the morphologically rich language, Korean, which can erode word semantics and lead to semantic confusion during training. This semant…

2024

Pairwise Distance Distillation for Unsupervised Real-World Image Super-Resolution

ECCV 2024poster

"Standard single-image super-resolution creates paired training data from high-resolution images through fixed downsampling kernels. However, real-world super-resolution (RWSR) faces unknown degradations in the low-resolution inputs, all the while lacking paired training data. Existing methods appro…

2024

Translation of Multifaceted Data without Re-Training of Machine Translation Systems

EMNLP 2024finding

Translating major language resources to build minor language resources becomes a widely-used approach. Particularly in translating complex data points composed of multiple components, it is common to translate each component separately. However, we argue that this practice often overlooks the interr…

Cited by 0SourcePDFScholar
2022

Consistency Learning via Decoding Path Augmentation for Transformers in Human Object Interaction Detection

CVPR 2022poster

Human-Object Interaction detection is a holistic visual recognition task that entails object detection as well as interaction classification. Previous works of HOI detection has been addressed by the various compositions of subset predictions, e.g., Image -> HO -> I, Image -> HI -> O. Recently, tran…

Cited by 31PDFcodeScholar
2022

Development of a cable-driven Growing Sling to assist patient transfer

IROS 2022poster

As the aging of society continues to accelerate, the number of elderly patients is increasing, as is the demand for manpower to care for them. In particular, there is an urgent need for bedridden patient care. However, limitations in the supply of human resources have caused an increase in the burde…

Cited by 0SourceScholar
2020

Development of a pneumatically-driven Growing Sling to assist patient transfer

IROS 2020poster

In this study, a new type of sling for assisting bedridden patients is developed using a pneumatic growing mechanism. Growing Sling focuses on minimizing the labor input of the caregivers by automating the sling insertion and retraction process while maintaining safety and comfort. Improvements over…

Cited by 8SourceScholar