← Search

Yulin Li

29 accepted papers

2026

A Reconfigured Wheel-Legged Robot for Enhanced Steering and Adaptability

RA-L 2026

Wheel-legged robots integrate leg agility on rough terrain with wheel efficiency on flat ground. However, most existing designs do not fully capitalize on the benefits of both legged and wheeled structures, which limits overall system flexibility and efficiency. We present FLORES, a novel wheel-legg

Cited by 1SourcecodeScholar
2026

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

ICML 2026poster

While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the structural mismatch between discrete textual intents and continuous acoustic realizations. Inspired by human cognitive decoupling, we introduce Agen…

Cited by 0SourceScholar
2026

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

ICML 2026poster

Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm limits their ability to handle complex prompts requiring iterative refinement. To enable multi-round Reflective Visual …

Cited by 0SourceScholar
2026

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

ICML 2026poster

While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds. These thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iteratio…

Cited by 0SourceScholar
2026

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling

ICML 2026poster

Recent advances in Large Reasoning Models (LRMs) demonstrate remarkable performance improvements by iteratively reflecting, exploring, and executing complex tasks, yet suffer from inefficiencies due to redundant reasoning, known as "overthinking". Existing methods to mitigate this issue either rely …

Cited by 0SourceScholar
2026

Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM

ICLR 2026poster

Diffusion Large Language Models (dLLMs) offer a promising alternative to autoregressive models, excelling in text generation tasks due to their bidirectional attention mechanisms. However, their computational complexity, scaling as $\mathcal{O}(L^3)$ with sequence length $L$, poses significant chall…

Cited by 0SourcecodeScholar
2026

Efficient Reasoning with Balanced Thinking

ICLR 2026poster

Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they often suffer from overthinking, expending redundant computational steps on simple problems, or underthinking, failing to explore sufficient reasoning paths despite inherent capabilities. These issues lead to ineffic…

Cited by 0SourcecodeScholar
2026

Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning

ICRA 2026poster

Whole-body manipulation (WBM) for humanoid robots presents a promising approach for executing embracing tasks involving bulky objects, where traditional grasping relying on end-effectors only remains limited in such scenarios due to inherent stability and payload constraints. This paper introduces a…

2026

FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

ICLR 2026oral

Although Video Large Language Models (VLLMs) have shown remarkable capabilities in video understanding, they are required to process high volumes of visual tokens, causing significant computational inefficiency. Existing VLLMs acceleration frameworks usually compress spatial and temporal redundancy…

Cited by 0SourcecodeScholar
2026

MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks. However, they remain brittle on mechanical engineering drawings, where high annotation density and weak domain knowledge, compounded by unreliable spatial relation re…

Cited by 0SourceScholar
2026

One-Step Model Predictive Path Integral for Manipulator Motion Planning Using Configuration Space Distance Fields

ICRA 2026poster

Motion planning for robotic manipulators is a fundamental problem in robotics. Classical optimization-based methods typically rely on the gradients of signed distance fields (SDF) to impose collision-avoidance constraints. However, these methods are susceptible to local minima and may fail when the …

2026

Online Trajectory Optimization for Arbitrary-Shaped Mobile Robots via Polynomial Separating Hypersurfaces

RA-L 2026

An emerging class of trajectory optimization methods enforces collision avoidance by jointly optimizing the robot's configuration and a separating hyperplane. However, as linear separators only apply to convex sets, these methods require convex approximations of both the robot and obstacles, which b

Cited by 1SourceScholar
2025

A Bio-inspired Spherical Soft Magnetic Millirobot for Gastrointestinal Applications

IROS 2025

Gastroscopy and colonoscopy have become the fundamental tools for gastrointestinal (GI) tract diagnosis and treatment. Conventional tethered devices usually lead to the use of anesthetic agents and patient discomfort. Capsule endoscopy is becoming an ideal alternative, however, the smooth capsule sh

Cited by 0SourceScholar
2025

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception

CVPR 2025poster

Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have shown promise in open-vocabulary tasks, their direct applicatio…

2025

FRTree Planner: Robot Navigation in Cluttered and Unknown Environments With Tree of Free Regions

RA-L 2025

In this work, we present FRTree planner, a novel robot navigation framework that leverages a tree structure of free regions, specifically designed for navigation in cluttered and unknown environments with narrow passages. The framework continuously incorporates real-time perceptive information to id

Cited by 6SourceScholar
2025

G2-SDF: Geometry-Guided Neural Signed Distance Fields for Scalable and Detailed Reconstruction

RA-L 2025

Effcient reconstruction methods, particularly capable of providing detailed information on obstacle distances across diverse environments, are crucial for effective robot motion planning. In this context, neural Signed Distance Fields (SDFs) offer a powerful solution by learning implicit representat

Cited by 1SourceScholar
2025

Interactive Navigation for Legged Manipulators with Learned Arm-Pushing Controller

IROS 2025

Interactive navigation is crucial in scenarios where proactively interacting with objects can yield shorter paths, thus significantly improving traversal efficiency. Existing methods primarily focus on using the robot body to relocate obstacles during navigation. However, they prove ineffective in n

Cited by 5SourcecodeScholar
2025

Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior

NeurIPS 2025poster

Recent advances in Video Large Language Models (VLLMs) have achieved remarkable video understanding capabilities, yet face critical efficiency bottlenecks due to quadratic computational growth with lengthy visual token sequences of long videos. While existing keyframe sampling methods can improve te…

Cited by 0SourceScholar
2025

Local Reactive Control for Mobile Manipulators With Whole-Body Safety in Complex Environments

RA-L 2025

Mobile manipulators typically encounter significant challenges in navigating narrow, cluttered environments due to their high-dimensional state spaces and complex kinematics. While reactive methods excel in dynamic settings, they struggle to efficiently incorporate complex, coupled constraints acros

Cited by 5SourceScholar
2025

Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models

IROS 2025

Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing for precise prediction of action trajectories. However, diffusion models typically rely on large parameter UNet backbones as policy networks, which can be cha

Cited by 22SourcecodeScholar
2025

OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision

AAAI 2025technical

Open-vocabulary detection aims to detect objects from novel categories beyond the base categories on which the detector is trained. However, existing open-vocabulary detectors trained on base category data tend to assign higher confidence to trained categories and confuse novel categories with the b…

2025

On the Surprising Robustness of Sequential Convex Optimization for Contact-Implicit Motion Planning

RSS 2025poster

Contact-implicit motion planning—embedding contact sequencing as implicit complementarity constraints—holds the promise of leveraging continuous optimization to discover new contact patterns online. Nevertheless, the resulting optimization, being an instance of Mathematical Programming with Compleme…

Cited by 0PDFScholar
2025

Robot Navigation in Unknown and Cluttered Workspace with Dynamical System Modulation in Starshaped Roadmap

ICRA 2025

Compared to conventional decomposition methods that use ellipses or polygons to represent free space, starshaped representation can better capture the natural distribution of sensor data, thereby exploiting a larger portion of traversable space. This paper introduces a novel motion planning and cont

Cited by 3SourcecodeScholar
2024

Chance-Aware Lane Change with High-Level Model Predictive Control Through Curriculum Reinforcement Learning

ICRA 2024poster

Lane change in dense traffic typically requires the recognition of an appropriate opportunity for maneuvers, which remains a challenging problem in self-driving. In this work, we propose a chance-aware lane-change strategy with high-level model predictive control (MPC) through curriculum reinforceme…

Cited by 7SourceScholar
2024

Collision-Free Trajectory Optimization in Cluttered Environments Using Sums-of-Squares Programming

RA-L 2024

In this work, we propose a trajectory optimization approach for robot navigation in cluttered 3D environments. We represent the robot's geometry as a semialgebraic set defined by polynomial inequalities such that robots with general shapes can be suitably characterized. We exploit the collision-free

Cited by 12SourcecodeScholar
2024

Geometry-Aware Safety-Critical Local Reactive Controller for Robot Navigation in Unknown and Cluttered Environments

RA-L 2024

This work proposes a safety-critical local reactive controller that enables the robot to navigate in unknown and cluttered environments. In particular, the trajectory tracking task is formulated as a constrained polynomial optimization problem. Then, safety constraints are imposed on the control var

Cited by 13SourceScholar
2023

Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding

IJCAI 2023poster

Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are challenging to be directly adapted to model document. They are unabl…

Cited by 8SourcePDFScholar
2023

StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

ICLR 2023poster

In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-training tasks: masked image modeling and masked language modeling, based on text region-level image masking. The proposed…

2021

Diverse Part Discovery: Occluded Person Re-Identification With Part-Aware Transformer

CVPR 2021poster

Occluded person re-identification (Re-ID) is a challenging task as persons are frequently occluded by various obstacles or other persons, especially in the crowd scenario. To address these issues, we propose a novel end-to-end Part-Aware Transformer (PAT) for occluded person Re-ID through diverse pa…

Cited by 433PDFScholar