← Search

Yuhan Li

23 accepted papers

2026

OmniXtreme: Breaking the Generality Barrier in High-Dynamic Humanoid Control

RSS 2026poster

High-fidelity motion tracking serves as the ultimate litmus test for generalizable, human-level motor skills. However, current policies often hit a “generality barrier”: as motion libraries scale in diversity, tracking fidelity inevitably collapses—especially for real-world deployment of high-dynami…

Cited by 0SourceScholar
2025

GraphArena: Evaluating and Exploring Large Language Models on Graph Computation

ICLR 2025poster

The ``arms race'' of Large Language Models (LLMs) demands new benchmarks to examine their progresses. In this paper, we introduce GraphArena, a benchmarking tool designed to evaluate LLMs on real-world graph computational problems. It offers a suite of four polynomial-time tasks (e.g., Shortest Dist…

2025

INFER: A Neural-symbolic Model For Extrapolation Reasoning on Temporal Knowledge Graph

ICLR 2025poster

Temporal Knowledge Graph(TKG) serves as an efficacious way to store dynamic facts in real-world. Extrapolation reasoning on TKGs, which aims at predicting possible future events, has attracted consistent research interest. Recently, some rule-based methods have been proposed, which are considered mo…

Cited by 0SourcePDFScholar
2025

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

ICCV 2025poster

Standard clothing asset generation involves restoring forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-world contexts, which presents significant challenges due to highly standardized structure sampling distributions and clot…

Cited by 0SourcePDFScholar
2025

Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps

ACL 2025finding

Low-Rank Adaptation (LoRA) has emerged as a prominent technique for fine-tuning large foundation models. Despite its successes, the substantial parameter redundancy, which limits the capacity and efficiency of LoRA, has been recognized as a bottleneck. In this work, we systematically investigate the…

Cited by 0SourcePDFScholar
2025

STORM-BORN: A Challenging Mathematical Derivations Dataset Curated via a Human-in-the-Loop Multi-Agent Framework

ACL 2025finding

High-quality math datasets are crucial for advancing the reasoning abilities of large language models (LLMs). However, existing datasets often suffer from three key issues: outdated and insufficient challenging content, neglecting human-like reasoning, and limited reliability due to single-LLM gener…

2025

ShoeFit: A New Dataset and Dual-image-stream DiT Framework for Virtual Footwear Try-On

NeurIPS 2025poster

Virtual footwear try-on (VFTON), a critical yet underexplored area in virtual try-on (VTON), aims to synthesize faithful try-on results given diverse footwear and model images while maintaining 3D consistency and texture authenticity. Unlike conventional garment-focused VTON methods, VFTON present…

Cited by 0SourceScholar
2025

SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic Priors

CVPR 2025poster

Despite significant advances in accurately estimating geometry in contemporary single-image 3D human reconstruction, creating a high-quality, efficient, and animatable 3D avatar remains an open challenge. Two key obstacles persist: incomplete observation and inconsistent 3D priors. To address these…

2025

TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-Control

ICCV 2025poster

In image editing tasks, high-quality text editing capabilities can significantly reduce both human and material resource costs. Existing methods, however, face significant limitations in terms of stroke accuracy for complex text and controllability of generated text styles. To address these challeng…

Cited by 0SourcePDFScholar
2024

A Survey of Graph Meets Large Language Model: Progress and Future Directions

IJCAI 2024poster

Graph plays a significant role in representing and analyzing complex relationships in real-world applications such as citation networks, social networks, and biological data. Recently, Large Language Models (LLMs), which have achieved tremendous success in various domains, have also been leveraged i…

2024

AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario

NeurIPS 2024poster

While image-based virtual try-on has made significant strides, emerging approaches still fall short of delivering high-fidelity and robust fitting images across various scenarios, as their models suffer from issues of ill-fitted garment styles and quality degrading during the training process, not t…

Cited by 7SourcePDFScholar
2024

FocalDreamer: Text-Driven 3D Editing via Focal-Fusion Assembly

AAAI 2024technical

While text-3D editing has made significant strides in leveraging score distillation sampling, emerging approaches still fall short in delivering separable, precise and consistent outcomes that are vital to content creation. In response, we introduce FocalDreamer, a framework that merges base shape w…

Cited by 56SourcePDFScholar
2024

Fundamental Limits of Direction Finding in Distributed Arrays Exploiting Auxiliary Sources

ICASSP 2024accepted

We consider the problem of estimating the directions of multiple target sources by exploiting auxiliary sources, focusing on a single snapshot obtained by the distributed array with position errors and angular offsets of subarrays. Former calibration methods generally assume the directions of auxili…

Cited by 0SourceScholar
2024

GLBench: A Comprehensive Benchmark for Graph with Large Language Models

NeurIPS 2024poster

The emergence of large language models (LLMs) has revolutionized the way we interact with graphs, leading to a new paradigm called GraphLLM. Despite the rapid development of GraphLLM methods in recent years, the progress and understanding of this field remain unclear due to the lack of a benchmark w…

2024

VS: Reconstructing Clothed 3D Human from Single Image via Vertex Shift

CVPR 2024poster

Various applications require high-fidelity and artifact-free 3D human reconstructions. However current implicit function-based methods inevitably produce artifacts while existing deformation methods are difficult to reconstruct high-fidelity humans wearing loose clothing. In this paper we propose a…

2023

Generalized Deep 3D Shape Prior via Part-Discretized Diffusion Process

CVPR 2023poster

We develop a generalized 3D shape generation prior model, tailored for multiple 3D tasks including unconditional shape generation, point cloud completion, and cross-modality shape generation, etc. On one hand, to precisely capture local fine detailed shape information, a vector quantized variational…

2023

Large Language Models Meet Harry Potter: A Dataset for Aligning Dialogue Agents with Characters

EMNLP 2023long findings

In recent years, Dialogue-style Large Language Models (LLMs) such as ChatGPT and GPT4 have demonstrated immense potential in constructing open-domain dialogue agents. However, aligning these agents with specific characters or individuals remains a considerable challenge due to the complexities of ch…

Cited by 0SourceScholar
2023

Sparse Bayesian Learning Based Three-Dimensional Imaging for Antenna Array Radar

ICASSP 2023accepted

In recent years, the development of compressed sensing and sparse representation provide us with a broader perspective of three-dimensional (3-D) imaging. In this work, we propose a 3-D imaging method based on a sparse Bayesian learning(SBL) framework for antenna array radar. It solves the problem o…

Cited by 0SourceScholar
2022

Community Question Answering Entity Linking via Leveraging Auxiliary Data

IJCAI 2022poster

Community Question Answering (CQA) platforms contain plenty of CQA texts (i.e., questions and answers corresponding to the question) where named entities appear ubiquitously. In this paper, we define a new task of CQA entity linking (CQAEL) as linking the textual entity mentions detected from CQA te…

2022

Sparse Modeling of The Early Part of Noisy Room Impulse Responses with Sparse Bayesian Learning

ICASSP 2022accepted

A model of a room impulse response (RIR) is useful for a wide range of applications. Typically, the early part of a RIR is sparse, and its sparse structure allows for accurate and simple modeling of the RIR. The existing ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w…

Cited by 0SourceScholar
2022

TIARA: Multi-grained Retrieval for Robust Question Answering over Large Knowledge Base

EMNLP 2022main

Pre-trained language models (PLMs) have shown their effectiveness in multiple scenarios. However, KBQA remains challenging, especially regarding coverage and generalization settings. This is due to two main factors: i) understanding the semantics of both questions and relevant knowledge from the KB;…