← Search

Yutong Liu

15 accepted papers

2026

Estimation of the Caged Object's Posture Under Forces Using Stepwise Geometric Calculations

RA-L 2026

In pin-hole assembly processes, precise alignment or compliance mechanisms are typically required. This paper proposes a method for connecting objects by utilizing caging to constrain their motion, enabling the insertion of a pin into a hole to adjust the allowable relative pose for assembly. This a

Cited by 0SourceScholar
2026

Estimation of the Caged Object's Posture under Forces Using Stepwise Geometric Calculations

ICRA 2026poster

In pin–hole assembly processes, precise alignment or compliance mechanisms are typically required. This paper proposes a method for connecting objects by utilizing caging to constrain their motion, enabling the insertion of a pin into a hole to adjust the allowable relative pose for assembly. This a…

Cited by 0SourceScholar
2026

GRASP: Hard-Label Black-Box Malware Evasion with Higher Success, Fewer Queries, and Smaller Perturbations

IJCAI 2026

Machine learning (ML)-based malware detectors are widely deployed but remain vulnerable to adversarial attacks. However, under hard-label black-box access, existing adversarial attacks on Windows Portable Executable (PE) malware are often query-inefficient and incur large file-size inflation. A comm

Cited by 0Scholar
2026

LSGQuant: Layer-Sensitivity Guided Quantization for One-Step Diffusion Real-World Video Super-Resolution

ICML 2026poster

One-Step Diffusion Models have demonstrated promising capability and fast inference in real-world Video Super-Resolution (VSR). However, the substantial model size and high computational cost of Diffusion Transformers (DiTs) hinder their practical deployment. While low-bit quantization is a common a…

Cited by 0SourceScholar
2026

TMD-TTS: A UNIFIED TIBETAN MULTI-DIALECT TEXT-TO-SPEECH FRAMEWORK FOR U-TSANG, AMDO AND KHAM SPEECH DATASET GENERATION

ICASSP 2026poster

Tibetan is a low-resource language with limited parallel speech corpora spanning its three major dialects (Ü-Tsang, Amdo, and Kham), limiting progress in speech modeling. To address this issue, we propose TMD-TTS, a unified Tibetan multi-dialect text-to-speech (TTS) framework that synthesizes parall…

Cited by 0SourcePDFScholar
2025

OSDFace: One-Step Diffusion Model for Face Restoration

CVPR 2025poster

Diffusion models have demonstrated impressive performance in face restoration. Yet, their multi-step inference process remains computationally intensive, limiting their applicability in real-world scenarios. Moreover, existing methods often struggle to generate face images that are harmonious, reali…

2025

S2D-LFE: Sparse-to-Dense Light Field Event Generation

CVPR 2025poster

In this paper, we present S2D-LFE, an innovative approach for sparse-to-dense light field event generation. For the first time to our knowledge, S2D-LFE enables controllable novel view synthesis only from sparse-view light field event (LFE) data, and addresses three critical challenges for the LFE g…

2025

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

ICCV 2025poster

We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQVAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabiliti…

2025

TLUE: A Tibetan Language Understanding Evaluation Benchmark

EMNLP 2025

Large language models have made tremendous progress in recent years, but low-resource languages, like Tibetan, remain significantly underrepresented in their evaluation. Despite Tibetan being spoken by over seven million people, it has largely been neglected in the development and assessment of LLMs

2025

uMedSum: A Unified Framework for Clinical Abstractive Summarization

ACL 2025long

Clinical abstractive summarization struggles to balance faithfulness and informativeness, sacrificing key information or introducing confabulations. Techniques like in-context learning and fine-tuning have improved overall summary quality orthogonally, without considering the above issue. Conversely…

Cited by 0SourcePDFScholar
2024

Rolling-Unet: Revitalizing MLP’s Ability to Efficiently Extract Long-Distance Dependencies for Medical Image Segmentation

AAAI 2024technical

Medical image segmentation methods based on deep learning network are mainly divided into CNN and Transformer. However, CNN struggles to capture long-distance dependencies, while Transformer suffers from high computational complexity and poor local feature learning. To efficiently extract and fuse l…

Cited by 30SourcePDFScholar
2023

CutMIB: Boosting Light Field Super-Resolution via Multi-View Image Blending

CVPR 2023poster

Data augmentation (DA) is an efficient strategy for improving the performance of deep neural networks. Recent DA strategies have demonstrated utility in single image super-resolution (SR). Little research has, however, focused on the DA strategy for light field SR, in which multi-view information ut…

2022

A Fine-grained Chinese Software Privacy Policy Dataset for Sequence Labeling and Regulation Compliant Identification

EMNLP 2022main

Privacy protection raises great attention on both legal levels and user awareness. To protect user privacy, countries enact laws and regulations requiring software privacy policies to regulate their behavior. However, privacy policies are written in professional languages with many legal terms and s…

2020

Kalman Filtering Attention for User Behavior Modeling in CTR Prediction

NeurIPS 2020spotlight

Click-through rate (CTR) prediction is one of the fundamental tasks for e-commerce search engines. As search becomes more personalized, it is necessary to capture the user interest from rich behavior data. Existing user behavior modeling algorithms develop different attention mechanisms to emphasize…