← Search

Yuqi Li

29 accepted papers

2026

A Comprehensive Survey of Interaction Techniques in 3D Scene Generation

IJCAI 2026

The rapid evolution of 3D scene generation has revolutionized content creation across domains such as gaming, film production, and architectural visualization. Within this landscape, interaction techniques serve as the pivotal bridge connecting user intent with generative models, enabling precise co

Cited by 0Scholar
2026

A Pseudo-Label Optimization Method Based on Polar Coordinate Modeling and Prior Constraints

AAAI 2026technical

Magnetic Resonance Imaging (MRI) and its automatic segmentation are pivotal in assisting physicians with clinical diagnosis. In recent years, with the scarcity of labeled data, significant advancements have been made in semi-supervised segmentation. However, the prediction of many current methods is

Cited by 0SourcePDFScholar
2026

DISTILLING TIME SERIES FOUNDATION MODELS FOR EFFICIENT FORECASTING

ICASSP 2026poster

Time Series foundation models (TSFMs) deliver strong forecasting performance through large-scale pretraining, but their large parameter sizes make deployment costly. While knowledge distillation offers a natural and effective approach for model compression, techniques developed for general machine l…

Cited by 0SourcePDFScholar
2026

Fast-SAM3D: 3Dfy Anything in Images but Faster

ICML 2026poster

SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this work, we conduct the **first systematic investigation** into its inference dynamics, revealing that generic acceleration strategies are brittle in this c…

Cited by 0SourceScholar
2026

KiRAS: Keyframe Guided Self-Imitation for Robust and Adaptive Skill Learning in Quadruped Robots

ICRA 2026poster

With advances in reinforcement learning and imitation learning, quadruped robots can acquire diverse skills within a single policy by imitating multiple skill-specific datasets. However, the lack of datasets on complex terrains limits the ability of such multi-skill policies to generalize effectivel…

2026

LEVERAGING LARGE MULTIMODAL MODELS FOR AUDIO-VIDEO DEEPFAKE DETECTION: A PILOT STUDY

ICASSP 2026oral

Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-specific models: they work well on curated tests but scale poorly and generalize weakly across domains. We introduce AV-LM…

Cited by 0SourcePDFScholar
2026

MUJICA: Multi-Skill Unified Joint Integration of Control Architecture for Wheeled-Legged Robots

ICRA 2026poster

Wheeled-legged robots hold promise for traversing complex terrains and offer superior mobility compared to legged robots. However, wheeled-legged robots must effectively balance both wheeled driving and legged control. Furthermore, due to noisy proprioceptive sensing and real-world motor constraints…

2026

QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification

ICLR 2026poster

Diffusion transformers exhibit remarkable video generation capability, yet their prohibitive computational and memory costs hinder practical deployment. Model quantization and attention sparsification are two promising directions for compression, but each alone suffers severe performance degradation…

Cited by 0SourcecodeScholar
2026

Quantized Visual Geometry Grounded Transformer

ICLR 2026poster

Learning-based 3D reconstruction models, represented by Visual Geometry Grounded Transformers (VGGTs), have achieved remarkable progress with large-scale transformers. Their prohibitive computational and memory costs severely hinder real-world deployment. Post-Training Quantization (PTQ) has emerged…

Cited by 0SourcecodeScholar
2026

SepPrune: Structured Pruning for Efficient Deep Speech Separation

AAAI 2026technical

Although deep learning has substantially advanced speech separation in recent years, most existing studies continue to prioritize separation quality while overlooking computational efficiency, an essential factor for low-latency speech processing in real-time applications. In this paper, we propose

Cited by 0SourcePDFScholar
2026

WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

ICML 2026poster

Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts. While feature caching can accelerate inference without training, we find that policies designed for single-modal diffus…

Cited by 0SourceScholar
2025

$\text{S}^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation

NeurIPS 2025poster

Diffusion transformers have emerged as the mainstream paradigm for video generation models. However, the use of up to billions of parameters incurs significant computational costs. Quantization offers a promising solution by reducing memory usage and accelerating inference. Nonetheless, we observe t…

Cited by 0SourcecodeScholar
2025

Cross-Layer Graph Knowledge Distillation for Image Recognition

ICASSP 2025accepted

Knowledge Distillation (KD) aims to improve a light-weight student network supervised by a large teacher network. The core idea of KD is to explore valuable knowledge from the teacher. Previous works often extract information from a single sample, but ignore relation modeling among multiple samples…

Cited by 0SourceScholar
2025

Enhancing Image Generation Fidelity via Progressive Prompts

ICASSP 2025accepted

Diffusion transformer (DiT) architecture catches much attention in image generation, which achieves better fidelity, performance, and diversity. However, most existing DiT-based image generation methods are global-aware synthesis and regional prompt control is less explored. In this paper, we propos…

Cited by 0SourceScholar
2025

Few-Shot Domain Adaptation for Learned Image Compression

AAAI 2025technical

Learned image compression (LIC) has achieved state-of-the-art rate-distortion performance, deemed promising for next-generation image compression techniques. However, pre-trained LIC models usually suffer from significant performance degradation when applied to out-of-training-domain images, implyin…

Cited by 0SourcePDFScholar
2025

Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting

ICCV 2025poster

Spatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation, which trans…

2025

Learned Image Compression with Hierarchical Progressive Context Modeling

ICCV 2025poster

Context modeling is essential in learned image compression for accurately estimating the distribution of latents. While recent advanced methods have expanded context modeling capacity, they still struggle to efficiently exploit long-range dependency and diverse context information across different c…

2025

Prototype-Driven Multi-Feature Generation for Visible-Infrared Person Re-identification

ICASSP 2025accepted

The primary challenges in visible-infrared person re-identification arise from the differences between visible (vis) and infrared (ir) images, including inter-modal and intra-modal variations. These challenges are further complicated by varying viewpoints and irregular movements. Existing methods of…

Cited by 0SourceScholar
2025

SFE-Net: Harnessing Biological Principles of Differential Gene Expression for Improved Feature Selection in Deep Learning Networks

ICASSP 2025accepted

In the realm of DeepFake detection, the challenge of adapting to various synthesis methodologies such as Faceswap, Deepfakes, Face2Face, and NeuralTextures significantly impacts the performance of traditional machine learning models. These models often suffer from static feature representation, whic…

Cited by 0SourceScholar
2025

SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech

ICASSP 2025accepted

This paper presents SpecWav-Attack, an adversarial model for detecting speakers in anonymized speech. It leverages Wav2Vec2 for feature extraction [1] and incorporates spectrogram resizing and incremental training for improved performance. Evaluated on librispeech-dev and librispeech-test, SpecWav-A…

Cited by 0SourceScholar
2025

Towards Zero-shot Cross-lingual SLU with Syntax-aware Multi-view Contrastive Learning

ICASSP 2025accepted

Recent state-of-the-art zero-shot cross-lingual spoken language understanding (SLU) models utilize contrastive learning to achieve multilingual semantics alignment between the original utterance and code-switched counterpart. Despite achieving promising results, we discover that they still suffer fr…

Cited by 0SourceScholar
2024

BatteryML: An Open-source Platform for Machine Learning on Battery Degradation

ICLR 2024spotlight

Battery degradation remains a pivotal concern in the energy storage domain, with machine learning emerging as a potent tool to drive forward insights and solutions. However, this intersection of electrochemical science and machine learning poses complex challenges. Machine learning experts often gra…

2024

Distill Vision Transformers to CNNs via Teacher Collaboration

ICASSP 2024accepted

The vision transformer (ViT) has recently emerged as a leading approach in various domains, outperforming other methods. Therefore, it is logical to explore the possibility of transferring the superior knowledge from ViT to more compact and cost-effective convolutional neural networks (CNNs). Howeve…

Cited by 0SourceScholar
2023

Scale-Adaptive Tiny Object Detection Enhanced by Across-Scale and Shape-Preserved Semantic Location

ICASSP 2023accepted

In tiny object detection, the main challenges are tiny objects’ weak feature responses and possible semantic disappearance in deep networks. To address the problems, we proposed an Instance-level, Scale-adaptive, Shape-preserved, and Semantic-consistent Supervision (I4S) module for better locating t…

Cited by 0SourceScholar
2021

Single-Shot Hyperspectral-Depth Imaging With Learned Diffractive Optics

ICCV 2021poster

Imaging depth and spectrum have been extensively studied in isolation from each other for decades. Recently, hyperspectral-depth (HS-D) imaging emerges to capture both information simultaneously by combining two different imaging systems; one for depth, the other for spectrum. While being accurate,…

Cited by 155PDFScholar
2019

GAN-Based Projector for Faster Recovery With Convergence Guarantees in Linear Inverse Problems

ICCV 2019poster

A Generative Adversarial Network (GAN) with generator G trained to model the prior of images has been shown to perform better than sparsity-based regularizers in ill-posed inverse problems. Here, we propose a new method of deploying a GAN-based prior to solve linear inverse problems using projected…

Cited by 77PDFScholar