← Search

Jing Lin

31 accepted papers

2026

ASTRAEA: A Token-wise Acceleration Framework for Video Diffusion Transformers

ICLR 2026poster

Video diffusion transformers (vDiTs) have made tremendous progress in text-to-video generation, but their high computational demands pose a major challenge for practical deployment. While existing studies propose acceleration methods to reduce workload at various granularities, they often rely on he…

Cited by 0SourceScholar
2026

EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer

ICLR 2026poster

Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This limitation stems from the intrinsic constraints of pixel-only training objectives, which inherently bias models toward app…

Cited by 0SourceScholar
2026

FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference

ICLR 2026poster

Large language models (LLMs) have been widely deployed with rapidly expanding context windows to support increasingly demanding applications. However, long contexts pose significant deployment challenges, primarily due to the KV cache whose size grows proportionally with context length. While KV cac…

Cited by 0SourcecodeScholar
2026

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

ICLR 2026poster

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most notably video generation (ViGen), have demonstrated remarkable generalization in…

Cited by 0SourcecodeScholar
2026

TimeRipples: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space

CVPR 2026

The recent surge in video generation has shown the growing demand for high-quality video synthesis using large vision models. Existing video generation models are predominantly based on the video diffusion transformer (vDiT), however, they suffer from substantial inference delay due to self-attentio

Cited by 0SourceScholar
2026

TopoMesh: High-Fidelity Mesh Autoencoding via Topological Unification

CVPR 2026

The dominant paradigm for high-fidelity 3D generation relies on a VAE-Diffusion pipeline, where the VAE's reconstruction capability sets a firm upper bound on generation quality. A fundamental challenge limiting existing VAEs is the representation mismatch between ground-truth meshes and network pre

Cited by 0SourceScholar
2026

UniMo: Unified Motion Generation and Understanding with Chain of Thought

AAAI 2026technical

Existing 3D human motion generation and understanding methods often exhibit limited interpretability, restricting effective mutual enhancement between these inherently related tasks. While current unified frameworks based on large language models (LLMs) leverage linguistic priors, they frequently en

Cited by 0SourcePDFScholar
2026

UniSketch: A Unified Framework for Parametric Sketch Generation and Constraint Prediction

AAAI 2026technical

In modern Computer-Aided Design (CAD), parametric sketches play a crucial role by capturing both the geometric structure and design intent through constraints. However, existing deep learning–based sketch methods remain restricted to simple geometric primitives and limited constraint types, hinderin

Cited by 0SourcePDFScholar
2026

ViLearn: Accelerating Training Convergence of Image-to-3D Generation via Visibility Learning

CVPR 2026

Single-image-to-3D shape generation has seen remarkable progress, driven by latent diffusion models trained on the compressed latent space of 3D VAEs. However, the task remains intrinsically ill-posed: recovering complete 3D geometry--especially occluded surfaces--from a single view is inherently am

Cited by 0SourceScholar
2025

DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior

ICCV 2025poster

We present DPoser-X, a diffusion-based prior model for 3D whole-body human poses. Building a versatile and robust full-body human pose prior remains challenging due to the inherent complexity of articulated human poses and the scarcity of high-quality whole-body pose datasets. To address these limit…

Cited by 0SourcePDFScholar
2025

DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization

NeurIPS 2025poster

Quantization plays a crucial role in accelerating the inference of large-scale models, and rotational matrices have been shown to effectively improve quantization performance by smoothing outliers. However, end-to-end fine-tuning of rotational optimization algorithms incurs high computational costs…

Cited by 0SourceScholar
2025

Dynamic Low-Rank Sparse Adaptation for Large Language Models

ICLR 2025poster

Despite the efficacy of network sparsity in alleviating the deployment strain of Large Language Models (LLMs), it endures significant performance degradation. Applying Low-Rank Adaptation (LoRA) to fine-tune the sparse LLMs offers an intuitive approach to counter this predicament, while it holds sho…

2025

HumanMM: Global Human Motion Recovery from Multi-shot Videos

CVPR 2025poster

In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to applications such as motion generation and motion understan…

2025

Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture

CVPR 2025poster

Existing methods for capturing multi-person holistic human motions from a monocular video usually involve integrating the detector, the tracker, and the human pose & shape estimator into a cascaded system. Differently, we develop a one-stage multi-person holistic human motion capture system, which 1…

2025

RazorAttention: Efficient KV Cache Compression Through Retrieval Heads

ICLR 2025poster

The memory and computational demands of Key-Value (KV) cache present significant challenges for deploying long-context language models. Previous approaches attempt to mitigate this issue by selectively dropping tokens, which irreversibly erases critical information that might be needed for future qu…

Cited by 23SourcePDFScholar
2025

Revisiting and Refining Lagunas' Beamforming for Acoustic Imaging

ICASSP 2025accepted

Lagunas et al proposed an adaptive beamforming method in 1986. However, this method never receives attention: until 2024, the citation of this paper is only 66. Actually, the Lagunas’ beamforming failed to identify sound sources if the required covariance matrix of array measurements is derived from…

Cited by 0SourceScholar
2025

SkillMimic: Learning Basketball Interaction Skills from Demonstrations

CVPR 2025highlight

Traditional reinforcement learning methods for human-object interaction (HOI) rely on labor-intensive, manually designed skill rewards that do not generalize well across different interactions. We introduce SkillMimic, a unified data-driven framework that fundamentally changes how agents learn inter…

2025

VaporTok: RL-Driven Adaptive Video Tokenizer with Prior & Task Awareness

NeurIPS 2025poster

Recent advances in visual tokenizers have demonstrated their effectiveness for multimodal large language models and autoregressive generative models. However, most existing visual tokenizers rely on a fixed downsampling rate at a given visual resolution, and consequently produce a constant number of…

Cited by 0SourceScholar
2024

ChatPose: Chatting about 3D Human Pose

CVPR 2024poster

We introduce ChatPose a framework employing Large Language Models (LLMs) to understand and reason about 3D human poses from images or textual descriptions. Our work is motivated by the human ability to intuitively understand postures from a single image or a brief description a process that intertwi…

2024

HumanTOMATO: Text-aligned Whole-body Motion Generation

ICML 2024poster

This work targets a novel text-driven **whole-body** motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and body motions simultaneously. Previous works on text-driven motion generation…

2023

Binarized Spectral Compressive Imaging

NeurIPS 2023poster

Existing deep learning models for hyperspectral image (HSI) reconstruction achieve good performance but require powerful hardwares with enormous memory and computational resources. Consequently, these methods can hardly be deployed on resource-limited mobile devices. In this paper, we propose a nove…

2023

One-Stage 3D Whole-Body Mesh Recovery With Component Aware Transformer

CVPR 2023poster

Whole-body mesh recovery aims to estimate the 3D human body, face, and hands parameters from a single image. It is challenging to perform this task with a single network due to resolution issues, i.e., the face and hands are usually located in extremely small regions. Existing works usually detect h…

2023

Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement

ICCV 2023poster

When enhancing low-light images, many deep learning algorithms are based on the Retinex theory. However, the Retinex model does not consider the corruptions hidden in the dark or introduced by the light-up process. Besides, these methods usually require a tedious multi-stage training pipeline and re…

Cited by 410PDFcodeScholar
2022

Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction

ECCV 2022poster

"Many learning-based algorithms have been developed to solve the inverse problem of coded aperture snapshot spectral imaging (CASSI). However, CNN-based methods show limitations in capturing long-range dependencies. Previous Transformer-based methods densely sample tokens, some of which are uninform…

2022

Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging

NeurIPS 2022accept

In coded aperture snapshot spectral compressive imaging (CASSI) systems, hyperspectral image (HSI) reconstruction methods are employed to recover the spatial-spectral signal from a compressed measurement. Among these algorithms, deep unfolding methods demonstrate promising performance but suffer fro…

2022

Flow-Guided Sparse Transformer for Video Deblurring

ICML 2022spotlight

Exploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Tra…

2022

HDNet: High-Resolution Dual-Domain Learning for Spectral Compressive Imaging

CVPR 2022poster

The rapid development of deep learning provides a better solution for the end-to-end reconstruction of hyperspectral image (HSI). However, existing learning-based methods have two major defects. Firstly, networks with self-attention usually sacrifice internal resolution to balance model performance…

Cited by 188PDFcodeScholar
2022

Mask-Guided Spectral-Wise Transformer for Efficient Hyperspectral Image Reconstruction

CVPR 2022poster

Hyperspectral image (HSI) reconstruction aims to recover the 3D spatial-spectral signal from a 2D measurement in the coded aperture snapshot spectral imaging (CASSI) system. The HSI representations are highly similar and correlated across the spectral dimension. Modeling the inter-spectra interactio…

Cited by 333PDFcodeScholar
2022

Unsupervised Flow-Aligned Sequence-to-Sequence Learning for Video Restoration

ICML 2022spotlight

How to properly model the inter-frame relation within the video sequence is an important but unsolved challenge for video restoration (VR). In this work, we propose an unsupervised flow-aligned sequence-to-sequence model (S2SVR) to address this problem. On the one hand, the sequence-to-sequence mode…

2019

Joint Torque Estimation toward Dynamic and Compliant Control for Gear-Driven Torque Sensorless Quadruped Robot

IROS 2019poster

This paper investigates dynamic and compliant control based on joint output torque estimation for electrically actuated quadruped robots with large-reduction-ratio harmonic gear. Compared with position control, force control exhibits better performance of dynamics and compliance for the robot's inte…

Cited by 29SourceScholar