← Search

Chengjiang Long

33 accepted papers

2026

FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation

CVPR 2026

We present FlexAvatar, a flexible large reconstruction model for high-fidelity 3D head avatars with detailed dynamic deformation from single or sparse images, without requiring camera poses or expression labels. It leverages a transformer-based reconstruction model with structured head query tokens

Cited by 0SourceScholar
2026

Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability

ICLR 2026poster

Temporal Video Grounding (TVG) aims to localize video segments corresponding to a given textual query, which often describes human actions. However, we observe that current methods, usually optimizing for high temporal Intersection-over-Union (IoU), frequently struggle to accurately recognize or und…

Cited by 0SourceScholar
2026

T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding

CVPR 2026

Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dynamics. Existing Vision-LMMs typically perceive temporal dynamics via positional encoding, text-based timestamps, or visu

Cited by 0SourceScholar
2025

DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse Rendering

ICCV 2025poster

Recent methods have shown that pre-trained diffusion models can be fine-tuned to enable generative inverse rendering by learning image-conditioned noise-to-intrinsic mapping. Despite their remarkable progress, they struggle to robustly produce high-quality results as the noise-to-intrinsic paradigm…

2025

PUMPS: Skeleton-Agnostic Point-based Universal Motion Pre-Training for Synthesis in Human Motion Tasks

ICCV 2025poster

Motion skeletons drive 3D character animation by transforming bone hierarchies, but differences in proportions or structure make motion data hard to transfer across skeletons, posing challenges for data-driven motion synthesis. Temporal Point Clouds (TPCs) offer an unstructured, cross-compatible mot…

2025

PhysSplat: Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting

ICCV 2025poster

Recent advancements in 3D generation models have opened new possibilities for simulating dynamic 3D object movements and customizing behaviors, yet creating this content remains challenging. Current methods often require manual assignment of precise physical properties for simulations or rely on vid…

Cited by 0SourcePDFScholar
2025

Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation

ICCV 2025poster

Text-to-image diffusion models have seen significant development recently due to increasing availability of paired 2D data. Although a similar trend is emerging in 3D generation, the limited availability of high-quality 3D data has resulted in less competitive 3D diffusion models compared to their 2…

Cited by 0SourcePDFScholar
2024

CoreRec: A Counterfactual Correlation Inference for Next Set Recommendation

AAAI 2024technical

Next set recommendation aims to predict the items that are likely to be bought in the next purchase. Central to this endeavor is the task of capturing intra-set and cross-set correlations among items. However, the modeling of cross-set correlations poses challenges due to specific issues. Primarily,…

Cited by 0SourcePDFScholar
2024

Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery

NeurIPS 2024poster

Human Mesh Recovery (HMR) is the task of estimating a parameterized 3D human mesh from an image. There is a kind of methods first training a regression model for this problem, then further optimizing the pretrained regression model for any specific sample individually at test time. However, the pret…

2024

Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection

ECCV 2024poster

"Video Anomaly Detection (VAD) has been extensively studied under the settings of One-Class Classification (OCC) and Weakly-Supervised learning (WS), which however both require laborious human-annotated normal/abnormal labels. In this paper, we study Unsupervised VAD (UVAD) that does not depend on a…

2024

Motion Keyframe Interpolation for Any Human Skeleton using Point Cloud-based Human Motion Data Homogenisation

ECCV 2024poster

"In the character animation field, modern supervised keyframe interpolation models have demonstrated exceptional performance in constructing natural human motions from sparse pose definitions. As supervised models, large motion datasets are necessary to facilitate the learning process; however, sinc…

Cited by 0SourcePDFScholar
2024

Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses

ECCV 2024poster

"Besides a 3D mesh, Human Mesh Recovery (HMR) methods usually need to estimate a camera for computing 2D reprojection loss. Previous approaches may encounter the following problem: both the mesh and camera are not correct but the combination of them can yield a low reprojection loss. To alleviate th…

2023

Continuous Intermediate Token Learning With Implicit Motion Manifold for Keyframe Based Motion Interpolation

CVPR 2023poster

Deriving sophisticated 3D motions from sparse keyframes is a particularly challenging problem, due to continuity and exceptionally skeletal precision. The action features are often derivable accurately from the full series of keyframes, and thus, leveraging the global context with transformers has b…

2023

Discriminative Active Learning for Robotic Grasping in Cluttered Scene

RA-L 2023

Robotic grasping is a challenging task due to the diversity of object shapes. A sufficiently labeled dataset is essential for the grasp pose detection methods based on deep learning. However, data annotation is a costly procedure. Active learning aims to mitigate the greedy need for massive labeled

Cited by 18SourceScholar
2023

Feature Representation Learning With Adaptive Displacement Generation and Transformer Fusion for Micro-Expression Recognition

CVPR 2023poster

Micro-expressions are spontaneous, rapid and subtle facial movements that can neither be forged nor suppressed. They are very important nonverbal communication clues, but are transient and of low intensity thus difficult to recognize. Recently deep learning based methods have been developed for micr…

Cited by 45SourcePDFScholar
2022

CPRAL: Collaborative Panoptic-Regional Active Learning for Semantic Segmentation

AAAI 2022technical

Acquiring the most representative examples via active learning (AL) can benefit many data-dependent computer vision tasks by minimizing efforts of image-level or pixel-wise annotations. In this paper, we propose a novel Collaborative Panoptic-Regional Active Learning framework (CPRAL) to address the…

Cited by 18SourcePDFScholar
2022

Complementary Attention Gated Network for Pedestrian Trajectory Prediction

AAAI 2022technical

Pedestrian trajectory prediction is crucial in many practical applications due to the diversity of pedestrian movements, such as social interactions and individual motion behaviors. With similar observable trajectories and social environments, different pedestrians may make completely different futu…

2022

Deep Image-Based Illumination Harmonization

CVPR 2022poster

Integrating a foreground object into a background scenewith illumination harmonization is an important but chal-lenging task in computer vision and augmented reality community. Existing methods mainly focus on foreground andbackground appearance consistency or the foreground object shadow generation…

Cited by 18PDFcodeScholar
2022

Progressively Generating Better Initial Guesses Towards Next Stages for High-Quality Human Motion Prediction

CVPR 2022poster

This paper presents a high-quality human motion prediction method that accurately predicts future human poses given observed ones. Our method is based on the observation that a good initial guess of the future poses is very helpful in improving the forecasting accuracy. This motivates us to propose…

Cited by 140PDFcodeScholar
2022

Social Interpretable Tree for Pedestrian Trajectory Prediction

AAAI 2022technical

Understanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social Interpretable Tree (SIT), to address this multi-modal prediction task, where a hand-crafted tree is built depending on th…

2022

Video Shadow Detection via Spatio-Temporal Interpolation Consistency Training

CVPR 2022poster

It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results. In this paper, we address these challenges by proposing a Sp…

Cited by 21PDFcodeScholar
2021

A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action Localization

AAAI 2021technical

Weakly supervised temporal action localization is a challenging vision task due to the absence of ground-truth temporal locations of actions in the training videos. With only video-level supervision during training, most existing methods rely on a Multiple Instance Learning (MIL) framework to predic…

2021

A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame Prediction

ICCV 2021poster

In this paper, we propose HF2-VAD, a Hybrid framework that integrates Flow reconstruction and Frame prediction seamlessly to handle Video Anomaly Detection. Firstly, we design the network of ML-MemAE-SC (Multi-Level Memory modules in an Autoencoder with Skip Connections) to memorize normal patterns…

Cited by 278PDFcodeScholar
2021

DRB-GAN: A Dynamic ResBlock Generative Adversarial Network for Artistic Style Transfer

ICCV 2021poster

In this work, we propose a Dynamic ResBlock Generative Adversarial Network (DRB-GAN) for artistic style transfer. The style code is modeled as the shared parameters for Dynamic ResBlocks connecting both the style encoding network and the style transfer network. In the style encoding network, a style…

Cited by 117PDFcodeScholar
2021

MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion Prediction

ICCV 2021poster

Human motion prediction is a challenging task due to the stochasticity and aperiodicity of future poses. Recently, graph convolutional network has been proven to be very effective to learn dynamic relations among pose joints, which is helpful for pose prediction. On the other hand, one can abstract…

Cited by 262PDFcodeScholar
2021

SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory Prediction

CVPR 2021poster

Pedestrian trajectory prediction is a key technology in autopilot, which remains to be very challenging due to complex interactions between pedestrians. However, previous works based on dense undirected interaction suffer from modeling superfluous interactions and neglect of trajectory motion tenden…

Cited by 327PDFcodeScholar
2020

ARShadowGAN: Shadow Generative Adversarial Network for Augmented Reality in Single Light Scenes

CVPR 2020poster

Generating virtual object shadows consistent with the real-world environment shading effects is important but challenging in computer vision and augmented reality applications. To address this problem, we propose an end-to-end Generative Adversarial Network for shadow generation named ARShadowGAN fo…

Cited by 106PDFcodeScholar
2020

DOA-GAN: Dual-Order Attentive Generative Adversarial Network for Image Copy-Move Forgery Detection and Localization

CVPR 2020poster

Images can be manipulated for nefarious purposes to hide content or to duplicate certain objects through copy-move operations. Discovering a well-crafted copy-move forgery in images can be very challenging for both humans and machines; for example, an object on a uniform background can be replaced b…

Cited by 163PDFScholar
2019

ARGAN: Attentive Recurrent Generative Adversarial Network for Shadow Detection and Removal

ICCV 2019poster

In this paper we propose an attentive recurrent generative adversarial network (ARGAN) to detect and remove shadows in an image. The generator consists of multiple progressive steps. At each step a shadow attention detector is firstly exploited to generate an attention map which specifies shadow reg…

Cited by 174PDFScholar