← Search

Hongsong Wang

16 accepted papers

2026

DreamCS: Geometry-Aware Text-to-3D Generation with Unpaired 3D Reward Supervision

ICLR 2026poster

While text-to-3D generation has attracted growing interest, existing methods often struggle to produce 3D assets that align well with human preferences. Current preference alignment techniques for 3D content typically rely on hardly-collected preference-paired multi-view 2D images to train 2D reward…

Cited by 0SourceScholar
2026

EasyTune: Efficient Step-Aware Fine-Tuning for Diffusion-Based Motion Generation

ICLR 2026poster

In recent years, motion generative models have undergone significant advancement, yet pose challenges in aligning with downstream objectives. Recent studies have shown that using differentiable rewards to directly align the preference of diffusion models yields promising results. However, these meth…

Cited by 0SourcecodeScholar
2026

HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal Grounding

ICLR 2026poster

Temporal grounding in long, untrimmed videos is critical for real-world video understanding, yet it remains a challenging task owing to complex temporal structures and pervasive visual redundancy. Existing methods rely heavily on supervised training with task-specific annotations, which inherently l…

Cited by 0SourceScholar
2026

Matting Anything 2: Towards Video Matting for Anything

ICLR 2026poster

Video matting is a crucial task for many applications, but existing methods face significant limitations. They are often domain-specific, focusing primarily on human portraits, and rely on the mask of first frame that is challenging to acquire for transparent or intricate objects like fire or smoke.…

Cited by 0SourceScholar
2026

ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided Alignment

AAAI 2026technical

Text-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more diversity and realistic motion. However, there exists a misalignment between text

Cited by 0SourcePDFScholar
2025

Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly Detection

AAAI 2025technical

Video Anomaly Detection (VAD) is essential for computer vision and multimedia research. Existing VAD methods utilize either reconstruction-based or prediction-based frameworks. The former excels at detecting irregular patterns or structures, whereas the latter is capable of spotting abnormal deviati…

2025

LOTA: Bit-Planes Guided AI-Generated Image Detection

ICCV 2025poster

The rapid advancement of GAN and Diffusion models makes it more difficult to distinguish AI-generated images from real ones. Recent studies often use image-based reconstruction errors as an important feature for determining whether an image is AI-generated. However, these approaches typically incur…

2025

SAM-Aware Graph Prompt Reasoning Network for Cross-Domain Few-Shot Segmentation

AAAI 2025technical

The primary challenge of cross-domain few-shot segmentation (CD-FSS) is the domain disparity between the training and inference phases, which can exist in either the input data or the target classes. Previous models struggle to learn feature representations that generalize to various unknown domains…

2025

SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization

NeurIPS 2025poster

Text-to-motion generation is essential for advancing the creative industry but often presents challenges in producing consistent, realistic motions. To address this, we focus on fine-tuning text-to-motion models to consistently favor high-quality, human-preferred motions—a critical yet largely unexp…

Cited by 0SourcecodeScholar
2025

USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature Decorrelation

AAAI 2025technical

Contrastive learning has achieved great success in skeleton-based representation learning recently. However, the prevailing methods are predominantly negative-based, necessitating additional momentum encoder and memory bank to get negative samples, which increases the difficulty of model training. F…

2017

Modeling Temporal Dynamics and Spatial Configurations of Actions Using Two-Stream Recurrent Neural Networks

CVPR 2017poster

Recently, skeleton based action recognition gains more popularity due to cost-effective depth sensors coupled with real-time skeleton estimation algorithms. Traditional approaches based on handcrafted features are limited to represent the complexity of motion patterns. Recent methods that use Recurr…

Cited by 511PDFcodeScholar