← Search

Shentong Mo

24 accepted papers

2025

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

ICASSP 2025accepted

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignment of text, video, and audio embeddings via a single contrastive loss. However,…

Cited by 0SourceScholar
2025

Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows

CVPR 2025poster

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames. Previous studies leverage a two-stage design where the AV encoders are firstly…

Cited by 0SourcePDFScholar
2024

Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning

NeurIPS 2024spotlight

In recent advancements in unsupervised visual representation learning, the Joint-Embedding Predictive Architecture (JEPA) has emerged as a significant method for extracting visual features from unlabeled imagery through an innovative masking strategy. Despite its success, two primary limitations hav…

Cited by 2SourcePDFScholar
2024

Continual Audio-Visual Sound Separation

NeurIPS 2024poster

In this paper, we introduce a novel continual audio-visual sound separation task, aiming to continuously separate sound sources for new classes while preserving performance on previously learned classes, with the aid of visual guidance. This problem is crucial for practical visually guided auditory…

2024

Fast Training of Diffusion Transformer with Extreme Masking for 3D Point Clouds Generation

ECCV 2024poster

"Diffusion Transformers have recently shown remarkable effectiveness in generating high-quality 3D point clouds. However, training voxel-based diffusion models for high-resolution 3D voxels remains prohibitively expensive due to the cubic complexity of attention operators, which arises from the addi…

Cited by 5SourcePDFScholar
2024

Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked Modeling

CVPR 2024poster

Humans possess a remarkable ability to integrate auditory and visual information enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues demonstrated through cognitive psychology and neuroscience research offers promising potential for developing mu…

2023

A Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition

ICML 2023poster

The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately, with several methods developed independently for each task. However, given the interconnected nature of source localizat…

2023

DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation

NeurIPS 2023poster

Recent Diffusion Transformers (i.e., DiT) have demonstrated their powerful effectiveness in generating high-quality 2D images. However, it is unclear how the Transformer architecture performs equally well in 3D shape generation, as previous 3D diffusion methods mostly adopted the U-Net architecture.…

Cited by 73SourcePDFScholar
2023

DiffComplete: Diffusion-based Generative 3D Shape Completion

NeurIPS 2023poster

We introduce a new diffusion-based approach for shape completion on 3D range scans. Compared with prior deterministic and probabilistic methods, we strike a balance between realism, multi-modality, and high fidelity. We propose DiffComplete by casting shape completion as a generative task conditione…

Cited by 31SourcePDFScholar
2022

"Unitail: Detecting, Reading, and Matching in Retail Scene"

ECCV 2022poster

"To make full use of computer vision technology in stores, it is required to consider the actual needs that fit the characteristics of the retail scene. Pursuing this goal, we introduce the United Retail Datasets (Unitail), a large-scale benchmark of basic visual tasks on products that challenges al…

2022

Multi-modal Grouping Network for Weakly-Supervised Audio-Visual Video Parsing

NeurIPS 2022accept

The audio-visual video parsing task aims to parse a video into modality- and category-aware temporal segments. Previous work mainly focuses on weakly-supervised approaches, which learn from video-level event labels. During training, they do not know which modality perceives and meanwhile which tempo…