← Search

Minghao Zhu

5 accepted papers

2026

GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding

CVPR 2026

Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid-LLMs) to broader applications. However, existing Vid-LLMs rely on uniform frame sampling to extract video information, resulting in a sparse distribution of ke

Cited by 0SourcecodeScholar
2025

CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation

ICCV 2025poster

In the effort to achieve robust and generalizable category-level object pose estimation, recent methods primarily focus on learning fundamental representations from data. However, the inherent biases within the data are often overlooked: the repeated training samples and similar environments may mis…

2024

MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer

NeurIPS 2024poster

Transferring visual-language knowledge from large-scale foundation models for video recognition has proved to be effective. To bridge the domain gap, additional parametric modules are added to capture the temporal information. However, zero-shot generalization diminishes with the increase in the num…

2024

SNF-Feat: Semantic-Guided Negative-Sample-Free Representation Learning for Local Feature Extraction

IROS 2024poster

Local feature extraction constitutes a foundational module crucial for numerous downstream tasks of computer vision. Its primary challenge lies in the generation of discriminative feature representations. Prior methodologies have employed contrastive learning within their pipelines, yet have encount…

Cited by 0SourceScholar
2022

Non-Autoregressive Neural Machine Translation with Consistency Regularization Optimized Variational Framework

NAACL 2022long

Variational Autoencoder (VAE) is an effective framework to model the interdependency for non-autoregressive neural machine translation (NAT). One of the prominent VAE-based NAT frameworks, LaNMT, achieves great improvements to vanilla models, but still suffers from two main issues which lower down t…