← Search

MengMeng Wang

54 accepted papers

2026

Deforming Videos to Masks: Flow Matching for Referring Video Segmentation

ICLR 2026poster

Referring Video Object Segmentation (RVOS) requires segmenting specific objects in a video guided by a natural language description. The core challenge of RVOS is to anchor abstract linguistic concepts onto a specific set of pixels and continuously segment them through the complex dynamics of a vide…

Cited by 0SourceScholar
2026

DynBridge: Bridging Imagination and Control through Interaction Dynamics for Robot Manipulation

CVPR 2026

Recent generative models allow robots to generate future visual outcomes for action guidance, yet most still address imagination and control independently, resulting in visually coherent rollouts but physically inconsistent behaviors. While structural priors enhance spatial grounding, these methods

Cited by 0SourceScholar
2026

PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation

ICRA 2026poster

Bimanual manipulation is essential for advanced robotic systems because it offers higher efficiency and flexibility compared to single-arm configurations. However, existing approaches either lack inter-arm interaction or ignore the need for a dynamic division of labor, treating the arms as functiona…

2026

PoliCon: Evaluating LLMs on Achieving Diverse Political Consensus Objectives

ICLR 2026poster

Achieving political consensus is crucial yet challenging for the effective functioning of social governance. However, although frontier AI systems represented by large language models (LLMs) have developed rapidly in recent years, their capabilities in this scope are still understudied. In this pape…

Cited by 0SourcecodeScholar
2026

Representation Alignment for Diffusion Transformers without External Components

ICLR 2026poster

Recent studies have demonstrated that learning a meaningful internal represen- tation can accelerate generative training. However, existing approaches necessi- tate to either introduce an off-the-shelf external representation task or rely on a large-scale, pre-trained external representation encoder…

Cited by 0SourcecodeScholar
2026

SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Training

CVPR 2026

Denoising-based diffusion transformers, despite their strong generation performance, suffer from inefficient training convergence. Existing methods addressing this issue, such as REPA (relying on external representation encoders) or SRA (requiring dual-model setups), inevitably incur heavy computati

Cited by 0SourceScholar
2025

Action Detail Matters: Refining Video Recognition with Local Action Queries

CVPR 2025poster

Video action recognition involves interpreting both global context and specific details to accurately identify actions. While previous models are effective at capturing spatiotemporal features, they often lack a focused representation of key action details. To address this, we introduce \nameo, a fr…

Cited by 0SourcePDFScholar
2025

Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective

ACL 2025finding

As large language models (LLMs) become increasingly integrated into critical applications, aligning their behavior with human values presents significant challenges. Current methods, such as Reinforcement Learning from Human Feedback (RLHF), typically focus on a limited set of coarse-grained values…

2025

Density-aware and Depth-aware Visual Representation for Zero-Shot Object Counting

ICASSP 2025accepted

Previous methods often utilize CLIP semantic classifiers with class names for zero-shot object counting. However, they ignore crucial density and depth knowledge for counting tasks. Thus, we propose a density-aware and depth-aware prompt counting model, which captures density information via learnin…

Cited by 0SourceScholar
2025

DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-Making

ICML 2025poster

Integrating natural language instructions and visual perception with decision-making is a critical challenge for embodied agents. Existing methods often struggle to balance the conciseness of language commands with the richness of video content. To bridge the gap between modalities, we propose extra…

Cited by 0SourcePDFScholar
2025

EchoGPT: An Interactive Cardiac Function Assessment Model for Echocardiogram Videos

IJCAI 2025

With the development of wearable cardiac ultrasound devices, it is no longer sufficient to solely rely on doctors for diagnosing long-term echocardiogram videos. Automated diagnosis of echocardiogram videos has now become a research hotspot. Existing studies only analyze echocardiogram video through

2025

Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia

NeurIPS 2025poster

Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing eval…

Cited by 0SourceScholar
2025

Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization

IJCAI 2025

Text-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by the need for manual prompt crafting, which can be time-consuming, introduce irrele

2025

LLM-TPF: Multiscale Temporal Periodicity-Semantic Fusion LLMs for Time Series Forecasting

IJCAI 2025

Large language models have demonstrated remarkable generalization capabilities and strong performance across various fields. Recent research has highlighted their significant potential in time series forecasting. However, time series data often exhibit complex periodic characteristics, posing a subs

2025

Low-Biased General Annotated Dataset Generation

CVPR 2025poster

Pre-training backbone networks on a general annotated dataset (e.g., ImageNet) that comprises numerous manually collected images with category annotations has proven to be indispensable for enhancing the generalization capacity of downstream visual tasks. However, those manually collected images oft…

2025

Manifold Constraint Reduces Exposure Bias in Accelerated Diffusion Sampling

ICLR 2025poster

Diffusion models have demonstrated significant potential for generating high-quality images, audio, and videos. However, their iterative inference process entails substantial computational costs, limiting practical applications. Recently, researchers have introduced accelerated sampling methods that…

Cited by 0SourcePDFScholar
2025

MonoLift: Learning 3D Manipulation Policies from Monocular RGB via Distillation

NeurIPS 2025spotlight

Although learning 3D manipulation policies from monocular RGB images is lightweight and deployment-friendly, the lack of structural information often leads to inaccurate action estimation. While explicit 3D inputs can mitigate this issue, they typically require additional sensors and introduce data…

Cited by 0SourcecodeScholar
2025

SpotActor: Training-Free Layout-Controlled Consistent Image Generation

AAAI 2025technical

Text-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of eac…

Cited by 2SourcePDFScholar
2025

TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking

ICCV 2025poster

3D LiDAR-based single object tracking (SOT) relies on sparse and irregular point clouds, posing challenges from geometric variations in scale, motion patterns, and structural complexity across object categories. Current category-specific approaches achieve good accuracy but are impractical for real-…

Cited by 0SourcePDFScholar
2025

VidEvo: Evolving Video Editing through Exhaustive Temporal Modeling

IJCAI 2025

Text-guided video editing (TGVE) has become a recent hotspot due to its entertainment value and practical applications. To reduce overhead, existing methods primarily extend from text-to-image diffusion models and typically involve reconstruction and editing phases. However, challenges persist, part

Cited by 0SourcePDFScholar
2024

A Multimodal, Multi-Task Adapting Framework for Video Action Recognition

AAAI 2024technical

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance a…

Cited by 17SourcePDFScholar
2024

A Robotic-centric Paradigm for 3D Human Tracking Under Complex Environments Using Multi-modal Adaptation

IROS 2024poster

The goal of this paper is to strike a feasible tracking paradigm that can make 3D human trackers applicable on robot platforms and enable more high-level tasks. Till now, two fundamental problems haven’t been adequately addressed. One is the computational cost lightweight enough for robotic deployme…

Cited by 0SourceScholar
2024

Decentralized Riemannian Conjugate Gradient Method on the Stiefel Manifold

ICLR 2024poster

The conjugate gradient method is a crucial first-order optimization method that generally converges faster than the steepest descent method, and its computational cost is much lower than that of second-order methods. However, while various types of conjugate gradient methods have been studied in Euc…

Cited by 11SourcePDFScholar
2024

Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category Discovery

NeurIPS 2024oral

Recent advancements have shown promise in applying traditional Semi-Supervised Learning strategies to the task of Generalized Category Discovery (GCD). Typically, this involves a teacher-student framework in which the teacher imparts knowledge to the student to classify categories, even in the absen…

Cited by 2SourcePDFScholar
2024

LangSuit·E: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments

ACL 2024findings

Recent advances in Large Language Models (LLMs) have shown inspiring achievements in constructing autonomous agents that rely onlanguage descriptions as inputs. However, it remains unclear how well LLMs can function as few-shot or zero-shot embodied agents in dynamic interactive environments. To add…

2024

LooGLE: Can Long-Context Language Models Understand Long Contexts?

ACL 2024long

Large language models (LLMs) are typically limited to processing texts within context window size, which has spurred significant research efforts into enhancing LLMs’ long-context understanding as well as developing high-quality benchmarks to evaluate the ability. However, prior datasets suffer from…

2024

Multi-modal 3D Human Tracking for Robots in Complex Environment with Siamese Point-Video Transformer

ICRA 2024poster

Tracking a specific person in 3D scene is gaining momentum due to its numerous applications in robotics. Currently, most 3D trackers focus on driving scenarios with neglected jitter and uncomplicated surroundings, which results in their severe degeneration in complex environments, especially on jolt…

Cited by 3SourceScholar
2024

OneActor: Consistent Subject Generation via Cluster-Conditioned Guidance

NeurIPS 2024poster

Text-to-image diffusion models benefit artists with high-quality image generation. Yet their stochastic nature hinders artists from creating consistent images of the same subject. Existing methods try to tackle this challenge and generate consistent content in various ways. However, they either depe…

2024

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

CVPR 2024poster

Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers which was inefficient and lacked generalized representation due to the scarcity of multimodal data. Therefore recent studies have ut…

2024

SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image Generation

AAAI 2024technical

Despite significant progress in Text-to-Image (T2I) generative models, even lengthy and complex text descriptions still struggle to convey detailed controls. In contrast, Layout-to-Image (L2I) generation, aiming to generate realistic and complex scene images from user-specified layouts, has risen to…

Cited by 16SourcePDFScholar
2024

Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image Editing

NeurIPS 2024poster

Text-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires inverting the source image into a latent space, a process often hindered by prediction errors inherent in DDIM inversion…

2023

Boosting Few-shot Action Recognition with Graph-guided Hybrid Matching

ICCV 2023poster

Class prototype construction and matching are core aspects of few-shot action recognition. Previous methods mainly focus on designing spatiotemporal relation modeling modules or complex temporal alignment algorithms. Despite the promising results, they ignored the value of class prototype constructi…

Cited by 37PDFcodeScholar
2023

PANet: LiDAR Panoptic Segmentation with Sparse Instance Proposal and Aggregation

IROS 2023poster

Reliable LiDAR panoptic segmentation (LPS), including both semantic and instance segmentation, is vital for many robotic applications, such as autonomous driving. This work proposes a new LPS framework named PANet to eliminate the dependency on the offset branch and improve the performance on large…

Cited by 5SourcecodeScholar
2023

RICO: Regularizing the Unobservable for Indoor Compositional Reconstruction

ICCV 2023poster

Recently, neural implicit surfaces have become popular for multi-view reconstruction. To facilitate practical applications like scene editing and manipulation, some works extend the framework with semantic masks input for the object-compositional reconstruction rather than the holistic perspective.…

Cited by 13PDFcodeScholar
2023

Revisiting the Spatial and Temporal Modeling for Few-Shot Action Recognition

AAAI 2023technical

Spatial and temporal modeling is one of the most core aspects of few-shot action recognition. Most previous works mainly focus on long-term temporal relation modeling based on high-level spatial representations, without considering the crucial low-level spatial features and short-term temporal relat…

Cited by 45SourcePDFScholar
2023

SSC-RS: Elevate LiDAR Semantic Scene Completion with Representation Separation and BEV Fusion

IROS 2023poster

Semantic scene completion (SSC) jointly predicts the semantics and geometry of the entire 3D scene, which plays an essential role in 3D scene understanding for autonomous driving systems. SSC has achieved rapid progress with the help of semantic context in segmentation. However, how to effectively e…

Cited by 23SourcecodeScholar
2023

Synchronize Feature Extracting and Matching: A Single Branch Framework for 3D Object Tracking

ICCV 2023poster

Siamese network has been a de facto benchmark framework for 3D LiDAR object tracking with a shared-parametric encoder extracting features from template and search region, respectively. This paradigm relies heavily on an additional matching network to model the cross-correlation/similarity of the tem…

Cited by 18PDFScholar
2022

E-NeRV: Expedite Neural Video Representation with Disentangled Spatial-Temporal Context

ECCV 2022poster

"Recently, the image-wise implicit neural representation of videos, NeRV, has gained popularity for its promising results and swift speed compared to regular pixel-wise implicit representations. However, the redundant parameters within the network structure can cause a large model size when scaling…

2021

FCFR-Net: Feature Fusion based Coarse-to-Fine Residual Learning for Depth Completion

AAAI 2021technical

Depth completion aims to recover a dense depth map from a sparse depth map with the corresponding color image as input. Recent approaches mainly formulate the depth completion as a one-stage end-to-end learning task, which outputs dense depth maps directly. However, the feature extraction and superv…

Cited by 138SourcePDFScholar
2021

HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation

AAAI 2021technical

Self-supervised learning shows great potential in monocular depth estimation, using image sequences as the only source of supervision. Although people try to use the high-resolution image for depth estimation, the accuracy of prediction has not been significantly improved. In this work…

2021

One-shot Face Reenactment Using Appearance Adaptive Normalization

AAAI 2021technical

The paper proposes a novel generative adversarial network for one-shot face reenactment, which can animate a single face image to a different pose-and-expression (provided by a driving image) while keeping its original appearance. The core of our network is a novel mechanism called appearance adapti…

Cited by 31SourcePDFScholar
2021

RFNet: Recurrent Forward Network for Dense Point Cloud Completion

ICCV 2021poster

Point cloud completion is an interesting and challenging task in 3D vision, aiming to recover complete shapes from sparse and incomplete point clouds. Existing learning-based methods often require vast computation cost to achieve excellent performance, which limits their practical applications. In t…

Cited by 48PDFScholar
2021

Self-Supervised Monocular Depth Estimation for All Day Images Using Domain Separation

ICCV 2021poster

Remarkable results have been achieved by DCNN based self-supervised depth estimation approaches. However, most of these approaches can only handle either day-time or night-time images, while their performance degrades for all-day images due to large domain shift and the variation of illumination bet…

Cited by 86PDFcodeScholar
2021

Structure-aware Person Image Generation with Pose Decomposition and Semantic Correlation

AAAI 2021technical

In this paper we tackle the problem of pose guided person image generation, which aims to transfer a person image from the source pose to a novel target pose while maintaining the source appearance. Given the inefficiency of standard CNNs in handling large spatial transformation, we propose a struct…

Cited by 23SourcePDFScholar
2020

DTVNet: Dynamic Time-lapse Video Generation via Single Still Image

ECCV 2020poster

This paper presents a novel end-to-end dynamic time-lapse video generation framework, named DTVNet, to generate diversified time-lapse videos from a single landscape image, which are conditioned on normalized motion vectors. The proposed DTVNet consists of two submodules: mph{Optical Flow Encoder} (…

2020

Semantic Graph Based Place Recognition for 3D Point Clouds

IROS 2020poster

Due to the difficulty in generating the effective descriptors which are robust to occlusion and viewpoint changes, place recognition for 3D point cloud remains an open issue. Unlike most of the existing methods that focus on extracting local, global, and statistical features of raw point clouds, our…

Cited by 147SourcecodeScholar
2019

Motion Artefact Removal in Functional Near-infrared Spectroscopy Signals Based on Robust Estimation

ICASSP 2019accepted

Functional Near-InfraRed Spectroscopy (fNIRS) has gained widespread acceptance as a non-invasive neuroimaging modality for monitoring functional brain activities. fNIRS uses light in the near infra-red spectrum (600-900 nm) to penetrate human brain tissues and estimates the oxygenation conditions ba…

Cited by 0SourceScholar
2019

STM: SpatioTemporal and Motion Encoding for Action Recognition

ICCV 2019poster

Spatiotemporal and motion features are two complementary and crucial information for video action recognition. Recent state-of-the-art methods adopt a 3D CNN stream to learn spatiotemporal features and another flow stream to learn motion features. In this work, we aim to efficiently encode these two…

Cited by 556PDFScholar
2017

Real-time 3D human tracking for mobile robots with multisensors

ICRA 2017poster

Acquiring the accurate 3-D position of a target person around a robot provides fundamental and valuable information that is applicable to a wide range of robotic tasks, including home service, navigation and entertainment. This paper presents a real-time robotic 3-D human tracking system which combi…

Cited by 34SourceScholar
2015

Automated tracking of cells from phase contrast images by multiple hypothesis Kalman filters

ICASSP 2015accepted

Cell migration is a fundamental process for the development and maintenance of all multicellular organisms. Accurate cell tracking may lead to better interpretations of long-term cell behaviours. This paper describes an automated system to track multiple cells from experimental phase contrast images…

Cited by 0SourceScholar