← Search

Meng Chen

32 accepted papers

2026

MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction

ICML 2026poster

Predicting street-level socioeconomic indicators from street view imagery is fundamental to urban planning. Existing methods typically extract visual features via pretrained encoders and propagate information through graph-based learning, but they fail to fully exploit the structured, task-relevant,…

Cited by 0SourceScholar
2026

POLY-SVC: POLYPHONY-AWARE SINGING VOICE CONVERSION WITH HARMONIC MODELING

ICASSP 2026poster

Singing Voice Conversion (SVC) aims to transform a source singing voice into a target singer while preserving lyrics and melody. Most existing SVC methods depend on F0 extractors to capture the lead melody from clean vocals. However, no existing method can reliably extract clean vocals from accompan…

Cited by 0SourcePDFScholar
2026

Tackling Model Bias via Game-theoretic Multi-agent Collaboration Framework for Hateful Meme Classification

CVPR 2026

Hateful meme classification aims to identify memes containing hateful content and has become increasingly important in the era of social media dominance. Large multimodal models (LMMs) have significantly enhanced the understanding of multimodal content, advancing this field. However, cognitive biase

Cited by 0SourcecodeScholar
2025

Accelerating Dense LLMs via L0-regularized Mixture-of-Experts

ACL 2025short

Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a ligh…

2025

AutoMV: An Autonomous Agent Framework for Real Estate Marketing Video Generation

AAAI 2025technical

In this paper, we introduce AutoMV, an autonomous agent framework designed for generating real estate marketing videos. The framework integrates a diverse set of existing models into a tool library, allowing the agent to intelligently select and execute the appropriate tools. Given property images a…

Cited by 0SourcePDFScholar
2025

Cross-City Latent Space Alignment for Consistency Region Embedding

ICML 2025poster

Learning urban region embeddings has substantially advanced urban analysis, but their typical focus on individual cities leads to disparate embedding spaces, hindering cross-city knowledge transfer and the reuse of downstream task predictors. To tackle this issue, we present Consistent Region Embedd…

Cited by 0SourcePDFScholar
2025

Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability

NAACL 2025system demonstrations

As large language models (LLMs) continue to evolve, leaderboards play a significant role in steering their development. Existing leaderboards often prioritize model capabilities while overlooking safety concerns, leaving a significant gap in responsible AI development. To address this gap, we introd…

2025

Mastering the Craft of Data Synthesis for CodeLLMs

NAACL 2025long

Large language models (LLMs) have shown impressive performance in code understanding and generation, making coding tasks a key focus for researchers due to their practical applications and value as a testbed for LLM evaluation. Data synthesis and filtering techniques have been widely adopted and sho…

2025

Results of the Big ANN: NeurIPS’23 competition

NeurIPS 2025poster

The 2023 Big ANN Challenge, held at NeurIPS 2023, focused on advancing the state-of-the-art in indexing data structures and search algorithms for practical variants of Approximate Nearest Neighbor (ANN) search that reflect its the growing complexity and diversity of workloads. Unlike prior challenge…

Cited by 0SourcecodeScholar
2025

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

NeurIPS 2025poster

Transformer-based Large Language Models (LLMs) have become increasingly important. However, scaling LLMs to longer contexts incurs slow inference speed and high GPU memory consumption for caching key-value (KV) vectors. This paper presents RetrievalAttention, a training-free approach to both acceler…

Cited by 0SourcecodeScholar
2025

Spatio-temporal Prototype-based Hierarchical Learning for OD Demand Prediction

IJCAI 2025

Origin-Destination (OD) demand prediction is a pivotal yet highly challenging task in intelligent transportation systems, aiming to accurately forecast cross-region ridership flows within urban networks. While previous studies have focused on modeling node-to-node relationships, most of them neglect

Cited by 0SourcePDFScholar
2025

Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation

IROS 2025

The significant advancements in embodied vision navigation have raised concerns about its susceptibility to adversarial attacks exploiting deep neural networks. Investigating the adversarial robustness of embodied vision navigation is crucial, especially given the threat of 3D physical attacks that

Cited by 7SourcecodeScholar
2024

Exploring Urban Semantics: A Multimodal Model for POI Semantic Annotation with Street View Images and Place Names

IJCAI 2024poster

Semantic annotation for points of interest (POIs) is the process of annotating a POI with a category label, which facilitates many services related to POIs, such as POI search and recommendation. Most of the existing solutions extract features related to POIs from abundant user-generated content dat…

2024

G^2SAM: Graph-Based Global Semantic Awareness Method for Multimodal Sarcasm Detection

AAAI 2024technical

Multimodal sarcasm detection, aiming to detect the ironic sentiment within multimodal social data, has gained substantial popularity in both the natural language processing and computer vision communities. Recently, graph-based studies by drawing sentimental relations to detect multimodal sarcasm ha…

2024

Learning Hierarchy-Enhanced POI Category Representations Using Disentangled Mobility Sequences

IJCAI 2024poster

Points of interest (POIs) carry a wealth of semantic information of varying locations in cities and thus have been widely used to enable various location-based services. To understand POI semantics, existing methods usually model contextual correlations of POI categories in users' check-in sequences…

2024

Urban Region Embedding via Multi-View Contrastive Prediction

AAAI 2024technical

Recently, learning urban region representations utilizing multi-modal data (information views) has become increasingly popular, for deep understanding of the distributions of various socioeconomic features in cities. However, previous methods usually blend multi-view information in a posteriors stag…

2023

Dialog-Post: Multi-Level Self-Supervised Objectives and Hierarchical Model for Dialogue Post-Training

ACL 2023long

Dialogue representation and understanding aim to convert conversational inputs into embeddings and fulfill discriminative tasks. Compared with free-form text, dialogue has two important characteristics, hierarchical semantic structure and multi-facet attributes. Therefore, directly applying the pret…

2023

DiffusEmp: A Diffusion Model-Based Framework with Multi-Grained Control for Empathetic Response Generation

ACL 2023long

Empathy is a crucial factor in open-domain conversations, which naturally shows one’s caring and understanding to others. Though several methods have been proposed to generate empathetic responses, existing works often lead to monotonous empathy that refers to generic and safe expressions. In this p…

Cited by 20SourcePDFScholar
2023

Enhancing Multimodal Alignment with Momentum Augmentation for Dense Video Captioning

ICASSP 2023accepted

Dense video captioning aims to localize multiple events from an untrimmed video and generate corresponding captions for each event. Fusing different modalities(e.g. rgb, flow, audio) via transformer structure is a promising way to improve the caption performance. However, it is challenging for the c…

Cited by 0SourceScholar
2023

Improving Disfluency Detection with Multi-Scale Self Attention and Contrastive Learning

ICASSP 2023accepted

Disfluency detection aims to recognize disfluencies in sentences. Existing works usually adopt a sequence labeling model to tackle this task. They also attempt to integrate into models the feature that the disfluencies are similar to the correct phrase, the so-called "rough copy". However, they heav…

Cited by 0SourceScholar
2023

MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query Grounding

AAAI 2023technical

Multimodal named entity recognition (MNER) is a critical step in information extraction, which aims to detect entity spans and classify them to corresponding entity types given a sentence-image pair. Existing methods either (1) obtain named entities with coarse-grained visual clues from attention me…

Cited by 53SourcePDFScholar
2023

Tackling Modality Heterogeneity with Multi-View Calibration Network for Multimodal Sentiment Detection

ACL 2023long

With the popularity of social media, detecting sentiment from multimodal posts (e.g. image-text pairs) has attracted substantial attention recently. Existing works mainly focus on fusing different features but ignore the challenge of modality heterogeneity. Specifically, different modalities with in…

2023

Towards an Integrated View of Semantic Annotation for POIs with Spatial and Textual Information

IJCAI 2023poster

Categories of Point of Interest (POI) facilitate location-based services from many aspects like location search and POI recommendation. However, POI categories are often incomplete and new POIs are being consistently generated, this rises the demand for semantic annotation for POIs, i.e., labeling t…

Cited by 9SourcePDFScholar
2023

UFO2: A Unified Pre-Training Framework for Online and Offline Speech Recognition

ICASSP 2023accepted

In this paper, we propose a Unified pre-training Framework for Online and Offline (UFO2) Automatic Speech Recognition (ASR), which 1) simplifies the two separate training workflows for online and offline modes into one process, and 2) improves the Word Error Rate (WER) performance with limited utter…

Cited by 0SourceScholar
2022

Building Robust Spoken Language Understanding by Cross Attention Between Phoneme Sequence and ASR Hypothesis

ICASSP 2022accepted

Building Spoken Language Understanding (SLU) robust to Automatic Speech Recognition (ASR) errors is an essential issue for various voice-enabled virtual assistants. Considering that most ASR errors are caused by phonetic confusion between similar-sounding expressions, intuitively, leveraging the pho…

Cited by 0SourceScholar
2022

Few-Shot Table Understanding: A Benchmark Dataset and Pre-Training Baseline

COLING 2022main

Few-shot table understanding is a critical and challenging problem in real-world scenario as annotations over large amount of tables are usually costly. Pre-trained language models (PLMs), which have recently flourished on tabular data, have demonstrated their effectiveness for table understanding t…

2022

Gated Multimodal Fusion with Contrastive Learning for Turn-Taking Prediction in Human-Robot Dialogue

ICASSP 2022accepted

Turn-taking, aiming to decide when the next speaker can start talking, is an essential component in building human-robot spoken dialogue systems. Previous studies indicate that multi-modal cues can facilitate this challenging task. However, due to the paucity of public multimodal datasets, current m…

Cited by 0SourceScholar
2022

Learning to Generate Poetic Chinese Landscape Painting with Calligraphy

IJCAI 2022poster

In this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generation, Polaca takes the classic poetry as input and outputs the artistic landscape painting image with the corresponding ca…

Cited by 10SourcePDFScholar
2021

Conversational Query Rewriting with Self-Supervised Learning

ICASSP 2021accepted

Context modeling plays a critical role in building multi-turn dialogue systems. Conversational Query Rewriting (CQR) aims to simplify the multi-turn dialogue modeling into a single-turn problem by explicitly rewriting the conversational query into a self-contained utterance. However, existing approa…

Cited by 0SourceScholar
2019

An autonomous exploration algorithm using environment-robot interacted traversability analysis

IROS 2019poster

Auto-exploration is a task for self-driving robots to explore unknown environments, which becomes much complicated when they move on irregular outdoor terrains. To improve the situation, a new frontier-based exploration algorithm is presented in this paper. It starts from original 3D cloud points of…

Cited by 27SourceScholar
2015

The calibration device and method of humanoid finger sensor based on multimodal perception

IROS 2015poster

Depending on a humanoid finger sensor with multimodal perception capability, physical properties such as pressure, temperature, texture features, surface roughness, micro vibration, etc., could be easily extracted when the sensor contacts with different objects. In this paper, an innovative calibrat…

Cited by 0SourceScholar