← Search

Gerard de Melo

48 accepted papers

2026

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes

AAAI 2026technical

Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such interactions from images involves interpreting subtle visual cues such as relations, proximity and co-movement – semantica

Cited by 0SourcePDFScholar
2026

Token Distillation: Attention-Aware Input Embeddings for New Tokens

ICLR 2026poster

Current language models rely on static vocabularies determined at pretraining time, which can lead to decreased performance and increased computational cost for domains underrepresented in the original vocabulary. New tokens can be added to solve this problem, when coupled with a good initialization…

Cited by 0SourcecodeScholar
2025

ACE-M3: Automatic Capability Evaluator for Multimodal Medical Models

COLING 2025main

As multimodal large language models (MLLMs) gain prominence in the medical field, the need for precise evaluation methods to assess their effectiveness has become critical. While benchmarks provide a reliable means to evaluate the capabilities of MLLMs, traditional metrics like ROUGE and BLEU employ…

Cited by 0SourcePDFScholar
2025

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation

ACL 2025long

With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, which rely on token overlaps to measure quality, significantly overlook the import…

Cited by 0SourcePDFScholar
2025

GraphLSS: Integrating Lexical, Structural, and Semantic Features for Long Document Extractive Summarization

NAACL 2025short

Heterogeneous graph neural networks have recently gained attention for long document summarization, modeling the extraction as a node classification task. Although effective, these models often require external tools or additional machine learning models to define graph components, producing highly…

2025

Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation

AAAI 2025technical

In the rapidly evolving landscape of large language models (LLMs) for medical applications, ensuring the reliability and accuracy of these models in clinical settings is paramount. Existing benchmarks often focus on fixed-format tasks like multiple-choice QA, which fail to capture the complexity of…

2025

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models

ACL 2025finding

With the increasing size of Large Vision-Language Models (LVLMs), network pruning techniques aimed at compressing models for deployment in resource-constrained environments have garnered significant attention. However, we observe that pruning often leads to a degradation in safety performance. To ad…

2025

Image Token Matters: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) with discrete image tokenizers unify multimodal representations by encoding visual inputs into a finite set of tokens. Despite their effectiveness, we find that these models still hallucinate non-existent objects. We hypothesize that one reason is due to visual p…

Cited by 0SourceScholar
2025

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning

AAAI 2025technical

The emergence of fine-tuning-as-a-service has revealed a new vulnerability in large language models (LLMs). A mere handful of malicious data uploaded by users can subtly manipulate the fine-tuning process, leading to a compromised alignment state. Existing methods to counteract fine-tuning attacks t…

2025

SLlama: Parameter-Efficient Language Model Architecture for Enhanced Linguistic Competence Under Strict Data Constraints

EMNLP 2025

Scaling data and model size has driven recent advances in language modeling, but this strategy falters under scenarios with strict data constraints, as in the BabyLM Challenge. However, insights from Chinchilla highlights that smaller models trained on more data outperform larger counterparts traine

2025

Vector Grimoire: Codebook-based Shape Generation under Raster Image Supervision

ICML 2025poster

Scalable Vector Graphics (SVG) is a popular format on the web and in the design industry. However, despite the great strides made in generative modeling, SVG has remained underexplored due to the discrete and complex nature of such data. We introduce GRIMOIRE, a text-guided SVG generative model that…

Cited by 0SourcePDFScholar
2024

CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios

EMNLP 2024main

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench, a comprehensive benchmark with 14 expert-guided core clinica…

2024

Generating Persona-Aware Empathetic Responses with Retrieval-Augmented Prompt Learning

ICASSP 2024accepted

Empathetic response generation requires perceiving and understanding the user’s emotion to deliver suitable responses. However, existing models generally lack an ability to respond in a persona-specific way, which has been shown to play a vital role in expressing appropriate empathy. To address this…

Cited by 0SourceScholar
2024

I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token

NeurIPS 2024poster

Large Language Models are known to capture real-world knowledge, allowing them to excel in many downstream tasks. Despite recent advances, these models are still prone to what are commonly known as hallucinations, causing them to emit unwanted and factually incorrect text. In this work, we propose a…

2024

LLMs Cannot (Yet) Match the Specificity and Simplicity of Online Communities in Long Form Question Answering

EMNLP 2024finding

Retail investing is on the rise, and a growing number of users is relying on online finance communities to educate themselves.However, recent years have positioned Large Language Models (LLMs) as powerful question answering (QA) tools, shifting users away from interacting in communities towards disc…

Cited by 0SourcePDFScholar
2024

MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

AAAI 2024technical

The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this issue, we introduce MedBench, a comprehensive benchmark for the…

2024

NextLevelBERT: Masked Language Modeling with Higher-Level Representations for Long Documents

ACL 2024long

While (large) language models have significantly improved over the last years, they still struggle to sensibly process long sequences found, e.g., in books, due to the quadratic scaling of the underlying attention mechanism. To address this, we propose NextLevelBERT, a Masked Language Model operatin…

2023

A Disentangled-Attention Based Framework with Persona-Aware Prompt Learning for Dialogue Generation

AAAI 2023technical

Endowing dialogue agents with personas is the key to delivering more human-like conversations. However, existing persona-grounded dialogue systems still lack informative details of human conversations and tend to reply with inconsistent and generic responses. One of the main underlying causes is tha…

Cited by 5SourcePDFScholar
2023

Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification using Graph Neural Networks?

EMNLP 2023long findings

Given the success of Graph Neural Networks (GNNs) for structure-aware machine learning, many studies have explored their use for text classification, but mostly in specific domains with limited data characteristics. Moreover, some strategies prior to GNNs relied on graph mining and classical machine…

Cited by 0SourcecodeScholar
2023

Disentangled CVAEs with Contrastive Learning for Explainable Recommendation

AAAI 2023technical

Modern recommender systems are increasingly expected to provide informative explanations that enable users to understand the reason for particular recommendations. However, previous methods struggle to interpret the input IDs of user--item pairs in real-world datasets, failing to extract adequate ch…

Cited by 7SourcePDFScholar
2023

FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models

EMNLP 2023long main

Using model weights pretrained on a high-resource language as a warm start can reduce the need for data and compute to obtain high-quality language models for other, especially low-resource, languages. However, if we want to use a new tokenizer specialized for the target language, we cannot transfer…

Cited by 0SourcecodeScholar
2023

Harnessing Neighborhood Modeling and Asymmetry Preservation for Digraph Representation Learning

IJCAI 2023poster

Digraph Representation Learning aims to learn representations for directed homogeneous graphs (digraphs). Prior work is largely constrained or has poor generalizability across tasks. Most Graph Neural Networks exhibit poor performance on digraphs due to the neglect of modeling neighborhoods and pres…

Cited by 0SourcePDFScholar
2022

Art Creation with Multi-Conditional StyleGANs

IJCAI 2022poster

Creating art is often viewed as a uniquely human endeavor. In this paper, we introduce a multi-conditional Generative Adversarial Network (GAN) approach trained on large amounts of human paintings to synthesize realistic-looking paintings that emulate human art. Our approach is based on the StyleGAN…

2022

Curriculum Prompt Learning with Self-Training for Abstractive Dialogue Summarization

EMNLP 2022main

Succinctly summarizing dialogue is a task of growing interest, but inherent challenges, such as insufficient training data and low information density impede our ability to train abstractive models. In this work, we propose a novel curriculum-based prompt learning method with self-training to addres…

2022

Fast-R2D2: A Pretrained Recursive Neural Network based on Pruned CKY for Grammar Induction and Text Representation

EMNLP 2022main

Chart-based models have shown great potential in unsupervised grammar induction, running recursively and hierarchically, but requiring O(n³) time-complexity. The Recursive Transformer based on Differentiable Trees (R2D2) makes it possible to scale to large language model pretraining even with a comp…

2022

Frozen CLIP Models Are Efficient Video Learners

ECCV 2022poster

"Video recognition has been dominated by the end-to-end learning paradigm - first initializing a video recognition model with weights of a pretrained image model and then conducting end-to-end training on videos. This enables the video network to benefit from the pretrained image model. However, thi…

2022

Improving Personalized Explanation Generation through Visualization

ACL 2022long

In modern recommender systems, there are usually comments or reviews from users that justify their ratings for different items. Trained on such textual corpus, explainable recommendation models learn to discover user interests and generate personalized explanations. Though able to provide plausible…

Cited by 37SourcePDFScholar
2022

Multi-Scale Distribution Deep Variational Autoencoder for Explanation Generation

ACL 2022findings

Generating explanations for recommender systems is essential for improving their transparency, as users often wish to understand the reason for receiving a specified recommendation. Previous methods mainly focus on improving the generation quality, but often produce generic explanations that fail to…

Cited by 5SourcePDFScholar
2021

Data Augmentation with Adversarial Training for Cross-Lingual NLI

ACL 2021long

Due to recent pretrained multilingual representation models, it has become feasible to exploit labeled data from one language to train a cross-lingual model that can then be applied to multiple new languages. In practice, however, we still face the problem of scarce labeled data, leading to subpar r…

2021

Faithfully Explainable Recommendation via Neural Logic Reasoning

NAACL 2021long

Knowledge graphs (KG) have become increasingly important to endow modern recommender systems with the ability to generate traceable reasoning paths to explain the recommendation process. However, prior research rarely considers the faithfulness of the derived explanations to justify the decision-mak…

2021

R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language Modeling

ACL 2021long

Human language understanding operates at multiple levels of granularity (e.g., words, phrases, and sentences) with increasing levels of abstraction that can be hierarchically combined. However, existing deep models with stacked layers do not explicitly model any sort of hierarchical process. In this…

2021

TIME: Text and Image Mutual-Translation Adversarial Networks

AAAI 2021technical

Focusing on text-to-image (T2I) generation, we propose Text and Image Mutual-Translation Adversarial Networks (TIME), a lightweight but effective model that jointly learns a T2I generator G and an image captioning discriminator D under the Generative Adversarial Network framework. While previous met…

Cited by 39SourcePDFScholar
2020

Cross-Lingual Emotion Lexicon Induction using Representation Alignment in Low-Resource Settings

COLING 2020main

Emotion lexicons provide information about associations between words and emotions. They have proven useful in analyses of reviews, literary texts, and posts on social media, among other things. We evaluate the feasibility of deriving emotion lexicons cross-lingually, especially for low-resource lan…

Cited by 8SourcePDFScholar
2020

Data Augmentation for Multiclass Utterance Classification – A Systematic Study

COLING 2020main

Utterance classification is a key component in many conversational systems. However, classifying real-world user utterances is challenging, as people may express their ideas and thoughts in manifold ways, and the amount of training data for some categories may be fairly limited, resulting in imbalan…

Cited by 23SourcePDFScholar
2020

HID: Hierarchical Multiscale Representation Learning for Information Diffusion

IJCAI 2020poster

Multiscale modeling has yielded immense success on various machine learning tasks. However, it has not been properly explored for the prominent task of information diffusion, which aims to understand how information propagates along users in online social networks. For a specific user, whether and w…

2020

Incorporating Pragmatic Reasoning Communication into Emergent Language

NeurIPS 2020spotlight

Emergentism and pragmatics are two research fields that study the dynamics of linguistic communication along quite different timescales and intelligence levels. From the perspective of multi-agent reinforcement learning, they correspond to stochastic games with reinforcement training and stage games…

Cited by 26SourcePDFScholar
2020

Interactive Question Clarification in Dialogue via Reinforcement Learning

COLING 2020industry

Coping with ambiguous questions has been a perennial problem in real-world dialogue systems. Although clarification by asking questions is a common form of human interaction, it is hard to define appropriate questions to elicit more specific intents from a user. In this work, we propose a reinforcem…

Cited by 8SourcePDFScholar
2020

Query Distillation: BERT-based Distillation for Ensemble Ranking

COLING 2020industry

Recent years have witnessed substantial progress in the development of neural ranking networks, but also an increasingly heavy computational burden due to growing numbers of parameters and the adoption of model ensembles. Knowledge Distillation (KD) is a common solution to balance the effectiveness…

Cited by 5SourcePDFScholar
2020

SCALOR: Generative World Models with Scalable Object Representations

ICLR 2020poster

Scalability in terms of object density in a scene is a primary challenge in unsupervised sequential object-oriented representation learning. Most of the previous models have been shown to work only on scenes with a few objects. In this paper, we propose SCALOR, a probabilistic generative world model…

Cited by 140SourceScholar
2018

Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification

CVPR 2018poster

Recently, substantial research effort has focused on how to apply CNNs or RNNs to better capture temporal patterns in videos, so as to improve the accuracy of video classification. In this paper, however, we show that temporal information, especially longer-term patterns, may not be necessary to ach…