← Search

Marcio Monteiro

3 accepted papers

2026

Landmark-Guided Policy Optimization for Multi-Objective Language Model Selection

ICML 2026poster

Selecting a pretrained large language model (LLM) to fine-tune for a task-specific dataset can be time-consuming and costly. With several candidate models available to choose from, varying in size, architecture, and pretraining data, finding the best model for a specific task often involves extensiv…

Cited by 0SourceScholar
2026

TORA: Train Once, Realign Anytime for Offline Multi-Objective Reinforcement Learning

AAAI 2026technical

Intelligent agents in real-world applications must adapt their behavior to changing contexts and user preferences. For example, planning a road trip requires considering both travel time and cost. Multi-objective reinforcement learning (MORL) provides a principled approach to navigate such trade-of

Cited by 0SourcePDFScholar
2024

Characterizing Text Datasets with Psycholinguistic Features

EMNLP 2024finding

Fine-tuning pretrained language models on task-specific data is a common practice in Natural Language Processing (NLP) applications. However, the number of pretrained models available to choose from can be very large, and it remains unclear how to select the optimal model without spending considerab…