← Search

Davide Modolo

17 accepted papers

2026

MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated promising capabilities in multimodal coding tasks such as chart-to-code generation. However, existing methods primarily rely on supervised fine-tuning (SFT), which requires the model to learn code patterns through chart-code pairs b

Cited by 0SourceScholar
2025

Enhancing Numerical Prediction of MLLMs with Soft Labeling

ICCV 2025poster

The optimality of using the de facto cross-entropy loss with one-hot target distribution (hard labeling) is questioned when training (Multimodal) Large Language Models (LLMs/MLLMs). Although it is reasonable for language token prediction, which is a typical multi-class classification problem in disc…

Cited by 0SourcePDFScholar
2024

Hyperbolic Learning with Synthetic Captions for Open-World Detection

CVPR 2024poster

Open-world detection poses significant challenges as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption datasets for training which are extremely expensive to collect. Instead we propose t…

Cited by 6SourcePDFScholar
2024

Self-Supervised Multi-Object Tracking with Path Consistency

CVPR 2024highlight

In this paper we propose a novel concept of path consistency to learn robust object matching without using manual object identity supervision. Our key idea is that to track a object through frames we can obtain multiple different association results from a model by varying the frames it can observe…

2023

ScaleDet: A Scalable Multi-Dataset Object Detector

CVPR 2023poster

Multi-dataset training provides a viable solution for exploiting heterogeneous large-scale datasets without extra annotation cost. In this work, we propose a scalable multi-dataset detector (ScaleDet) that can scale up its generalization across datasets when increasing the number of training dataset…

Cited by 23SourcePDFScholar
2023

SkeleTR: Towards Skeleton-based Action Recognition in the Wild

ICCV 2023poster

We present SkeleTR, a new framework for skeleton-based action recognition. In contrast to prior work, which focuses mainly on controlled environments, we target in-the-wild scenarios that typically involve a variable number of people and various forms of interaction between people. SkeleTR works wit…

Cited by 35PDFScholar
2022

Hierarchical Self-Supervised Representation Learning for Movie Understanding

CVPR 2022poster

Most self-supervised video representation learning approaches focus on action recognition. In contrast, in this paper we focus on self-supervised video learning for movie understanding and propose a novel hierarchical self-supervised pretraining strategy that separately pretrains each level of our h…

Cited by 29PDFcodeScholar
2022

MaCLR: Motion-Aware Contrastive Learning of Representations for Videos

ECCV 2022poster

"We present MaCLR, a novel method to explicitly perform cross-modal self-supervised video representations learning from visual and motion modalities. Compared to previous video representation learn- ing methods that mostly focus on learning motion cues implicitly from RGB inputs, MaCLR enriches stan…

2022

Semi-supervised Vision Transformers at Scale

NeurIPS 2022accept

We study semi-supervised learning (SSL) for vision transformers (ViT), an under-explored topic despite the wide adoption of the ViT architectures to different tasks. To tackle this problem, we use a SSL pipeline, consisting of first un/self-supervised pre-training, followed by supervised fine-tuning…

2022

TubeR: Tubelet Transformer for Video Action Detection

CVPR 2022oral

We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an off-line actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in video by simulta…

Cited by 95PDFScholar
2022

What To Look at and Where: Semantic and Spatial Refined Transformer for Detecting Human-Object Interactions

CVPR 2022oral

We propose a novel one-stage Transformer-based semantic and spatial refined transformer (SSRT) to solve the Human-Object Interaction detection task, which requires to localize humans and objects, and predicts their interactions. Differently from previous Transformer-based HOI approaches, which mostl…

Cited by 66PDFcodeScholar
2021

Selective Feature Compression for Efficient Activity Recognition Inference

ICCV 2021poster

Most action recognition solutions rely on dense sampling to precisely cover the informative temporal clip. Extensively searching temporal region is expensive for a real-world application. In this work, we focus on improving the inference efficiency of current action recognition backbones on trimmed…

Cited by 11PDFScholar
2019

Action Recognition With Spatial-Temporal Discriminative Filter Banks

ICCV 2019poster

Action recognition has seen a dramatic performance improvement in the last few years. Most of the current state-of-the-art literature either aims at improving performance through changes to the backbone CNN network, or exploring different trade-offs between computational efficiency and performance,…

Cited by 90PDFScholar
2015

Joint Calibration of Ensemble of Exemplar SVMs

CVPR 2015poster

We present a method for calibrating the Ensemble of Exemplar SVMs model. Unlike the standard approach, which calibrates each SVM independently, our method optimizes their joint performance as an ensemble. We formulate joint calibration as a constrained optimization problem and devise an efficient op…

Cited by 14SourcePDFScholar