← Search

Sridha Sridharan

20 accepted papers

2025

AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification

CVPR 2025poster

We introduce AG-VPReID, a new large-scale dataset for aerial-ground video-based person re-identification (ReID) that comprises 6,632 subjects, 32,321 tracklets and over 9.6 million frames captured by drones (altitudes ranging from 15-120m), CCTV, and wearable cameras. This dataset offers a real-worl…

2025

Event2Tracking: Reconstructing Multi-Agent Soccer Trajectories Using Long-Term Multimodal Context

AAAI 2025technical

Soccer is a rich testbed for studying multi-agent adversarial systems. In this work we focus on the task of reconstructing the noisy trajectories of soccer agents (players and the ball). Previous works that model the behaviours of agents in soccer are limited in two respects: (i) they only focus on…

Cited by 0SourcePDFScholar
2024

GeoAdapt: Self-Supervised Test-Time Adaptation in LiDAR Place Recognition Using Geometric Priors

RA-L 2024

LiDAR place recognition approaches based on deep learning suffer from significant performance degradation when there is a shift between the distribution of training and test datasets, often requiring re-training the networks to achieve peak performance. However, obtaining accurate ground truth data

Cited by 11SourceScholar
2023

Spectral Geometric Verification: Re-Ranking Point Cloud Retrieval for Metric Localization

RA-L 2023

In large-scale metric localization, an incorrect result during retrieval will lead to an incorrect pose estimate or loop closure. Re-ranking methods propose to take into account all the top retrieval candidates and re-order them to increase the likelihood of the top candidate being correct. However,

Cited by 35SourcecodeScholar
2023

Wild-Places: A Large-Scale Dataset for Lidar Place Recognition in Unstructured Natural Environments

ICRA 2023poster

Many existing datasets for lidar place recognition are solely representative of structured urban environments, and have recently been saturated in performance by deep learning based approaches. Natural and unstructured environments present many additional challenges for the tasks of long-term locali…

Cited by 53SourcecodeScholar
2022

InCloud: Incremental Learning for Point Cloud Place Recognition

IROS 2022poster

Place recognition is a fundamental component of robotics, and has seen tremendous improvements through the use of deep learning models in recent years. Networks can experience significant drops in performance when deployed in unseen or highly dynamic environments, and require additional training on…

Cited by 32SourcecodeScholar
2022

LoGG3D-Net: Locally Guided Global Descriptor Learning for 3D Place Recognition

ICRA 2022poster

Retrieval-based place recognition is an efficient and effective solution for re-localization within a pre-built map, or global data association for Simultaneous Localization and Mapping (SLAM). The accuracy of such an approach is heavily dependant on the quality of the extracted scene-level represen…

Cited by 98SourcecodeScholar
2022

SESS: Saliency Enhancing with Scaling and Sliding

ECCV 2022poster

"High-quality saliency maps are essential in several machine learning application areas including explainable AI and weakly supervised object detection and segmentation. Many techniques have been developed to generate better saliency using neural networks. However, they are often limited to specific…

2021

Locus: LiDAR-based Place Recognition using Spatiotemporal Higher-Order Pooling

ICRA 2021poster

Place Recognition enables the estimation of a globally consistent map and trajectory by providing non-local constraints in Simultaneous Localisation and Mapping (SLAM). This paper presents Locus, a novel place recognition method using 3D LiDAR point clouds in large-scale environments. We propose a m…

Cited by 91SourcecodeScholar
2020

Attention Driven Fusion for Multi-Modal Emotion Recognition

ICASSP 2020accepted

Deep learning has emerged as a powerful alternative to hand-crafted methods for emotion recognition on combined acoustic and text modalities. Baseline systems model emotion information in text and acoustic modes independently using Deep Convolutional Neural Networks (DCNN) and Recurrent Neural Netwo…

Cited by 0SourceScholar
2020

Spatiotemporal Camera-LiDAR Calibration: A Targetless and Structureless Approach

RA-L 2020

The demand for multimodal sensing systems for robotics is growing due to the increase in robustness, reliability and accuracy offered by these systems. These systems also need to be spatially and temporally co-registered to be effective. In this letter, we propose a targetless and structureless spat

Cited by 99SourceScholar
2019

Investigating Domain Sensitivity of DNN Embeddings for Speaker Recognition Systems

ICASSP 2019accepted

A speaker embeddings framework achieves state-of-the-art speaker recognition performance by modeling speaker discriminant information directly using deep neural networks (DNNs). After the introduction of neural network based speaker embeddings, researchers have explored the requirements for training…

Cited by 0SourceScholar
2019

Predicting the Future: A Jointly Learnt Model for Action Anticipation

ICCV 2019poster

Inspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current state-of-the-art methods which first learn a model to predict…

Cited by 116PDFScholar
2019

Robust Photogeometric Localization Over Time for Map-Centric Loop Closure

RA-L 2019

Map-centric Simultaneous Localization And Mapping (SLAM) is emerging as an alternative of conventional graph-based SLAM for its accuracy and efficiency in long-term mapping problems. However, in map-centric SLAM, the process of loop closure differs from that of conventional SLAM and the result of in

Cited by 14SourceScholar
2018

Calibrating Cameras in Poor-Conditioned Pitch-Based Sports Games

ICASSP 2018accepted

Camera calibration is a preliminary step in sports analytics which enables us to transform player positions to standard playing area coordinates. While many camera calibration systems work well when the visual content contains sufficient clues, such as a key frame, calibrating without such informati…

Cited by 0SourceScholar
2018

Elastic LiDAR Fusion: Dense Map-Centric Continuous-Time SLAM

ICRA 2018poster

The concept of continuous-time trajectory representation has brought increased accuracy and efficiency to multi-modal sensor fusion in modern SLAM. However, regardless of these advantages, its offline property caused by the requirement of global batch optimization is critically hindering its relevan…

Cited by 120SourceScholar
2015

A cluster-voting approach for speaker diarization and linking of Australian broadcast news recordings

ICASSP 2015accepted

We present a clustering-only approach to the problem of speaker diarization to eliminate the need for the commonly employed and computationally expensive Viterbi segmentation and realignment stage. We use multiple linear segmentations of a recording and carry out complete-linkage clustering within e…

Cited by 0SourceScholar
2015

Improving out-domain PLDA speaker verification using unsupervised inter-dataset variability compensation approach

ICASSP 2015accepted

Experimental studies have found that when the state-of-the-art probabilistic linear discriminant analysis (PLDA) speaker verification systems are trained using out-domain data, it significantly affects speaker verification performance due to the mismatch between development data and evaluation data.…

Cited by 0SourceScholar
2015

Searching for semantic person queries using channel representations

ICASSP 2015accepted

It is not uncommon to hear a person of interest described by their height, build, and clothing (i.e. type and colour). These semantic descriptions are commonly used by people to describe others, as they are quick to relate and easy to understand. However such queries are not easily utilised within i…

Cited by 0SourceScholar