← Search

Adrian Hilton

18 accepted papers

2025

Improving Gaussian Splatting with Localized Points Management

CVPR 2025highlight

Point management is critical for optimizing 3D Gaussian Splatting models, as point initiation (e.g., via structure from motion) is often distributionally inappropriate. Typically, Adaptive Density Control (ADC) algorithm is adopted, leveraging view-averaged gradient magnitude thresholding for point…

Cited by 0SourcePDFScholar
2025

NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

ICLR 2025poster

Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by characters or agents. This lack of narrative restricts models’ ability to generate text descriptions that capture the causal…

2024

ANIM: Accurate Neural Implicit Model for Human Reconstruction from a single RGB-D Image

CVPR 2024poster

Recent progress in human shape learning shows that neural implicit models are effective in generating 3D human surfaces from limited number of views and even from a single RGB image. However existing monocular approaches still struggle to recover fine geometric details such as face hands or cloth wr…

Cited by 8SourcePDFScholar
2024

CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing

ECCV 2024poster

"Weakly supervised audio-visual video parsing (AVVP) methods aim to detect audible-only, visible-only, and audible-visible events using only video-level labels. Existing approaches tackle this by leveraging unimodal and cross-modal contexts. However, we argue that while cross-modal learning is benef…

2022

Super-Resolution 3D Human Shape from a Single Low-Resolution Image

ECCV 2022poster

"We propose a novel framework to reconstruct super-resolution human shape from a single low-resolution input image. The approach overcomes limitations of existing approaches that reconstruct 3D human shape from a single image, which require high-resolution images together with auxiliary data such as…

2018

Acoustic Reflector Localization and Classification

ICASSP 2018accepted

The process of understanding acoustic properties of environments is important for several applications, such as spatial audio, augmented reality and source separation. In this paper, multichannel room impulse responses are recorded and transformed into their direction of arrival (DOA)-time domain, b…

Cited by 0SourceScholar
2018

Deep Autoencoder for Combined Human Pose Estimation and Body Model Upscaling

ECCV 2018poster

We present a method for simultaneously estimating 3D human pose and body shape from a sparse set of wide-baseline camera views. We train a symmetric convolutional autoencoder with a dual loss that enforces learning of a latent representation that encodes skeletal joint positions, and at the same tim…

Cited by 75SourcePDFScholar
2018

Non-Zero Diffusion Particle Flow SMC-PHD Filter for Audio-Visual Multi-Speaker Tracking

ICASSP 2018accepted

The sequential Monte Carlo probability hypothesis density (SMC-PHD) filter has been shown to be promising for audio-visual multi-speaker tracking. Recently, the zero diffusion particle flow (ZPF) has been used to mitigate the weight degeneracy problem in the SMC-PHD filter. However, this leads to a…

Cited by 0SourceScholar
2018

Volumetric performance capture from minimal camera viewpoints

ECCV 2018poster

We present a convolutional autoencoder that enables high fidelity volumetric reconstructions of human performance to be captured from multi-view video comprising only a small set of camera views. Our method yields similar end-to-end reconstruction error to that of a probabilistic visual hull compute…

Cited by 68SourcePDFScholar
2016

Identity association using PHD filters in multiple head tracking with depth sensors

ICASSP 2016accepted

The work on 3D human pose estimation has been through a significant amount of progress in recent years, particularly due to the widespread availability of commodity depth sensors. However, most pose estimation methods follow a tracking-as-detection approach which does not explicitly handle occlusion…

Cited by 0SourceScholar
2016

Temporally Coherent 4D Reconstruction of Complex Dynamic Scenes

CVPR 2016oral

This paper presents an approach for reconstruction of 4D temporally coherent models of complex dynamic scenes. No prior knowledge is required of scene structure or camera calibration allowing reconstruction from multiple moving cameras. Sparse-to-dense temporal correspondence is integrated with join…

Cited by 91PDFScholar
2015

FaceDirector: Continuous Control of Facial Performance in Video

ICCV 2015poster

We present a method to continuously blend between multiple facial performances of an actor, which can contain different facial expressions or emotional states. As an example, given sad and angry video takes of a scene, our method empowers the movie director to specify arbitrary weighted combinations…

Cited by 19PDFScholar