← Search

Hailin Jin

42 accepted papers

2025

Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval

AAAI 2025technical

Video moment retrieval (VMR) aims to locate the most likely video moment(s) corresponding to a text query in untrimmed videos. Training of existing methods is limited by the lack of diverse and generalisable VMR datasets, hindering their ability to generalise moment-text associations to queries cont…

Cited by 0SourcePDFScholar
2023

Efficient Adaptive Human-Object Interaction Detection with Concept-guided Memory

ICCV 2023poster

Human Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare classes and the high computational cost and time required to…

Cited by 24PDFcodeScholar
2023

Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence Localization

ACL 2023long

Video sentence localization aims to locate moments in an unstructured video according to a given natural language query. A main challenge is the expensive annotation costs and the annotation bias. In this work, we study video sentence localization in a zero-shot setting, which learns with only video…

2023

Moment Detection in Long Tutorial Videos

ICCV 2023poster

Tutorial videos play an increasingly important role in professional development and self-directed education. For users to realise the full benefits of this medium, tutorial videos must be efficiently searchable. In this work, we focus on the task of moment detection, in which the goal is to localise…

Cited by 4PDFcodeScholar
2023

SCCS: Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment

ACL 2023findings

Multimedia summarization with multimodal output (MSMO) is a recently explored application in language grounding. It plays an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or providing introductions to online videos. However, exist…

Cited by 9SourcePDFScholar
2023

Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-Training

CVPR 2023poster

The correlation between the vision and text is essential for video moment retrieval (VMR), however, existing methods heavily rely on separate pre-training feature extractors for visual and textual understanding. Without sufficient temporal boundary annotations, it is non-trivial to learn universal v…

Cited by 40SourcePDFScholar
2022

Cross Modal Retrieval With Querybank Normalisation

CVPR 2022poster

Profiting from large-scale training datasets, advances in neural architecture design and efficient inference, joint embeddings have become the dominant approach for tackling cross-modal retrieval. In this work we first show that, despite their effectiveness, state-of-the-art joint embeddings suffer…

Cited by 99PDFcodeScholar
2022

StyleBabel: Artistic Style Tagging and Captioning

ECCV 2022poster

"We present StyleBabel, a unique open access dataset of natural language captions and free-form tags describing the artistic style of over 135K digital artworks, collected via a novel participatory method from experts studying at specialist art and design schools. StyleBabel was collected via an ite…

Cited by 15SourcePDFScholar
2022

Video Activity Localisation with Uncertainties in Temporal Boundary

ECCV 2022poster

"Current methods for video activity localisation over time assume implicitly that activity temporal boundaries labelled for model training are determined and precise. However, in unscripted natural videos, different activities mostly transit smoothly, so that it is intrinsically ambiguous to determi…

Cited by 31SourcePDFScholar
2021

A Multi-Implicit Neural Representation for Fonts

NeurIPS 2021poster

Fonts are ubiquitous across documents and come in a variety of styles. They are either represented in a native vector format or rasterized to produce fixed resolution images. In the first case, the non-standard representation prevents benefiting from latest network architectures for neural represen…

Cited by 29SourcePDFScholar
2021

ALADIN: All Layer Adaptive Instance Normalization for Fine-Grained Style Similarity

ICCV 2021poster

We present ALADIN (All Layer AdaIN); a novel architecture for searching images based on the similarity of their artistic style. Representation learning is critical to visual search, where distance in the learned search embedding reflects image similarity. Learning an embedding that discriminates fin…

Cited by 32PDFScholar
2021

Cross-Sentence Temporal and Semantic Relations in Video Activity Localisation

ICCV 2021poster

Video activity localisation has recently attained increasing attention due to its practical values in automatically localising the most salient visual segments corresponding to their language descriptions (sentences) from untrimmed and unstructured videos. For supervised model training, a temporal a…

Cited by 82PDFScholar
2021

Look at What I’m Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

NeurIPS 2021spotlight

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompanying transcribed narrations. To achieve this goal, we propose a multilayer cro…

Cited by 26SourcePDFScholar
2021

Magic Layouts: Structural Prior for Component Detection in User Interface Designs

CVPR 2021poster

We present Magic Layouts; a method for parsing screenshots or hand-drawn sketches of user interface (UI) layouts. Our core contribution is to extend existing detectors to exploit a learned structural prior for UI designs, enabling robust detection of UI components; buttons, text boxes and similar. S…

Cited by 11PDFScholar
2021

StreamHover: Livestream Transcript Summarization and Annotation

EMNLP 2021main

With the explosive growth of livestream broadcasting, there is an urgent need for new summarization technology that enables us to create a preview of streamed content and tap into this wealth of knowledge. However, the problem is nontrivial due to the informal nature of spoken language. Further, the…

2021

TeachText: CrossModal Generalized Distillation for Text-Video Retrieval

ICCV 2021poster

In recent years, considerable progress on the task of text-video retrieval has been achieved by leveraging large-scale pretraining on visual and audio datasets to construct powerful video encoders. By contrast, despite the natural symmetry, the design of effective algorithms for exploiting large-sca…

Cited by 167PDFcodeScholar
2020

Geo-PIFu: Geometry and Pixel Aligned Implicit Functions for Single-view Human Reconstruction

NeurIPS 2020poster

We propose Geo-PIFu, a method to recover a 3D mesh from a monocular color image of a clothed person. Our method is based on a deep implicit function-based representation to learn latent voxel features using a structure-aware 3D U-Net, to constrain the model in two ways: first, to resolve feature a…

2019

An Internal Learning Approach to Video Inpainting

ICCV 2019poster

We propose a novel video inpainting algorithm that simultaneously hallucinates missing appearance and motion (optical flow) information, building upon the recent 'Deep Image Prior' (DIP) that exploits convolutional network architectures to enforce plausible texture in static images. In extending DIP…

Cited by 101PDFcodeScholar
2019

Large-Scale Tag-Based Font Retrieval With Generative Feature Learning

ICCV 2019poster

Font selection is one of the most important steps in a design workflow. Traditional methods rely on ordered lists which require significant domain knowledge and are often difficult to use even for trained professionals. In this paper, we address the problem of large-scale tag-based font retrieval wh…

Cited by 36PDFScholar
2018

Disentangling Structure and Aesthetics for Style-Aware Image Completion

CVPR 2018poster

Content-aware image completion or in-painting is a fundamental tool for the correction of defects or removal of objects in images. We propose a non-parametric in-painting algorithm that enforces both structural and aesthetic (style) consistency within the resulting image. Our contributions are two…

Cited by 15SourcePDFScholar
2018

Interactive Boundary Prediction for Object Selection

ECCV 2018poster

Interactive image segmentation is critical for many image editing tasks. While recent advanced methods on interactive segmentation focus on the region-based paradigm, more traditional boundary-based methods such as Intelligent Scissor are still popular in practice as they allow users to have active…

Cited by 65SourcePDFScholar
2018

Multi-Task Adversarial Network for Disentangled Feature Learning

CVPR 2018poster

We address the problem of image feature learning for the applications where multiple factors exist in the image generation process and only some factors are of our interest. We present a novel multi-task adversarial network based on an encoder-discriminator-generator architecture. The encoder extrac…

Cited by 77SourcePDFScholar
2018

Synthetically Supervised Feature Learning for Scene Text Recognition

ECCV 2018poster

We address the problem of image feature learning for scene text recognition. The image features in the state-of-the-art methods are learned from large-scale synthetic image datasets. However, most methods only rely on outputs of the synthetic data generation process, namely realistically looking ima…

Cited by 109SourcePDFScholar
2018

Towards Privacy-Preserving Visual Recognition via Adversarial Training: A Pilot Study

ECCV 2018poster

This paper aims to improve privacy-preserving visual recognition, an increasingly demanded feature in smart camera applications, by formulating a unique adversarial training framework. The proposed framework explicitly learns a degradation transform for the original video inputs, in order to optimiz…

2018

What do I Annotate Next? An Empirical Study of Active Learning for Action Localization

ECCV 2018poster

Despite tremendous progress achieved in temporal action localization, state-of-the-art methods still struggle to train accurate models when annotated data is scarce. In this paper, we introduce a novel active learning framework for temporal localization that aims to mitigate this data dependency iss…

Cited by 51SourcePDFScholar
2018

``Factual'' or ``Emotional'': Stylized Image Captioning with Adaptive Learning and Attention

ECCV 2018poster

Generating stylized captions for an image is an emerging topic in image captioning. Given an image as input, it requires the system to generate a caption that has a specific style (e.g., humorous, romantic, positive, and negative) while describing the image content semantically accurately. In this p…

Cited by 95SourcePDFScholar
2017

BAM! The Behance Artistic Media Dataset for Recognition Beyond Photography

ICCV 2017poster

Computer vision systems are designed to work well within the context of everyday photography. However, artists often render the world around them in ways that do not resemble photographs. Artwork produced by people is not constrained to mimic the physical world, making it more challenging for machin…

Cited by 191PDFScholar
2017

Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks

CVPR 2017poster

Indoor scene understanding is central to applications such as robot navigation and human companion assistance. Over the last years, data-driven deep neural networks have outperformed many traditional approaches thanks to their representation learning capabilities. One of the bottlenecks in training…

Cited by 329PDFScholar
2017

Sketching With Style: Visual Search With Sketches and Aesthetic Context

ICCV 2017poster

We propose a novel measure of visual similarity for image retrieval that incorporates both structural and aesthetic (style) constraints. Our algorithm accepts a query as sketched shape, and a set of one or more contextual images specifying the desired visual aesthetic. A triplet network is used to l…

Cited by 77PDFScholar
2017

Spatial-Semantic Image Search by Visual Feature Synthesis

CVPR 2017spotlight

The performance of image retrieval has been improved tremendously in recent years through the use of deep feature representations. Most existing methods, however, aim to retrieve images that are visually similar or semantically relevant to the query, irrespective of spatial configuration. In this pa…

Cited by 52PDFcodeScholar