← Search

Takayuki Okatani

22 accepted papers

2025

Action-Agnostic Point-Level Supervision for Temporal Action Detection

AAAI 2025technical

We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who…

2025

MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval

IJCAI 2025

Result diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasing the diversity metric of image appearances. However, the diversity metric and its desired value vary depending on the

2025

Self-Supervised Learning of Intertwined Content and Positional Features for Object Detection

ICML 2025poster

We present a novel self-supervised feature learning method using Vision Transformers (ViT) as the backbone, specifically designed for object detection and instance segmentation. Our approach addresses the challenge of extracting features that capture both class and positional information, which are…

Cited by 0SourcePDFScholar
2024

Globalizing Local Features: Image Retrieval Using Shared Local Features with Pose Estimation for Faster Visual Localization

ICRA 2024poster

Visual localization is an important sub-task in SfM and visual SLAM that involves estimating a 6-DoF camera pose for an input query image relative to a given 3D model of the environment. The most accurate approach is a hierarchical one that splits the task into two stages: image retrieval and camera…

Cited by 0SourceScholar
2022

Bridging the Gap from Asymmetry Tricks to Decorrelation Principles in Non-contrastive Self-supervised Learning

NeurIPS 2022accept

Recent non-contrastive methods for self-supervised representation learning show promising performance. While they are attractive since they do not need negative samples, it necessitates some mechanism to avoid collapsing into a trivial solution. Currently, there are two approaches to collapse preven…

Cited by 13SourcePDFScholar
2022

GRIT: Faster and Better Image Captioning Transformer Using Dual Visual Features

ECCV 2022poster

"Current state-of-the-art methods for image captioning employ region-based features, as they provide object-level information that is essential to describe the content of images; they are usually extracted by an object detector such as Faster R-CNN. However, they have several issues, such as lack of…

2021

Learning To Bundle-Adjust: A Graph Network Approach to Faster Optimization of Bundle Adjustment for Vehicular SLAM

ICCV 2021poster

Bundle adjustment (BA) occupies a large portion of SfM and visual SLAM's total execution time. Local BA over the latest several keyframes plays a crucial role in visual SLAM. Its execution time should be sufficiently short for robust tracking; this is especially critical for embedded systems with a…

Cited by 9PDFcodeScholar
2021

Look Wide and Interpret Twice: Improving Performance on Interactive Instruction-following Tasks

IJCAI 2021poster

There is a growing interest in the community in making an embodied AI agent perform a complicated task while interacting with an environment following natural language directives. Recent studies have tackled the problem using ALFRED, a well-designed dataset for the task, but achieved only very low a…

Cited by 37SourcePDFScholar
2021

Matching in the Dark: A Dataset for Matching Image Pairs of Low-Light Scenes

ICCV 2021poster

This paper considers matching images of low-light scenes, aiming to widen the frontier of SfM and visual SLAM applications. Recent image sensors can record the brightness of scenes with more than eight-bit precision, available in their RAW-format image. We are interested in making full use of such h…

Cited by 21PDFcodeScholar
2020

Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs

ECCV 2020poster

It has been a primary concern in recent studies of vision and language tasks to design an effective attention mechanism dealing with interactions between the two modalities. The Transformer has recently been extended and applied to several bi-modal tasks, yielding promising results. For visual dialo…

2019

A Generative Model of Underwater Images for Active Landmark Detection and Docking

IROS 2019poster

Underwater active landmarks (UALs) are widely used for short-range underwater navigation in underwater robotics tasks. Detection of UALs is challenging due to large variance of underwater illumination, water quality and change of camera viewpoint. Moreover, improvement of detection accuracy relies u…

Cited by 8SourceScholar
2019

Attention-Based Adaptive Selection of Operations for Image Restoration in the Presence of Unknown Combined Distortions

CVPR 2019poster

Many studies have been conducted so far on image restoration, the problem of restoring a clean image from its distorted version. There are many different types of distortion affecting image quality. Previous studies have focused on single types of distortion, proposing methods for removing them. How…

Cited by 112PDFcodeScholar
2019

Dual Residual Networks Leveraging the Potential of Paired Operations for Image Restoration

CVPR 2019poster

In this paper, we study design of deep neural networks for tasks of image restoration. We propose a novel style of residual connections dubbed "dual residual connection", which exploits the potential of paired operations, e.g., up- and down-sampling or convolution with large- and small-size kernels.…

Cited by 295PDFcodeScholar
2018

Exploiting the Potential of Standard Convolutional Autoencoders for Image Restoration by Evolutionary Search

ICML 2018oral

Researchers have applied deep neural networks to image restoration tasks, in which they proposed various network architectures, loss functions, and training methods. In particular, adversarial training, which is employed in recent studies, seems to be a key ingredient to success. In this paper, we s…

2018

Feature Quantization for Defending Against Distortion of Images

CVPR 2018poster

In this work, we address the problem of improving robustness of convolutional neural networks (CNNs) to image distortion. We argue that higher moment statistics of feature distributions can be shifted due to image distortion, and the shift leads to performance decrease and cannot be reduced by ordin…

Cited by 35SourcePDFScholar
2018

Improved Fusion of Visual and Language Representations by Dense Symmetric Co-Attention for Visual Question Answering

CVPR 2018poster

A key solution to visual question answering (VQA) exists in how to fuse visual and language features extracted from an input image and question. We show that an attention mechanism that enables dense, bi-directional interactions between the two modalities contributes to boost accuracy of prediction…

Cited by 379SourcePDFScholar
2017

Self-Calibration-Based Approach to Critical Motion Sequences of Rolling-Shutter Structure From Motion

CVPR 2017poster

In this paper we consider critical motion sequences (CMSs) of rolling-shutter (RS) SfM. Employing an RS camera model with linearized pure rotation, we show that the RS distortion can be approximately expressed by two internal parameters of an "imaginary" camera plus one-parameter nonlinear transform…

Cited by 36PDFScholar