← Search

Hamid Rezatofighi

47 accepted papers

2026

An Empirical Study on How Video-LLMs Answer Video Questions

CVPR 2026

Taking advantage of large-scale data and pretrained language models, Video Large Language Models (Video-LLMs) have shown strong capabilities in answering video questions. However, most existing efforts focus on improving performance, with limited attention to understanding their internal mechanisms.

Cited by 0SourceScholar
2026

DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities

CVPR 2026

The performance of Vision-Language Transformers drops sharply when an input modality (e.g., image) is missing, because the model is forced to make predictions using incomplete information. Existing missing-aware prompt methods help reduce this degradation, but they still rely on conventional predict

Cited by 0SourcecodeScholar
2026

JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics

AAAI 2026technical

Recent advances in Vision-Language Models (VLMs) and large language models (LLMs) have greatly enhanced visual reasoning, a key capability for embodied AI agents like robots. However, existing visual reasoning benchmarks often suffer from several limitations: they lack a clear definition of reasonin

Cited by 0SourcePDFScholar
2026

MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning

ICLR 2026poster

Recent vision-language models have strong perceptual ability but their implicit reasoning is hard to explain and easily generates hallucinations on complex queries. Compositional methods improve interpretability, but most rely on a single agent or hand-crafted pipeline and cannot decide when to coll…

Cited by 0SourcecodeScholar
2026

Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Convex Parametric Shapes

AAAI 2026technical

Optimizing the similarity between parametric shapes is crucial for numerous computer vision tasks, where Intersection over Union (IoU) stands as the canonical measure. However, existing optimization methods exhibit significant shortcomings: regression-based losses like L1/L2 lack correlation with Io

Cited by 0SourcePDFScholar
2026

NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions

ICRA 2026poster

This paper addresses the problem of autonomous UAV search missions, where a UAV must locate specific Entities of Interest (EOIs) within a time limit, based on brief descriptions in large, hazard-prone environments with keep-out zones. The UAV must perceive, reason, and make decisions with limited an…

2026

VQ-VA World: Towards High-Quality Visual Question-Visual Answering

CVPR 2026

This paper studies Visual Question-Visual Answering (VQ-VA): generating an image, rather than text, in response to a visual question---an ability that has recently emerged in proprietary systems such as NanoBanana and GPT-Image. To also bring this capability to open-source models, we introduce VQ-VA

Cited by 0SourcecodeScholar
2025

Acoustic-to-Hyper-Spectral: Hyper-Spectral Image Construction from Frequency Spectrums Through Simulated Annealing (Student Abstract)

AAAI 2025technical

This abstract presents a simulated annealing based approach that constructs hyper-spectral images from the frequency spectrums of a distributed acoustic sensing system and iteratively improves them through the training of learnable filters. The aim is to construct an image that represents features o…

Cited by 0SourcePDFScholar
2025

DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning

ICCV 2025poster

Visual reasoning (VR), which is crucial in many fields for enabling human-like visual understanding, remains highly challenging. Recently, compositional visual reasoning approaches, which leverage the reasoning abilities of large language models (LLMs) with integrated tools to solve problems, have s…

Cited by 0SourcePDFScholar
2025

DrVideo: Document Retrieval Based Long Video Understanding

CVPR 2025poster

Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos poses two main challenges: difficulty in locating key information and performing…

2025

Hier-SLAM: Scaling-Up Semantics in SLAM with a Hierarchically Categorical Gaussian Splatting

ICRA 2025

We propose Hier-SLAM, a semantic 3D Gaussian Splatting SLAM method featuring a novel hierarchical categorical representation, which enables accurate global 3D semantic mapping, scaling-up capability, and explicit semantic label prediction in the 3D world. The parameter usage in semantic SLAM systems

Cited by 19SourcecodeScholar
2025

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

ICCV 2025poster

Visual Grounding (VG) tasks, such as referring expression detection and segmentation tasks are important for linking visual entities to context, especially in complex reasoning tasks that require detailed query interpretation. This paper explores VG beyond basic perception, highlighting challenges f…

2025

NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions

RA-L 2025

This paper addresses the problem of autonomous UAV search missions, where a UAV must locate specific Entities of Interest (EOIs) within a time limit, based on brief descriptions in large, hazard-prone environments with keep-out zones. The UAV must perceive, reason, and make decisions with limited an

Cited by 14SourceScholar
2024

Diffusion Model for Robust Multi-Sensor Fusion in 3D Object Detection and BEV Segmentation

ECCV 2024poster

"Diffusion models have recently gained prominence as powerful deep generative models, demonstrating unmatched performance across various domains. However, their potential in multi-sensor fusion remains largely unexplored. In this work, we introduce “DifFUSER”, a novel approach that leverages diffusi…

Cited by 2SourcePDFScholar
2024

HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning

ECCV 2024poster

"Recent advances in visual reasoning (VR), particularly with the aid of Large Vision-Language Models (VLMs), show promise but require access to large-scale datasets and face challenges such as high computational costs and limited generalization capabilities. Compositional visual reasoning approaches…

2024

JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments

CVPR 2024poster

Autonomous robot systems have attracted increasing research attention in recent years where environment understanding is a crucial step for robot navigation human-robot interaction and decision. Real-world robot systems usually collect visual data from multiple sensors and are required to recognize…

Cited by 2SourcePDFScholar
2024

JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups

CVPR 2024poster

Understanding human social behaviour is crucial in computer vision and robotics. Micro-level observations like individual actions fall short necessitating a comprehensive approach that considers individual behaviour intra-group dynamics and social group levels for a thorough understanding. To addres…

Cited by 5SourcePDFScholar
2024

Normal-GS: 3D Gaussian Splatting with Normal-Involved Rendering

NeurIPS 2024poster

Rendering and reconstruction are long-standing topics in computer vision and graphics. Achieving both high rendering quality and accurate geometry is a challenge. Recent advancements in 3D Gaussian Splatting (3DGS) have enabled high-fidelity novel view synthesis at real-time speeds. However, the noi…

Cited by 2SourcePDFScholar
2024

Shape-prior Free Space-time Neural Radiance Field for 4D Semantic Reconstruction of Dynamic Scene from Sparse-View RGB Videos

IROS 2024

Many applications in Augmented/Virtual Reality or robotics require precise geometry modeling of individual elements in a dynamic scene under a sparse-view camera setup, without any prior information about their semantic labels or shapes. In our research, we introduce a 3D shape prior-free Neural Rad

Cited by 0SourceScholar
2024

TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene

NeurIPS 2024poster

Despite advancements in Neural Implicit models for 3D surface reconstruction, handling dynamic environments with interactions between arbitrary rigid, non-rigid, or deformable entities remains challenging. The generic reconstruction methods adaptable to such dynamic scenes often require additional i…

2023

Accurate and Real-Time 3D Pedestrian Detection Using an Efficient Attentive Pillar Network

RA-L 2023

Efficiently and accurately detecting people from 3D point cloud data is of great importance in many robotic and autonomous driving applications. This fundamental perception task is still very challenging due to <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/199

Cited by 29SourcecodeScholar
2023

Energy-based Self-Training and Normalization for Unsupervised Domain Adaptation

ICCV 2023poster

We propose an Unsupervised Domain Adaptation (UDA) method by making use of Energy-Based Learning (EBL) and demonstrate 1. EBL can be used to improve the instance selection for a self-training task on the unlabelled target domain, and 2. alignment and normalizing energy scores can learn domain-invari…

Cited by 9PDFScholar
2023

JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking

CVPR 2023poster

Autonomous robotic systems operating in human environments must understand their surroundings to make accurate and safe decisions. In crowded human scenes with close-up human-robot interaction and robot navigation, a deep understanding of surrounding people requires reasoning about human motion and…

Cited by 36SourcePDFScholar
2023

Knowledge Combination To Learn Rotated Detection Without Rotated Annotation

CVPR 2023poster

Rotated bounding boxes drastically reduce output ambiguity of elongated objects, making it superior to axis-aligned bounding boxes. Despite the effectiveness, rotated detectors are not widely employed. Annotating rotated bounding boxes is such a laborious process that they are not provided in many d…

2023

MARLIN: Masked Autoencoder for Facial Video Representation LearnINg

CVPR 2023poster

This paper proposes a self-supervised approach to learn universal facial representations from videos, that can transfer across a variety of facial analysis tasks such as Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). O…

2023

Physically Plausible 3D Human-Scene Reconstruction From Monocular RGB Image Using an Adversarial Learning Approach

RA-L 2023

Holistic 3D human-scene reconstruction is a crucial and emerging research area in robot perception. A key challenge in holistic 3D human-scene reconstruction is to generate a physically plausible 3D scene from a single monocular RGB image. The existing research mainly proposes optimization-based app

Cited by 4SourceScholar
2023

ProtoCon: Pseudo-Label Refinement via Online Clustering and Prototypical Consistency for Efficient Semi-Supervised Learning

CVPR 2023highlight

Confidence-based pseudo-labeling is among the dominant approaches in semi-supervised learning (SSL). It relies on including high-confidence predictions made on unlabeled data as additional targets to train the model. We propose ProtoCon, a novel SSL method aimed at the less-explored label-scarce SSL…

Cited by 29SourcePDFScholar
2023

Tracking Different Ant Species: An Unsupervised Domain Adaptation Framework and a Dataset for Multi-object Tracking

IJCAI 2023poster

Tracking individuals is a vital part of many experiments conducted to understand collective behaviour. Ants are the paradigmatic model system for such experiments but their lack of individually distinguishing visual features and their high colony densities make it extremely difficult to perform reli…

2022

GMFlow: Learning Optical Flow via Global Matching

CVPR 2022oral

Learning-based optical flow estimation has been dominated with the pipeline of cost volume with convolutions for flow regression, which is inherently limited to local correlations and thus is hard to address the long-standing challenge of large displacements. To alleviate this, the state-of-the-art…

Cited by 459PDFcodeScholar
2022

JRDB-Act: A Large-Scale Dataset for Spatio-Temporal Action, Social Group and Activity Detection

CVPR 2022poster

The availability of large-scale video action understanding datasets has facilitated advances in the interpretation of visual scenes containing people. However, learning to recognise human actions and their social interactions in an unconstrained real-world environment comprising numerous people, wit…

Cited by 46PDFScholar
2021

MOLTR: Multiple Object Localization, Tracking and Reconstruction From Monocular RGB Videos

RA-L 2021

Semantic aware reconstruction is more advantageous than geometric-only reconstruction for future robotic and AR/VR applications because it represents not only where things are, but also what things are. Object-centric mapping is a task to build an object-level reconstruction where objects are separa

Cited by 25SourceScholar
2021

ODAM: Object Detection, Association, and Mapping Using Posed RGB Video

ICCV 2021poster

Localizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system…

Cited by 34PDFcodeScholar
2021

Probabilistic Tracklet Scoring and Inpainting for Multiple Object Tracking

CVPR 2021poster

Despite the recent advances in multiple object tracking (MOT), achieved by joint detection and tracking, dealing with long occlusions remains a challenge. This is due to the fact that such techniques tend to ignore the long-term motion information. In this paper, we introduce a probabilistic autoreg…

Cited by 110PDFcodeScholar
2021

TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild

ICCV 2021poster

Joint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems. Predicting body dynamics requires capturing subtle information embedded in the humans' interactions with each other an…

Cited by 64PDFScholar
2021

Unsupervised Image Segmentation by Mutual Information Maximization and Adversarial Regularization

RA-L 2021

Semantic segmentation is one of the basic, yet essential scene understanding tasks for an autonomous agent. The recent developments in supervised machine learning and neural networks have enjoyed great success in enhancing the performance of the state-of-the-art techniques for this task. However, th

Cited by 36SourceScholar
2020

JRMOT: A Real-Time 3D Multi-Object Tracker and a New Large-Scale Dataset

IROS 2020poster

Robots navigating autonomously need to perceive and track the motion of objects and other agents in its surroundings. This information enables planning and executing robust and safe trajectories. To facilitate these processes, the motion should be perceived in 3D Cartesian space. However, most recen…

Cited by 106SourcecodeScholar
2020

Joint Learning of Social Groups, Individuals Action and Sub-group Activities in Videos

ECCV 2020poster

Individuals Action and Sub-group Activities in Videos","The state-of-the art solutions for human activity understanding from a video stream formulate the task as a spatio-temporal problem which requires joint localization of all individuals in the scene and classification of their actions or group a…

2020

LAVAPilot: Lightweight UAV Trajectory Planner with Situational Awareness for Embedded Autonomy to Track and Locate Radio-tags

IROS 2020poster

Tracking and locating radio-tagged wildlife is a labor-intensive and time-consuming task necessary in wildlife conservation. In this article, we focus on the problem of achieving embedded autonomy for a resource-limited aerial robot for the task capable of avoiding undesirable disturbances to wildli…

Cited by 31SourcecodeScholar
2020

Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation

IROS 2020poster

Video object segmentation plays a vital role to many robotic tasks, beyond the satisfied accuracy, quickly adapt to the new scenario with very limited annotations and conduct a quick inference are also important. In this paper, we are specifically concerned with the task of fast segmenting all pixel…

Cited by 14SourceScholar
2020

Socially and Contextually Aware Human Motion and Pose Forecasting

RA-L 2020

Smooth and seamless robot navigation while interacting with humans depends on predicting human movements. Forecasting such human dynamics often involves modeling human trajectories (global motion) or detailed body joint movements (local motion). Prior work typically tackled local and global human mo

Cited by 94SourceScholar
2019

Approximating the Permanent by Sampling from Adaptive Partitions

NeurIPS 2019poster

Computing the permanent of a non-negative matrix is a core problem with practical applications ranging from target tracking to statistical thermodynamics. However, this problem is also #P-complete, which leaves little hope for finding an exact solution that can be computed efficiently. While the pr…

Cited by 11SourcePDFScholar
2019

Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression

CVPR 2019poster

Intersection over Union (IoU) is the most popular evaluation metric used in the object detection benchmarks. However, there is a gap between optimizing the commonly used distance losses for regressing the parameters of a bounding box and maximizing this metric value. The optimal objective for a metr…

Cited by 6789PDFcodeScholar
2019

SoPhie: An Attentive GAN for Predicting Paths Compliant to Social and Physical Constraints

CVPR 2019poster

This paper addresses the problem of path prediction for multiple interacting agents in a scene, which is a crucial step for many autonomous platforms such as self-driving cars and social robots. We present SoPhie; an interpretable framework based on Generative Adversarial Network (GAN), which levera…

Cited by 1245PDFcodeScholar
2019

Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks

NeurIPS 2019poster

Predicting the future trajectories of multiple interacting pedestrians in a scene has become an increasingly important problem for many different applications ranging from control of autonomous vehicles and social robots to security and surveillance. This problem is compounded by the presence of soc…

Cited by 838SourcePDFScholar