← Search

Hossein Rahmani

35 accepted papers

2026

DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation

CVPR 2026

The rapid growth of the text-to-image (T2I) community has fostered a thriving online ecosystem of expert models, which are variants of pretrained diffusion models specialized for diverse generative capabilities. Yet, existing model merging methods remain limited in fully leveraging abundant online e

Cited by 0SourceScholar
2026

SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition

ICML 2026poster

Skeleton-based action recognition aims to understand human behaviors from body joint sequences and is especially challenging in the one-shot setting, where only a single labeled exemplar is available for each novel action. A key challenge is learning representations that capture the hierarchical and…

Cited by 0SourceScholar
2026

Translating Signals to Languages for sEMG-Based Activity Recognition

CVPR 2026

Surface electromyography (sEMG) signal-based activity recognition has attracted increasing research attention in recent years. To develop accurate sEMG signal-based activity recognizers, numerous approaches have been proposed. Some studies focus on designing larger and more expressive model architec

Cited by 0SourceScholar
2025

An Image-like Diffusion Method for Human-Object Interaction Detection

CVPR 2025poster

Human-object interaction (HOI) detection often faces high levels of ambiguity and indeterminacy, as the same interaction can appear vastly different across different human-object pairs. Additionally, the indeterminacy can be further exacerbated by issues such as occlusions and cluttered backgrounds.…

Cited by 0SourcePDFScholar
2025

Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation

NeurIPS 2025poster

We introduce Skeleton-Cache, the first training-free test-time adaptation framework for skeleton-based zero-shot action recognition (SZAR), aimed at improving model generalization to unseen actions during inference. Skeleton-Cache reformulates inference as a lightweight retrieval process over a non-…

Cited by 0SourcecodeScholar
2025

Boundary Probing for Input Privacy Protection When Using LMM Services

ICCV 2025poster

Alongside the rapid development of Large Multimodal Models (LMMs) like GPT-4V, privacy concerns also rise. As LMMs are commonly deployed as cloud services, users are typically required to upload their personal images and videos to the cloud to access these services, raising great concerns about visu…

Cited by 0SourcePDFScholar
2025

DiffIP: Representation Fingerprints for Robust IP Protection of Diffusion Models

ICCV 2025poster

Intellectual property (IP) protection for diffusion models is a critical concern, given the significant resources and time required for their development. To effectively safeguard the IP of diffusion models, a key step is enabling the comparison of unique identifiers (fingerprints) between suspect a…

Cited by 0SourcePDFScholar
2025

GaussianBlock: Building Part-Aware Compositional and Editable 3D Scene by Primitives and Gaussians

ICLR 2025poster

Recently, with the development of Neural Radiance Fields and Gaussian Splatting, 3D reconstruction techniques have achieved remarkably high fidelity. However, the latent representations learnt by these methods are highly entangled and lack interpretability. In this paper, we propose a novel part-awa…

Cited by 0SourcePDFScholar
2025

MonoDiff9D: Monocular Category-Level 9D Object Pose Estimation via Diffusion Model

ICRA 2025

Object pose estimation is a core means for robots to understand and interact with their environment. For this task, monocular category-level methods are attractive as they require only a single RGB camera. However, current methods rely on shape priors or CAD models of the intra-class known objects.

Cited by 5SourcecodeScholar
2025

Performing Defocus Deblurring by Modeling its Formation Process

ICCV 2025poster

Single image defocus deblurring (SIDD) is a challenging task that aims to recover an all-in-focus image from a defocused one. In this paper, we make the observation that a defocused image can be viewed as a blend of illuminated blobs based on fundamental imaging principles, and the defocus blur in t…

Cited by 0SourcePDFScholar
2024

Class-Agnostic Object Counting with Text-to-Image Diffusion Model

ECCV 2024poster

"Class-agnostic object counting aims to count objects of arbitrary classes with limited information (, a few exemplars or the class names) provided. It requires the model to effectively acquire the characteristics of the target objects and accurately perform counting, which can be challenging. In th…

Cited by 7SourcePDFScholar
2024

Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers

ECCV 2024poster

"We introduce Diff-Tracker, a novel approach for the challenging unsupervised visual tracking task leveraging the pre-trained text-to-image diffusion model. Our main idea is to leverage the rich knowledge encapsulated within the pre-trained diffusion model, such as the understanding of image semanti…

Cited by 12SourcePDFScholar
2024

Weakly Supervised Co-training with Swapping Assignments for Semantic Segmentation

ECCV 2024poster

"Class activation maps (CAMs) are commonly employed in weakly supervised semantic segmentation (WSSS) to produce pseudo-labels. Due to incomplete or excessive class activation, existing studies often resort to offline CAM refinement, introducing additional stages or proposing offline modules. This c…

2023

A Probabilistic Attention Model With Occlusion-Aware Texture Regression for 3D Hand Reconstruction From a Single RGB Image

CVPR 2023poster

Recently, deep learning based approaches have shown promising results in 3D hand reconstruction from a single RGB image. These approaches can be roughly divided into model-based approaches, which are heavily dependent on the model's parameter space, and model-free approaches, which require large num…

2023

DiffPose: Toward More Reliable 3D Pose Estimation

CVPR 2023poster

Monocular 3D human pose estimation is quite challenging due to the inherent ambiguity and occlusion, which often lead to high uncertainty and indeterminacy. On the other hand, diffusion models have recently emerged as an effective tool for generating high-quality images from noise. Inspired by their…

2023

Reinforced Learning for Label-Efficient 3D Face Reconstruction

ICRA 2023poster

3D face reconstruction plays a major role in many human-robot interaction systems, from automatic face authentication to human-computer interface-based entertainment. To improve robustness against occlusions and noise, 3D face reconstruction networks are often trained on a set of in-the-wild face im…

Cited by 1SourceScholar
2023

Token Boosting for Robust Self-Supervised Visual Transformer Pre-Training

CVPR 2023poster

Learning with large-scale unlabeled data has become a powerful tool for pre-training Visual Transformers (VTs). However, prior works tend to overlook that, in real-world scenarios, the input data may be corrupted and unreliable. Pre-training VTs on such corrupted data can be challenging, especially…

Cited by 6SourcePDFScholar
2022

Dynamic Spatio-Temporal Specialization Learning for Fine-Grained Action Recognition

ECCV 2022poster

"The goal of fine-grained action recognition is to successfully discriminate between action categories with subtle differences. To tackle this, we derive inspiration from the human visual system which contains specialized regions in the brain that are dedicated towards handling specific tasks. We de…

Cited by 30SourcePDFScholar
2022

ERA: Expert Retrieval and Assembly for Early Action Prediction

ECCV 2022poster

"Early action prediction aims to successfully predict the class label of an action before it is completely performed. This is a challenging task because the beginning stages of different actions can be very similar, with only minor subtle differences for discrimination. In this paper, we propose a n…

Cited by 29SourcePDFScholar
2022

GradAuto: Energy-Oriented Attack on Dynamic Neural Networks

ECCV 2022poster

"Dynamic neural networks could adapt their structures or parameters based on different inputs. By reducing the computation redundancy for certain samples, it can greatly improve the computational efficiency without compromising the accuracy. In this paper, we investigate the robustness of dynamic ne…

2022

Graph-Context Attention Networks for Size-Varied Deep Graph Matching

CVPR 2022poster

Deep learning for graph matching has received growing interest and developed rapidly in the past decade. Although recent deep graph matching methods have shown excellent performance on matching between graphs of equal size in the computer vision area, the size-varied graph matching problem, where th…

Cited by 29PDFcodeScholar
2022

IGFormer: Interaction Graph Transformer for Skeleton-Based Human Interaction Recognition

ECCV 2022poster

"Human interaction recognition is very important in many applications. One crucial cue in recognizing an interaction is the interactive body parts. In this work, we propose a novel Interaction Graph Transformer (IGFormer) network for skeleton-based interaction recognition via modeling the interactiv…

Cited by 48SourcePDFScholar
2022

REMOTE: Reinforced Motion Transformation Network for Semi-supervised 2D Pose Estimation in Videos

AAAI 2022technical

Existing approaches for 2D pose estimation in videos often require a large number of dense annotations, which are costly and labor intensive to acquire. In this paper, we propose a semi-supervised REinforced MOtion Transformation nEtwork (REMOTE) to leverage a few labeled frames and temporal pose va…

Cited by 13SourcePDFScholar
2021

Else-Net: Elastic Semantic Network for Continual Action Recognition From Skeleton Data

ICCV 2021poster

We address continual action recognition from skeleton sequence, which aims to learn a recognition model over time from a continuous stream of skeleton data. This task is very important in changing environment. Due to catastrophic forgetting problems of deep neural networks and large discrepancies be…

Cited by 56PDFScholar