← Search

Ajmal Saeed Mian

16 accepted papers

2026

3DSMT: A Hybrid Spiking Mamba-Transformer for Point Cloud Analysis

ICLR 2026poster

The sparse unordered structure of point clouds causes unnecessary computation and energy consumption in deep models. Conventionally, the Transformer architecture is leveraged to model global relationships in point clouds, however, its quadratic complexity restricts scalability. Although the Mamba a…

Cited by 0SourceScholar
2026

ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation

CVPR 2026

3D human reaction generation faces three main challenges: (1) high motion fidelity, (2) real-time inference, and (3) autoregressive adaptability for online scenarios. Existing methods fail to meet all three simultaneously. We propose ARMFlow, a MeanFlow-based autoregressive framework that models tem

Cited by 0SourcecodeScholar
2026

Class-Partitioned VQ-VAE and Latent Flow Matching for Point Cloud Scene Generation

AAAI 2026technical

Most 3D scene generation methods are limited to only generating object bounding box parameters while newer diffusion methods also generate class labels and latent features. Using object size or latent feature, they then retrieve objects from a predefined database. For complex scenes of varied, multi

Cited by 0SourcePDFScholar
2026

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

ICLR 2026poster

Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single latent representation, limiting their ability to capture fine-grained actions an…

Cited by 0SourcecodeScholar
2026

NatADiff: Adversarial Boundary Guidance for Natural Adversarial Diffusion

ICLR 2026poster

Adversarial samples exploit irregularities in the manifold "learned" by deep learning models to cause misclassifications. The study of these adversarial samples provides insight into the features a model uses to classify inputs, which can be leveraged to improve robustness against future attacks. Ho…

Cited by 0SourceScholar
2025

Auto-Regressive Diffusion for Generating 3D Human-Object Interactions

AAAI 2025technical

Text-driven Human-Object Interaction (Text-to-HOI) generation is an emerging field with applications in animation, video games, virtual reality, and robotics. A key challenge in HOI generation is maintaining interaction consistency in long sequences. Existing Text-to-Motion-based approaches, such as…

2025

Beyond Human Perception: Understanding Multi-Object World from Monocular View

CVPR 2025poster

Language and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often availab…

2025

CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation

NeurIPS 2025poster

Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, th…

Cited by 0SourceScholar
2025

Mono3DVLT: Monocular-Video-Based 3D Visual Language Tracking

CVPR 2025poster

Visual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal wi…

2025

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

CVPR 2025poster

In this paper, we introduce Motion-Grounded Video Reasoning, a new motionunderstanding task that requires generating visual answers (video segmentationmasks) according to the input question, and hence needs implicit spatiotemporalreasoning and grounding. This task extends existing spatiotemporal gro…

Cited by 3SourcePDFScholar
2025

Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-Localization

AAAI 2025technical

Visual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing…

2025

Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion

AAAI 2025technical

3D semantic scene completion is critical for multiple downstream tasks in autonomous systems. It estimates missing geometric and semantic information in the acquired scene data. Due to the challenging real-world conditions, this task usually demands complex models that process multi-modal data to ac…

2024

Referring Human Pose and Mask Estimation In the Wild

NeurIPS 2024poster

We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to pr…

2023

Re-calibrating Feature Attributions for Model Interpretation

ICLR 2023top-25%

The ability to interpret machine learning models is critical for high-stakes applications. Due to its desirable theoretical properties, path integration is a widely used scheme for feature attribution to interpret model predictions. However, the methods implementing this scheme currently rely on abs…

Cited by 13SourcePDFScholar
2021

Object-to-Scene: Learning to Transfer Object Knowledge to Indoor Scene Recognition

IROS 2021poster

Accurate perception of the surrounding scene is helpful for robots to make reasonable judgments and behaviours. Therefore, developing effective scene representation and recognition methods are of significant importance in robotics. Currently, a large body of research focuses on developing novel auxi…

Cited by 34SourcecodeScholar