← Search

Mohan Kankanhalli

45 accepted papers

2026

Aggregating Diverse Cue Experts for AI-Generated Image Detection

AAAI 2026technical

The rapid emergence of image synthesis models poses challenges to the generalization of AI-generated image detectors. However, existing methods often rely on model-specific features, leading to overfitting and poor generalization. In this paper, we introduce the Multi-Cue Aggregation Network (MCAN),

Cited by 0SourcePDFScholar
2026

MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer

CVPR 2026

3D pose transfer aims to transfer the pose-style of a source mesh to a target character while preserving both the target's geometry and the source's pose characteristic. Existing methods are largely restricted to characters with similar structures and fail to generalize to category-free settings (e.

Cited by 0SourceScholar
2026

Mitigating Noise-Induced Layout Priors for Object Counting in Diffusion Models

ICML 2026poster

Despite remarkable progress in text-to-image diffusion models, accurately generating the specified number of objects remains a persistent challenge. We identify the initial noise as a primary determinant of spatial layout formation, with early-stage cross-attention serving as the key mechanism that …

Cited by 0SourceScholar
2026

Object-Centric Framework for Video Moment Retrieval

AAAI 2026technical

Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing mo

Cited by 0SourcePDFScholar
2025

Fair Deepfake Detectors Can Generalize

NeurIPS 2025poster

Deepfake detection models face two critical challenges: generalization to unseen manipulations and demographic fairness among population groups. However, existing approaches often demonstrate that these two objectives are inherently conflicting, revealing a trade-off between them. In this paper, we,…

Cited by 0SourceScholar
2025

Multi-Modal Recommendation Unlearning for Legal, Licensing, and Modality Constraints

AAAI 2025technical

User data spread across multiple modalities has popularized multi-modal recommender systems (MMRS). They recommend diverse content such as products, social media posts, TikTok reels, etc., based on a user-item interaction graph. With rising data privacy demands, recent methods propose unlearning pri…

2025

Nine Ways to Break Copyright Law and Why Our LLM Won’t: A Fair Use Aligned Generation Framework

EMNLP 2025

Large language models (LLMs) commonly risk copyright infringement by reproducing protected content verbatim or with insufficient transformative modifications, posing significant ethical, legal, and practical concerns. Current inference-time safeguards predominantly rely on restrictive refusal-based

Cited by 0SourcePDFScholar
2025

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

NeurIPS 2025poster

The vulnerability of Vision Large Language Models (VLLMs) to jailbreak attacks appears as no surprise. However, recent defense mechanisms against these attacks have reached near-saturation performance on benchmark evaluations, often with minimal effort. This dual high performance in both attack and…

Cited by 0SourceScholar
2025

Tree-of-Evolution: Tree-Structured Instruction Evolution for Code Generation in Large Language Models

ACL 2025long

Data synthesis has become a crucial research area in large language models (LLMs), especially for generating high-quality instruction fine-tuning data to enhance downstream performance. In code generation, a key application of LLMs, manual annotation of code instruction data is costly. Recent method…

2025

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

CVPR 2025poster

Large multimodal models (LMMs) with advanced video analysis capabilities have recently garnered significant attention. However, most evaluations rely on traditional methods like multiple-choice question answering in benchmarks such as VideoMME and LongVideoBench, which are prone to lack the depth ne…

Cited by 6SourcePDFScholar
2024

An LLM can Fool Itself: A Prompt-Based Adversarial Attack

ICLR 2024poster

The wide-ranging applications of large language models (LLMs), especially in safety-critical domains, necessitate the proper evaluation of the LLM’s adversarial robustness. This paper proposes an efficient tool to audit the LLM’s adversarial robustness via a prompt-based adversarial attack (PromptAt…

2024

Bilateral Adaptation for Human-Object Interaction Detection with Occlusion-Robustness

CVPR 2024poster

Human-Object Interaction (HOI) Detection constitutes an important aspect of human-centric scene understanding which requires precise object detection and interaction recognition. Despite increasing advancement in detection recognizing subtle and intricate interactions remains challenging. Recent met…

Cited by 6SourcePDFScholar
2024

Finetuning Text-to-Image Diffusion Models for Fairness

ICLR 2024oral

The rapid adoption of text-to-image diffusion models in society underscores an urgent need to address their biases. Without interventions, these biases could propagate a skewed worldview and restrict opportunities for minority groups. In this work, we frame fairness as a distributional alignment pro…

2024

Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition Learning

ICML 2024poster

Previous efforts using frozen Large Language Models (LLMs) for visual understanding, via image captioning or image-text retrieval tasks, face challenges when dealing with complex multimodal scenarios. In order to enhance the capabilities of Multimodal Large Language Models (MLLM) in comprehending th…

2024

MCM: Multi-condition Motion Synthesis Framework

IJCAI 2024poster

Conditional human motion synthesis (HMS) aims to generate human motion sequences that conform to specific conditions. Text and audio represent the two predominant modalities employed as HMS control conditions. While existing research has primarily focused on single conditions, the multi-condition hu…

Cited by 1SourcePDFScholar
2024

PELA: Learning Parameter-Efficient Models with Low-Rank Approximation

CVPR 2024poster

Applying a pre-trained large model to downstream tasks is prohibitive under resource-constrained conditions. Recent dominant approaches for addressing efficiency issues involve adding a few learnable parameters to the fixed backbone model. This strategy however leads to more challenges in loading la…

2024

Perplexity-aware Correction for Robust Alignment with Noisy Preferences

NeurIPS 2024poster

Alignment techniques are critical in ensuring that large language models (LLMs) output helpful and harmless content by enforcing the LLM-generated content to align with human preferences. However, the existence of noisy preferences (NPs), where the responses are mistakenly labelled as chosen or rej…

2024

TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

NeurIPS 2024spotlight

Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability of substantial web video-text data. This difficulty primarily arises from the inherent complexity of videos and the inef…

2023

Can Bad Teaching Induce Forgetting? Unlearning in Deep Networks Using an Incompetent Teacher

AAAI 2023technical

Machine unlearning has become an important area of research due to an increasing need for machine learning (ML) applications to comply with the emerging data privacy regulations. It facilitates the provision for removal of certain set or class of data from an already trained ML model without requiri…

2023

Continuous-Discrete Convolution for Geometry-Sequence Modeling in Proteins

ICLR 2023poster

The structure of proteins involves 3D geometry of amino acid coordinates and 1D sequence of peptide chains. The 3D structure exhibits irregularity because amino acids are distributed unevenly in Euclidean space and their coordinates are continuous variables. In contrast, the 1D structure is regular…

Cited by 50SourcePDFScholar
2023

DSFNet: Dual Space Fusion Network for Occlusion-Robust 3D Dense Face Alignment

CVPR 2023poster

Sensitivity to severe occlusion and large view angles limits the usage scenarios of the existing monocular 3D dense face alignment methods. The state-of-the-art 3DMM-based method, directly regresses the model's coefficients, underutilizing the low-level 2D spatial and semantic information, which can…

2023

Efficient Adversarial Contrastive Learning via Robustness-Aware Coreset Selection

NeurIPS 2023spotlight

Adversarial contrastive learning (ACL) does not require expensive data annotations but outputs a robust representation that withstands adversarial attacks and also generalizes to a wide range of downstream tasks. However, ACL needs tremendous running time to generate the adversarial variants of all…

2023

Enhancing Adversarial Contrastive Learning via Adversarial Invariant Regularization

NeurIPS 2023poster

Adversarial contrastive learning (ACL) is a technique that enhances standard contrastive learning (SCL) by incorporating adversarial data to learn a robust representation that can withstand adversarial attacks and common corruptions without requiring costly annotations. To improve transferability, t…

2023

Text to Point Cloud Localization with Relation-Enhanced Transformer

AAAI 2023technical

Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, we focus on a text-to-point-cloud cross-modal localization problem. Given a textual query, it aims to identify the descri…

2022

Adversarial Attack and Defense for Non-Parametric Two-Sample Tests

ICML 2022spotlight

Non-parametric two-sample tests (TSTs) that judge whether two sets of samples are drawn from the same distribution, have been widely used in the analysis of critical data. People tend to employ TSTs as trusted basic tools and rarely have any doubt about their reliability. This paper systematically u…

2022

Chairs Can Be Stood On: Overcoming Object Bias in Human-Object Interaction Detection

ECCV 2022poster

"Detecting Human-Object Interaction (HOI) in images is an important step towards high-level visual comprehension. Existing work often shed light on improving either human and object detection, or interaction recognition. However, due to the limitation of datasets, these methods tend to fit well on f…

2022

Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action Segmentation

NeurIPS 2022accept

We propose Differentiable Temporal Logic (DTL), a model-agnostic framework that introduces temporal constraints to deep networks. DTL treats the outputs of a network as a truth assignment of a temporal logic formula, and computes a temporal logic loss reflecting the consistency between the output an…

Cited by 40SourcePDFScholar
2022

Learning Realistic Patterns from Visually Unrealistic Stimuli: Generalization and Data Anonymization (Extended Abstract)

IJCAI 2022poster

Good training data is a prerequisite to develop useful Machine Learning applications. However, in many domains existing data sets cannot be shared due to privacy regulations (e.g., from medical studies). This work investigates a simple yet unconventional approach for anonymized data synthesis to en…

Cited by 4SourcePDFScholar
2022

Self-Supervised Global-Local Structure Modeling for Point Cloud Domain Adaptation With Reliable Voted Pseudo Labels

CVPR 2022poster

In this paper, we propose an unsupervised domain adaptation method for deep point cloud representation learning. To model the internal structures in target point clouds, we first propose to learn the global representations of unlabeled data by scaling up or down point clouds and then predicting the…

Cited by 68PDFScholar
2021

Geometry-aware Instance-reweighted Adversarial Training

ICLR 2021oral

In adversarial machine learning, there was a common belief that robustness and accuracy hurt each other. The belief was challenged by recent studies where we can maintain the robustness and improve the accuracy. However, the other direction, whether we can keep the accuracy and improve the robustnes…

Cited by 339SourcePDFScholar
2021

Learning Causal Representation for Training Cross-Domain Pose Estimator via Generative Interventions

ICCV 2021poster

3D pose estimation has attracted increasing attention with the availability of high-quality benchmark datasets. However, prior works show that deep learning models tend to learn spurious correlations, which fail to generalize beyond the specific dataset they are trained on. In this work, we take a s…

Cited by 39PDFScholar
2021

Learning to Predict Trustworthiness with Steep Slope Loss

NeurIPS 2021poster

Understanding the trustworthiness of a prediction yielded by a classifier is critical for the safe and effective use of AI models. Prior efforts have been proven to be reliable on small-scale datasets. In this work, we study the problem of predicting trustworthiness on real-world large-scale dataset…

2021

PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences

ICLR 2021poster

Point cloud sequences are irregular and unordered in the spatial dimension while exhibiting regularities and order in the temporal dimension. Therefore, existing grid based convolutions for conventional video processing cannot be directly applied to spatio-temporal modeling of raw point cloud sequen…

2021

Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud Videos

CVPR 2021poster

Point cloud videos exhibit irregularities and lack of order along the spatial dimension where points emerge inconsistently across different frames. To capture the dynamics in point cloud videos, point tracking is usually employed. However, as points may flow in and out across frames, computing accur…

Cited by 214PDFcodeScholar
2021

Unsupervised Motion Representation Learning with Capsule Autoencoders

NeurIPS 2021poster

We propose the Motion Capsule Autoencoder (MCAE), which addresses a key challenge in the unsupervised learning of motion representations: transformation invariance. MCAE models motion in a two-level hierarchy. In the lower level, a spatio-temporal motion signal is divided into short, local, and sema…

2020

Attacks Which Do Not Kill Training Make Adversarial Learning Stronger

ICML 2020poster

Adversarial training based on the minimax formulation is necessary for obtaining adversarial robustness of trained models. However, it is conservative or even pessimistic so that it sometimes hurts the natural generalization. In this paper, we raise a fundamental question{—}do we have to trade off n…

Cited by 505SourcePDFScholar
2020

Inferring DQN structure for high-dimensional continuous control

ICML 2020poster

Despite recent advancements in the field of Deep Reinforcement Learning, Deep Q-network (DQN) models still show lackluster performance on problems with high-dimensional action spaces. The problem is even more pronounced for cases with high-dimensional continuous action spaces due to a combinatorial…

Cited by 11SourcePDFScholar
2019

Sublinear Time Nearest Neighbor Search over Generalized Weighted Space

ICML 2019oral

Nearest Neighbor Search (NNS) over generalized weighted space is a fundamental problem which has many applications in various fields. However, to the best of our knowledge, there is no sublinear time solution to this problem. Based on the idea of Asymmetric Locality-Sensitive Hashing (ALSH), we intr…

Cited by 24SourcePDFScholar