← Search

Cheng Deng

98 accepted papers

2026

Backtrace Mamba: Reviving Critical Temporal Contexts via Hierarchical Memory Compression for Online Action Detection

AAAI 2026technical

Online Action Detection (OAD) requires real-time prediction of ongoing actions without access to future frames, posing challenges in balancing computational efficiency and long-term dependencies modeling.Existing methods either suffer from slow training and limited temporal receptive fields, or face

Cited by 0SourcePDFScholar
2026

Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning

ICML 2026spotlight

As a foundational task in human-centric cross-modal intelligence, motion-language retrieval aims to bridge the semantic gap between natural language and human motion, enabling intuitive motion analysis, yet existing approaches predominantly focus on aligning entire motion sequences with global textu…

Cited by 0SourceScholar
2026

Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillation

CVPR 2026

Cloud-Edge Continual Test-Time Adaptation (CTTA)--with edge devices processing real-time data and the cloud offering strong computing power--is a critical paradigm for models that adapt to dynamic data distributions in real-world scenarios. However, most existing frameworks assume architectural homo

Cited by 0SourceScholar
2026

Editing Is a Bargaining Game: Balanced Knowledge Editing in Large Language Models

AAAI 2026technical

Large Language Models (LLMs) are prone to generating incorrect or outdated information, thereby necessitating efficient and precise mechanisms for knowledge updates. Existing knowledge editing approaches, however, often encounter conflicts between two competing objectives: maintaining existing knowl

Cited by 0SourcePDFScholar
2026

Learning Attribute–Affordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance Grounding

ICML 2026poster

This paper pays attention to open-vocabulary 3D object affordance grounding (OVAG), which aims to localize affordance regions on 3D objects by leveraging interaction images or textual instructions. Most existing methods treat interaction images as sources of external affordance knowledge and align t…

Cited by 0SourceScholar
2026

Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation

CVPR 2026

The latest progress in text-to-3D generative models makes it possible to generate high-quality 3D content. Recent text-to-3D large model have achieved remarkable breakthroughs in multi-view consistency. However, their effectiveness is often affected by inherent biases, resulting in sensitivity to de

Cited by 0SourceScholar
2026

Raise One and Infer Three: Toward Reasoning- and Memory-Augmented Diffusion Policy Generalization

IJCAI 2026

Diffusion policy has shown impressive performance in robotic manipulation tasks while struggling with out-of-distribution shifts and limited demonstrations. Recent advances primarily focus on improving geometric or perceptual representations for diffusion policy. However, these approaches rely heavi

Cited by 0Scholar
2026

SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs

ICLR 2026poster

Humans can imagine and manipulate visual images mentally, a capability known as \textit{spatial visualization}. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through spatial visualization remains insufficiently evaluated…

Cited by 0SourcecodeScholar
2026

Trajectory-Stabilized Inference for Diffusion-Based Video Inpainting

ICML 2026poster

Video inpainting aims to restore missing regions while preserving spatial and temporal coherence. Diffusion-based methods achieve strong per-frame reconstruction, but their sampling implicitly generates temporally coupled latent trajectories whose long-horizon stability is not explicitly modeled, le…

Cited by 0SourceScholar
2026

VoG: Enhancing LLM Reasoning through Stepwise Verification on Knowledge Graphs

ICLR 2026poster

Large Language Models (LLMs) excel at various reasoning tasks but still encounter challenges such as hallucination and factual inconsistency in knowledge-intensive tasks, primarily due to a lack of external knowledge and factual verification. These challenges could be mitigated by leveraging knowled…

Cited by 0SourceScholar
2025

AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing

ICASSP 2025accepted

With the development of data-centric AI, the focus has shifted from model-driven approaches to improving data quality. Academic literature, as one of the crucial types, is predominantly stored in PDF formats and needs to be parsed into texts before further processing. However, parsing diverse struct…

Cited by 0SourceScholar
2025

CFD: Learning Generalized Molecular Representation via Concept-Enhanced Feedback Disentanglement

ICLR 2025poster

To accelerate biochemical research, e.g., drug and protein discovery, molecular representation learning (MRL) has attracted much attention. However, most existing methods follow the closed-set assumption that training and testing data share identical distribution, which limits their generalization a…

2025

Compress to One Point: Neural Collapse for Pre-Trained Model-Based Class-Incremental Learning

AAAI 2025technical

Class-Incremental Learning (CIL) requires an artificial intelligence system to learn different tasks without class overlaps continually. To achieve CIL, some methods introduce the Pre-Trained Model (PTM) and leverage the generalized feature representation of PTM to learn downstream incremental tasks…

2025

Dual-Space Semantic Synergy Distillation for Continual Learning of Unlabeled Streams

NeurIPS 2025poster

Continual learning from unlabeled data streams while effectively combating catastrophic forgetting poses an intractable challenge. Traditional methods predominantly rely on visual clustering techniques to generate pseudo labels, which are frequently plagued by problems such as noise and suboptimal q…

Cited by 0SourceScholar
2025

Energy vs. Noise: Towards Robust Temporal Action Localization in Open-World

AAAI 2025technical

Temporal Action Localization (TAL) aims to accurately identify the start and end times of actions in untrimmed videos and classify them according to specific labels. However, the complexity and imbalance between target actions and background in video data make this task particularly challenging. Alt…

2025

In-context Learning Demonstration Generation with Text Distillation

IJCAI 2025

In-context learning (ICL), a paradigm derived from large language models (LLMs), holds significant promise but is notably sensitive to the choice of input demonstrations. While numerous methodologies have been developed to select the optimal demonstrations from existing datasets, our work alternativ

2025

MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework

NeurIPS 2025poster

Simulating collective decision-making involves more than aggregating individual behaviors; it emerges from dynamic interactions among individuals. While large language models (LLMs) offer strong potential for social simulation, achieving quantitative alignment with real-world data remains a key chal…

Cited by 0SourcecodeScholar
2025

Meta-Learning Dynamic Center Distance: Hard Sample Mining for Learning with Noisy Labels

ICCV 2025poster

The sample selection approach is a widely adopted strategy for learning with noisy labels, where examples with lower losses are effectively treated as clean during training. However, this clean set often becomes dominated by easy examples, limiting the model's meaningful exposure to more challenging…

Cited by 0SourcePDFScholar
2025

NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens

ICLR 2025poster

Recent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of these models' long-context abilities remains a challenge due to the limitations of current benchmarks. To address this g…

2025

Percept, Memory, and Imagine: World Feature Simulating for Open-Domain Unknown Object Detection

CVPR 2025poster

To accelerate the safe deployment of object detectors, we focus on reducing the impact of both covariate and semantic shifts. And we consider a realistic yet challenging scenario, namely Open-Domain Unknown Object Detection (ODU-OD), which aims to detect unknown objects in unseen target domains with…

2025

Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance Grounding

CVPR 2025poster

This paper pays attention to Weakly Supervised Affordance Grounding (WSAG) task that aims to train model to identify affordance regions using human-object interaction images and egocentric images without the need for costly pixel-level annotations. Most existing methods usually consider the affordan…

Cited by 0SourcePDFScholar
2025

Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative Modeling

NeurIPS 2025poster

In dance performances, choreographers define the visual expression of movement, while cinematographers shape its final presentation through camera work. Consequently, the synthesis of camera movements informed by both music and dance has garnered increasing research interest. While recent advancemen…

Cited by 0SourceScholar
2025

Tackling Long-Tailed Data Challenges in Spiking Neural Networks via Heterogeneous Knowledge Distillation

IJCAI 2025

Spiking Neural Networks (SNNs), inspired by the behavior of biological neurons, have gained significant research interest for resource-constrained edge devices and neuromorphic hardware due to their use of binary spike signals for inter-unit communication with low power consumption. However, the abs

Cited by 0SourcePDFScholar
2025

Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization

ICLR 2025poster

Recently, the comprehensive understanding of human motion has been a prominent area of research due to its critical importance in many fields. However, existing methods often prioritize specific downstream tasks and roughly align text and motion features within a CLIP-like framework. This results in…

Cited by 1SourcePDFScholar
2025

VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual Grounding

ICCV 2025poster

As an important direction of embodied intelligence, 3D Visual Grounding has attracted much attention, aiming to identify 3D objects matching the given language description. Most existing methods often follow a two-stage process, i.e., first detecting proposal objects and identifying the right object…

Cited by 0SourcePDFScholar
2025

Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph Generation

ICCV 2025poster

To promote the deployment of scenario understanding in the real world, Open-Vocabulary Scene Graph Generation (OV-SGG) has attracted much attention recently, aiming to generalize beyond the limited number of relation categories labeled during training and detect those unseen relations during inferen…

2024

A Versatile Framework for Continual Test-Time Domain Adaptation: Balancing Discriminability and Generalizability

CVPR 2024poster

Continual test-time domain adaptation (CTTA) aims to adapt the source pre-trained model to a continually changing target domain without additional data acquisition or labeling costs. This issue necessitates an initial performance enhancement within the present domain without labels while concurrentl…

Cited by 3SourcePDFScholar
2024

Asymmetric Mutual Alignment for Unsupervised Zero-Shot Sketch-Based Image Retrieval

AAAI 2024technical

In recent years, many methods have been proposed to address the zero-shot sketch-based image retrieval (ZS-SBIR) task, which is a practical problem in many applications. However, in real-world scenarios, on the one hand, we can not obtain training data with the same distribution as the test data, an…

Cited by 6SourcePDFScholar
2024

DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning

ICML 2024poster

In this work, we investigate the potential of large language models (LLMs) based agents to automate data science tasks, with the goal of comprehending task requirements, then building and training the best-fit machine learning models. Despite their widespread success, existing LLM agents are hindere…

2024

Exploiting Intrinsic Multilateral Logical Rules for Weakly Supervised Natural Language Video Localization

ACL 2024long

Weakly supervised natural language video localization (WS-NLVL) aims to retrieve the moment corresponding to a language query in a video with only video-language pairs utilized during training. Despite great success, existing WS-NLVL methods seldomly consider the complex temporal relations enclosing…

2024

LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures

ACL 2024long

In response to the escalating demand for digital human representations, progress has been made in the generation of realistic human gestures from given speeches. Despite the remarkable achievements of recent research, the generation process frequently includes unintended, meaningless, or non-realist…

Cited by 3SourcePDFScholar
2024

Long-Tail Class Incremental Learning via Independent Sub-prototype Construction

CVPR 2024poster

Long-tail class incremental learning (LT-CIL) is designed to perpetually acquire novel knowledge from an imbalanced and perpetually evolving data stream while ensuring the retention of previously acquired knowledge. The existing method only re-balances data distribution and ignores exploring the pot…

Cited by 5SourcePDFScholar
2024

Modulated Phase Diffusor: Content-Oriented Feature Synthesis for Detecting Unknown Objects

ICLR 2024poster

To promote the safe deployment of object detectors, a task of unsupervised out-of-distribution object detection (OOD-OD) is recently proposed, aiming to detect unknown objects during training without reliance on any auxiliary OOD data. To alleviate the impact of lacking OOD data, for this task, one…

2024

Navigating Continual Test-time Adaptation with Symbiosis Knowledge

IJCAI 2024poster

Continual test-time domain adaptation seeks to adapt the source pre-trained model to a continually changing target domain without incurring additional data acquisition or labeling costs. Unfortunately, existing mainstream methods may result in a detrimental cycle. This is attributed to noisy pseudo-…

Cited by 0SourcePDFScholar
2024

Retrieval Across Any Domains via Large-scale Pre-trained Model

ICML 2024poster

In order to enhance the generalization ability towards unseen domains, universal cross-domain image retrieval methods require a training dataset encompassing diverse domains, which is costly to assemble. Given this constraint, we introduce a novel problem of data-free adaptive cross-domain retrieval…

Cited by 0SourcePDFScholar
2024

Robust Noisy Correspondence Learning with Equivariant Similarity Consistency

CVPR 2024poster

The surge in multi-modal data has propelled cross-modal matching to the forefront of research interest. However the challenge lies in the laborious and expensive process of curating a large and accurately matched multimodal dataset. Commonly sourced from the Internet these datasets often suffer from…

Cited by 6SourcePDFScholar
2024

Unveiling the Unknown: Unleashing the Power of Unknown to Known in Open-Set Source-Free Domain Adaptation

CVPR 2024poster

Open-Set Source-Free Domain Adaptation aims to transfer knowledge in realistic scenarios where the target domain has additional unknown classes compared to the limited-access source domain. Due to the absence of information on unknown classes existing methods mainly transfer knowledge of known class…

2023

Bootstrap Your Own Prior: Towards Distribution-Agnostic Novel Class Discovery

CVPR 2023poster

Novel Class Discovery (NCD) aims to discover unknown classes without any annotation, by exploiting the transferable knowledge already learned from a base set of known classes. Existing works hold an impractical assumption that the novel class distribution prior is uniform, yet neglect the imbalanced…

2023

Discriminating Known From Unknown Objects via Structure-Enhanced Recurrent Variational AutoEncoder

CVPR 2023poster

Discriminating known from unknown objects is an important essential ability for human beings. To simulate this ability, a task of unsupervised out-of-distribution object detection (OOD-OD) is proposed to detect the objects that are never-seen-before during model training, which is beneficial for pro…

2023

Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus

EMNLP 2023long main

Large Language Models (LLMs) have gained significant popularity for their impressive performance across diverse fields. However, LLMs are prone to hallucinate untruthful or nonsensical outputs that fail to meet user expectations in many real-world applications. Existing works for detecting hallucina…

Cited by 0SourcecodeScholar
2023

Environment-Invariant Curriculum Relation Learning for Fine-Grained Scene Graph Generation

ICCV 2023poster

The scene graph generation (SGG) task is designed to identify the predicates based on the subject-object pairs. However, existing datasets generally include two imbalance cases: one is the class imbalance from the predicted predicates and another is the context imbalance from the given subject-objec…

Cited by 13PDFcodeScholar
2023

Hierarchical Prompt Learning for Compositional Zero-Shot Recognition

IJCAI 2023poster

Compositional Zero-Shot Learning (CZSL) aims to imitate the powerful generalization ability of human beings to recognize novel compositions of known primitive concepts that correspond to a state and an object, e.g., purple apple. To fully capture the intra- and inter-class correlations between compo…

Cited by 23SourcePDFScholar
2023

Similarity Distribution Based Membership Inference Attack on Person Re-identification

AAAI 2023technical

While person Re-identification (Re-ID) has progressed rapidly due to its wide real-world applications, it also causes severe risks of leaking personal information from training data. Thus, this paper focuses on quantifying this risk by membership inference (MI) attack. Most of the existing MI attack…

2022

Attention-guided Contrastive Hashing for Long-tailed Image Retrieval

IJCAI 2022poster

Image hashing is to represent an image using a binary code for efficient storage and accurate retrieval. Recently, deep hashing methods have shown great improvements on ideally balanced datasets, however, long-tailed data is more common due to rare samples or data collection costs in the real world.…

2022

Divide and Conquer: Compositional Experts for Generalized Novel Class Discovery

CVPR 2022poster

In response to the explosively-increasing requirement of annotated data, Novel Class Discovery (NCD) has emerged as a promising alternative to automatically recognize unknown classes without any annotation. To this end, a model makes use of a base set to learn basic semantic discriminability that ca…

Cited by 51PDFcodeScholar
2022

MetricFormer: A Unified Perspective of Correlation Exploring in Similarity Learning

NeurIPS 2022accept

Similarity learning can be significantly advanced by informative relationships among different samples and features. The current methods try to excavate the multiple correlations in different aspects, but cannot integrate them into a unified framework. In this paper, we provide to consider the multi…

Cited by 9SourcePDFScholar
2022

Noise Is Also Useful: Negative Correlation-Steered Latent Contrastive Learning

CVPR 2022poster

How to effectively handle label noise has been one of the most practical but challenging tasks in Deep Neural Networks (DNNs). Recent popular methods for training DNNs with noisy labels mainly focus on directly filtering out samples with low confidence or repeatedly mining valuable information from…

Cited by 27PDFScholar
2022

Not Just Selection, but Exploration: Online Class-Incremental Continual Learning via Dual View Consistency

CVPR 2022poster

Online class-incremental continual learning aims to learn new classes continually from a never-ending and single-pass data stream, while not forgetting the learned knowledge of old classes. Existing replay-based methods have shown promising performance by storing a subset of old class data. Unfortun…

Cited by 104PDFcodeScholar
2022

RSA: Reducing Semantic Shift from Aggressive Augmentations for Self-supervised Learning

NeurIPS 2022accept

Most recent self-supervised learning methods learn visual representation by contrasting different augmented views of images. Compared with supervised learning, more aggressive augmentations have been introduced to further improve the diversity of training pairs. However, aggressive augmentations may…

2022

Single-Domain Generalized Object Detection in Urban Scene via Cyclic-Disentangled Self-Distillation

CVPR 2022poster

In this paper, we are concerned with enhancing the generalization capability of object detectors. And we consider a realistic yet challenging scenario, namely Single-Domain Generalized Object Detection (Single-DGOD), which aims to learn an object detector that performs well on many unseen target dom…

Cited by 101PDFcodeScholar
2021

Domain-Smoothing Network for Zero-Shot Sketch-Based Image Retrieval

IJCAI 2021poster

Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a novel cross-modal retrieval task, where abstract sketches are used as queries to retrieve natural images under zero-shot scenario. Most existing methods regard ZS-SBIR as a traditional classification problem and employ a cross-entropy or triplet-…

2021

Generalized and Discriminative Few-Shot Object Detection via SVD-Dictionary Enhancement

NeurIPS 2021poster

Few-shot object detection (FSOD) aims to detect new objects based on few annotated samples. To alleviate the impact of few samples, enhancing the generalization and discrimination abilities of detectors on new objects plays an important role. In this paper, we explore employing Singular Value Decomp…

2021

Graph Debiased Contrastive Learning with Joint Representation Clustering

IJCAI 2021poster

By contrasting positive-negative counterparts, graph contrastive learning has become a prominent technique for unsupervised graph representation learning. However, existing methods fail to consider the class information and will introduce false-negative samples in the random negative sampling, causi…

Cited by 179SourcePDFScholar
2021

Secure Bilevel Asynchronous Vertical Federated Learning with Backward Updating

AAAI 2021technical

Vertical federated learning (VFL) attracts increasing attention due to the emerging demands of multi-party collaborative modeling and concerns of privacy leakage. In the real VFL applications, usually only one or partial parties hold labels, which makes it challenging for all parties to collaborativ…

Cited by 90SourcePDFScholar
2021

SelfSAGCN: Self-Supervised Semantic Alignment for Graph Convolution Network

CVPR 2021poster

Graph convolution networks (GCNs) are a powerful deep learning approach and have been successfully applied to representation learning on graphs in a variety of real-world applications. Despite their success, two fundamental weaknesses of GCNs limit their ability to represent graph-structured data: p…

Cited by 42PDFcodeScholar
2020

Fewer is More: A Deep Graph Metric Learning Perspective Using Fewer Proxies

NeurIPS 2020spotlight

Deep metric learning plays a key role in various machine learning tasks. Most of the previous works have been confined to sampling from a mini-batch, which cannot precisely characterize the global geometry of the embedding space. Although researchers have developed proxy- and classification-based me…

2020

Learning Unseen Concepts via Hierarchical Decomposition and Composition

CVPR 2020poster

Composing and recognizing new concepts from known sub-concepts has been a fundamental and challenging vision task, mainly due to 1) the diversity of sub-concepts and 2) the intricate contextuality between sub-concepts and their corresponding visual features. However, most of the current methods simp…

Cited by 66PDFScholar
2020

Lifelong Zero-Shot Learning

IJCAI 2020poster

Zero-Shot Learning (ZSL) handles the problem that some testing classes never appear in training set. Existing ZSL methods are designed for learning from a fixed training set, which do not have the ability to capture and accumulate the knowledge of multiple training sets, causing them infeasible to m…

Cited by 0SourcePDFScholar
2020

Multi-Task Collaborative Network for Joint Referring Expression Comprehension and Segmentation

CVPR 2020oral

Referring expression comprehension (REC) and segmentation (RES) are two highly-related tasks, which both aim at identifying the referent according to a natural language expression. In this paper, we propose a novel Multi-task Collaborative Network (MCN) to achieve a joint learning of REC and RES for…

Cited by 348PDFcodeScholar
2020

Projection & Probability-Driven Black-Box Attack

CVPR 2020poster

Generating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer from the need for excessive queries, as it is non-trivial to find an appropriate direction to optimize in the high-dimens…

Cited by 58PDFcodeScholar
2019

Adversarial Fine-Grained Composition Learning for Unseen Attribute-Object Recognition

ICCV 2019poster

Recognizing unseen attribute-object pairs never appearing in the training data is a challenging task, since an object often refers to a specific entity while an attribute is an abstract semantic description. Besides, attributes are highly correlated to objects, i.e., an attribute tends to describe d…

Cited by 116PDFScholar
2019

Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language Query

ICCV 2019poster

Actor and action video segmentation from natural language query aims to selectively segment the actor and its action in a video based on an input textual description. Previous works mostly focus on learning simple correlation between two heterogeneous features of vision and language via dynamic conv…

Cited by 96PDFcodeScholar
2019

Balanced Self-Paced Learning for Generative Adversarial Clustering Network

CVPR 2019oral

Clustering is an important problem in various machine learning applications, but still a challenging task when dealing with complex real data. The existing clustering algorithms utilize either shallow models with insufficient capacity for capturing the non-linear nature of data, or deep models with…

Cited by 126PDFScholar
2019

DistillHash: Unsupervised Deep Hashing by Distilling Data Pairs

CVPR 2019poster

Due to storage and search efficiency, hashing has become significantly prevalent for nearest neighbor search. Particularly, deep hashing methods have greatly improved the search performance, typically under supervised scenarios. In contrast, unsupervised deep hashing models can hardly achieve satis…

Cited by 185PDFScholar
2019

Pyramidal Person Re-IDentification via Multi-Loss Dynamic Training

CVPR 2019poster

Most existing Re-IDentification (Re-ID) methods are highly dependent on precise bounding boxes that enable images to be aligned with each other. However, due to the challenging practical scenarios, current detection models often produce inaccurate bounding boxes, which inevitably degenerate the perf…

Cited by 502PDFcodeScholar
2018

Direct Shape Regression Networks for End-to-End Face Alignment

CVPR 2018poster

Face alignment has been extensively studied in computer vision community due to its fundamental role in facial analysis, but it remains an unsolved problem. The major challenges lie in the highly nonlinear relationship between face images and associated facial shapes, which is coupled by underlying…

2018

Faster Derivative-Free Stochastic Algorithm for Shared Memory Machines

ICML 2018oral

Asynchronous parallel stochastic gradient optimization has been playing a pivotal role to solve large-scale machine learning problems in big data applications. Zeroth-order (derivative-free) methods estimate the gradient only by two function evaluations, thus have been applied to solve the problems…

Cited by 28SourcePDFScholar
2018

Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval

CVPR 2018poster

Thanks to the success of deep learning, cross-modal retrieval has made significant progress recently. However, there still remains a crucial bottleneck: how to bridge the modality gap to further enhance the retrieval accuracy. In this paper, we propose a self-supervised adversarial hashing (SSAH) ap…

2018

Unsupervised Deep Generative Adversarial Hashing Network

CVPR 2018poster

Unsupervised deep hash functions have not shown satisfactory improvements against the shallow alternatives, and usually, require supervised pretraining to avoid getting stuck in bad local minima. In this paper, we propose a deep unsupervised hashing function, called HashGAN, which outperforms unsupe…

Cited by 144SourcePDFScholar
2017

Deep Clustering via Joint Convolutional Autoencoder Embedding and Relative Entropy Minimization

ICCV 2017poster

In this paper, we propose a new clustering model, called DEeP Embedded RegularIzed ClusTering (DEPICT), which efficiently maps data into a discriminative embedding subspace and precisely predicts cluster assignments. DEPICT generally consists of a multinomial logistic regression function stacked on…

Cited by 713PDFcodeScholar
2017

Learning A Structured Optimal Bipartite Graph for Co-Clustering

NeurIPS 2017poster

Co-clustering methods have been widely applied to document clustering and gene expression analysis. These methods make use of the duality between features and samples such that the co-occurring structure of sample and feature clusters can be extracted. In graph based co-clustering methods, a biparti…

Cited by 176SourcePDFScholar
2015

Coupled fisher discrimination dictionary learning for single image super-resolution

ICASSP 2015accepted

Image Super-resolution (SR) reconstruction techniques based on sparse representation have attracted ever-increasing attentions in recent years, where the choice of over-complete dictionary is of prime important for reconstruction quality. However, most of the image SR methods based on sparse represe…

Cited by 0SourceScholar