← Search

Miao Zhang

73 accepted papers

2026

Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation

ICLR 2026poster

View transformers process multi-view observations to predict actions and have shown impressive performance in robotic manipulation. Existing methods typically extract static visual representations in a view-specific manner, leading to inadequate 3D spatial reasoning ability and a lack of dynamic ada…

Cited by 0SourceScholar
2026

DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation

AAAI 2026technical

Recent advances in self-supervised learning (SSL) have shown tremendous potential for learning 3D point cloud representations without human annotations. However, SSL for 3D point clouds still faces critical challenges due to irregular geometry, shortcut-prone reconstruction, and unbalanced semantics

Cited by 0SourcePDFScholar
2026

Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration

ICML 2026poster

Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like heal…

Cited by 0SourceScholar
2026

HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering

IJCAI 2026

Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge. Most prior methods adapt by updating a largely shared parameter set. This often leads to cross-level task interference, hindering accurate adaptatio

Cited by 0Scholar
2026

Instance-wise Adaptive Scheduling via Derivative-Free Meta-Learning

ICLR 2026poster

Deep Reinforcement Learning has achieved remarkable progress in solving NP-hard scheduling problems. However, existing methods primarily focus on optimizing average performance over training instances, overlooking the core objective of solving each individual instance with high quality. While severa…

Cited by 0SourceScholar
2026

KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing

AAAI 2026technical

Knowledge Tracing (KT) aims to dynamically model a student’s mastery of knowledge concepts based on their historical learning interactions. Most current methods rely on single-point estimates, which cannot distinguish true ability from outburst or carelessness, creating ambiguity in judging mastery.

Cited by 0SourcePDFScholar
2026

LECDPR:LLM Enhancement and Concept-Document Interactive Modeling for Prerequisite Relation Prediction

IJCAI 2026

Accurate prediction of prerequisite relations among concepts is important for course planning and intelligent tutoring systems. Existing text-based methods are frequently contaminated by noise such as redundant phrasing, ambiguous sentences, and domain-specific colloquialisms. Moreover, previous met

Cited by 0Scholar
2026

MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering

AAAI 2026technical

Visual Question Answering (VQA) requires models to reason over multimodal information, combining visual and textual data. With the development of continual learning, significant progress has been made in retaining knowledge and adapting to new information in the VQA domain. However, current methods

Cited by 0SourcePDFScholar
2026

MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment

AAAI 2026technical

Multi-modal entity alignment aims to identify equivalent entities between two multi-modal Knowledge graphs by integrating multi-modal data, such as images and text, to enrich the semantic representations of entities. However, existing methods may overlook the structural contextual information within

Cited by 0SourcePDFScholar
2026

SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting

ICLR 2026poster

Continual Learning (CL) requires a model to learn multiple tasks in sequence while maintaining both stability—preserving knowledge from previously learned tasks, and plasticity—effectively learning new tasks. Orthogonal projection has emerged as an effective and popular paradigm in CL, where it part…

Cited by 0SourcecodeScholar
2026

StructMamPose: From Sequential Perception to Structural Reasoning for 3D Human Pose Estimation

ICML 2026poster

Accurately modeling complex temporal and topological dependencies and depth information is critical for monocular 3D human pose estimation, yet existing Mamba-based approaches struggle to fulfill these demands, suffering from internal state update confusion induced by forced sequence flattening and …

Cited by 0SourceScholar
2026

Sustainable Intelligence for the Wild: Democratizing Ecological Monitoring via Knowledge-Adaptive Edge Expert Agents

IJCAI 2026

Rapid biodiversity loss underscore the urgency of effective monitoring, yet manual surveys remain resource-intensive. While on-device AI offers a scalable alternative, its performance in the wild is often challenged by environmental variability. Current methods rely heavily on cloud resource, which

Cited by 0Scholar
2026

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) inherit the auto-regressive generation paradigm and cache the keys and values (KV) of all previous tokens to accelerate inference, resulting in memory consumption that scales linearly with context length. This issue is particularly pronounced in VLMs due to substantial …

Cited by 0SourceScholar
2026

TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models

CVPR 2026

Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for their safe deployment to prevent the creation of harmful content.This has fostered a dynamic interplay between the development of erasure defenses and the adversarial probes desi

Cited by 0SourcecodeScholar
2026

Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds

CVPR 2026

Recent advances in self-supervised learning (SSL) for point clouds have substantially improved 3D scene understanding without human annotations. Existing approaches emphasize semantic awareness by enforcing feature consistency across augmented views or by masked scene modeling. However, the resultin

Cited by 0SourceScholar
2025

APKGC: Noise-enhanced Multi-Modal Knowledge Graph Completion with Attention Penalty

AAAI 2025technical

Multimodal knowledge graphs (MMKG) store structured world knowledge enriched with multimodal descriptive information. However, MMKG often faces the challenge of incompleteness. The primary objective of multimodal knowledge graph completion (MMKGC) is to predict missing entities within MMKG. Current…

2025

AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration

ACL 2025long

Multi-agent systems (MAS) based on large language models (LLMs) have demonstrated significant potential in collaborative problem-solving. However, they still face substantial challenges of low communication efficiency and suboptimal task performance, making the careful design of the agents’ communic…

2025

AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration

EMNLP 2025

Proper initialization is crucial for any system, particularly in multi-agent systems (MAS), where it plays a pivotal role in determining both the system’s efficiency and effectiveness. However, existing MAS initialization methods do not fully account for the collaborative needs of the generated agen

2025

ArchiSet: Benchmarking Editable and Consistent Single-View 3D Reconstruction of Buildings with Specific Window-to-Wall Ratios

ICCV 2025poster

Image-based 3D Genetation has made significant progress in typical scenarios, achieving high fidelity in capturing intricate textures. However, in the Architecture, Engineering, and Construction (AEC) design stages, existing technologies still face considerable challenges, particularly in handling s…

Cited by 0SourcePDFScholar
2025

CARD: Cross-modal Agent Framework for Generative and Editable Residential Design

EMNLP 2025

In recent years, architectural design automation has made significant progress, but the complexity of open-world environments continues to make residential design a challenging task, often requiring experienced architects to perform multiple iterations and human-computer interactions. Therefore, ass

Cited by 0SourcePDFScholar
2025

Class-Aware PillarMix: Can Mixed Sample Data Augmentation Enhance 3D Object Detection with Radar Point Clouds?

IROS 2025

Due to the significant effort required for data collection and annotation in 3D perception tasks, mixed sample data augmentation (MSDA) has been widely studied to generate diverse training samples by mixing existing data. Among these methods, MixUp is a prominent approach that generates new samples

Cited by 1SourceScholar
2025

Common Sense Bias Modeling for Classification Tasks

AAAI 2025technical

Machine learning model bias can arise from dataset composition: correlated sensitive features can distort the downstream classification model's decision boundary and lead to performance differences along these features. Existing de-biasing works tackle the most prominent bias features, such as color…

Cited by 0SourcePDFScholar
2025

Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation

CVPR 2025poster

Dataset distillation (DD) excels in synthesizing a small number of images per class (IPC) but struggles to maintain its effectiveness in high-IPC settings. Recent works on dataset distillation demonstrate that combining distilled and real data can mitigate the effectiveness decay. However, our analy…

2025

DGCPL: Dual Graph Distillation for Concept Prerequisite Relation Learning

IJCAI 2025

Concept prerequisite relations determine the learning order of knowledge concepts in one domain, which has an important impact on teachers' course design and students' personalized learning. Current research usually predicts concept prerequisite relations from the perspective of knowledge, and rarel

2025

DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecture

CVPR 2025poster

Diffusion models (DMs) have demonstrated exceptional generative capabilities across various domains, including image, video, and so on. A key factor contributing to their effectiveness is the high quantity and quality of data used during training. However, mainstream DMs now consume increasingly lar…

2025

DefMamba: Deformable Visual State Space Model

CVPR 2025poster

Recently, state space models (SSM), particularly Mamba, have attracted significant attention from scholars due to their ability to effectively balance computational efficiency and performance. However, most existing visual Mamba methods flatten images into 1D sequences using predefined scan orders,…

Cited by 1SourcePDFScholar
2025

Enhancing Diversity for Data-free Quantization

CVPR 2025poster

Model quantization is an effective way to compress deep neural networks and accelerate the inference time on edge devices. Existing quantization methods usually require original data for calibration during the compressing process, which may be inaccessible due to privacy issues. A common way is to g…

Cited by 1SourcePDFScholar
2025

Enhancing GUI Agent with Uncertainty-Aware Self-Trained Evaluator

NeurIPS 2025poster

Benefiting from the availability of extensive navigation trajectories, both manually and automatically annotated, current graphical user interface (GUI) agents have achieved remarkable advancements in performance. However, these annotated datasets often contain substantial noise, which impedes effec…

Cited by 0SourceScholar
2025

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

ICCV 2025poster

The incorporation of high-resolution visual input equips multimodal large language models (MLLMs) with enhanced visual perception capabilities for real-world tasks. However, most existing high-resolution MLLMs rely on a cropping-based approach to process images, which leads to fragmented visual enco…

2025

FloorPlan-LLaMa: Aligning Architects’ Feedback and Domain Knowledge in Architectural Floor Plan Generation

ACL 2025long

Floor plans serve as a graphical language through which architects sketch and communicate their design ideas. Actually, in the Architecture, Engineering, and Construction (AEC) design stages, generating floor plans is a complex task requiring domain expertise and alignment with user requirements. Ho…

Cited by 0SourcePDFScholar
2025

FuncGenFoil: Airfoil Generation and Editing Model in Function Space

NeurIPS 2025poster

Aircraft manufacturing is the jewel in the crown of industry, in which generating high-fidelity airfoil geometries with controllable and editable representations remains a fundamental challenge. Existing deep learning methods, which typically rely on predefined parametric representations (e.g., Bézi…

Cited by 0SourcecodeScholar
2025

Image Stitching in Adverse Condition: A Bidirectional-Consistency Learning Framework and Benchmark

NeurIPS 2025poster

Deep learning-based image stitching methods have achieved promising performance on conventional stitching datasets. However, real-world scenarios may introduce challenges such as complex weather conditions, illumination variations, and dynamic scene motion, which severely degrade image quality and l…

Cited by 0SourceScholar
2025

Learning Concept Prerequisite Relation via Global Knowledge Relation Optimization

AAAI 2025technical

Learning concept prerequisite relations helps better master and build a logically coherent knowledge structure. Many studies use graph neural networks to create heterogeneous knowledge networks that enhance concept representations. However, different types of relations in these networks can influenc…

2025

PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models

ACL 2025long

Large Language Models (LLMs) suffer severe performance degradation when facing extremely low-bit (sub 2-bit) quantization. Several existing sub 2-bit post-training quantization (PTQ) methods utilize a mix-precision scheme by leveraging an unstructured fine-grained mask to explicitly distinguish sali…

2025

Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning

NeurIPS 2025poster

Continual Learning (CL) aims to enable models to continuously acquire new knowledge from a sequence of tasks with avoiding the forgetting of learned information. However, existing CL methods only rely on the parameters of the most recent task for inference, which makes them susceptible to catastroph…

Cited by 0SourcecodeScholar
2025

Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression

NeurIPS 2025poster

State-space models (SSMs), particularly Mamba, emerge as an efficient Transformer alternative with linear complexity for long-sequence modeling. Recent empirical works demonstrate Mamba's in-context learning (ICL) capabilities competitive with Transformers, a critical capacity for large foundation m…

Cited by 0SourceScholar
2025

Understanding the Forgetting of (Replay-based) Continual Learning via Feature Learning: Angle Matters

ICML 2025poster

Continual learning (CL) is crucial for advancing human-level intelligence, but its theoretical understanding, especially regarding factors influencing forgetting, is still relatively limited. This work aims to build a unified theoretical framework for understanding CL using feature learning theory.…

Cited by 0SourcePDFScholar
2025

Unveiling the Power of Multiple Gossip Steps: A Stability-Based Generalization Analysis in Decentralized Training

NeurIPS 2025spotlight

Decentralized training removes the centralized server, making it a communication-efficient approach that can significantly improve training efficiency, but it often suffers from degraded performance compared to centralized training. Multi-Gossip Steps (MGS) serve as a simple yet effective bridge bet…

Cited by 0SourceScholar
2025

Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language Models

EMNLP 2025

Activation sparsity provides a dynamic, input-dependent alternative to weight pruning for accelerating inference in large language models (LLMs), effectively reducing unnecessary computations and memory accesses during the forward pass. Despite its promise, existing activation sparsification methods

2024

A Novel Surgical Robotic System for Cochlear Implantation via the Tympanic Antrum Approach

RA-L 2024

With the advancement of surgical robots, researches on the cochlear implantation (CI) using robotic systems has become increasingly robust. However, due to the narrowness of the human facial recess, even the precision of the most advanced robotic systems for CI cannot completely eliminate risks. The

Cited by 2SourceScholar
2024

A Retinex Structure-based Low-light Enhancement Model Guided by Spatial Consistency

ICRA 2024poster

Images captured by robotics under low-light conditions are often plagued by several challenges, including diminished contrast, increased noise, loss of fine details, and unnatural color reproduction. These factors can significantly hinder the performance of computer vision tasks such as object detec…

Cited by 11SourceScholar
2024

AFBench: A Large-scale Benchmark for Airfoil Design

NeurIPS 2024poster

Data-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale benchmarks in this field. It is mainly the case for airfoil inverse…

2024

Domain-Aware k-Nearest-Neighbor Knowledge Distillation for Machine Translation

ACL 2024findings

kNN-MT has utilized neighborhood knowledge for auxiliary decoding, significantly improving translation performance. Subsequently, kNN-KD transitions the use of neighborhood knowledge from the decoding phase to the training phase, to address the temporal and spatial inefficiencies inherent in kNN-MT.…

2024

FLHetBench: Benchmarking Device and State Heterogeneity in Federated Learning

CVPR 2024poster

Federated learning (FL) is a powerful technology that enables collaborative training of machine learning models without sharing private data among clients. The fundamental challenge in FL lies in learning over extremely heterogeneous data distributions device capacities and device state availabiliti…

Cited by 6SourcePDFScholar
2024

Gaussian Adaptive Strategy Based Multi-Objective Evolutionary Optimization for Path Planning on Uneven Terrains

RA-L 2024

To enable mobile robots to safely and effectively accomplish path planning tasks on uneven terrains, this letter proposes a Gaussian adaptive strategy-based multi-objective evolutionary optimization. Firstly, path solutions are generated by employing spline interpolation on the digital elevation mod

Cited by 10SourceScholar
2024

LRQuant: Learnable and Robust Post-Training Quantization for Large Language Models

ACL 2024long

Post-training quantization (PTQ) for large language models (LLMs) significantly accelerates model inference and relieves memory constraints, without incurring model training. A “smoothing paradigm” is commonly used in LLM quantization, which transfers the quantization difficulty of activation to wei…

2024

Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization

NeurIPS 2024poster

Transformers have demonstrated great power in the recent development of large foundational models. In particular, the Vision Transformer (ViT) has brought revolutionary changes to the field of vision, achieving significant accomplishments on the experimental side. However, their theoretical capabili…

Cited by 5SourcePDFScholar
2024

Voltage Regulation in Polymer Electrolyte Fuel Cell Systems Using Gaussian Process Model Predictive Control

IROS 2024poster

This study presents a novel approach using Gaussian process model predictive control (MPC) to stabilize the output voltage of a polymer electrolyte fuel cell (PEFC) by regulating hydrogen and airflow rates. Two Gaussian process models capture PEFC dynamics, accounting for constraints like hydrogen p…

Cited by 3SourceScholar
2023

GNNEvaluator: Evaluating GNN Performance On Unseen Graphs Without Labels

NeurIPS 2023poster

Evaluating the performance of graph neural networks (GNNs) is an essential task for practical GNN model deployment and serving, as deployed GNNs face significant performance uncertainty when inferring on unseen and unlabeled test graphs, due to mismatched training-test graph distributions. In this p…

Cited by 15SourcePDFScholar
2023

Structure-free Graph Condensation: From Large-scale Graphs to Condensed Graph-free Data

NeurIPS 2023spotlight

Graph condensation, which reduces the size of a large-scale graph by synthesizing a small-scale condensed graph as its substitution, has immediate benefits for various graph learning tasks. However, existing graph condensation methods rely on the joint optimization of nodes and structures in the con…

2022

Adaptive Co-Teaching for Unsupervised Monocular Depth Estimation

ECCV 2022poster

"Unsupervised depth estimation using photometric losses suffers from local minimum and training instability. We address this issue by proposing an adaptive co-teaching framework to distill the learned knowledge from unsupervised teacher networks to a student network. We design an ensemble architectu…

2022

BaLeNAS: Differentiable Architecture Search via the Bayesian Learning Rule

CVPR 2022poster

Differentiable Architecture Search (DARTS) has received massive attention in recent years, mainly because it significantly reduces the computational cost through weight sharing and continuous relaxation. However, more recent works find that existing differentiable NAS techniques struggle to outperfo…

Cited by 24PDFScholar
2022

DRLK: Dynamic Hierarchical Reasoning with Language Model and Knowledge Graph for Question Answering

EMNLP 2022main

In recent years, Graph Neural Network (GNN) approaches with enhanced knowledge graphs (KG) perform well in question answering (QA) tasks. One critical challenge is how to effectively utilize interactions between the QA context and KG. However, existing work only adopts the identical QA context repre…

Cited by 15SourcePDFScholar
2022

Interpreting Operation Selection in Differentiable Architecture Search: A Perspective from Influence-Directed Explanations

NeurIPS 2022accept

The Differentiable ARchiTecture Search (DARTS) has dominated the neural architecture search community due to its search efficiency and simplicity. DARTS leverages continuous relaxation to convert the intractable operation selection problem into a continuous magnitude optimization problem which can b…

Cited by 3SourcePDFScholar
2022

Robust Speaker Verification with Joint Self-Supervised and Supervised Learning

ICASSP 2022accepted

Supervised learning and self-supervised learning address different facets. Supervised learning achieves high accuracy, but it requires numerous expensive labeled data indeed. Correspondingly, self-supervised learning, makes use of abundant unlabeled data to learn, but the performance lags behind tha…

Cited by 0SourceScholar
2022

Semi-Supervised Video Salient Object Detection Based on Uncertainty-Guided Pseudo Labels

NeurIPS 2022accept

Semi-Supervised Video Salient Object Detection (SS-VSOD) is challenging because of the lack of temporal information in video sequences caused by sparse annotations. Most works address this problem by generating pseudo labels for unlabeled data. However, error-prone pseudo labels negatively affect th…

Cited by 13SourcePDFScholar
2022

Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective

ICLR 2022poster

Graph convolutional networks (GCNs) and their variants have achieved great success in dealing with graph-structured data. Nevertheless, it is well known that deep GCNs suffer from the over-smoothing problem, where node representations tend to be indistinguishable as more layers are stacked up. The t…

Cited by 33SourcePDFScholar
2022

Weighted Mutual Learning with Diversity-Driven Model Compression

NeurIPS 2022accept

Online distillation attracts attention from the community as it simplifies the traditional two-stage knowledge distillation process into a single stage. Online distillation collaboratively trains a group of peer models, which are treated as students, and all students gain extra knowledge from each o…

Cited by 10SourcePDFScholar
2021

Dynamic Context-Sensitive Filtering Network for Video Salient Object Detection

ICCV 2021poster

The ability to capture inter-frame dynamics has been critical to the development of video salient object detection (VSOD). While many works have achieved great success in this field, a deeper insight into its dynamic nature should be developed. In this work, we aim to answer the following questions:…

Cited by 128PDFcodeScholar
2021

Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection

NeurIPS 2021poster

Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when on…

2021

MFNet: Multi-Filter Directive Network for Weakly Supervised Salient Object Detection

ICCV 2021poster

Weakly supervised salient object detection (WSOD) targets to train a CNNs-based saliency network using only low-cost annotations. Existing WSOD methods take various techniques to pursue single "high-quality" pseudo label from low-cost annotations and then develop their saliency networks. Though thes…

Cited by 82PDFcodeScholar
2021

iDARTS: Differentiable Architecture Search with Stochastic Implicit Gradients

ICML 2021spotlight

Differentiable ARchiTecture Search(DARTS) has recently become the mainstream in the neural architecture search (NAS) due to its efficiency and simplicity. With a gradient-based bi-level optimization, DARTS alternately optimizes the inner model weights and the outer architecture parameter in a weight…

2020

A2dele: Adaptive and Attentive Depth Distiller for Efficient RGB-D Salient Object Detection

CVPR 2020poster

Existing state-of-the-art RGB-D salient object detection methods explore RGB-D data relying on a two-stream architecture, in which an independent subnetwork is required to process depth data. This inevitably incurs extra computational costs and memory consumption, and using depth data during testing…

Cited by 281PDFcodeScholar
2020

Accurate RGB-D Salient Object Detection via Collaborative Learning

ECCV 2020poster

Benefiting from the spatial cues embedded in depth images, recent progress on RGB-D saliency detection shows impressive ability on some challenge scenarios. However, there are still two limitations. One hand is that the pooling and upsampling operations in FCNs might cause blur object boundaries. On…

2020

Asymmetric Two-Stream Architecture for Accurate RGB-D Saliency Detection

ECCV 2020poster

Most existing RGB-D saliency detection methods adopt symmetric two-stream architectures for learning discriminative RGB and depth representations. In fact, there is another level of ambiguity that is often overlooked: if RGB and depth data are necessary to fit into the same network. In this paper, w…

2020

Differentiable Neural Architecture Search in Equivalent Space with Exploration Enhancement

NeurIPS 2020poster

Recent works on One-Shot Neural Architecture Search (NAS) mostly adopt a bilevel optimization scheme to alternatively optimize the supernet weights and architecture parameters after relaxing the discrete search space into a differentiable space. However, the non-negligible incongruence in their rela…

Cited by 42SourcePDFScholar
2020

One-Shot Neural Architecture Search via Novelty Driven Sampling

IJCAI 2020poster

One-Shot Neural architecture search (NAS) has received wide attentions due to its computational efficiency. Most state-of-the-art One-Shot NAS methods use the validation accuracy based on inheriting weights from the supernet as the stepping stone to search for the best performing architecture, adopt…

2020

Overcoming Multi-Model Forgetting in One-Shot NAS With Diversity Maximization

CVPR 2020poster

One-Shot Neural Architecture Search (NAS) significantly improves the computational efficiency through weight sharing. However, this approach also introduces multi-model forgetting during the supernet training (architecture search phase), where the performance of previous architectures degrade when s…

Cited by 103PDFcodeScholar
2020

Select, Supplement and Focus for RGB-D Saliency Detection

CVPR 2020poster

Depth data containing a preponderance of discriminative power in location have been proven beneficial for accurate saliency prediction. However, RGB-D saliency detection methods are also negatively influenced by randomly distributed erroneous or missing regions on the depth map or along the object b…

Cited by 266PDFcodeScholar
2019

Depth-Induced Multi-Scale Recurrent Attention Network for Saliency Detection

ICCV 2019poster

In this work, we propose a novel depth-induced multi-scale recurrent attention network for saliency detection. It achieves dramatic performance especially in complex scenarios. There are three main contributions of our network that are experimentally demonstrated to have significant practical merits…

Cited by 526PDFScholar
2019

Memory-oriented Decoder for Light Field Salient Object Detection

NeurIPS 2019poster

Light field data have been demonstrated in favor of many tasks in computer vision, but existing works about light field saliency detection still rely on hand-crafted features. In this paper, we present a deep-learning-based method where a novel memory-oriented decoder is tailored for light field sal…

2018

Human and Machine Speaker Recognition Based on Short Trivial Events

ICASSP 2018accepted

Human speech often has events that we will call trivial events, e.g., cough, laugh and sniff. Compared to regular speech, these trivial events are usually short and variable, thus generally regarded as not speaker discriminative and so are largely ignored by present speaker recognition research. How…

Cited by 0SourceScholar