← Search

Cenk Baykal

12 accepted papers

2023

Alternating Updates for Efficient Transformers

NeurIPS 2023spotlight

It has been well established that increasing scale in deep transformer networks leads to improved quality and performance. However, this increase in scale often comes with prohibitive increases in compute cost and inference latency. We introduce Alternating Updates (AltUp), a simple-to-implement met…

Cited by 6SourcePDFScholar
2023

SLaM: Student-Label Mixing for Distillation with Unlabeled Examples

NeurIPS 2023poster

Knowledge distillation with unlabeled examples is a powerful training paradigm for generating compact and lightweight student models in applications where the amount of labeled data is limited but one has access to a large pool of unlabeled data. In this setting, a large teacher model generates "sof…

Cited by 10SourcePDFScholar
2022

A Theoretical View on Sparsely Activated Networks

NeurIPS 2022accept

Deep and wide neural networks successfully fit very complex functions today, but dense models are starting to be prohibitively expensive for inference. To mitigate this, one promising research direction is networks that activate a sparse subgraph of the network. The subgraph is chosen by a data-depe…

Cited by 9SourcePDFScholar
2022

Weighted Distillation with Unlabeled Examples

NeurIPS 2022accept

Distillation with unlabeled examples is a popular and powerful method for training deep neural networks in settings where the amount of labeled data is limited: A large “teacher” neural network is trained on the labeled data available, and then it is used to generate labels on an unlabeled dataset (…

Cited by 14SourcePDFScholar
2020

Provable Filter Pruning for Efficient Neural Networks

ICLR 2020poster

We present a provable, sampling-based approach for generating compact Convolutional Neural Networks (CNNs) by identifying and removing redundant filters from an over-parameterized network. Our algorithm uses a small batch of input data points to assign a saliency score to each filter and constructs…

Cited by 199SourcecodeScholar
2019

Data-Dependent Coresets for Compressing Neural Networks with Applications to Generalization Bounds

ICLR 2019poster

We present an efficient coresets-based neural network compression algorithm that sparsifies the parameters of a trained fully-connected neural network in a manner that provably approximates the network's output. Our approach is based on an importance sampling scheme that judiciously defines a sampli…

Cited by 98SourcePDFScholar
2018

Kinematic Design Optimization of a Parallel Surgical Robot to Maximize Anatomical Visibility via Motion Planning

ICRA 2018poster

We introduce a method to optimize on a patient-specific basis the kinematic design of the Continuum Reconfigurable Incisionless Surgical Parallel (CRISP) robot, a needle-diameter medical robot based on a parallel structure that is capable of performing minimally invasive procedures. Our objective is…

Cited by 18SourceScholar
2018

Sampling-Based Approximation Algorithms for Reachability Analysis with Provable Guarantees

RSS 2018poster

The successful deployment of many autonomous systems in part hinges on providing rigorous guarantees on their performance and safety through a formal verification method, such as reachability analysis. In this work, we present a simple-to-implement, sampling-based algorithm for reachability analysis…

Cited by 30SourcePDFScholar
2017

Persistent surveillance of events with unknown, time-varying statistics

ICRA 2017poster

We consider the problem of monitoring stochastic, time-varying events occurring at discrete locations. Our problem formulation extends prior work in persistent surveillance by considering the objective of maximizing event detections in unknown, dynamic environments where the rates of events are time…

Cited by 11SourceScholar
2015

Optimizing design parameters for sets of concentric tube robots using sampling-based motion planning

IROS 2015poster

Concentric tube robots are tentacle-like medical robots that can bend around anatomical obstacles to access hard-to-reach clinical targets. The component tubes of these robots can be swapped prior to performing a task in order to customize the robot's behavior and reachable workspace. Optimizing a r…

Cited by 41SourceScholar