← Search

Erkut Erdem

13 accepted papers

2026

EeveeDark: A Binary Neural Framework for Low-Light Video Enhancement via Event-Guided Sensor-Level Fusion

RA-L 2026

Enhancing videos under extreme low-light conditions remains challenging due to the difficulty of balancing restoration quality and computational efficiency in resource-constrained settings. This paper introduces EeveeDark, a low-light video enhancement framework that combines the spatial richness of

Cited by 0SourceScholar
2026

LAMP: Language-Assisted Motion Planning for Controllable Video Generation

CVPR 2026

Video generation has achieved remarkable progress in visual fidelity and controllability, enabling conditioning on text, layout, or motion. Among these, motion control -- specifying object dynamics and camera trajectories -- is essential for composing complex, cinematic scenes, yet existing interfac

Cited by 0SourcecodeScholar
2025

GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting

ICCV 2025poster

Efficient neural representations for dynamic video scenes are critical for applications ranging from video compression to interactive simulations. Yet, existing methods often face challenges related to high memory usage, lengthy training times, and temporal consistency. To address these issues, we i…

Cited by 0SourcePDFScholar
2024

CLIPAway: Harmonizing focused embeddings for removing objects via diffusion models

NeurIPS 2024poster

Advanced image editing techniques, particularly inpainting, are essential for seamlessly removing unwanted elements while preserving visual integrity. Traditional GAN-based methods have achieved notable success, but recent advancements in diffusion models have produced superior results due to their…

2024

Sequential Compositional Generalization in Multimodal Models

NAACL 2024long

The rise of large-scale multimodal models has paved the pathway for groundbreaking advances in generative modeling and reasoning, unlocking transformative applications in a variety of complex tasks. However, a pressing question that remains is their genuine capability for stronger forms of generaliz…

Cited by 1SourcePDFScholar
2024

ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

ICLR 2024poster

With the ever-increasing popularity of pretrained Video-Language Models (VidLMs), there is a pressing need to develop robust evaluation methodologies that delve deeper into their visio-linguistic capabilities. To address this challenge, we present ViLMA (Video Language Model Assessment), a task-agno…

Cited by 13SourcePDFScholar
2023

Harnessing Dataset Cartography for Improved Compositional Generalization in Transformers

EMNLP 2023long findings

Neural networks have revolutionized language modeling and excelled in various downstream tasks. However, the extent to which these models achieve compositional generalization comparable to human cognitive abilities remains a topic of debate. While existing approaches in the field have mainly focused…

Cited by 0SourcecodeScholar
2023

ST360IQ: No-Reference Omnidirectional Image Quality Assessment With Spherical Vision Transformers

ICASSP 2023accepted

Omnidirectional images, aka 360° images, can deliver immersive and interactive visual experiences. As their popularity has increased dramatically in recent years, evaluating the quality of 360° images has become a problem of interest since it provides insights for capturing, transmitting, and consum…

Cited by 0SourceScholar
2023

VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs

ICCV 2023poster

We propose VidStyleODE, a spatiotemporally continuous disentangled video representation based upon StyleGAN and Neural-ODEs. Effective traversal of the latent space learned by Generative Adversarial Networks (GANs) has been the basis for recent breakthroughs in image editing. However, the applicabil…

Cited by 2PDFScholar
2022

CRAFT: A Benchmark for Causal Reasoning About Forces and inTeractions

ACL 2022findings

Humans are able to perceive, understand and reason about causal events. Developing models with similar physical and causal understanding capabilities is a long-standing goal of artificial intelligence. As a step towards this direction, we introduce CRAFT, a new video question answering dataset that…

2020

Belief Regulated Dual Propagation Nets for Learning Action Effects on Groups of Articulated Objects

ICRA 2020poster

In recent years, graph neural networks have been successfully applied for learning the dynamics of complex and partially observable physical systems. However, their use in the robotics domain is, to date, still limited. In this paper, we introduce Belief Regulated Dual Propagation Networks (BRDPN),…

Cited by 16SourceScholar