← Search

Piotr Koniusz

62 accepted papers

2026

Artemis: Structured Visual Reasoning for Perception Policy Learning

ICML 2026poster

Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indicate that such purely linguistic intermediate reasoning often reduces performance on perception tasks. We argue that the c…

Cited by 0SourceScholar
2026

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection

ICLR 2026poster

Open-Vocabulary Object Detection (OVOD) plays a critical role in autonomous driving and human-computer interaction by enabling perception beyond closed-set categories. However, current approaches predominantly rely on multimodal fusion, facing dual limitations: multimodal fusion methods incur heavy…

Cited by 0SourceScholar
2026

Does a Hybrid Space-Aware Randomized Defense Improve Empirical and Certified Adversarial Robustness?

ICML 2026poster

We introduce Hybrid Space-aware Stochastic Convolution Attention Noise (HySCAN), a hybrid randomized defense that helps close the long-standing gap between provable robustness under ℓ2 certificates and empirical robustness against strong ℓ∞ attacks, while maintaining strong generalization across div…

Cited by 0SourceScholar
2026

Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting

CVPR 2026

Conventional few-shot medical image segmentation (FSMIS) approaches face performance bottlenecks that hinder broader clinical applicability. Although the Segment Anything Model (SAM) exhibits strong category-agnostic segmentation capabilities, its direct application to medical images often leads to

Cited by 0SourcecodeScholar
2026

Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs

ICLR 2026poster

Diagrams represent a form of visual language that encodes abstract concepts and relationships through structured symbols and their spatial arrangements. Unlike natural images, they are inherently symbolic, and entirely artificial. They thus pose unique challenges for Multimodal Large Language Model…

Cited by 0SourceScholar
2026

Tug-of-War No More: Harmonizing Accuracy and Robustness in Vision-Language Models via Stability-Aware Task Vector Merging

ICLR 2026poster

Foundation Vision-Language Models (VLMs) excel across benchmarks yet remain vulnerable to adversarial attacks. While adversarial fine-tuning improves robustness, attaining a desirable clean–robust performance trade-off typically requires costly hyperparameter searches with multiple retraining runs.…

Cited by 0SourceScholar
2025

Amortized Active Generation of Pareto Sets

NeurIPS 2025poster

We introduce active generation of Pareto sets (A-GPS), a new framework for online discrete black-box multi-objective optimization (MOO). A-GPS learns a generative model of the Pareto set that supports a-posteriori conditioning on user preferences. The method employs a class probability estimator (CP…

Cited by 0SourceScholar
2025

BiLoRA: Almost-Orthogonal Parameter Spaces for Continual Learning

CVPR 2025poster

Continual learning requires models to learn tasks sequentially while maintaining a delicate balance between stability (retaining knowledge of previous tasks) and plasticity (adapting to new tasks). A key challenge is preventing interference between tasks - where learning new tasks degrades performan…

2025

CrossSpectra: Exploiting Cross-Layer Smoothness for Parameter-Efficient Fine-Tuning

NeurIPS 2025poster

Parameter-efficient fine-tuning (PEFT) is essential for adapting large foundation models without excessive storage cost. However, current approaches such as LoRA treat each layer’s adaptation independently, overlooking correlations across layers. This independence causes the number of trainable para…

Cited by 0SourceScholar
2025

Improving Zero-Shot Adversarial Robustness in Vision-Language Models by Closed-form Alignment of Adversarial Path Simplices

ICML 2025spotlight

Vision-Language Models (VLMs) such as CLIP excel at zero-shot classification due to large-scale pre-training but are vulnerable to adversarial examples. Adversarial fine-tuning robustifies zero-shot models by aligning prediction scores of individual adversaries with their clean counterparts, which t…

Cited by 0SourcePDFScholar
2025

Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion

ICLR 2025poster

In computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing these features is essential for robust performance, especially with the availability of powerful pre-trained models like…

Cited by 1SourcePDFScholar
2025

Machine Unlearning via Task Simplex Arithmetic

NeurIPS 2025poster

As foundation Vision-Language Models (VLMs) unlock fine-tuning on smaller datasets while leveraging large-scale pre-training data, machine unlearning becomes critical in addressing privacy concerns and regulatory compliance. Task vector, representing the difference between parameters of models fine-…

Cited by 0SourceScholar
2025

Open-World Objectness Modeling Unifies Novel Object Detection

CVPR 2025poster

The challenge in open-world object detection, similarly to few- and zero-shot learning, is to generalize beyond the class distribution of the training data. In this paper, we propose a general class-agnostic objectness measure to limit bias toward labeled samples. One issue in open-world detection…

Cited by 1SourcePDFScholar
2025

Primitive Vision: Improving Diagram Understanding in MLLMs

ICML 2025poster

Mathematical diagrams have a distinctive structure. Standard feature transforms designed for natural images (e.g., CLIP) fail to process them effectively, limiting their utility in multimodal large language models (MLLMs). Current efforts to improve MLLMs have primarily focused on scaling mathematic…

2025

Robust SuperAlignment: Weak-to-Strong Robustness Generalization for Vision-Language Models

NeurIPS 2025spotlight

Numerous well-established studies have demonstrated the superhuman capabilities of modern Vision-Language Models (VLMs) across a wide range of tasks. However, growing is the doubt about the continuing availability of reliable high-quality labeling (supervision) from human annotators, leading to stag…

Cited by 0SourceScholar
2025

Robustifying Zero-Shot Vision Language Models by Subspaces Alignment

ICCV 2025poster

Vision-Language Models (VLMs) enjoy strong zero-shot performance but are vulnerable to adversarial attacks posing security risks. Adversarially robust fine-tuning enhances zero-shot robustness on new datasets while preserving the natural performance of pre-trained VLMs. However, prior methods use sa…

Cited by 0SourcePDFScholar
2024

Adversarially Robust Few-shot Learning via Parameter Co-distillation of Similarity and Class Concept Learners

CVPR 2024poster

Few-shot learning (FSL) facilitates a variety of computer vision tasks yet remains vulnerable to adversarial attacks. Existing adversarially robust FSL methods rely on either visual similarity learning or class concept learning. Our analysis reveals that these two learning paradigms are complementar…

Cited by 3SourcePDFScholar
2024

CHAIN: Enhancing Generalization in Data-Efficient GANs via lipsCHitz continuity constrAIned Normalization

CVPR 2024poster

Generative Adversarial Networks (GANs) significantly advanced image generation but their performance heavily depends on abundant training data. In scenarios with limited data GANs often struggle with discriminator overfitting and unstable training. Batch Normalization (BN) despite being known for en…

2024

PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularization

NeurIPS 2024spotlight

Parameter-Efficient Fine-Tuning (PEFT) effectively adapts pre-trained transformers to downstream tasks. However, the optimization of tasks performance often comes at the cost of generalizability in fine-tuned models. To address this issue, we theoretically connect smaller weight gradient norms durin…

2024

Pre-training with Random Orthogonal Projection Image Modeling

ICLR 2024spotlight

Masked Image Modeling (MIM) is a powerful self-supervised strategy for visual pre-training without the use of labels. MIM applies random crops to input images, processes them with an encoder, and then recovers the masked inputs with a decoder, which encourages the network to capture and learn struct…

2024

Robust Distillation via Untargeted and Targeted Intermediate Adversarial Samples

CVPR 2024poster

Adversarially robust knowledge distillation aims to compress large-scale models into lightweight models while preserving adversarial robustness and natural performance on a given dataset. Existing methods typically align probability distributions of natural and adversarial samples between teacher an…

Cited by 5SourcePDFScholar
2023

Distilling Self-Supervised Vision Transformers for Weakly-Supervised Few-Shot Classification & Segmentation

CVPR 2023poster

We address the task of weakly-supervised few-shot image classification and segmentation, by leveraging a Vision Transformer (ViT) pretrained with self-supervision. Our proposed method takes token representations from the self-supervised ViT and leverages their correlations, via self-attention, to pr…

Cited by 39SourcePDFScholar
2023

Learning Partial Correlation Based Deep Visual Representation for Image Classification

CVPR 2023poster

Visual representation based on covariance matrix has demonstrates its efficacy for image classification by characterising the pairwise correlation of different channels in convolutional feature maps. However, pairwise correlation will become misleading once there is another channel correlating with…

2023

Learning Spatial-context-aware Global Visual Feature Representation for Instance Image Retrieval

ICCV 2023poster

In instance image retrieval, considering local spatial information within an image has proven effective to boost retrieval performance, as demonstrated by local visual descriptor based geometric verification. Nevertheless, it will be highly valuable to make ordinary global image representations spat…

Cited by 9PDFcodeScholar
2023

Mitigating the Popularity Bias of Graph Collaborative Filtering: A Dimensional Collapse Perspective

NeurIPS 2023spotlight

Graph-based Collaborative Filtering (GCF) is widely used in personalized recommendation systems. However, GCF suffers from a fundamental problem where features tend to occupy the embedding space inefficiently (by spanning only a low-dimensional subspace). Such an effect is characterized in GCF by th…

Cited by 27SourcePDFScholar
2023

Spectral Feature Augmentation for Graph Contrastive Learning and Beyond

AAAI 2023technical

Although augmentations (e.g., perturbation of graph edges, image crops) boost the efficiency of Contrastive Learning (CL), feature level augmentation is another plausible, complementary yet not well researched strategy. Thus, we present a novel spectral feature argumentation for contrastive learni…

2023

Transductive Few-Shot Learning With Prototype-Based Label Propagation by Iterative Graph Refinement

CVPR 2023poster

Few-shot learning (FSL) is popular due to its ability to adapt to novel classes. Compared with inductive few-shot learning, transductive models typically perform better as they leverage all samples of the query set. The two existing classes of methods, prototype-based and graph-based, have the disad…

2022

Kernelized Few-Shot Object Detection With Efficient Integral Aggregation

CVPR 2022poster

We design a Kernelized Few-shot Object Detector by leveraging kernelized matrices computed over multiple proposal regions, which yield expressive non-linear representations whose model complexity is learned on the fly. Our pipeline contains several modules. An Encoding Network encodes support and qu…

Cited by 79PDFcodeScholar
2022

Time-rEversed diffusioN tEnsor Transformer: A New TENET of Few-Shot Object Detection

ECCV 2022poster

"In this paper, we tackle the challenging problem of Few-shot Object Detection. Existing FSOD pipelines (i) use average-pooled representations that result in information loss; and/or (ii) discard position information that can help detect object instances. Consequently, such pipelines are sensitive t…

2021

Rethinking Class Relations: Absolute-Relative Supervised and Unsupervised Few-Shot Learning

CVPR 2021poster

The majority of existing few-shot learning methods describe image relations with binary labels. However, such binary relations are insufficient to teach the network complicated real-world relations, due to the lack of decision smoothness. Furthermore, current few-shot learning models capture only th…

Cited by 81PDFScholar
2020

Few-shot Action Recognition with Permutation-invariant Attention

ECCV 2020poster

Many few-shot learning models focus on recognising images. In contrast, we tackle a challenging task of few-shot action recognition from videos. We build on a C3D encoder for spatio-temporal video blocks to capture short-range action patterns. Such encoded blocks are aggregated by permutation-invari…

Cited by 219SourcePDFScholar
2020

On Modulating the Gradient for Meta-Learning

ECCV 2020poster

Inspired by optimization techniques, we propose a novel meta-learning algorithm with gradient modulation to encourage fast-adaptation of neural networks in the absence of abundant data. Our method, termed ModGrad, is designed to circumvent the noisy nature of the gradients which is prevalent in low-…

2019

Hallucinating IDT Descriptors and I3D Optical Flow Features for Action Recognition With CNNs

ICCV 2019poster

In this paper, we revive the use of old-fashioned handcrafted video representations for action recognition and put new life into these techniques via a CNN-based hallucination step. Despite of the use of RGB and optical flow frames, the I3D model (amongst others) thrives on combining its output with…

Cited by 119PDFScholar
2018

Museum Exhibit Identification Challenge for the Supervised Domain Adaptation and Beyond

ECCV 2018poster

We study an open problem of artwork identification and propose a new dataset dubbed Open Museum Identification Challenge (Open MIC). It contains photos of exhibits captured in 10 distinct exhibition spaces of several museums which showcase paintings, timepieces, sculptures, glassware, relics, scienc…

Cited by 54SourcePDFScholar
2017

Domain Adaptation by Mixture of Alignments of Second- or Higher-Order Scatter Tensors

CVPR 2017poster

In this paper, we propose an approach to the domain adaptation, dubbed Second- or Higher-order Transfer of Knowledge (So-HoT), based on the mixture of alignments of second- or higher-order scatter statistics between the source and target domains. The human ability to learn from few labeled samples i…

Cited by 163PDFScholar
2016

Sparse Coding for Third-Order Super-Symmetric Tensor Descriptors With Application to Texture Recognition

CVPR 2016spotlight

Super-symmetric tensors - a higher-order extension of scatter matrices - are becoming increasingly popular in machine learning and computer vision for modeling data statistics, co-occurrences, or even as visual descriptors. They were shown recently to outperform second-order approaches, however, the…

Cited by 41PDFScholar