← Search

Dimitris N. Metaxas

56 accepted papers

2026

K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model

ICLR 2026poster

Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented. They are usually trained on single knowledge sources and specific to individual tasks, modalities, or organs. This fragmentation contrasts sharply with clinical practice, where experts seaml…

Cited by 0SourcecodeScholar
2026

Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization

ICRA 2026poster

Solving complex, long-horizon robotic manipulation tasks requires a deep understanding of physical interactions, reasoning about their long-term consequences, and precise high-level planning. Vision-Language Models (VLMs) offer a general perceive-reason-act framework for this goal. However, previous…

2026

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

ICLR 2026poster

While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to identify trustworthy responses, especially in complex tasks requiring multi-step reasoning. In this paper, we propose a…

Cited by 0SourcecodeScholar
2025

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction

CVPR 2025poster

Prevailing Multimodal Large Language Models (MLLMs) encode the input image(s) as vision tokens and feed them into the language backbone, similar to how Large Language Models (LLMs) process the text tokens. However, the number of vision tokens increases quadratically as the image resolutions, leading…

2025

AutoEdit: Automatic Hyperparameter Tuning for Image Editing

NeurIPS 2025poster

Recent advances in diffusion models have revolutionized text-guided image editing, yet existing editing methods face critical challenges in hyperparameter identification. To get the reasonable editing performance, these methods often require the user to brute-force tune multiple interdependent hyper…

Cited by 0SourceScholar
2025

FlowChef: Steering of Rectified Flow Models for Controlled Generations

ICCV 2025poster

Despite recent advances in Rectified Flow Models (RFMs), unlocking their full potential for controlled generation tasks--such as inverse problems and image editing--remains a significant hurdle. Although RFMs and Diffusion Models (DMs) represent state-of-the-art approaches in generative modeling, th…

2025

Improved Training Technique for Latent Consistency Models

ICLR 2025poster

Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demonstrated impressive performance, achieving results on par with diffusion models in the pixel space. However, the success…

2025

LUCAS: Layered Universal Codec Avatars

CVPR 2025poster

Photorealistic 3D head avatar reconstruction faces critical challenges in modeling dynamic face-hair interactions and achieving cross-identity generalization, particularly during expressions and head movements. We present LUCAS, a novel Universal Prior Model (UPM) for codec avatar modeling that dise…

Cited by 0SourcePDFScholar
2025

LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation

ICLR 2025poster

Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual p…

2025

MLLM-as-a-Judge for Image Safety without Human Labeling

CVPR 2025highlight

Image content safety has become a significant challenge with the rise of visual media on online platforms. Meanwhile, in the age of AI-generated content (AIGC), many image generation models are capable of producing harmful content, such as images containing sexual or violent material. Thus, it becom…

Cited by 2SourcePDFScholar
2025

Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Image Generation

AAAI 2025technical

Flow matching has emerged as a promising framework for training generative models, demonstrating impressive empirical performance while offering relative ease of training compared to diffusion-based models. However, this method still requires numerous function evaluations in the sampling process. To…

2025

Show and Segment: Universal Medical Image Segmentation via In-Context Learning

CVPR 2025poster

Medical image segmentation remains challenging due to the vast diversity of anatomical structures, imaging modalities, and segmentation tasks. While deep learning has made significant advances, current approaches struggle to generalize as they require task-specific training or fine-tuning on unseen…

Cited by 0SourcePDFScholar
2025

SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

CVPR 2025poster

We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image g…

Cited by 2SourcePDFScholar
2025

Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models

NeurIPS 2025poster

Vision–language models (VLMs) excel at zero-shot inference but often degrade under test-time domain shifts. For this reason, episodic test-time adaptation strategies have recently emerged as powerful techniques for adapting VLMs to a single unlabeled image. However, existing adaptation strategies, s…

Cited by 0SourcecodeScholar
2025

The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information Steering

ICML 2025poster

Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded contents. In this paper, we investigate the internal dynamics of hallucination by examining the tokens logits rankings througho…

2025

VISIAR: Empower MLLM for Visual Story Ideation

ACL 2025finding

Ideation, the process of forming ideas from concepts, is a big part of the content creation process. However, the noble goal of helping visual content creators by suggesting meaningful sequences of visual assets from a limited collection is challenging. It requires a nuanced understanding of visual…

2024

BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models

NeurIPS 2024poster

Large Language Models (LLMs) often suffer from overconfidence during inference, particularly when adapted to downstream domain-specific tasks with limited data. Previous work addresses this issue by employing approximate Bayesian estimation after the LLMs are trained, enabling them to quantify uncer…

2024

DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models

ICLR 2024poster

Recent text-to-image diffusion models have shown surprising performance in generating high-quality images. However, concerns have arisen regarding the unauthorized data usage during the training or fine-tuning process. One example is when a model trainer collects a set of images created by a particu…

2024

DiMSUM: Diffusion Mamba - A Scalable and Unified Spatial-Frequency Method for Image Generation

NeurIPS 2024poster

We introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in input images for image generation tasks. While state-space networks, including Mamba, a revolutionary advancement in re…

2024

Generating Enhanced Negatives for Training Language-Based Object Detectors

CVPR 2024poster

The recent progress in language-based open-vocabulary object detection can be largely attributed to finding better ways of leveraging large-scale data with free-form text annotations. Training such models with a discriminative objective function has proven successful but requires good positive and n…

2024

How to Trace Latent Generative Model Generated Images without Artificial Watermark?

ICML 2024poster

Latent generative models (e.g., Stable Diffusion) have become more and more popular, but concerns have arisen regarding potential misuse related to images generated by these models. It is, therefore, necessary to analyze the origin of images by inferring if a particular image was generated by a spec…

2024

Instantaneous Perception of Moving Objects in 3D

CVPR 2024poster

The perception of 3D motion of surrounding traffic participants is crucial for driving safety. While existing works primarily focus on general large motions we contend that the instantaneous detection and quantification of subtle motions is equally important as they indicate the nuances in driving b…

Cited by 1SourcePDFScholar
2024

Layout-Agnostic Scene Text Image Synthesis with Diffusion Models

CVPR 2024poster

While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-based methods for scene text generation are typically limited by their reliance on…

Cited by 5SourcePDFScholar
2024

Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate

NeurIPS 2024poster

Generalization remains a central challenge in machine learning. In this work, we propose *Learning from Teaching* (**LoT**), a novel regularization technique for deep neural networks to enhance generalization. Inspired by the human ability to capture concise and abstract patterns, we hypothesize tha…

2024

Rethinking Deep Unrolled Model for Accelerated MRI Reconstruction

ECCV 2024oral

"Magnetic Resonance Imaging (MRI) is a widely used imaging modality for clinical diagnostics and the planning of surgical interventions. Accelerated MRI seeks to mitigate the inherent limitation of long scanning time by reducing the amount of raw k-space data required for image reconstruction. Recen…

2024

SF-V: Single Forward Video Generation Model

NeurIPS 2024poster

Diffusion-based video generation models have demonstrated remarkable success in obtaining high-fidelity videos through the iterative denoising process. However, these models require multiple denoising steps during sampling, resulting in high computational costs. In this work, we propose a novel appr…

2024

Taming Self-Training for Open-Vocabulary Object Detection

CVPR 2024poster

Recent studies have shown promising performance in open-vocabulary object detection (OVD) by utilizing pseudo labels (PLs) from pretrained vision and language models (VLMs). However teacher-student self-training a powerful and widely used paradigm to leverage PLs is rarely explored for OVD. This wor…

2023

DeFormer: Integrating Transformers with Deformable Models for 3D Shape Abstraction from a Single Image

ICCV 2023poster

Explicit 3D shape abstraction from a single 2D image is a long-standing problem in computer vision and graphics. By leveraging a set of primitives to represent the target shape, recent methods have achieved promising results. However, these methods either use a relatively larger number of primitives…

Cited by 8PDFScholar
2023

LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction

NeurIPS 2023poster

Reconstructing the 3D articulated shape of an animal from a single in-the-wild image is a challenging task. We propose LEPARD, a learning-based framework that discovers semantically meaningful 3D parts and reconstructs 3D shapes in a part-based manner. This is advantageous as 3D parts are robust to…

Cited by 13SourcePDFScholar
2023

Learning Articulated Shape With Keypoint Pseudo-Labels From Web Images

CVPR 2023poster

This paper shows that it is possible to learn models for monocular 3D reconstruction of articulated objects (e.g. horses, cows, sheep), using as few as 50-150 images labeled with 2D keypoints. Our proposed approach involves training category-specific keypoint estimators, generating 2D keypoint pseud…

Cited by 8SourcePDFScholar
2023

Revisiting Multimodal Representation in Contrastive Learning: From Patch and Token Embeddings to Finite Discrete Tokens

CVPR 2023poster

Contrastive learning-based vision-language pre-training approaches, such as CLIP, have demonstrated great success in many vision-language tasks. These methods achieve cross-modal alignment by encoding a matched image-text pair with similar feature embeddings, which are generated by aggregating infor…

2023

SINE: SINgle Image Editing With Text-to-Image Diffusion Models

CVPR 2023poster

Recent works on diffusion models have demonstrated a strong capability for conditioning image generation, e.g., text-guided image synthesis. Such success inspires many efforts trying to use large-scale pre-trained diffusion models for tackling a challenging problem--real image editing. Works conduct…

2022

A Manifold View of Adversarial Risk

AISTATS 2022poster

The adversarial risk of a machine learning model has been widely studied. Most previous works assume that the data lies in the whole ambient space. We propose to take a new angle and take the manifold assumption into consideration. Assuming data lies in a manifold, we investigate two new types of ad…

Cited by 4SourcePDFScholar
2022

Exploiting Unlabeled Data with Vision and Language Models for Object Detection

ECCV 2022poster

"Building robust and generic object detection frameworks requires scaling to larger label spaces and bigger training datasets. However, it is prohibitively costly to acquire annotations for thousands of categories at a large scale. We propose a novel method that leverages the rich semantics availabl…

2022

Hierarchically Self-Supervised Transformer for Human Skeleton Representation Learning

ECCV 2022poster

"Despite the success of fully-supervised human skeleton sequence modeling, utilizing self-supervised pre-training for skeleton sequence representation learning has been an active field because acquiring task-specific skeleton annotations at large scales is difficult. Recent studies focus on learning…

2022

Learning Transferable Reward for Query Object Localization with Policy Adaptation

ICLR 2022poster

We propose a reinforcement learning based approach to query object localization, for which an agent is trained to localize objects of interest specified by a small exemplary set. We learn a transferable reward signal formulated using the exemplary set by ordinal metric learning. Our proposed method…

2022

Social ODE: Multi-agent Trajectory Forecasting with Neural Ordinary Differential Equations

ECCV 2022poster

"Multi-agent trajectory forecasting has recently attracted a lot of attention due to its widespread applications including autonomous driving. Most previous methods use RNNs or Transformers to model agent dynamics in the temporal dimension and social pooling or GNNs to model interactions with other…

Cited by 37SourcePDFScholar
2021

A Good Image Generator Is What You Need for High-Resolution Video Synthesis

ICLR 2021spotlight

Image and video synthesis are closely related areas aiming at generating content from noise. While rapid progress has been demonstrated in improving image-based models to handle large resolutions, high-quality renderings, and wide variations in image content, achieving comparable video generation re…

2021

CrossNorm and SelfNorm for Generalization Under Distribution Shifts

ICCV 2021poster

Traditional normalization techniques (e.g., Batch Normalization and Instance Normalization) generally and simplistically assume that training and test data follow the same distribution. As distribution shifts are inevitable in real-world applications, well-trained models with previous normalization…

Cited by 73PDFcodeScholar
2021

Dual Projection Generative Adversarial Networks for Conditional Image Generation

ICCV 2021poster

Conditional Generative Adversarial Networks (cGANs) extend the standard unconditional GAN framework to learning joint data-label distributions from samples, and have been established as powerful generative models capable of generating high-fidelity imagery. A challenge of training such a model lies…

Cited by 25PDFcodeScholar
2021

Improved Transformer for High-Resolution GANs

NeurIPS 2021poster

Attention-based models, exemplified by the Transformer, can effectively model long range dependency, but suffer from the quadratic complexity of self-attention operation, making them difficult to be adopted for high-resolution image generation based on Generative Adversarial Networks (GANs). In this…

2021

Semantic Aware Data Augmentation for Cell Nuclei Microscopical Images With Artificial Neural Networks

ICCV 2021poster

There exists many powerful architectures for object detection and semantic segmentation of both biomedical and natural images. However, a difficulty arises in the ability to create training datasets that are large and well-varied. The importance of this subject is nested in the amount of training da…

Cited by 7PDFcodeScholar
2021

Stochastic Transformer Networks With Linear Competing Units: Application To End-to-End SL Translation

ICCV 2021poster

Automating sign language translation (SLT) is a challenging real-world application. Despite its societal importance, though, research progress in the field remains rather poor. Crucially, existing methods that yield viable performance necessitate the availability of laborious to obtain gloss sequenc…

Cited by 61PDFcodeScholar
2020

Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior Knowledge

CVPR 2020poster

Cross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student). Existing approaches require paired training examples exist in both modalities. However, accessing the data from superi…

Cited by 92PDFScholar
2020

Learning Trailer Moments in Full-Length Movies with Co-Contrastive Attention

ECCV 2020poster

A movie's key moments stand out of the screenplay to grab an audience's attention and make movie browsing efficient. But a lack of annotations makes the existing approaches not applicable to movie key moment detection. To get rid of human annotations, we leverage the officially-released trailers as…

Cited by 65SourcePDFScholar
2020

MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps

CVPR 2020poster

The ability to reliably perceive the environmental states, particularly the existence of objects and their motion behavior, is crucial for autonomous driving. In this work, we propose an efficient deep model, called MotionNet, to jointly perform perception and motion prediction from 3D point clouds.…

Cited by 204PDFcodeScholar
2020

Synthetic Learning: Learn From Distributed Asynchronized Discriminator GAN Without Sharing Medical Image Data

CVPR 2020poster

In this paper, we propose a data privacy-preserving and communication efficient distributed GAN learning framework named Distributed Asynchronized Discriminator GAN (AsynDGAN). Our proposed framework aims to train a central generator learns from distributed discriminator, and use the generated synth…

Cited by 113PDFcodeScholar
2019

Distributed Inexact Newton-type Pursuit for Non-convex Sparse Learning

AISTATS 2019poster

In this paper, we present a sample distributed greedy pursuit method for non-convex sparse learning under cardinality constraint. Given the training samples uniformly randomly partitioned across multiple machines, the proposed method alternates between local inexact sparse minimization of a Newton-t…

2019

Semantic Graph Convolutional Networks for 3D Human Pose Regression

CVPR 2019poster

In this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convolution filters and shared transformation matrix for each node. To address these limitations, we propose Semantic Graph Co…

Cited by 694PDFcodeScholar
2019

Sharpen Focus: Learning With Attention Separability and Consistency

ICCV 2019poster

Recent developments in gradient-based attention modeling have seen attention maps emerge as a powerful tool for interpreting convolutional neural networks. Despite good localization for an individual class of interest, these techniques produce attention maps with substantially overlapping responses…

Cited by 41PDFScholar
2017

Dual Iterative Hard Thresholding: From Non-convex Sparse Minimization to Non-smooth Concave Maximization

ICML 2017poster

Iterative Hard Thresholding (IHT) is a class of projected gradient descent methods for optimizing sparsity-constrained minimization models, with the best known efficiency and scalability in practice. As far as we know, the existing IHT-style methods are designed for sparse minimization in primal for…

Cited by 20SourcePDFScholar
2017

Reconstruction-Based Disentanglement for Pose-Invariant Face Recognition

ICCV 2017poster

Deep neural networks (DNNs) trained on large-scale datasets have recently achieved impressive improvements in face recognition. But a persistent challenge remains to develop methods capable of handling large pose variations that are relatively under-represented in training data. This paper presents…

Cited by 189PDFScholar
2017

StackGAN: Text to Photo-Realistic Image Synthesis With Stacked Generative Adversarial Networks

ICCV 2017oral

Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image approaches can roughly reflect the meaning of the given descriptions, but they fail to contain necessary details and vi…

Cited by 2957PDFcodeScholar