CVPR 2026 Accepted Papers
The full list of 4,068 papers accepted at CVPR 2026 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
accepted: 4,068
- $L^{2}DGS$: Low-Light Dynamic Gaussian Splattingaccepted
- $\alpha$Matte4K & $\mu$Matting: Dataset and Model for Ultra-Micro Precision Alpha Video Mattingaccepted
- $\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimizationaccepted
- $\phi$-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Modelsaccepted
- 2-Shots in the Dark: Low-Light Denoising with Minimal Data Acquisitionaccepted
- 240FPS Stereo Vision from Monocular Mixed Spikesaccepted
- 2D-LFM: Lifting Foundation Model without 3D Supervisionaccepted
- 2ndMatch: Finetuning Pruned Diffusion Models via Second-Order Jacobian Matchingaccepted
- 3D Gaussian Splatting at Arbitrary Resolutions with Compact Proxy Anchorsaccepted
- 3D Gaussian Splatting from Unposed Spike Streamaccepted
- 3D Gaussian Splatting with Self-Constrained Priors for High Fidelity Surface Reconstructionaccepted
- 3D Space as a Scratchpad for Editable Text-to-Image Generationaccepted
- 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Cloudsaccepted
- 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generationaccepted
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understandingaccepted
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Imageaccepted
- 3D-IDE: 3D Implicit Depth Emergentaccepted
- 3D-LATTE: Latent Space 3D Editing from Textual Instructionsaccepted
- 3D-Object Perception Transformer (3PT)accepted
- 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decodingaccepted
- 3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objectsaccepted
- 3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experienceaccepted
- 3M-TI: High-Quality Mobile Thermal Imaging via Calibration-free Multi-Camera Cross-Modal Diffusionaccepted
- 4C4D: 4 Camera 4D Gaussian Splattingaccepted
- 4D Local Modeling Toward Dynamic Global Perception for Ambiguity-free Rotation-Invariant Point Cloud Analysisaccepted
- 4D Primitive-Mache: Glueing Primitives for Persistent 4D Scene Reconstructionaccepted
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillationaccepted
- 4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Videoaccepted
- 4DP-QA: Scalable QA for 4D Perception in Vision Language Modelsaccepted
- 4DSurf: High-Fidelity Dynamic Scene Surface Reconstructionaccepted
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Modelsaccepted
- A Bit is All You Need! Efficient Video Capture via Single Bit Imagingaccepted
- A Causal Marriage between VLM and IRM from Understanding to Reasoningaccepted
- A Closed-Form Solution for Debiasing Vision-Language Models with Utility Guarantees Across Modalities and Tasksaccepted
- A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helpsaccepted
- A Combination of Noise and Bilateral Filters Achieve Supralinear and Scalable Adversarial Robustness in CNNsaccepted
- A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimationaccepted
- A Debiased Reconstruction-based Framework for Training-Free Detection of AI-Generated Imagesaccepted
- A Difference-in-Difference Approach to Detecting AI-Generated Imagesaccepted
- A Faster Path to Continual Learningaccepted
- A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokensaccepted
- A Geometric Algebra-Informed 3DGS Framework for Wireless Channel Predictionaccepted
- A Mixed Diet Makes DINO An Omnivorous Vision Encoderaccepted
- A More Word-like Image Tokenization for MLLMsaccepted
- A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoningaccepted
- A Polynomial Chaos Framework for Causal Discovery in Nonlinear Uncertain Systemsaccepted
- A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Modelsaccepted
- A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real Worldaccepted
- A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generationaccepted
- A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detectionaccepted
- A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Rankingaccepted
- A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Spaceaccepted
- A Supervised Multi-task Framework for Joint cryo-ET Restoration Enabled by Generative Physical Simulationaccepted
- A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image Generationaccepted
- A Training-Free Style-Personalization via SVD-Based Feature Decompositionaccepted
- A Unified Framework for Knowledge Transfer in Bidirectional Model Scalingaccepted
- A Unified Perspective on Adversarial Membership Manipulation in Vision Modelsaccepted
- A2GC: Asymmetric Aggregation with Geometric Constraints for Locally Aggregated Descriptorsaccepted
- A3: Towards Advertising Aesthetic Assessmentaccepted
- ACE-Merging: Data-Free Model Merging with Adaptive Covariance Estimationaccepted
- ACPV-Net: All-Class Polygonal Vectorization for Seamless Vector Map Generation from Aerial Imageryaccepted
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Modelsaccepted
- AD-GBC: Anisotropic Granular-Ball Skip-Connection Refiner for UNet-Based Medical Image Segmentationaccepted
- ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoningaccepted
- AE2VID: Event-based Video Reconstruction via Aperture Modulationaccepted
- AERGS-SLAM: Auto-Exposure-Robust Stereo 3D Gaussian Splatting SLAMaccepted
- AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Modelsaccepted
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructionsaccepted
- AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Modelsaccepted
- AGiLe: Learning Robust Long-Horizon Manipulation via Affordance-Grounded Bidirectional Latent Planningaccepted
- AHS: Adaptive Head Synthesis via Synthetic Data Augmentationsaccepted
- AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimationaccepted
- AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detectionaccepted
- ALLNet: Multi-task Dense Prediction for Degraded Imagesaccepted
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backendaccepted
- AMap: Distilling Future Priors for Ahead-Aware Online HD Map Constructionaccepted
- AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understandingaccepted
- ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoningaccepted
- APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigationaccepted
- APPO: Attention-guided Perception Policy Optimization for Video Reasoningaccepted
- AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videosaccepted
- ARC Is a Vision Problem!accepted
- AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidanceaccepted
- ARES: Unifying Asymmetric RGB-Event Stereo for Probabilistic Scene Flow Estimationaccepted
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavioraccepted
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoningaccepted
- ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generationaccepted
- ART: Articulated Reconstruction Transformeraccepted
- AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Modelsaccepted
- AToken: A Unified Tokenizer for Visionaccepted
- AURA: Multi-modal Shared Autonomy for Urban Navigationaccepted
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsaccepted
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Modelsaccepted
- AVA-VLA: Improving Vision-Language-Action models with Active Visual Attentionaccepted
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over Videoaccepted
- AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMsaccepted
- AVGGT: Rethinking Global Attention for Accelerating VGGTaccepted
- AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Networkaccepted
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Modelsaccepted
- Abstract 3D Perception for Spatial Intelligence in Vision-Language Modelsaccepted
- AcTTA: Rethinking Test-Time Adaptation via Dynamic Activationaccepted
- Accelerating Autoregressive Video Diffusion via History-Guided Cache and Residual Correctionaccepted
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspectiveaccepted
- Accelerating Diffusion via Hybrid Data-Pipeline Parallelism Based on Conditional Guidance Schedulingaccepted
- Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestepaccepted
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compressionaccepted
- AceTone: Bridging Words and Colors for Conditional Image Gradingaccepted
- Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoningaccepted
- Act2See: Emergent Active Visual Perception for Video Reasoningaccepted
- ActAvatar: Temporally-Aware Precise Action Control for Talking Avatarsaccepted
- Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movementsaccepted
- Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulationaccepted
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic Manipulationaccepted
- ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusionaccepted
- Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Modelsaccepted
- Active Inference for Micro-Gesture Recognition: EFE-Guided Temporal Sampling and Adaptive Learningaccepted
- Active Intelligence in Video Avatars via Closed-loop World Modelingaccepted
- Active Perceptual Inference: A Corticothalamic-Inspired Dynamic Nested Recurrent Network for Multimodal Sentiment Analysis with Incomplete Dataaccepted
- ActiveAD: Planning-Oriented Active Learning for End-to-End Autonomous Drivingaccepted
- ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Modelaccepted
- ActivePolicy: Active Gaussian Reconstruction and Optimization Strategy Based on Global-Local Information Gainaccepted
- ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulationaccepted
- ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videosaccepted
- AdaBet: Gradient-free Layer Selection for Efficient Training of Deep Neural Networksaccepted
- AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generationaccepted
- AdaDexTrack: Dynamic Modulation for Adaptive and Generalizable Dexterous Manipulation Trackingaccepted
- AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLMaccepted
- AdaPrior: Bayesian-Inspired Adaptive Prior Correction for Long-Tailed Continual Learningaccepted
- AdaRadar: Rate Adaptive Spectral Compression for Radar-based Perceptionaccepted
- AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor Environmentsaccepted
- AdaSVD: Singular Value Decomposition with Adaptive Mechanisms for Large Multimodal Modelsaccepted
- AdaSpark: Adaptive Sparsity for Efficient Long-Video Understandingaccepted
- AdaSpot: Spend Resolution Where It Matters for Precise Event Spottingaccepted
- AdapAction: Adaptive Target Action Backdoor Attack against GUI Agentsaccepted
- AdapTok: Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Spaceaccepted
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisitionaccepted
- Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generationaccepted
- Adapting In-context Generation for Enhanced Composed Image Retrievalaccepted
- Adapting Lightweight Image-based Counting Models for Video Crowd Countingaccepted
- Adapting Point Cloud Analysis via Multimodal Bayesian Distribution Learningaccepted
- Adapting a Pre-trained Single-Cell Foundation Model to Spatial Gene Expression Generation from Histology Imagesaccepted
- Adaptive 3D Perception for Small Aerial Targets Under Sparse Sampling via Reinforcement Learningaccepted
- Adaptive Action Chunking at Inference-time for Vision-Language-Action Modelsaccepted
- Adaptive Anisotropic Gaussian Splatting for Multi-contrast MRI Arbitrary-Scale Super-Resolution with Anatomy Guidanceaccepted
- Adaptive Auxiliary Prompt Blending for Target-Faithful Diffusion Generationaccepted
- Adaptive Bayesian Early-Exit Networks for Efficient Non-Transferable Learningaccepted
- Adaptive Capacity Autoregressive Visual Trackingaccepted
- Adaptive Confidence Regularization for Multimodal Failure Detectionaccepted
- Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognitionaccepted
- Adaptive Depth Lightweight RGB-T Tracking with Holistic Token Routingaccepted
- Adaptive Learned Image Compression with Graph Neural Networksaccepted
- Adaptive Spatial-Temporal Window: Unlocking the Potential of Event Cameras in Heterogeneous Velocity Scenariosaccepted
- Adaptive Spectral Feature Forecasting for Diffusion Sampling Accelerationaccepted
- Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generationaccepted
- Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learningaccepted
- AdvFM: Lookahead Flow-Matching Velocity-Field Attacks for Imperceptible and Transferable Adversarial Examplesaccepted
- Advancing Cancer Prognosis with Hierarchical Fusion of Genomic, Proteomic and Pathology Imaging Data from a Systems Biology Perspectiveaccepted
- Advancing Image Classification with Discrete Diffusion Classification Modelingaccepted
- Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimizationaccepted
- AeroAgent: A Vision-Physics-Decision Framework for Aerodynamic Vehicle Designaccepted
- AeroDGS: Physically Consistent Dynamic Gaussian Splatting for Single-Sequence Aerial 4D Reconstructionaccepted
- AeroGS: Scale-Aware Gaussian Splatting for Pose-Free Dynamic UAV Scene Reconstructionaccepted
- Aesthetic Camera Viewpoint Suggestion with 3D Aesthetic Fieldaccepted
- Affine Perspective-Three-Point Problemaccepted
- AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondenceaccepted
- AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesisaccepted
- AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiersaccepted
- Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulationaccepted
- Affordance-First Decomposition for Continual Learning in Video-Language Understandingaccepted
- Affostruction: 3D Affordance Grounding with Generative Reconstructionaccepted
- Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detectionaccepted
- AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detectionaccepted
- Agentic Retoucher for Text-To-Image Generationaccepted
- Agentic Video Summarization via Self-Reflecting Multimodal Understandingaccepted
- Agile Deliberation: Concept Deliberation for Subjective Visual Classificationaccepted
- Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrievalaccepted
- AirSim360: A Panoramic Simulation Platform within Drone Viewaccepted
- AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generationaccepted
- Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detectionaccepted
- Align Images Before You Generateaccepted
- Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision Transformersaccepted
- Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied Agentsaccepted
- AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignmentaccepted
- Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferencesaccepted
- Aligning Text, Images and 3D Structure Token-by-Tokenaccepted
- Aligning What Vision-Language Models See and Perceive with Adaptive Information Flowaccepted
- All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Modelsaccepted
- All Vehicles Can Lie: Efficient Adversarial Defense in Fully Untrusted-Vehicle Collaborative Perception via Pseudo-Random Bayesian Inferenceaccepted
- All in One: Unifying Deepfake Detection, Tampering Localization, and Source Tracing with a Robust Landmark-Identity Watermarkaccepted
- All-in-One Slider for Attribute Manipulation in Diffusion Modelsaccepted
- An Efficient Token Compression Framework for Visual Object Trackingaccepted
- An Empirical Study on How Video-LLMs Answer Video Questionsaccepted
- An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Drivingaccepted
- An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learningaccepted
- Anatomica: Localized Control over Geometric and Topological Properties for Anatomical Diffusion Modelsaccepted
- Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Predictionaccepted
- Anchor-Guided Gradient Alignment for Incomplete Multimodal Learningaccepted
- AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flowsaccepted
- AnchorSplat: Feed-Forward 3D Gaussian Splatting With 3D Geometric Priorsaccepted
- Anchoring and Rescaling Attention for Semantically Coherent Inbetweeningaccepted
- Anchoring the Mind of Multimodal Reasoners: Cognitive Bias as a Vector for Jailbreak Attacksaccepted
- Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Samplingaccepted
- AniMimic: Imitating 3D Animation from Video Priorsaccepted
- Animator-Centric Skeleton Generation on Objects with Fine-Grained Detailsaccepted
- Annotation-Efficient Coreset Selection for Context-dependent Segmentationaccepted
- Anomaly as Non-Conformity via Training-Free Graph Laplacian Energy Minimizationaccepted
- Anomaly-Related Residual Fields for Cross-domain Anomaly Detectionaccepted
- AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectorsaccepted
- AnthroTAP: Learning Point Tracking with Real-World Motionaccepted
- Anti-Degradation Lifelong Multi-View Clusteringaccepted
- Anti-I2V: Safeguarding your Photos from Malicious Image-to-video Generationaccepted
- AntiStyler: Defending Object Detection Models Against Adversarial Patch Attacks Using Style Removalaccepted
- Any Resolution Any Geometry: From Multi-View To Multi-Patchaccepted
- Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Studyaccepted
- Any4D: Unified Feed-Forward Metric 4D Reconstructionaccepted
- AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimizationaccepted
- AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual Referencesaccepted
- AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusionaccepted
- AnyPcc: Compressing Any Point Cloud with a Single Universal Modelaccepted
- ApET: Approximation-Error Guided Token Compression for Efficient VLMsaccepted
- Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generationaccepted
- Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learningaccepted
- ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wildaccepted
- Archon: A Unified Multimodal Model for Holistic Digital Human Generationaccepted
- Are Image-to-Video Models Good Zero-Shot Image Editors?accepted
- Are We Ready for RL in Text-to-3D Generation? A Progressive Investigationaccepted
- ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactionsaccepted
- ArtLLM: Generating Articulated Assets via 3D LLMaccepted
- ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility Proposalsaccepted
- ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understandingaccepted
- Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objectsaccepted
- Asking like Socrates: Socrates helps VLMs understand remote sensing imagesaccepted
- AssemblyBench: Physics-Aware Assembly of Complex Industrial Objectsaccepted
- Assignment-Driven Hash Learning in a Hyper-Semantic Space for On-the-Fly Category Discoveryaccepted
- AstraNav-Memory: Contexts Compression for Long Memoryaccepted
- AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localizationaccepted
- Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video Captioningaccepted
- AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robotsaccepted
- Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Modelaccepted
- Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazingaccepted
- Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformeraccepted
- Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Modelsaccepted
- Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decodingaccepted
- Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swappingaccepted
- Attribution as Retrieval: Model-Agnostic AI-Generated Image Attributionaccepted
- Attribution-Guided Model Rectification of Unreliable Neural Network Behaviorsaccepted
- Audio-sync Video Instance Editing with Granularity-Aware Mask Refineraccepted
- AudioAvatar: Personalized Audio-driven Whole-body Talking Avatarsaccepted
- AudioStory: Generating Long-Form Narrative Audio with Large Language Modelsaccepted
- Authorize-on-Demand: Dynamic Authorization with Legality-Aware Intellectual Property Protection for VLMsaccepted
- AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generationaccepted
- AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Modelsaccepted
- AutoRegressive Generation with B-rep Holistic Token Sequence Representationaccepted
- AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language Modelsaccepted
- Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversationaccepted
- AvatarPointillist: AutoRegressive 4D Gaussian Avatarizationaccepted
- AviaSafe: A Physics-Informed Data-Driven Model for Aviation Safety-Critical Cloud Forecastsaccepted
- AwareVLN: Reasoning with Self-awareness for Vision-Language Navigationaccepted
- B$^3$-Seg: Camera-Free, Training-Free 3DGS Segmentation via Analytic EIG and Beta-Bernoulli Bayesian Updatesaccepted
- BA-GS: Bayesian Adaptive Gaussian Splatting for SFM-Free 3D Reconstructionaccepted
- BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing Ratesaccepted
- BAMI: Training-Free Bias Mitigation in GUI Groundingaccepted
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Modelsaccepted
- BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learningaccepted
- BDNet:Bio-Inspired Dual-Backbone Small Object Detection Networkaccepted
- BEA-GS: BEyond RAdiance Supervision in 3DGS for Precise Object Extractionaccepted
- BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware Rasterizationaccepted
- BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Imagesaccepted
- BHCast: Unlocking Black Hole Plasma Dynamics from a Single Blurry Image with Long-Term Forecastingaccepted
- BIT: Matching-based Bi-directional Interaction Transformation Network for Visible-Infrared Person Re-Identificationaccepted
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Modelsaccepted
- BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship Detectionaccepted
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Modelsaccepted
- Back to Basics: Let Denoising Generative Models Denoiseaccepted
- Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detectionaccepted
- Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensationaccepted
- Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanationsaccepted
- BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentationaccepted
- Balanced Dataset Distillation via Modeling Multiple Visual Pattern Distributionaccepted
- Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Imagesaccepted
- BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognitionaccepted
- Basis-Oriented Low-rank Transfer for Few-Shot and Test-Time Adaptationaccepted
- Batch Loss Score for Dynamic Data Pruningaccepted
- Batman: Benign Knowledge Alignment Through Malicious Null Space in Federated Backdoor Attackaccepted
- Bayesian Decomposition and Semantic Completion for Few-shot Semantic Segmentationaccepted
- BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference Modelingaccepted
- Benchmarking Endoscopic Surgical Image Restoration and Beyondaccepted
- Benchmarking PhD-Level Coding in 3D Geometric Computer Visionaccepted
- Benchmarking Single-Factor Physical Video-to-Audio Generationaccepted
- Best Segmentation Buddies for Image-Shape Correspondenceaccepted
- Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream Performanceaccepted
- Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Predictionaccepted
- Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoningaccepted
- Beyond Appearance: Camouflaged Object Detection via Geometric Structureaccepted
- Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchorsaccepted
- Beyond Caption-Based Queries in Video Moment Retrievalaccepted
- Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object Detectionaccepted
- Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extractionaccepted
- Beyond Euclidean Gossip: KL-Barycentric Consensus on Heterogeneous and Imbalanced Imagesaccepted
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object Trackingaccepted
- Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Modelsaccepted
- Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3Daccepted
- Beyond Global Similarity: Multi-Conditional Retrieval for Fine-Grained Cross-Modal Understandingaccepted
- Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapteraccepted
- Beyond Ground-Truth: Leveraging Image Quality Priors for Real-World Image Restorationaccepted
- Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognitionaccepted
- Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language Modelsaccepted
- Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigationaccepted
- Beyond Mimicry: Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrationsaccepted
- Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognitionaccepted
- Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFTaccepted
- Beyond Myopic Alignment: Lookahead Optimization for Online Class-Incremental Learningaccepted
- Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classificationaccepted
- Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generationaccepted
- Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMsaccepted
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Controlaccepted
- Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detectionaccepted
- Beyond Reassembly: Fractured Object Recovery with Missing Partsaccepted
- Beyond Rule-Based Agents: Active Markov Games for Realistic Multi-Agent Interaction in Autonomous Drivingaccepted
- Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenesaccepted
- Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrievalaccepted
- Beyond Sequential Tools: A Unified VLM Agent System for Photographic Post-Processing via Dynamic Multi-Expert Fusionaccepted
- Beyond Single Images: A Comprehensive Benchmark for Album-Level Vision-Language Understandingaccepted
- Beyond Single Solution: Multi-Hypothesis Deep Unfolding Network for Image Compressive Sensingaccepted
- Beyond Single-View Sufficiency: CVBench for Cross-View Human Understandingaccepted
- Beyond Soft Label: Dataset Distillation via Orthogonal Gradient Matchingaccepted
- Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimationaccepted
- Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusionaccepted
- Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Interventionaccepted
- Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaborationaccepted
- Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentationaccepted
- Beyond Tie Points: Satellite Image Block Adjustment based on Dense Feature Consistencyaccepted
- Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretabilityaccepted
- Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detectionaccepted
- Beyond What's Shared: Recovering Lost Unique Information from Intermediate Layers to Boost Multimodal Geo-Foundation Modelsaccepted
- Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detectionaccepted
- Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinationsaccepted
- Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Trainingaccepted
- Beyond the Ground Truth: Enhanced Supervision for Image Restorationaccepted
- Beyond the Static World: Continual Category Discovery under Visual Driftaccepted
- Beyond the Static-World: Lifelong Learning for All-in-One Medical Image Restorationaccepted
- Bezier Degradation Modeling for LiDAR-based Human Motion Captureaccepted
- Bi-Bridge: Bidirectional Diffusion Bridges for Low-Light Image Enhancementaccepted
- Bi-directional Autoregressive Diffusion for Large Complex Motion Interpolationaccepted
- BiEvLight: Bi-level Learning of Task-Aware Event Refinement for Low-Light Image Enhancementaccepted
- BiFM: Bidirectional Flow Matching for Few-Step Image Editing and Generationaccepted
- BiGMINT: Biologically-guided Hierarchical Multimodal Integration for Modeling Multiple Compound Activities in Drug Discoveryaccepted
- BiGain: Unified Token Compression for Joint Generation and Classificationaccepted
- BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generationaccepted
- BiOTPrompt: Bidirectional Optimal Transport Guided Prompting for Disease Evolution-aware Radiology Report Generationaccepted
- BiPA: Bilevel Prompt Adaptation for Underwater Instance Segmentationaccepted
- BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaborationaccepted
- BiProLoRA: Bilevel Prompt LoRA for Real Scene Recoveryaccepted
- Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla Modelsaccepted
- Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Modelsaccepted
- Bias at the End of the Scoreaccepted
- Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereoaccepted
- Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detectionaccepted
- Bidirectional Normalizing Flow: From Data to Noise and Backaccepted
- Bidirectional Query-Driven Generation of Parametric CAD Sketchaccepted
- Bilevel Layer-Positioning LoRA for Real Image Dehazingaccepted
- BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformersaccepted
- BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignmentaccepted
- BiomedCCPL: Causal Conditional Prompt Learning for Biomedical Vision-Language Modelsaccepted
- Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compressionaccepted
- Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Modelsaccepted
- BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviationaccepted
- Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understandingaccepted
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformersaccepted
- Block-based Learned Image Compression without Blocking Artifactsaccepted
- BluRef: Unsupervised Image Deblurring with Dense-Matching Referencesaccepted
- BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic Enhanceraccepted
- Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processingaccepted
- Boosting Quantitive and Spatial Awareness for Zero-Shot Object Countingaccepted
- Boosting Reasoning in Large Multimodal Models via Activation Replayaccepted
- Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learningaccepted
- Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detectionaccepted
- Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prioraccepted
- Boosting Visual Reprogramming for CLIP with Dual Granularity Alignmentaccepted
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learningaccepted
- Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentationaccepted
- Bootstrapping Multi-view Learning for Test-time Noisy Correspondenceaccepted
- Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptationaccepted
- Boundary-Responsive Differentiable Gating for Superpixel-Based Segmentationaccepted
- Breaking Multimodal LLM Safety via Video-Driven Promptingaccepted
- Breaking Semantic Boundaries: Distribution-Guided Semantic Exploration for Creative Generationaccepted
- Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditionsaccepted
- Breaking Spurious Correlations: Uncertainty-Driven Causal Transformers for AU Detectionaccepted
- Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Graspingaccepted
- Breaking the Continuum: Discrete Distribution Learning for Structural MRI Reconstructionaccepted
- Breaking the Illusion: When Positive Meets Negative in Multimodal Decodingaccepted
- Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning Frameworkaccepted
- Breaking the Scalability Limit of Multi-Projector Calibration with Embedded Camerasaccepted
- BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splattingaccepted
- BrepVGAE: Variational Graph Autoencoder with Unified Latent Representation for B-repaccepted
- Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observationaccepted
- BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessmentaccepted
- BrickNet: Graph-Backed Generative Brick Assemblyaccepted
- Bridge: Basis-Driven Causal Inference Marries VFMs for Domain Generalizationaccepted
- BridgeEQA: Virtual Embodied Agents for Real Bridge Inspectionsaccepted
- Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstructionaccepted
- Bridging Domain Expertise and Generalization for Performance Estimationaccepted
- Bridging Domains through Subspace-Aware Model Mergingaccepted
- Bridging Facial Understanding and Animation via Language Modelsaccepted
- Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolutionaccepted
- Bridging Human Evaluation to Infrared and Visible Image Fusionaccepted
- Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verificationaccepted
- Bridging Privacy and Provenance: Traceable Virtual Identity Generationaccepted
- Bridging RGB and Hematoxylin Components: An Interleaved Guidance and Fusion Framework for Point Supervised Nuclei Segmentationaccepted
- Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigationaccepted
- Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory Bankaccepted
- Bridging the Perception Gap in Image Super-Resolution Evaluationaccepted
- Bringing Your Portrait to 3D Presenceaccepted
- BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point Cloudsaccepted
- Building Robust Vision Encoders for Cross-Dataset Evaluation in Immunofluorescent Microscopyaccepted
- Building a Precise Video Language with Human-AI Oversightaccepted
- BuildingGPT: Auto-Regressive Building Wireframe Reconstruction Model with Reinforcement Learningaccepted
- Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial Transcriptomicsaccepted
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generationaccepted
- Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal Transportaccepted
- C-GenReg: Training-Free 3D Point Cloud Registration by Multi-View-Consistent Geometry-to-Image Generation with Probabilistic Modalities Fusionaccepted
- C-LaV: Conditional Latent Velocity Field Denoising for Weather-Robust LiDAR Place Recognitionaccepted
- CAD-Refiner: A Unified Framework for CAD Generation and Iterative Editingaccepted
- CADC: Content Adaptive Diffusion-Based Generative Image Compressionaccepted
- CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Modelsaccepted
- CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language Misalignmentaccepted
- CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Modelaccepted
- CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topographyaccepted
- CARD: Correlation Aware Restoration with Diffusionaccepted
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoningaccepted
- CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editingaccepted
- CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysisaccepted
- CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interactionaccepted
- CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scaleaccepted
- CASPA: Graph-Structured Concept Anchors for Modality-Agnostic Adaptation in Vision-Language Modelsaccepted
- CASR: A Robust Cyclic Framework for Arbitrary Large-Scale Super-Resolution with Distribution Alignment and Self-Similarity Awarenessaccepted
- CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image Retrievalaccepted
- CATNet: Collaborative Alignment and Transformation Network for Cooperative Perceptionaccepted
- CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answeringaccepted
- CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioningaccepted
- CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detectionaccepted
- CD-Buffer: Complementary Dual-Buffer Framework for Test-Time Adaptation in Adverse Weather Object Detectionaccepted
- CDICS: Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased Decouplingaccepted
- CF-IPT: Cross-Modal Fusion Interactive Prompt Tuning of Vision-Language Pre-Trained Model for Multisource Remote Sensing Data Classificationaccepted
- CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidanceaccepted
- CG-Floor: Centroid-Guided Diffusion for Large-Scale Floorplan Generationaccepted
- CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency Metricaccepted
- CGHair: Compact Gaussian Hair Reconstruction with Card Clusteringaccepted
- CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuningaccepted
- CGU-Bayes: Causal Graph Uncertainty-Guided Bayesian Inference for Domain Generalizationaccepted
- CHAL: Causal-guided Hierarchical Anomaly-aware Learning for Moving Infrared Small Target Detectionaccepted
- CHEEM: Continual Learning by Reuse, New, Adapt and Skip - A Hierarchical Exploration-Exploitation Approachaccepted
- CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data Selectionaccepted
- CHIRP dataset: towards long-term, individual-level, behavioral monitoring of bird populations in the wildaccepted
- CI-VID: A Coherent Interleaved Text-Video Datasetaccepted
- CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment Analysisaccepted
- CIGMA: Causal Information-Gain Mechanistic Attribution of Attention Heads in Vision Transformersaccepted
- CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimationaccepted
- CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Spaceaccepted
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learningaccepted
- CLEP: Contrastive Language-Pose Pretrainingaccepted
- CLEX: Complementary Label Exchange Learning for Noisy Facial Expression Recognitionaccepted
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentenceaccepted
- CLIP-like Model as a Foundational Density Ratio Estimatoraccepted
- CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptationaccepted
- CLP: A Real-World Dataset of Contaminated Lens Protectors for Robust Semantic Segmentationaccepted
- CLaD: Planning with Grounded Foresight via Cross-Modal Latent Dynamicsaccepted
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoningaccepted
- CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generationaccepted
- CMR-RD: Long-Tailed Adaptive VLM for Explainable CMR Diagnosisaccepted
- COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimationaccepted
- COPE: Consistent Occlusion and Prompt Enhancement Network for Occluded Person Re-identificationaccepted
- COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMsaccepted
- COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learningaccepted
- CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMsaccepted
- COT-FM: Cluster-wise Optimal Transport Flow Matchingaccepted
- CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free Fusionaccepted
- CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Thinkaccepted
- CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructionsaccepted
- CREward: A Type-Specific Creativity Reward Modelaccepted
- CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registrationaccepted
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop Reasoningaccepted
- CROWn: A Unified Framework for Anti-Aliased Downsampling and Phase-Calibrated Fusion in 3D Medical Segmentationaccepted
- CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Modelsaccepted
- CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibrationaccepted
- CUBic: Coordinated Unified Bimanual Perception and Control Frameworkaccepted
- CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learningaccepted
- CUPID: Generative 3D Reconstruction via Joint Object and Pose Modelingaccepted
- CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generationaccepted
- CURVE: A Benchmark for Cultural and Multilingual Long Video Reasoningaccepted
- CVA: Context-aware Video-text Alignment for Video Temporal Groundingaccepted
- C^2FG: Control Classifier-Free Guidance via Score Discrepancy Analysisaccepted
- CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusionaccepted
- CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answeringaccepted
- CaT-GS: Efficient 3DGS Rendering for Large-Scale Scenes with Inter-frame Caching and Tile Schedulingaccepted
- CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenizationaccepted
- CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generationaccepted
- CamDirector: Towards Long-Term Coherent Video Trajectory Editingaccepted
- CamPI: Physical Adversarial Examples through Camera Power Signal Injectionaccepted
- Camera Control for Text-to-Image Generation via Learning Viewpoint Tokensaccepted
- Camouflage-aware Image-Text Retrieval via Expert Collaborationaccepted
- Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?accepted
- Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matchingaccepted
- Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformersaccepted
- Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detectionaccepted
- CanonCGT: Reference-Based Color Grading via Canonical Pivot Representationaccepted
- CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigationaccepted
- Captain Safari: A World Engine with Pose-Aligned 3D Memoryaccepted
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objectsaccepted
- CaptionQA: Is Your Caption as Useful as the Image Itself?accepted
- CaricHarmony: Contrastive Diffusion Paths for Identity-Preserving Caricature Synthesisaccepted
- Catalyst4D: High-Fidelity 3D-to-4D Scene Editing via Dynamic Propagationaccepted
- Causal Motion Diffusion Models for Autoregressive Motion Generationaccepted
- CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language Modelsaccepted
- CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Interventionaccepted
- Causality in Video Diffusers is Separable from Denoisingaccepted
- Cell-Type Prototype-Informed Neural Network for Gene Expression Estimation from Pathology Imagesaccepted
- ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Controlaccepted
- Chain of Event-Centric Causal Thought for Physically Plausible Video Generationaccepted
- Chain of World: World Model Thinking in Latent Motionaccepted
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoningaccepted
- Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Modelsaccepted
- Chain-of-Thought Guided Multi-Modal Object Re-Identificationaccepted
- ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Senisngaccepted
- Changes in Real Time: Online Scene Change Detection with Multi-View Fusionaccepted
- Charge: A Comprehensive Novel View Synthesis Benchmark and Dataset to Bind Them Allaccepted
- Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Chartsaccepted
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understandingaccepted
- ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question Answeringaccepted
- ChimeraLoRA: Multi-Head LoRA-Guided Synthetic Datasetsaccepted
- ChordEdit: One-Step Low-Energy Transport for Image Editingaccepted
- Choreographing a World of Dynamic Objectsaccepted
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encodingaccepted
- ChronoGS: Disentangling Invariants and Changes in Multi-Period Scenesaccepted
- CineBrain: A Large-Scale Multi-Modal Audiovisual Brain Dataset for Brain-Conditioned Video Generationaccepted
- CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarizationaccepted
- CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generationaccepted
- Cinematic Audio Source Separation Using Visual Cuesaccepted
- Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformersaccepted
- Circular-DPO: Aligning Multi-Stage 3D Generative Models via Preference Feedback Loopaccepted
- Clair Obscur: an Illumination-Aware Method for Real-World Image Vectorizationaccepted
- Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation Modelingaccepted
- Cleaning the Pool: Progressive Filtering of Unlabeled Pools in Deep Active Learningaccepted
- ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Dataaccepted
- Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosisaccepted
- ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstructionaccepted
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Modelsaccepted
- Closed-Form Concept Erasure via Double Projectionsaccepted
- Clothe and Poseaccepted
- Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Modelsaccepted
- Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretrainingaccepted
- Cluster-aware Anchor Learning for Multi-View Clusteringaccepted
- ClusterMark: Towards Robust Watermarking for Autoregressive Image Generators with Visual Token Clusteringaccepted
- Co-Me: Confidence Guided Token Merging for Visual Geometric Transformersaccepted
- CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detectionaccepted
- CoD: A Diffusion Foundation Model for Image Compressionaccepted
- CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inferenceaccepted
- CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detectionaccepted
- CoIn: Coverage and Informativeness-Guided Token Reduction for Efficient Large Multimodal Modelsaccepted
- CoLC: Communication-Efficient Collaborative Perception with LiDAR Completionaccepted
- CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generationaccepted
- CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual Localizationaccepted
- CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learningaccepted
- CoRiM: Conflict-driven Risk Minimization for Dynamic Multimodal Fusionaccepted
- CoRoGS: Contextual Gaussian Splatting for Robust Large-Deviation View Synthesisaccepted
- CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modelingaccepted
- CoT-Edit: Let CoT Guide Instruction Video Editingaccepted
- CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics Priorityaccepted
- CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Modelsaccepted
- CoWTracker: Tracking by Warping instead of Correlationaccepted
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoningaccepted
- CodeMMR: Bridging Natural Language, Code, and Image for Unified Retrievalaccepted
- CodePercept: Code-Grounded Visual STEM Perception for MLLMsaccepted
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimizationaccepted
- Coded-E2LF: Coded Aperture Light Field Imaging from Eventsaccepted
- CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Drivingaccepted
- CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editingaccepted
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoningaccepted
- ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Drivingaccepted
- Collaborative Multi-Mode Pruning for Vision-Language Modelsaccepted
- Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensingaccepted
- Color-Encoded Illumination for High-Speed Volumetric Scene Reconstructionaccepted
- ColorFLUX: A Structure-Color Decoupling Framework for Old Photo Colorizationaccepted
- ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimationaccepted
- Common Inpainted Objects In-N-Out of Contextaccepted
- CompBench: Benchmarking Complex Instruction-guided Image Editingaccepted
- CompetitorFormer: Mitigating Query Conflicts for 3D Instance Segmentation via Competitive Strategyaccepted
- Complementary Prototype Mapping for Efficient Multimodal Anomaly Detectionaccepted
- Complet4R: Geometric Complete 4D Reconstructionaccepted
- Composing Concepts from Images and Videos via Concept-prompt Bindingaccepted
- Composite-Attribute Person Re-Identification via Pose-Guided Disentanglementaccepted
- Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimizationaccepted
- Compositional Transformation Reasoning for Composed Video Retrievalaccepted
- Compressed-Domain-Aware Online Video Super-Resolutionaccepted
- Computation and Communication Efficient Federated Unlearning via On-server Gradient Conflict Mitigation and Expressionaccepted
- Computational Speckle Pattern Interferometryaccepted
- Computer Vision with a Superpixelation Cameraaccepted
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidenceaccepted
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approachaccepted
- Concept-Aware Batch Sampling Improves Language-Image Pretrainingaccepted
- Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generationaccepted
- Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustnessaccepted
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectorsaccepted
- ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimizationaccepted
- Condensed Test-Time Adaptation of VLMs for Action Recognitionaccepted
- Conditional Factuality Controlled LLMs with Generalization Certificates via Conformal Samplingaccepted
- ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrievalaccepted
- Confidence-Guided Multi-Scale Aggregation for Sparse-View High-Resolution 3D Gaussian Splattingaccepted
- Conflict-Aware Adaptive Cross-Reconstruction for Multimodal Sentiment Analysisaccepted
- Confusion-Aware Spectral Regularizer for Long-Tailed Recognitionaccepted
- ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generationaccepted
- Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCRaccepted
- Consensus vs. Controversy: Mapping the Decision Space Where Architectures Divergeaccepted
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulationaccepted
- ConsistCompose: Unified Multimodal Layout Control for Image Compositionaccepted
- Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learningaccepted
- Consistent Instance Field for Dynamic Scene Understandingaccepted
- Contact-Aware Neural Dynamicsaccepted
- Content-Adaptive Hierarchical Hyperprior for Neural Video Codingaccepted
- Content-Aware Dynamic Patchification for Efficient Video Diffusionaccepted
- Content-Aware Frequency Encoding for Implicit Neural Representations with Fourier-Chebyshev Featuresaccepted
- Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigationaccepted
- Continual Distillation of Teachers from Different Domainsaccepted
- Continual Learning for fMRI-Based Brain Disorder Diagnosis via Functional Connectivity Matrices Generative Replayaccepted
- Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesisaccepted
- Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classificationaccepted
- Controllable Federated Prompt Learning at Test Timeaccepted
- Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervisionaccepted
- Convexity-Aware Noise Calibration: A Self-Supervised Framework for Noise-Level-Unknown Image Denoisingaccepted
- Convolutional Neural Networks Driven by Content Similarityaccepted
- CoopDiff: A Diffusion-Guided Approach for Cooperation under Corruptionsaccepted
- CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generationaccepted
- Coordinate Denoising for Non-Equilibrium Molecular Representation Learningaccepted
- Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraintsaccepted
- Correspondence-Attention Alignment for Multi-View Diffusion Modelsaccepted
- CountGD++: Generalized Prompting for Open-World Countingaccepted
- Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoningaccepted
- Coupled Diffusion Sampling for Training-Free Multi-View Image Editingaccepted
- Coupling Liquid Time-Constant Encoders with Modern Hopfield Memoryaccepted
- Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimationaccepted
- Coverage Optimization for Camera View Selectionaccepted
- CrackSSM: Reviving SSMs for Crack Segmentation via Dynamic Scanningaccepted
- CraftMesh: High-Fidelity Generative Mesh Manipulation via Poisson Seamless Fusionaccepted
- Critical Patch-Aware Sparse Prompting with Decoupled Training for Continual Learning on the Edgeaccepted
- Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigationaccepted
- Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillationaccepted
- Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editingaccepted
- Cross-Domain Demo-to-Code via Neurosymbolic Counterfactual Reasoningaccepted
- Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptationaccepted
- Cross-Hand Latent Representation for Vision-Language-Action Modelsaccepted
- Cross-Instance Gaussian Splatting Registration via Geometry-Aware Feature-Guided Alignmentaccepted
- Cross-Modal Attention Calibration for LVLM Hallucination Mitigationaccepted
- Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Videoaccepted
- Cross-Modal Guided Visual Synthesis for Data-Efficient Multimodal Depression Recognitionaccepted
- Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmarkaccepted
- Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inferenceaccepted
- Cross-Subject EEG-to-Video Reconstruction and Beyondaccepted
- Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classificationaccepted
- Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Imagesaccepted
- Cross-domain Dual-stream Feature Disentanglement for Brain Disorder Prediction with Sparsely Labeled PETaccepted
- Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmarkaccepted
- Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learningaccepted
- Cross-modal Representation Learning for Diffusion-generated Image Detectionaccepted
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentationaccepted
- CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methodsaccepted
- CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstructionaccepted
- CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detectionaccepted
- CrowdGaussian: Reconstructing High-Fidelity 3D Gaussians for Human Crowd from a Single Imageaccepted
- CryoHype: Reconstructing a thousand cryo-EM structures with transformer-based hypernetworksaccepted
- CryoKRAQEN: Kernel-Regularized Annealing for Quantized Embedding Networks in Cryo-EM Heterogeneous Reconstructionaccepted
- CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective Videoaccepted
- Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokensaccepted
- Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generationaccepted
- Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understandingaccepted
- Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptationaccepted
- CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference Customizationaccepted
- Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusionaccepted
- Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Eventsaccepted
- Cycle-Consistent Tuning for Layered Image Decompositionaccepted
- CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird's-Eye-View Semantic Segmentationaccepted
- CycleManip: Enabling Cycle-based Manipulation via Effective History Perception and Understandingaccepted
- D$^2$-FOSA: Dual-Diffusion Guided EEG-to-Image Reconstruction with Frequency-Oriented Semantic Alignmentaccepted
- D-Convexity: A Unified Differentiable Convex Shape Prior via Quasi-Concavity for Data-driven Image Segmentationaccepted
- D-Prism: Differentiable Primitives for Structured Dynamic Modelingaccepted
- D2Cache: Second-Order Delta Caching for Higher Video Diffusion Accelerationaccepted
- D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarpingaccepted
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Networkaccepted
- D2T2 - Multimodal Automated Planning for Brachytherapyaccepted
- D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigationaccepted
- DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detectionaccepted
- DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignmentaccepted
- DABO: Difficulty-Aware Bayesian Optimization with Diffusion-Learned Priorsaccepted
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimationaccepted
- DARC: Dual Adjustment Reasoning with Counterfactuals for Trustworthy Chest X-ray Classificationaccepted
- DASH: A Meta-Attack Framework for Synthesizing Effective and Stealthy Adversarial Examplesaccepted
- DBMSolver: A Training-free Diffusion Bridge Sampler for High-Quality Image-to-Image Translationaccepted
- DC-Merge: Improving Model Merging with Directional Consistencyaccepted
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generationaccepted
- DDSF: Robust Few-Shot Learning via Disentangled Subspaces with Determinantal Point Processaccepted
- DDT: Decoupled Diffusion Transformeraccepted
- DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformersaccepted
- DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARsaccepted
- DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learningaccepted
- DEVA: Fine-tuning Multimodal Large Language Models for Visual Perception Tasksaccepted
- DFD-HR: Generalizable Deepfake Detection via Hierarchical Routing Learningaccepted
- DF^2-VB: Dual-level Fuzzy Fusion with View-specific Boosting for Multi-view Multi-label Classificationaccepted
- DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Imagesaccepted
- DGS: Dual Gradient and Semantic-Shift Guided Low-Rank Adaptation for Class Incremental Learningaccepted
- DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spacesaccepted
- DIMOS: Disentangling Instance-level Moving Object Segmentationaccepted
- DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrievalaccepted
- DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical Imagingaccepted
- DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual Promptingaccepted
- DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Drivingaccepted
- DMAligner: Enhancing Image Alignment via Diffusion Model Based View Synthesisaccepted
- DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Modelsaccepted
- DNF-SR: Dual-Input and Negative-Aware Feature Fine-Tuning for Real-World Image Super-Resolutionaccepted
- DP-FedAdamW: An Efficient Optimizer for Differentially Private Federated Large Modelsaccepted
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generationaccepted
- DPGF-Net: Dual-Prior Guided Fusion Network for Joint Assessment of Perceptual Quality and Semantic Consistency in AI-Generated Imagesaccepted
- DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalitiesaccepted
- DRAMA: Next-Gen Dynamic Orchestration for Resilient Multi-Agent Ecosystems in Fluxaccepted
- DREAM: Document Recognition with Explicit Adaptive Memoryaccepted
- DRM: Diffusion-based Reward Model With Step-wise Guidanceaccepted
- DROID-SLAM in the Wildaccepted
- DRS-GUI: Dynamic Region Search for Training-Free GUI Groundingaccepted
- DRiffusion: Draft-and-Refine Process Parallelizes Diffusion Models with Easeaccepted
- DSCA: Dynamic Subspace Concept Alignment for Lifelong VLM Editingaccepted
- DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDARaccepted
- DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtimeaccepted
- DSO: Direct Steering Optimization for Bias Mitigationaccepted
- DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolutionaccepted
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inferenceaccepted
- DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolutionaccepted
- DVAR: Dynamic Visual Autoregressive Modeling for Image Super-Resolutionaccepted
- DVGT: Driving Visual Geometry Transformeraccepted
- D^3FER: Dual Channel and Dual Branch Network for Robust Facial Expression Recognition under Dual Challengesaccepted
- Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridgingaccepted
- Dark3R: Learning Structure from Motion in the Darkaccepted
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action Recognitionaccepted
- DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera Conditionsaccepted
- Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasetsaccepted
- Data-Centric Meta-Learning for Robust Few-Shot Generalizationaccepted
- Dataset Distillation by Influence Matchingaccepted
- DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Rolesaccepted
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generationaccepted
- DeDelayed: Deleting Remote Inference Delay via On-Device Correctionaccepted
- DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentationaccepted
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representationsaccepted
- Debiased Sample Selection for Learning with Noisy Labelsaccepted
- Deciphering Genotype-Phenotype Mechanisms from High-Content Profiling via Knowledge-Guided Multi-modal Graph Learningaccepted
- Decision Boundary-aware Generation for Long-tailed Learningaccepted
- DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigationaccepted
- Decoding 3D Perception via BrainSSD: Synergistic Fusion of EEG Representations from Static and Dynamic Visual Streamsaccepted
- Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detectionaccepted
- Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compressionaccepted
- Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosisaccepted
- Decouple Your Discovery and Memory in Continual Generalized Category Discoveryaccepted
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoningaccepted
- Decoupled Generative Modeling for Human-Object Interaction Synthesisaccepted
- Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translationaccepted
- Decoupled and Reusable Adaptation for Efficient Cross-Modal Transferaccepted
- Decoupling Bias, Aligning Distributions: Synergistic Fairness Optimization for Deepfake Detectionaccepted
- Decoupling Defense Strategies for Robust Image Watermarkingaccepted
- Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptationaccepted
- Decoupling Vision and Language: Codebook Anchored Visual Adaptationaccepted
- Deep Feature Deformation Weightsaccepted
- DeepAlign: Mitigating Modality Conflict through Modality-Specific Alignmentaccepted
- DeepProtect: Proactive Face-Swapping Defense using Identity Blending and Attribute Distortionaccepted
- DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Modelsaccepted
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Modelsaccepted
- DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detectionaccepted
- Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detectionaccepted
- Defending Unauthorized Model Merging via Dual-Stage Weight Protectionaccepted
- Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillationaccepted
- Deformation-based In-Context Learning for Point Cloud Understandingaccepted
- Degradation-Consistent Test-Time Adaptation for All-in-One Image Restorationaccepted
- Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenariosaccepted
- Dehallu3D: Hallucination-Mitigated 3D Generation from a Single Image via Cyclic View Consistency Refinementaccepted
- Dejavu: Towards Experience Feedback Learning for Embodied Intelligenceaccepted
- Delta Rectified Flow Sampling for Text-to-Image Editingaccepted
- DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta Smoothingaccepted
- Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Modelsaccepted
- Demo2Tutorial: From Human Experience to Multimodal Software Tutorialsaccepted
- DemoFunGrasp: Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement Learningaccepted
- Den-TP: A Density-Balanced Data Curation and Evaluation Framework for Trajectory Predictionaccepted
- Denoise and Align: Towards Source-Free UDA for Robust Panoramic Semantic Segmentationaccepted
- Denoising as Path Planning: Training-Free Acceleration of Diffusion Models with DPCacheaccepted
- Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generationaccepted
- Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensorsaccepted
- Depth Any Endoscopy: Towards Self-Supervised Generalizable Depth Estimation in Monocular Endoscopyaccepted
- Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimationaccepted
- Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimationaccepted
- Depth Peeling for High-Fidelity Gaussian-Enhanced Surfel Renderingaccepted
- DepthFocus: Controllable Depth Estimation for See-Through Scenesaccepted
- Describe Anything Anywhere At Any Momentaccepted
- Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Modelsaccepted
- Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkageaccepted
- Designing to Forget: Deep Semi-parametric Models for Unlearningaccepted
- DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Videoaccepted
- Detect Any AI-Counterfeited Text Imageaccepted
- Detect Anything via Next Point Predictionaccepted
- DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imagingaccepted
- Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplificationaccepted
- Detecting Compressed AI-Generated Images via Phase Spectrum Robustnessaccepted
- Detecting Unknown Objects via Energy-based Separation for Open World Object Detectionaccepted
- DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoningaccepted
- Dexterous World Modelsaccepted
- DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Modelsaccepted
- DiGraphHal-Bench: Evaluating Multimodal Large Language Models on Complex Directed Graphsaccepted
- DiP: Taming Diffusion Models in Pixel Spaceaccepted
- DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledgeaccepted
- DiT-IC: Aligned Diffusion Transformer for Efficient Image Compressionaccepted
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Trainingaccepted
- Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbolsaccepted
- Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projectionaccepted
- Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline Diagramsaccepted
- DialogueVPR: Towards Conversational Visual Place Recognitionaccepted
- Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMsaccepted
- Diff-SemiER: Transparency-Aware Adaptive Fusion Diffusion Model with Generative Prior for Semi-Transparent Eyeglasses Removalaccepted
- Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene Generationaccepted
- DiffBMP: Differentiable Rendering with Bitmap Primitivesaccepted
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformersaccepted
- DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generationaccepted
- DiffSoup: Direct Differentiable Rasterization of Triangle Soup for Extreme Radiance Field Simplificationaccepted
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectificationaccepted
- Differentiable Adaptive 4D Structured Illumination for Joint Capture of Shape and Reflectanceaccepted
- Differentiable Laplacian Matrix Guided Superpixel Segmentationaccepted
- Differentiable Stroke Planning with Dual Parameterization for Efficient and High-Fidelity Painting Creationaccepted
- Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image Compressionaccepted
- Differentially Private 2D Human Pose Estimationaccepted
- DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulationaccepted
- Diffusion Forcing Planner: History-Annealed Planning with Time-Dependent Guidance for Autonomous Drivingaccepted
- Diffusion Guided Chain-of-Vision for Large Autoregressive Vision Modelsaccepted
- Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)accepted
- Diffusion Mental Averagesaccepted
- Diffusion Probe: Generated Image Result Prediction Using CNN Probesaccepted
- Diffusion Sampling Path Tells More: An Efficient Plug-and-Play Strategy for Sample Filteringaccepted
- Diffusion with a Linguistic Compass: Steering the Generation of Clinically Plausible Future sMRI Representations for Early MCI Conversion Predictionaccepted
- Diffusion-Based Makeup Transfer with Facial Region-Aware Makeup Featuresaccepted
- Diffusion-Based Native Adversarial Synthesis for Enhanced Medical Segmentation Generalizationaccepted
- Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learningaccepted
- DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localizationaccepted
- DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhanceraccepted
- Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentationaccepted
- DirectFisheye-GS: Enabling Native Fisheye Input in Gaussian Splatting with Cross-View Joint Optimizationaccepted
- Direction-aware 3D Large Multimodal Modelsaccepted
- DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Cachingaccepted
- Disco-GS: Gaussian Splatting in Dynamic Color Lightingaccepted
- Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentationaccepted
- Discovering Adaptive Task Dependencies for Efficient Multi-Task Representation Compressionaccepted
- Discriminative Perception via Anchored Description for Reasoning Segmentationaccepted
- Disentangle-then-Align: Non-Iterative Hybrid Multimodal Image Registration via Cross-Scale Feature Disentanglementaccepted
- Disentangled Textual Priors for Diffusion-based Image Super-Resolutionaccepted
- Disentanglement-wise Image Dehazing through Cross-Domain Manifold Consensusaccepted
- Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generationaccepted
- Distilling Balanced Knowledge from a Biased Teacheraccepted
- Distilling Quasi-Conformal Mapping: A Generalizable and Efficient Solution for Wide-Angle Correctionaccepted
- Distilling Unsigned Distance Function for Surface Reconstruction from 3D Gaussian Splattingaccepted
- Distributed Image Compression with Multimodal Side Information at Extremely Low Bitratesaccepted
- Distribution-Aligned Multimodal Fusion for Robust Object Detectionaccepted
- Diverse Video Generation with Determinantal Point Process-Guided Policy Optimizationaccepted
- DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformersaccepted
- DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPOaccepted
- Diversity over Uniformity: Rethinking Representation in Generated Image Detectionaccepted
- Divide, Conquer, and Aggregate: Asymmetric Experts for Class-Imbalanced Semi-Supervised Medical Image Segmentationaccepted
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understandingaccepted
- Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?accepted
- Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusionsaccepted
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacksaccepted
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBenchaccepted
- Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Controlaccepted
- Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answeringaccepted
- DocPrune: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruningaccepted
- DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understandingaccepted
- Does YOLO Really Need to See Every Training Image in Every Epoch?accepted
- Domain Sensitive Federated Learning with Fisher-Informed Pruningaccepted
- Domain-Skewed Federated Learning with Feature Decoupling and Calibrationaccepted
- Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programsaccepted
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Modelsaccepted
- Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Designaccepted
- Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Drivingaccepted
- Draft and Refine with Visual Expertsaccepted
- Drainage: A Unifying Framework for Addressing Class Uncertaintyaccepted
- DreamOmni2: Multimodal Instruction-based Generation and Editingaccepted
- DreamSAC: Learning Hamiltonian World Models via Symmetry Explorationaccepted
- DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion Transformeraccepted
- DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prioraccepted
- DreamStereo: Towards Real-Time Stereo Inpainting for HD Videosaccepted
- DreamStyle: A Unified Framework for Video Stylizationaccepted
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Modelsaccepted
- Drift-Resilient Temporal Priors for Visual Trackingaccepted
- Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Drivingaccepted
- DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Drivingaccepted
- DriveLaW: Unifying Planning and Video Generation in a Latent Driving Worldaccepted
- DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Drivingaccepted
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planningaccepted
- DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generationaccepted
- DriveVLN: Towards Mapless Vision-and-Language Navigation in Autonomous Drivingaccepted
- DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidanceaccepted
- Driving on Registersaccepted
- Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splattingaccepted
- Dual Ascent Diffusion for Inverse Problemsaccepted
- Dual Band Thermal Videography: Separating Time-Varying Reflection and Emission Near Ambient Conditionsaccepted
- Dual Graph Regularized Deep Unfolding Network for Guided Depth Map Super-resolutionaccepted
- Dual-Agent Reinforcement Learning for Adaptive and Cost-Aware Visual-Inertial Odometryaccepted
- Dual-Estimator: Decoupling Global and Local Semantic Shift for Drift Compensation in Class-Incremental Learningaccepted
- Dual-Granularity Memory for Efficient Video Generationaccepted
- Dual-Level Confidence based Implicit Self-Refinement for Medical Visual Question Answeringaccepted
- Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classificationaccepted
- Dual-Prototype-Guided Multi-task Learning for Unsupervised Anomaly Detection and Classificationaccepted
- Dual-branch Distilled Transformer for Efficient Asymmetric UAV Trackingaccepted
- Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and Intuitionaccepted
- Dual-level Adapter Boosting Prompt-free Curvilinear Structure Segmentationaccepted
- DualMirage: Hunting Stealthy Multimodal LLM Agents via CAPTCHAs with Contour and Adversarial Illusionsaccepted
- DualPrim: Compact 3D Reconstruction with Positive and Negative Primitivesaccepted
- DualReg: Dual-Space Filtering and Reinforcement for Rigid Registrationaccepted
- DualSplat: Robust 3D Gaussian Splatting via Pseudo-Mask Bootstrapping from Reconstruction Failuresaccepted
- Duala: Dual-Level Alignment of Subjects and Stimuli for Cross-Subject fMRI Decodingaccepted
- DuetMerging: Synergizing Dynamic and Static Strategies for Mitigating Task Interference in Model Mergingaccepted
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidanceaccepted
- DuoGen: Towards Autonomous Interleaved Multimodal Generationaccepted
- DuoMo: Dual Motion Diffusion for World-Space Human Reconstructionaccepted
- DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detectionaccepted
- DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generationaccepted
- DynBridge: Bridging Imagination and Control through Interaction Dynamics for Robot Manipulationaccepted
- DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image Generationaccepted
- Dynamic Black-hole Emission Tomography with Physics-informed Neural Fieldsaccepted
- Dynamic Exposure Burst Image Restorationaccepted
- Dynamic Important Example Mining for Reinforcement Finetuningaccepted
- Dynamic Label Noise Suppression with Optimal Teacher Pool for Facial Expression Recognitionaccepted
- Dynamic Logits Adjustment and Exploration for Test-Time Adaptation in Vision Language Modelsaccepted
- Dynamic Magic: Unleashing Restricted Knowledge for Lifelong Person Re-Identificationaccepted
- Dynamic Momentum Recalibration in Online Gradient Learningaccepted
- Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registrationaccepted
- Dynamic Token Reweighting for Robust Vision-Language Modelsaccepted
- Dynamic Visual SLAM using a General 3D Prioraccepted
- Dynamic-Static Decomposition for Novel View Synthesis of Dynamic Scenes with Spiking Neuronsaccepted
- Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformeraccepted
- DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAsaccepted
- DynamicTree: Interactive Real Tree Animation via Sparse Voxel Spectrumaccepted
- DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Drivingaccepted
- Dynamics-Aware Preference Optimization for Vision-Language Modelsaccepted
- Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videosaccepted
- DynamicsBoost: Dynamic Plausible Video Generation via Annotation-Free Continuation Preference Optimizationaccepted
- E$^2$-SCI: Elastic Edge-Cloud Speculative Decoding via Credit Inertiaaccepted
- E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimationaccepted
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingaccepted
- E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thoughtaccepted
- E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstructionaccepted
- E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Drivingaccepted
- EDGS: Eliminating Densification for Efficient Convergence of 3DGSaccepted
- EE-RL: Vision Language Guided Reinforcement Learning with Explorer and Expert model for End-to-End Autonomous Drivingaccepted
- EEGiT: Teaching Vision Transformers to Understand the EEG signalaccepted
- EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Groundingaccepted
- EI-Part:Explode for Completion and Implode for Refinementaccepted
- ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and Test-time Generative Adaptationaccepted
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understandingaccepted
- ELVIS: Enhance Low-Light for Video Instance Segmentation in the Darkaccepted
- ELiC: Efficient LiDAR Geometry Compression via Cross-Bit-depth Feature Propagation and Bag-of-Encodersaccepted
- EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's Diseaseaccepted
- EMGauss: Continuous Slice-to-3D Reconstruction via Dynamic Gaussian Modeling in Volume Electron Microscopyaccepted
- EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categoriesaccepted
- EMMA: Extracting Multiple physical parameters from Multimodal Dataaccepted
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Modelsaccepted
- EMR-Diff: Edge-aware Multimodal Residual Diffusion Model for Hyperspectral Image Super-resolutionaccepted
- ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understandingaccepted
- ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specializationaccepted
- EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detectionaccepted
- ESAM++: Efficient Online 3D Perception on the Edgeaccepted
- EV-CGNet: Co-visible Focused 3D-guided 2D Event Keypoint Detection Networkaccepted
- EVA: Efficient Reinforcement Learning for End-to-End Video Agentaccepted
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generationaccepted
- EVLF: Early Vision-Language Fusion for Generative Dataset Distillationaccepted
- EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformeraccepted
- EXOTIC: External Vision-driven Incomplete Multi-view Classificationaccepted
- EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrievalaccepted
- EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligenceaccepted
- EarlyTom: Early Token Compression Completes Fast Video Understandingaccepted
- Easy2Hard: From Partially to Fully Unmatched Modalities as Negative Samples in Contrastive Learningaccepted
- Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flowaccepted
- EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decompositioaccepted
- EasyV2V: A High-quality Instruction-based Video Editing Frameworkaccepted
- EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generationaccepted
- EchoPOSE: 6D Pose Estimation of Sparse Echocardiograms for Left-Ventricular 3D Shape Reconstructionaccepted
- EchoVDiff: Cardiac-Cycle Echocardiography Video Generation from Arbitrary Single Frameaccepted
- Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Modelsaccepted
- Echoes of Ownership: Adversarial-Guided Dual Injection for Copyright Protection in MLLMsaccepted
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignmentaccepted
- EcoSplat: Efficiency-controllable Feed-forward 3D Gaussian Splatting from Multi-view Imagesaccepted
- Edge-Focused Super-Resolution for Omnidirectional Images with Spherical Geometric Augmentationaccepted
- Edge-RecViT: Efficient Vision Transformer via Semantic-Refined Dynamic Recursionaccepted
- Edges Compete for Trust: Group Relative Edge Optimization for Building Reconstruction from Point Cloudsaccepted
- Edit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene Editingaccepted
- Edit-aware RAW reconstructionaccepted
- Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceiversaccepted
- EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editingaccepted
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editingaccepted
- Editprint: General Digital Image Forensics via Editing Fingerprint with Self-Augmentation Trainingaccepted
- EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal Modelsaccepted
CVPR accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.