ICCV 2025 Accepted Papers
The full list of 2,620 papers accepted at ICCV 2025 (IEEE/CVF International Conference on Computer Vision). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Poster: 2,595
- "Principal Components" Enable A New Language of ImagesPoster
- 2.5 Years in Class: A Multimodal Textbook for Vision-Language PretrainingPoster
- 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene UpdatePoster
- 2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human VideosPoster
- 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language NavigationPoster
- 3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage GenerationPoster
- 3D Mesh Editing using Masked LRMsPoster
- 3D Test-time Adaptation via Graph Spectral Driven Point ShiftPoster
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object DetectionPoster
- 3DGS-LM: Faster Gaussian-Splatting Optimization with Levenberg-MarquardtPoster
- 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene UnderstandingPoster
- 3DRealCar: An In-the-wild RGB-D Car Dataset with 360-degree ViewsPoster
- 3DSRBench: A Comprehensive 3D Spatial Reasoning BenchmarkPoster
- 4D Gaussian Splatting SLAMPoster
- 4D Visual Pre-training for Robot LearningPoster
- 4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object UnderstandingPoster
- 4DSegStreamer: Streaming 4D Panoptic Segmentation via Dual ThreadsPoster
- 6DOPE-GS: Online 6D Object Pose Estimation using Gaussian SplattingPoster
- 7DGS: Unified Spatial-Temporal-Angular Gaussian SplattingPoster
- A Conditional Probability Framework for Compositional Zero-shot LearningPoster
- A Constrained Optimization Approach for Gaussian Splatting from Coarsely-posed Images and Noisy Lidar Point CloudsPoster
- A Differentiable Wave Optics Model for End-to-End Computational Imaging System OptimizationPoster
- A Framework for Double-Blind Federated Adaptation of Foundation ModelsPoster
- A Good Teacher Adapts Their Knowledge for DistillationPoster
- A Hidden Stumbling Block in Generalized Category Discovery: Distracted AttentionPoster
- A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors For Visual Place RecognitionPoster
- A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D SupervisionPoster
- A Linear N-Point Solver for Structure and Motion from Asynchronous TracksPoster
- A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty MotionsPoster
- A Quality-Guided Mixture of Score-Fusion Experts Framework for Human RecognitionPoster
- A Real-world Display Inverse Rendering DatasetPoster
- A Recipe for Generating 3D Worlds from a Single ImagePoster
- A Simple yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision TasksPoster
- A Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with MambaPoster
- A Tiny Change, A Giant Leap: Long-Tailed Class-Incremental Learning via Geometric Prototype AlignmentPoster
- A Token-level Text Image Foundation Model for Document UnderstandingPoster
- A Unified Framework for Motion Reasoning and Generation in Human InteractionPoster
- A Unified Framework to BRIDGE Complete and Incomplete Deep Multi-View Clustering under Non-IID Missing PatternsPoster
- A Unified Interpretation of Training-Time Out-of-Distribution DetectionPoster
- A View-consistent Sampling Method for Regularized Training of Neural Radiance FieldsPoster
- A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual SetsPoster
- A3GS: Arbitrary Artistic Style into Arbitrary 3D Gaussian SplattingPoster
- AAA-Gaussians: Anti-Aliased and Artifact-Free 3D Gaussian RenderingPoster
- ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge DistillationPoster
- ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-TrainingPoster
- AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous DrivingPoster
- ADCD-Net: Robust Document Image Forgery Localization via Adaptive DCT Feature and Hierarchical Content DisentanglementPoster
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing EvaluationPoster
- AG2aussian: Anchor-Graph Structured Gaussian Splatting for Instance-Level 3D Scene Understanding and EditingPoster
- AGO: Adaptive Grounding for Open World 3D Occupancy PredictionPoster
- AHCPTQ: Accurate and Hardware-Compatible Post-Training Quantization for Segment Anything ModelPoster
- AIComposer: Any Style and Content Image Composition via Feature IntegrationPoster
- AID: Adapting Image2Video Diffusion Models for Instruction-guided Video PredictionPoster
- AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language ModelsPoster
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningPoster
- AIM: Amending Inherent Interpretability via Self-Supervised MaskingPoster
- AIRA: Activation-Informed Low-Rank Adaptation for Large ModelsPoster
- AJAHR: Amputated Joint Aware 3D Human Mesh RecoveryPoster
- ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow PredictionsPoster
- AM-Adapter: Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-WildPoster
- AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory PredictionPoster
- AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and SegmentationPoster
- AR-1-to-3: Single Image to Consistent 3D Object via Next-View PredictionPoster
- AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical ReasoningPoster
- ARGUS: Hallucination and Omission Evaluation in Video-LLMsPoster
- ARIG: Autoregressive Interactive Head Generation for Real-time ConversationsPoster
- ARMO: Autoregressive Rigging for Multi-Category ObjectsPoster
- ART: Adaptive Relation Tuning for Generalized Relation PredictionPoster
- ASCENT: Annotation-free Self-supervised Contrastive Embeddings for 3D Neuron Tracking in Fluorescence MicroscopyPoster
- ASGS: Single-Domain Generalizable Open-Set Object Detection via Adaptive Subgraph SearchingPoster
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense PredictionPoster
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingPoster
- ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human ModelingPoster
- AU-Blendshape for Fine-grained Stylized 3D Facial Expression ManipulationPoster
- AURELIA: Test-time Reasoning Distillation in Audio-Visual LLMsPoster
- AV-Flow: Transforming Text to Audio-Visual Human-like InteractionsPoster
- AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video GenerationPoster
- AVAM: a Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question AnsweringPoster
- AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMsPoster
- AcZeroTS: Active Learning for Zero-shot Tissue Segmentation in Pathology ImagesPoster
- Accelerate 3D Object Detection Models via Zero-Shot Attention Key PruningPoster
- Accelerating Diffusion Sampling via Exploiting Local Transition CoherencePoster
- AccidentalGS: 3D Gaussian Splatting from Accidental Camera MotionPoster
- Achieving More with Less: Additive Prompt Tuning for Rehearsal-Free Class-Incremental LearningPoster
- Acknowledging Focus Ambiguity in Visual QuestionsPoster
- Activation Subspaces for Out-of-Distribution DetectionPoster
- Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object DetectionPoster
- Active Membership Inference Test (aMINT): Enhancing Model Auditability with Multi-Task Learning.Poster
- Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex EnvironmentsPoster
- AdaDCP: Learning an Adapter with Discrete Cosine Prior for Clear-to-Adverse Domain GeneralizationPoster
- AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous DrivingPoster
- AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview DiffusionPoster
- Adapt Foundational Segmentation Models with Heterogeneous Searching SpacePoster
- Adapting In-Domain Few-Shot Segmentation to New Domains without Source Domain RetrainingPoster
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak SupervisionPoster
- Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part GroundingPoster
- Adaptive Caching for Faster Video Generation with Diffusion TransformersPoster
- Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time ShiftsPoster
- Adaptive Hyper-Graph Convolution Network for Skeleton-based Human Action Recognition with Virtual ConnectionsPoster
- Adaptive Learning of High-Value Regions for Semi-Supervised Medical Image SegmentationPoster
- Adaptive Prompt Learning via Gaussian Outlier Synthesis for Out-of-distribution DetectionPoster
- Adaptive Routing of Text-to-Image Generation Requests Between Large Cloud Model and Light-Weight Edge ModelPoster
- AdaptiveAE: An Adaptive Exposure Strategy for HDR Capturing in Dynamic ScenesPoster
- Adding Additional Control to One-Step Diffusion with Joint Distribution MatchingPoster
- Addressing Representation Collapse in Vector Quantized Models with One Linear LayerPoster
- Addressing Text Embedding Leakage in Diffusion-based Image Editing
- AdsQA: Towards Advertisement Video UnderstandingPoster
- AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?Poster
- Advancing Text-to-3D Generation with Linearized Lookahead Variational Score DistillationPoster
- Advancing Textual Prompt Learning with Anchored AttributesPoster
- Advancing Visual Large Language Model for Multi-granular Versatile PerceptionPoster
- Adversarial Attention Perturbations for Large Object Detection TransformersPoster
- Adversarial Data Augmentation for Single Domain Generalization via Lyapunov Exponent-Guided OptimizationPoster
- Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video SynthesisPoster
- Adversarial Exploitation of Data Diversity Improves Visual LocalizationPoster
- Adversarial Purification via Super-Resolution and DiffusionPoster
- Adversarial Reconstruction Feedback for Robust Fine-grained GeneralizationPoster
- Adversarial Robust Memory-Based Continual LearnerPoster
- Adversarial Robustness of Discriminative Self-Supervised Learning in VisionPoster
- Adversarial Training for Probabilistic RobustnessPoster
- AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional RelationsPoster
- Aether: Geometric-Aware Unified World ModelingPoster
- AffordDexGrasp: Open-set Language-guided Dexterous Grasp with Generalizable-Instructive AffordancePoster
- After the Party: Navigating the Mapping From Color to Ambient LightingPoster
- Agreement aware and dissimilarity oriented GLOMPoster
- AgroBench: Vision-Language Model Benchmark in AgriculturePoster
- AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model InferencePoster
- Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature RepresentationPoster
- AlignDiff: Learning Physically-Grounded Camera Alignment via DiffusionPoster
- AlignGuard: Scalable Safety Alignment for Text-to-Image GenerationPoster
- Aligning Constraint Generation with Design Intent in Parametric CADPoster
- Aligning Effective Tokens with Video Anomaly in Large Language ModelsPoster
- Aligning Global Semantics and Local Textures in Generative Video EnhancementPoster
- Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text MatchingPoster
- Aligning Moments in Time using Video QueriesPoster
- Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs ReasoningPoster
- All Parts Matter: A Unified Mask-Free Virtual Try-On FrameworkPoster
- All in One: Visual-Description-Guided Unified Point Cloud SegmentationPoster
- AllGCD: Leveraging All Unlabeled Data for Generalized Category DiscoveryPoster
- AllTracker: Efficient Dense Point Tracking at High ResolutionPoster
- Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic ApproximationPoster
- Allowing Oscillation Quantization: Overcoming Solution Space Limitation in Low Bit-Width QuantizationPoster
- Always Skip AttentionPoster
- Amodal Depth Anything: Amodal Depth Estimation in the WildPoster
- Amodal3R: Amodal 3D Reconstruction from Occluded 2D ImagesPoster
- An Efficient Hybrid Vision Transformer for TinyML ApplicationsPoster
- An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image RetrievalPoster
- An Empirical Study of Autoregressive Pre-training from VideosPoster
- An Information-Theoretic Regularizer for Lossy Neural Image CompressionPoster
- An Inversion-based Measure of Memorization for Diffusion ModelsPoster
- An OpenMind for 3D Medical Vision Self-supervised LearningPoster
- Analyzing Finetuning Representation Shift for Multimodal LLMs SteeringPoster
- Anchor Token Matching: Implicit Structure Locking for Training-free AR Image EditingPoster
- AnimalClue: Recognizing Animals by their TracesPoster
- Animate Anyone 2: High-Fidelity Character Image Animation with Environment AffordancePoster
- AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh AnimationPoster
- AnimeGamer: Infinite Anime Life Simulation with Next Game State PredictionPoster
- AnnofreeOD: Detecting All Classes at Low Frame Rates Without Human AnnotationsPoster
- Anomaly Detection of Integrated Circuits Package Substrates Using the Large Vision Model SAIC: Dataset Construction, Methodology, and ApplicationPoster
- Anti-Tamper Protection for Unauthorized Individual Image GenerationPoster
- Any-SSR: How Recursive Least Squares Works in Continual Learning of Large Language ModelPoster
- Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing TasksPoster
- AnyBimanual: Transferring Unimanual Policy for General Bimanual ManipulationPoster
- AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera CalibrationPoster
- AnyI2V: Animating Any Conditional Image with Motion ControlPoster
- AnyPortal: Zero-Shot Consistent Video Background ReplacementPoster
- ArchiSet: Benchmarking Editable and Consistent Single-View 3D Reconstruction of Buildings with Specific Window-to-Wall RatiosPoster
- Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsPoster
- Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data and Metric PerspectivesPoster
- ArgMatch: Adaptive Refinement Gathering for Efficient Dense MatchingPoster
- ArgoTweak: Towards Self-Updating HD Maps through Structured PriorsPoster
- ArtEditor: Learning Customized Instructional Image Editor from Few-Shot ExamplesPoster
- Arti-PG: A Toolbox for Procedurally Synthesizing Large-Scale and Diverse Articulated Objects with Rich AnnotationsPoster
- Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene DescriptionPoster
- Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question AnsweringPoster
- AstroLoc: Robust Space to Ground Image LocalizerPoster
- Asynchronous Event Error-Minimizing Noise for Safeguarding Event DatasetPoster
- Att-Adapter: A Robust and Precise Domain-Specific Multi-Attributes T2I Diffusion Adapter via Conditional Variational AutoencoderPoster
- Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!Poster
- Attention to Trajectory: Trajectory-Aware Open-Vocabulary TrackingPoster
- Attention to the Burstiness in Visual Prompt Tuning!Poster
- Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationPoster
- Augmented and Softened Matching for Unsupervised Visible-Infrared Person Re-IdentificationPoster
- Augmenting Moment Retrieval: Zero-Dependency Two-Stage LearningPoster
- Authentic 4D Driving Simulation with a Video Generation ModelPoster
- Auto-Controlled Image Perception in MLLMs via Visual Perception TokensPoster
- Auto-Regressive Transformation for Image AlignmentPoster
- Auto-Regressively Generating Multi-View Consistent ImagesPoster
- Auto-Vocabulary Semantic SegmentationPoster
- AutoComPose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsPoster
- AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian SplattingPoster
- AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial PromptsPoster
- AutoScape: Geometry-Consistent Long-Horizon Scene GenerationPoster
- Automated Model Evaluation for Object Detection via Prediction Consistency and ReliabilityPoster
- Automated Red Teaming for Text-to-Image Models through Feedback-Guided Prompt Iteration with Vision-Language ModelsPoster
- Autoregressive Denoising Score Matching is a Good Video Anomaly DetectorPoster
- Auxiliary Prompt Tuning of Vision-Language Models for Few-Shot Out-of-Distribution DetectionPoster
- Avat3r: Large Animatable Gaussian Reconstruction Model for High-fidelity 3D Head AvatarsPoster
- Axis-level Symmetry Detection with Group-Equivariant RepresentationPoster
- B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal TokensPoster
- BANet: Bilateral Aggregation Network for Mobile Stereo MatchingPoster
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language ModelsPoster
- BATCLIP: Bimodal Online Test-Time Adaptation for CLIPPoster
- BUFFER-X: Towards Zero-Shot Point Cloud Registration in Diverse ScenesPoster
- BVINet: Unlocking Blind Video Inpainting with Zero AnnotationsPoster
- BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant LearningPoster
- Back on Track: Bundle Adjustment for Dynamic Scene ReconstructionPoster
- Backdoor Attacks on Neural Networks via One-Bit FlipPoster
- Backdoor Defense via Enhanced Splitting and Trap IsolationPoster
- Backdoor Mitigation by Distance-Driven DetoxificationPoster
- Backdooring Self-Supervised Contrastive Learning by Noisy AlignmentPoster
- Background Invariance Testing According to Semantic ProximityPoster
- BadVideo: Stealthy Backdoor Attack against Text-to-Video GenerationPoster
- Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and ReconstructionPoster
- Balanced Image Stylization with Style Matching ScorePoster
- Balanced Sharpness-Aware Minimization for Imbalanced RegressionPoster
- Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot SegmentationPoster
- Balancing Task-invariant Interaction and Task-specific Adaptation for Unified Image FusionPoster
- Bayesian-Inspired Space-Time SuperpixelsPoster
- Benchmarking Burst Super-Resolution for Polarization Images: Noise Dataset and AnalysisPoster
- Benchmarking Egocentric Visual-Inertial SLAM at City ScalePoster
- Benchmarking Multimodal CoT Reward Model Stepwise by Visual ProgramPoster
- Benchmarking Multimodal Large Language Models Against Image CorruptionsPoster
- Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-to-3D GenerationPoster
- Benefit From Seen: Enhancing Open-Vocabulary Object Detection by Bridging Visual and Textual Co-Occurrence KnowledgePoster
- Beyond Blur: A Fluid Perspective on Generative Diffusion ModelsPoster
- Beyond Brain Decoding: Visual-Semantic Reconstructions to Mental Creation Extension Based on fMRIPoster
- Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line GenerationPoster
- Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action RecognitionPoster
- Beyond Losses Reweighting: Empowering Multi-Task Learning via the Generalization PerspectivePoster
- Beyond Low-Rank Tuning: Model Prior-Guided Rank Allocation for Effective Transfer in Low-Data and Large-Gap Regimes.Poster
- Beyond Next-Token: Next-X Prediction for Autoregressive Visual GenerationPoster
- Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR RepresentationsPoster
- Beyond Perspective: Neural 360-Degree Video CompressionPoster
- Beyond RGB: Adaptive Parallel Processing for RAW Object DetectionPoster
- Beyond Simple Edits: Composed Video Retrieval with Dense ModificationsPoster
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object DetectionPoster
- Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video DetectionPoster
- Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMsPoster
- Beyond Training: Dynamic Token Merging for Zero-Shot Video UnderstandingPoster
- Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly SearchPoster
- Beyond [cls]: Exploring the True Potential of Masked Image Modeling RepresentationsPoster
- Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question AnsweringPoster
- Beyond the Frame: Generating 360deg Panoramic Videos from Perspective VideosPoster
- Beyond the Limits: Overcoming Negative Correlation of Activation-Based Training-Free NASPoster
- BezierGS: Dynamic Urban Scene Reconstruction with Bezier Curve Gaussian SplattingPoster
- Bi-Level Optimization for Self-Supervised AI-Generated Face DetectionPoster
- Bias in Gender Bias Benchmarks: How Spurious Features Distort EvaluationPoster
- Bias-Resilient Weakly Supervised Semantic Segmentation Using Normalizing FlowsPoster
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video RetrievalPoster
- Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction DetectionPoster
- BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View SynthesisPoster
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoPoster
- Blended Point Cloud Diffusion for Localized Text-guided Shape EditingPoster
- Blind Noisy Image Deblurring Using Residual Guidance StrategyPoster
- Blind Video Super-Resolution based on Implicit KernelsPoster
- Blind2Sound: Self-Supervised Image Denoising without Residual NoisePoster
- BlinkTrack: Feature Tracking over 80 FPS via Events and ImagesPoster
- BlueNeg: A 35mm Negative Film Dataset for Restoring Channel-Heterogeneous DeteriorationPoster
- BokehDiff: Neural Lens Blur with One-Step DiffusionPoster
- Bokehlicious: Photorealistic Bokeh Rendering with Controllable AperturesPoster
- Bolt3D: Generating 3D Scenes in SecondsPoster
- Boost 3D Reconstruction using Diffusion-based Monocular Camera CalibrationPoster
- Boosting Adversarial Transferability via Negative Hessian Trace RegularizationPoster
- Boosting Adversarial Transferability via Residual Perturbation AttackPoster
- Boosting Class Representation via Semantically Related Instances for Robust Long-Tailed Learning with Noisy LabelsPoster
- Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and TransferabilityPoster
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer FeaturesPoster
- Boosting MLLM Reasoning with Text-Debiased Hint-GRPOPoster
- Boosting Multi-View Indoor 3D Object Detection via Adaptive 3D Volume ConstructionPoster
- Boosting Multimodal Learning via Disentangled Gradient LearningPoster
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-trainingPoster
- Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic DataPoster
- Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model AdaptationPoster
- Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection via Cross-Domain Retrieval AugmentationPoster
- Boundary Probing for Input Privacy Protection When Using LMM ServicesPoster
- BoxDreamer: Dreaming Box Corners for Generalizable Object Pose EstimationPoster
- Breaking Grid Constraints: Dynamic Graph Reconstruction Network for Multi-organ SegmentationPoster
- Breaking Rectangular Shackles: Cross-View Object Segmentation for Fine-Grained Object Geo-LocalizationPoster
- Breaking the Encoder Barrier for Seamless Video-Language UnderstandingPoster
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent AlignmentPoster
- Bridging 3D Anomaly Localization and Repair via High-Quality Continuous Geometric RepresentationPoster
- Bridging Class Imbalance and Partial Labeling via Spectral-Balanced Energy Propagation for Skeleton-based Action RecognitionPoster
- Bridging Continuous and Discrete Tokens for Autoregressive Visual GenerationPoster
- Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution FrameworkPoster
- Bridging Domain Generalization to Multimodal Domain Generalization via Unified RepresentationsPoster
- Bridging Local Inductive Bias and Long-Range Dependencies with Pixel-Mamba for End-to-end Whole Slide Image AnalysisPoster
- Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging ScenariosPoster
- Bridging the Gap between Brain and Machine in Interpreting Visual Semantics: Towards Self-adaptive Brain-to-Text DecodingPoster
- Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action RecognitionPoster
- Bridging the Sky and Ground: Towards View-Invariant Feature Learning for Aerial-Ground Person Re-IdentificationPoster
- Bring Your Rear Cameras for Egocentric 3D Human Pose EstimationPoster
- Bringing RNNs Back to Efficient Open-Ended Video UnderstandingPoster
- C2MIL: Synchronizing Semantic and Topological Causalities in Multiple Instance Learning for Robust and Interpretable Survival AnalysisPoster
- C4D: 4D Made from 3D through Dual CorrespondencesPoster
- CA-I2P: Channel-Adaptive Registration Network with Global Optimal SelectionPoster
- CA2C: A Prior-Knowledge-Free Approach for Robust Label Noise Learning via Asymmetric Co-learning and Co-trainingPoster
- CABLD: Contrast-Agnostic Brain Landmark Detection with Consistency-Based RegularizationPoster
- CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task SolversPoster
- CAD-Recode: Reverse Engineering CAD Code from Point CloudsPoster
- CAFA: a Controllable Automatic Foley ArtistPoster
- CAP: Evaluation of Persuasive and Creative Image GenerationPoster
- CAPTURE: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object CountingPoster
- CARIM: Caption-Based Autonomous Driving Scene Retrieval via Inclusive Text MatchingPoster
- CARL: Causality-guided Architecture Representation Learning for an Interpretable Performance PredictorPoster
- CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
- CAT: A Unified Click-and-Track Framework for Realistic TrackingPoster
- CATP-LLM: Empowering Large Language Models for Cost-Aware Tool PlanningPoster
- CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View ImagePoster
- CAVIS: Context-Aware Video Instance SegmentationPoster
- CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in LiteracyPoster
- CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian SplattingPoster
- CCMNet: Leveraging Calibrated Color Correction Matrices for Cross-Camera Color ConstancyPoster
- CE-FAM: Concept-Based Explanation via Fusion of Activation MapsPoster
- CF3: Compact and Fast 3D Feature FieldsPoster
- CHARM3R: Towards Unseen Camera Height Robust Monocular 3D DetectorPoster
- CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE SolversPoster
- CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single ImagePoster
- CIARD: Cyclic Iterative Adversarial Robustness DistillationPoster
- CL-Splats: Continual Learning of Gaussian Splatting with Local OptimizationPoster
- CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic SegmentationPoster
- CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian SplattingPoster
- CLIPSym: Delving into Symmetry Detection with CLIPPoster
- CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic SegmentationPoster
- CLOT: Closed Loop Optimal Transport for Unsupervised Action SegmentationPoster
- CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing ModalitiesPoster
- CMB-ML: A Cosmic Microwave Background Dataset for the Oldest Possible Computer Vision TaskPoster
- CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD GenerationPoster
- CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance ShiftsPoster
- CO2-Net: A Physics-Informed Spatio-Temporal Model for Global Surface CO2 ReconstructionPoster
- CODA: Repurposing Continuous VAEs for Discrete TokenizationPoster
- CODE-CL: Conceptor-Based Gradient Projection for Deep Continual LearningPoster
- COIN: Confidence Score-Guided Distillation for Annotation-Free Cell SegmentationPoster
- COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound DatasetsPoster
- COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language NavigationPoster
- COSTARR: Consolidated Open Set Technique with Attenuation for Robust RecognitionPoster
- CObL: Toward Zero-Shot Ordinal Layering without User PromptingPoster
- CRAM: Large Scale Video Continual Learning with Bootstrapped CompressionPoster
- CSD-VAR: Content-Style Decomposition in Visual Autoregressive ModelsPoster
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath ModelingPoster
- CULTURE3D: A Large-Scale and Diverse Dataset of Cultural Landmarks and Terrains for Gaussian-Based Scene RenderingPoster
- CVPT: Cross Visual Prompt TuningPoster
- CWNet: Causal Wavelet Network for Low-Light Image EnhancementPoster
- CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset DistillationPoster
- CaliMatch: Adaptive Calibration for Improving Safe Semi-supervised LearningPoster
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt EnsemblesPoster
- CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language ModelPoster
- CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion ModelsPoster
- Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?Poster
- Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge EditingPoster
- Can We Achieve Efficient Diffusion Without Self-Attention? Distilling Self-Attention into ConvolutionsPoster
- Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D GaussiansPoster
- CanFields: Consolidating Diffeomorphic Flows for Non-Rigid 4D Interpolation from Arbitrary-Length SequencesPoster
- CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space ModulationPoster
- CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language ModelsPoster
- CaptionSmiths: Flexibly Controlling Language Pattern in Image CaptioningPoster
- Capturing head avatar with hand contacts from a monocular videoPoster
- CarGait: Cross-Attention based Re-ranking for Gait recognitionPoster
- CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for GuidancePoster
- Cassic: Towards Content-Adaptive State-Space Models for Learned Image CompressionPoster
- Category-Specific Selective Feature Enhancement for Long-Tailed Multi-Label Image ClassificationPoster
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot LearningPoster
- Causal-Entity Reflected Egocentric Traffic Accident Video SynthesisPoster
- Causality-guided Prompt Learning for Vision-language Models via Visual GranulationPoster
- Certifiably Optimal Anisotropic Rotation AveragingPoster
- CharaConsist: Fine-Grained Consistent Character GenerationPoster
- ChartCap: Mitigating Hallucination of Dense Chart CaptioningPoster
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart ReasoningPoster
- ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language ModelsPoster
- Chimera: Improving Generalist Model with Domain-Specific ExpertsPoster
- CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene ReconstructionPoster
- CityNav: A Large-Scale Dataset for Real-World Aerial NavigationPoster
- ClaraVid: A Holistic Scene Reconstruction Benchmark From Aerial Perspective With Delentropy-Based Complexity ProfilingPoster
- Class-Wise Federated Averaging for Efficient PersonalizationPoster
- CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge DistillationPoster
- ClearSight: Human Vision-Inspired Solutions for Event-Based Motion DeblurringPoster
- Client2Vec: Improving Federated Learning by Distribution Shifts Aware Client IndexingPoster
- Clink! Chop! Thud! - Learning Object Sounds from Real-World InteractionsPoster
- Closed-Loop Transfer for Weakly-supervised Affordance GroundingPoster
- Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive InjectionPoster
- CoA-VLA: Improving Vision-Language-Action Models via Visual-Text Chain-of-AffordancePoster
- CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous DrivingPoster
- CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression SegmentationPoster
- CoLMDriver: LLM-based Negotiation Benefits Cooperative Autonomous DrivingPoster
- CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion ModelsPoster
- CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image MatchingPoster
- CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred ImagesPoster
- CoSMIC: Continual Self-supervised Learning for Multi-Domain Medical Imaging via Conditional Mutual Information MaximizationPoster
- CoST: Efficient Collaborative Perception From Unified Spatiotemporal PerspectivePoster
- CoStoDet-DDPM: Collaborative Training of Stochastic and Deterministic Models Improves Surgical Workflow Anticipation and RecognitionPoster
- CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real VideosPoster
- CogCM: Cognition-Inspired Contextual Modeling for Audio-Visual Speech EnhancementPoster
- CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMsPoster
- Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent DialoguesPoster
- Color Matching Using Hypernetwork-Based Kolmogorov-Arnold NetworksPoster
- Colors See Colors Ignore: Clothes Changing ReID with Color DisentanglementPoster
- CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing GamesPoster
- Combinative Matching for Geometric Shape AssemblyPoster
- Communication-Efficient Multi-Vehicle Collaborative Semantic Segmentation via Sparse 3D Gaussian SharingPoster
- CompCap: Improving Multimodal Large Language Models with Composite CaptionsPoster
- CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image GenerationPoster
- Competitive Distillation: A Simple Learning Strategy for Improving Visual ClassificationPoster
- CompleteMe: Reference-based Human Image CompletionPoster
- Compression of 3D Gaussian Splatting with Optimized Feature Planes and Standard Video CodecsPoster
- Compression-Aware One-Step Diffusion Model for JPEG Artifact RemovalPoster
- ConceptSplit: Decoupled Multi-Concept Personalization of Diffusion Models via Token-wise Adaptation and Attention DisentanglementPoster
- Conditional Latent Diffusion Models for Zero-Shot Instance SegmentationPoster
- Conditional Visual Autoregressive Modeling for Pathological Image RestorationPoster
- ConformalSAM: Unlocking the Potential of Foundational Segmentation Models in Semi-Supervised Semantic Segmentation with Conformal PredictionPoster
- Confound from All Sides, Distill with Resilience: Multi-Objective Adversarial Paths to Zero-Shot RobustnessPoster
- ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using ConstraintsPoster
- Consensus-Driven Active Model SelectionPoster
- Consistency Trajectory Matching for One-Step Generative Super-ResolutionPoster
- ConsistentCity: Semantic Flow-guided Occupancy DiT for Temporally Consistent Driving Scene SynthesisPoster
- ConstStyle: Robust Domain Generalization with Unified Style TransformationPoster
- Constraint-Aware Feature Learning for Parametric Point CloudPoster
- Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain ReasoningPoster
- Contact-Aware Amodal Completion for Human-Object Interaction via Multi-Regional InpaintingPoster
- Contact-Aware Refinement of Human Pose Pseudo-Ground Truth via Bioimpedance SensingPoster
- Context Guided Transformer Entropy Modeling for Video CompressionPoster
- Context-Aware Academic Emotion Dataset and BenchmarkPoster
- ContextFace: Generating Facial Expressions from Emotional ContextsPoster
- Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic ScenariosPoster
- Continual Multiple Instance Learning with Enhanced Localization for Histopathological Whole Slide Image AnalysisPoster
- Continual Personalization for Diffusion ModelsPoster
- Continuous-Time Human Motion Field from Event CamerasPoster
- ContraGS: Codebook-Condensed and Trainable Gaussian Splatting for Fast, Memory-Efficient Reconstruction
- Contrastive Flow MatchingPoster
- Contrastive Test-Time Composition of Multiple LoRA Models for Image GenerationPoster
- Controllable 3D Outdoor Scene Generation via Scene GraphsPoster
- Controllable Feature Whitening for Hyperparameter-Free Bias MitigationPoster
- Controllable Latent Space Augmentation for Digital PathologyPoster
- Controllable Weather Synthesis and Removal with Video Diffusion ModelsPoster
- Controllable and Expressive One-Shot Video Head SwappingPoster
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-ExpertsPoster
- Controlling Multimodal LLMs via Reward-guided DecodingPoster
- CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential PerceptionPoster
- Cooperative Pseudo Labeling for Unsupervised Federated ClassificationPoster
- Coordinate-based Speed of Sound Recovery for Aberration-Corrected Photoacoustic Computed TomographyPoster
- CopyrightShield: Enhancing Diffusion Model Security Against Copyright Infringement AttacksPoster
- CoralSRT: Revisiting Coral Reef Semantic Segmentation by Feature Rectification via Self-supervised GuidancePoster
- CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic SegmentationPoster
- Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the WildPoster
- Correspondence-Free Fast and Robust Spherical Point Pattern RegistrationPoster
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought ReasoningPoster
- CountSE: Soft Exemplar Open-set Object CountingPoster
- CounterPC: Counterfactual Feature Realignment for Unsupervised Domain Adaptation on Point CloudsPoster
- Counting Stacked ObjectsPoster
- Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image AnalysisPoster
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationPoster
- Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLMsPoster
- Cross-Architecture Distillation Made Simple with Redundancy SuppressionPoster
- Cross-Category Subjectivity Generalization for Style-Adaptive Sketch Re-IDPoster
- Cross-Subject Mind Decoding from Inaccurate RepresentationsPoster
- Cross-View Isolated Sign Language Recognition via View Synthesis and Feature DisentanglementPoster
- Cross-modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and MethodPoster
- CryoFastAR: Fast Cryo-EM Ab initio Reconstruction Made EasyPoster
- CuMPerLay: Learning Cubical Multiparameter Persistence VectorizationsPoster
- CuRe: Cultural Gaps in the Long Tail of Text-to-Image SystemsPoster
- Curve-Aware Gaussian Splatting for 3D Parametric Curve ReconstructionPoster
- Customizing Domain Adapters for Domain GeneralizationPoster
- CutS3D: Cutting Semantics in 3D for 2D Unsupervised Instance SegmentationPoster
- Cycle Consistency as Reward: Learning Image-Text Alignment without Human PreferencesPoster
- Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object DetectionPoster
- CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image TranslationPoster
- D-Attn: Decomposed Attention for Large Vision-and-Language ModelPoster
- D2ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action RecognitionPoster
- D3: Training-Free AI-Generated Video Detection Using Second-Order FeaturesPoster
- D3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image DetectionPoster
- DAA*: Deep Angular A Star for Image-based Path PlanningPoster
- DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference ImagesPoster
- DADM: Dual Alignment of Domain and Modality for Face Anti-spoofingPoster
- DADet: Safeguarding Image Conditional Diffusion Models against Adversarial and Backdoor Attacks via Diffusion Anomaly DetectionPoster
- DALIP: Distribution Alignment-based Language-Image Pre-Training for Domain-Specific DataPoster
- DAMap: Distance-aware MapNet for High Quality HD Map ConstructionPoster
- DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain LearningPoster
- DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene RenderingPoster
- DASH: Detection and Assessment of Systematic Hallucinations of VLMsPoster
- DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative PerceptionPoster
- DAViD: Data-efficient and Accurate Vision Models from Synthetic DataPoster
- DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion ModelsPoster
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent SpacePoster
- DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid TokenizerPoster
- DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion ModelsPoster
- DC-TTA: Divide-and-Conquer Framework for Test-Time Adaptation of Interactive SegmentationPoster
- DCHM: Depth-Consistent Human Modeling for Multiview DetectionPoster
- DCT-Shield: A Robust Frequency Domain Defense against Malicious Image EditingPoster
- DDB: Diffusion Driven Balancing to Address Spurious CorrelationsPoster
- DEPTHOR: Depth Enhancement from a Practical Light-Weight dToF Sensor and RGB ImagePoster
- DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking HeadsPoster
- DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video GenerationPoster
- DIA: The Adversarial Exposure of Deterministic Inversion in Diffusion ModelsPoster
- DICE: Staleness-Centric Optimizations for Parallel Diffusion MoE InferencePoster
- DIH-CLIP: Unleashing the Diversity of Multi-Head Self-Attention for Training-Free Open-Vocabulary Semantic SegmentationPoster
- DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative ModelsPoster
- DIMO: Diverse 3D Motion Generation for Arbitrary ObjectsPoster
- DIP: Unsupervised Dense In-Context Post-training of Visual RepresentationsPoster
- DISTA-Net: Dynamic Closely-Spaced Infrared Small Target UnmixingPoster
- DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent DiffusionPoster
- DIVE: Taming DINO for Subject-Driven Video EditingPoster
- DLF: Extreme Image Compression with Dual-generative Latent FusionPoster
- DLFR-Gen: Diffusion-based Video Generation with Dynamic Latent Frame RatePoster
- DM-EFS: Dynamically Multiplexed Expanded Features Set Form for Robust and Efficient Small Object DetectionPoster
- DMQ: Dissecting Outliers of Diffusion Models for Post-Training QuantizationPoster
- DMesh++: An Efficient Differentiable Mesh for Complex ShapesPoster
- DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse RenderingPoster
- DOGR: Towards Versatile Visual Document Grounding and ReferringPoster
- DOLLAR: Few-Step Video Generation via Distillation and Latent Reward OptimizationPoster
- DONUT: A Decoder-Only Model for Trajectory PredictionPoster
- DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose PriorPoster
- DRaM-LHM: A Quaternion Framework for Iterative Camera Pose EstimationPoster
- DSO: Aligning 3D Generators with Simulation Feedback for Physical SoundnessPoster
- DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking TuningPoster
- DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary DescriptionsPoster
- Dark-ISP: Enhancing RAW Image Processing for Low-Light Object DetectionPoster
- Dataset Distillation as Data Compression: A Rate-Utility PerspectivePoster
- Dataset Distillation via Vision-Language Category PrototypePoster
- Dataset Distillation via the Wasserstein MetricPoster
- Dataset Ownership Verification for Pre-trained Masked ModelsPoster
- DeFSS: Image-to-Mask Denoising Learning for Few-shot SegmentationPoster
- DeGauss: Dynamic-Static Decomposition with Gaussian Splatting for Distractor-free 3D ReconstructionPoster
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback SynergyPoster
- DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity UnderstandingPoster
- Debiased Curriculum Adaptation for Safe Transfer Learning in Chest X-ray ClassificationPoster
- Debiasing Trace Guidance: Top-down Trace Distillation and Bottom-up Velocity Alignment for Unsupervised Anomaly DetectionPoster
- DecAD: Decoupling Anomalies in Latent Space for Multi-Class Unsupervised Anomaly DetectionPoster
- Deciphering Cross-Modal Alignment in Large Vision-Language Models via Modality Integration RatePoster
- Decoding Correlation-Induced Misalignment in the Stable Diffusion Workflow for Text-to-Image GenerationPoster
- Decouple and Track: Benchmarking and Improving Video Diffusion Transformers For Motion TransferPoster
- Decouple to Reconstruct: High Quality UHD Restoration via Active Feature Disentanglement and Reversible FusionPoster
- Decoupled Diffusion Sparks Adaptive Scene GenerationPoster
- Decoupled Multi-Predictor Optimization for Inference-Efficient Model TuningPoster
- Deep Adaptive Unfolded Network via Spatial Morphology Stripping and Spectral Filtration for Pan-sharpeningPoster
- Deep Incomplete Multi-view Clustering with Distribution Dual-Consistency Recovery GuidancePoster
- Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength ImagesPoster
- DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement LearningPoster
- DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery AnalysisPoster
- Deeply Supervised Flow-Based Generative ModelsPoster
- Degradation-Modeled Multipath Diffusion for Tunable Metalens PhotographyPoster
- Demeter: A Parametric Model of Crop Plant Morphology from the Real WorldPoster
- Democratizing High-Fidelity Co-Speech Gesture Video GenerationPoster
- Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional TokensPoster
- Denoising Token Prediction in Masked Autoregressive ModelsPoster
- Dense Policy: Bidirectional Autoregressive Learning of ActionsPoster
- Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image GenerationPoster
- DepR: Depth Guided Single-view Scene Reconstruction with Instance-level DiffusionPoster
- Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationPoster
- Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth EstimationPoster
- DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationPoster
- Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for DermatologyPoster
- Describe Anything: Detailed Localized Image and Video CaptioningPoster
- Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object RetrievalPoster
- Describe, Don't Dictate: Semantic Image Editing with Natural Language IntentPoster
- Details Matter for Indoor Open-vocabulary 3D Instance SegmentationPoster
- Detect Anything 3D in the WildPoster
- Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous CirclePoster
- Deterministic Object Pose Confidence Region EstimationPoster
- Devil is in the Uniformity: Exploring Diverse Learners within Transformer for Image RestorationPoster
- DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-to-Robot HandoverPoster
- DexVLG: Dexterous Vision-Language-Grasp Model at ScalePoster
- DiGA3D: Coarse-to-Fine Diffusional Propagation of Geometry and Appearance for Versatile 3D InpaintingPoster
- DiMPLe - Disentangled Multi-Modal Prompt Learning: Enhancing Out-Of-Distribution Alignment with Invariant and Spurious Feature SeparationPoster
- DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-vocabulary Queries in NeRFPoster
- DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationPoster
- DiT4SR: Taming Diffusion Transformer for Real-World Image Super-ResolutionPoster
- DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion TransformersPoster
- DiTaiListener: Controllable High Fidelity Listener Video Generation with DiffusionPoster
- Di[M]O: Distilling Masked Diffusion Models into One-step GeneratorPoster
- Diagnosing Pretrained Models for Out-of-distribution DetectionPoster
- DialNav: Multi-turn Dialog Navigation with a Remote GuidePoster
- DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary LookupPoster
- Diff2I2P: Differentiable Image-to-Point Cloud Registration with Diffusion PriorPoster
- DiffDoctor: Diagnosing Image Diffusion Models Before TreatingPoster
- DiffIP: Representation Fingerprints for Robust IP Protection of Diffusion ModelsPoster
- DiffRefine: Diffusion-based Proposal Specific Point Cloud Densification for Cross-Domain Object DetectionPoster
- DiffSim: Taming Diffusion Models for Evaluating Visual SimilarityPoster
- DiffTell: A High-Quality Dataset for Describing Image Manipulation ChangesPoster
- DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex DegradationsPoster
- Differentiable Room Acoustic Rendering with Multi-View Vision PriorsPoster
- Differential-informed Sample Selection Accelerates Multimodal Contrastive LearningPoster
- Differentially Private Fine-Tuning of Diffusion ModelsPoster
- DiffuMatch: Category-Agnostic Spectral Diffusion Priors for Robust Non-rigid Shape MatchingPoster
- Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion ModelsPoster
- Diffusion Curriculum: Synthetic-to-Real Data Curriculum via Image-Guided DiffusionPoster
- Diffusion Epistemic Uncertainty with Asymmetric Learning for Diffusion-Generated Image DetectionPoster
- Diffusion Guided Adaptive Augmentation for Generalization in Visual Reinforcement LearningPoster
- Diffusion Image PriorPoster
- Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-ResolutionPoster
- Diffusion-Based Extreme High-speed Scenes Reconstruction with the Complementary Vision SensorPoster
- Diffusion-Based Imaginative Coordination for Bimanual ManipulationPoster
- Diffusion-based 3D Hand Motion Recovery with Intuitive PhysicsPoster
- Diffusion-based Source-biased Model for Single Domain Generalized Object DetectionPoster
- DimensionX: Create Any 3D and 4D Scenes from a Single Image with Decoupled Video DiffusionPoster
- Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene ModelingPoster
- Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face RestorationPoster
- DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMsPoster
- DisCoPatch: Taming Adversarially-driven Batch Statistics for Improved Out-of-Distribution DetectionPoster
- DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow DecodingPoster
- DisTime: Distribution-based Time Representation for Video Large Language ModelsPoster
- Discontinuity-aware Normal Integration for Generic Central Camera ModelsPoster
- Discovering Divergent Representations between Text-to-Image ModelsPoster
- Discretized Gaussian Representation for Tomographic ReconstructionPoster
- DisenQ: Disentangling Q-Former for Activity-BiometricsPoster
- Disentangled Clothed Avatar Generation with Layered RepresentationPoster
- Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningPoster
- Disentangling Instance and Scene Contexts for 3D Semantic Scene CompletionPoster
- Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing AccuracyPoster
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-DeconstructionPoster
- DistillDrive: End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning ModelPoster
- Distilling Diffusion Models to Efficient 3D LiDAR Scene CompletionPoster
- Distilling Parallel Gradients for Fast ODE Solvers of Diffusion ModelsPoster
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyPoster
- Diversity-Enhanced Distribution Alignment for Dataset DistillationPoster
- Divide-and-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-supervised Continual LearningPoster
- Diving into the Fusion of Monocular Priors for Generalized Stereo MatchingPoster
- Do It Yourself: Learning Semantic Correspondence from Pseudo-LabelsPoster
- DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document UnderstandingPoster
- Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?Poster
- Domain Generalizable Portrait Style TransferPoster
- Domain-aware Category-level Geometry Learning Segmentation for 3D Point CloudsPoster
- Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint DetectionPoster
- DoppDrive: Doppler-Driven Temporal Aggregation for Improved Radar Object DetectionPoster
- Doppler-Aware LiDAR-RADAR Fusion for Weather-Robust 3D DetectionPoster
- Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion ModelsPoster
- Drawing Developmental Trajectory from Cortical Surface ReconstructionPoster
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
- DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid GuidancePoster
- DreamCube: RGB-D Panorama Generation via Multi-plane SynchronizationPoster
- DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesPoster
- DreamFuse: Adaptive Image Fusion with Diffusion TransformerPoster
- DreamLayer: Simultaneous Multi-Layer Generation via Diffusion ModelPoster
- DreamRelation: Relation-Centric Video CustomizationPoster
- DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image ModelsPoster
- DriveArena: A Closed-loop Generative Simulation Platform for Autonomous DrivingPoster
- DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous DrivingPoster
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive TransformersPoster
- DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video GenerationPoster
- DuCos: Duality Constrained Depth Super-Resolution via Foundation ModelPoster
- DuET: Dual Incremental Object Detection via Exemplar-Free Task ArithmeticPoster
- Dual Domain Control via Active Learning for Remote Sensing Domain Incremental Object DetectionPoster
- Dual Reciprocal Learning of Language-based Human Motion Understanding and GenerationPoster
- Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image DiffusionPoster
- Dual-Expert Consistency Model for Efficient and High-Quality Video GenerationPoster
- Dual-Process Image GenerationPoster
- Dual-S3D: Hierarchical Dual-Path Selective SSM-CNN for High-Fidelity Implicit ReconstructionPoster
- Dual-Temporal Exemplar Representation Network for Video Semantic SegmentationPoster
- Dual-level Prototype Learning for Composite Degraded Image RestorationPoster
- DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video CustomizationPoster
- DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-based Human Action SegmentationPoster
- DuoLoRA : Cycle-consistent and Rank-disentangled Content-Style PersonalizationPoster
- DyGS-SLAM: Real-Time Accurate Localization and Gaussian Reconstruction for Dynamic ScenesPoster
- DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile ManipulationPoster
- DynFaceRestore: Balancing Fidelity and Quality in Diffusion-Guided Blind Face Restoration with Dynamic Blur-Level Mapping and GuidancePoster
- DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video UnderstandingPoster
- Dynamic Dictionary Learning for Remote Sensing Image SegmentationPoster
- Dynamic Group Detection using VLM-augmented Temporal Groupness GraphPoster
- Dynamic Multi-Layer Null Space Projection for Vision-Language Continual LearningPoster
- Dynamic Multimodal Prototype Learning in Vision-Language ModelsPoster
- Dynamic Point Maps: A Versatile Representation for Dynamic 3D ReconstructionPoster
- Dynamic Reconstruction of Hand-Object Interaction with Distributed Force-aware Contact RepresentationPoster
- Dynamic Typography: Bringing Text to Life via Video Diffusion PriorPoster
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object DetectionPoster
- Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLMPoster
- DynamicFace: High-Quality and Consistent Face Swapping for Image and Video using Composable 3D Facial PriorsPoster
- DynamicID: Zero-Shot Multi-ID Image Personalization with Flexible Facial EditabilityPoster
- E-NeMF: Event-based Neural Motion Field for Novel Space-time View Synthesis of Dynamic ScenesPoster
- E-SAM: Training-Free Segment Every Entity ModelPoster
- EA-KD: Entropy-based Adaptive Knowledge DistillationPoster
- EA-Vit: Efficient Adaptation for Elastic Vision TransformerPoster
- EAMamba: Efficient All-Around Vision State Space Model for Image RestorationPoster
- EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric FlowPoster
- EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image RegistrationPoster
- EDM: Efficient Deep Feature MatchingPoster
- EDiT: Efficient Diffusion Transformers with Linear Compressed AttentionPoster
- EEdit : Rethinking the Spatial and Temporal Redundancy for Efficient Image EditingPoster
- EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained ClientsPoster
- EMD: Explicit Motion Modeling for High-Quality Street Gaussian SplattingPoster
- EMatch: A Unified Framework for Event-based Optical Flow and Stereo MatchingPoster
- EMoTive: Event-guided Trajectory Modeling for 3D Motion EstimationPoster
- ERNet: Efficient Non-Rigid Registration Network for Point SequencesPoster
- ESCNet:Edge-Semantic Collaborative Network for Camouflaged Object DetectionPoster
- ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental LearningPoster
- ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large ModelsPoster
- ETA: Energy-based Test-time Adaptation for Depth CompletionPoster
- ETCH: Generalizing Body Fitting to Clothed Humans via Equivariant TightnessPoster
- ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and AnsweringPoster
- EVDM: Event-based Real-world Video Deblurring with MambaPoster
- EVER: Exact Volumetric Ellipsoid Rendering for Real-time View SynthesisPoster
- EVEv2: Improved Baselines for Encoder-Free Vision-Language ModelsPoster
- EVOLVE: Event-Guided Deformable Feature Transfer and Dual-Memory Refinement for Low-Light Video Object SegmentationPoster
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionPoster
- EYE3:Turn Anything into Naked-eye 3DPoster
- Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image EditingPoster
- Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingPoster
- Easy3D: A Simple Yet Effective Method for 3D Interactive SegmentationPoster
- EasyControl: Adding Efficient and Flexible Control for Diffusion TransformerPoster
- Edicho: Consistent Image Editing in the WildPoster
- Edit360: 2D Image Edits to 3D Assets from Any AnglePoster
- EditCLIP: Representation Learning for Image EditingPoster
- Effective Training Data Synthesis for Improving MLLM Chart UnderstandingPoster
- Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning StrategyPoster
- Efficient Autoregressive Shape Generation via Octree-Based Adaptive TokenizationPoster
- Efficient Concertormer for Image Deblurring and BeyondPoster
- Efficient Event Camera Data Pretraining with Adaptive Prompt FusionPoster
- Efficient Fine-Tuning of Large Models via Nested Low-Rank AdaptationPoster
- Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation VariationPoster
- Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal InteractionsPoster
- Efficient Spiking Point Mamba for Point Cloud AnalysisPoster
- Efficient Track AnythingPoster
- Efficient Unsupervised Shortcut Learning Detection and Mitigation in TransformersPoster
- Efficient Visual Place Recognition Through Multimodal Semantic Knowledge IntegrationPoster
- EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion ModelsPoster
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric PerceptionPoster
- EgoAgent: A Joint Predictive Agent Model in Egocentric WorldsPoster
- EgoM2P: Egocentric Multimodal Multitask Pretraining
- EgoMusic-driven Human Dance Motion Estimation with Skeleton MambaPoster
- Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language GuidancePoster
- Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image DescriptionsPoster
- Embodied Navigation with Auxiliary Task of Action Description PredictionPoster
- Embodied Representation Alignment with Mirror NeuronsPoster
- Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene UnderstandingPoster
- EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene UnderstandingPoster
- EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile DevicePoster
- EmotiCrafter: Text-to-Emotional-Image Generation based on Valence-Arousal ModelPoster
- Emulating Self-attention with Convolution for Efficient Image Super-ResolutionPoster
- End-to-End Driving with Online Trajectory Evaluation via BEV World ModelPoster
- End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationPoster
- End-to-End Multi-Modal Diffusion MambaPoster
- Engage for All: Making Ordinary Image Descriptions Appealing Again!Poster
- Enhanced Event-based Dense Stereo via Cross-Sensor Knowledge DistillationPoster
- Enhanced Pansharpening via Quaternion Spatial-Spectral InteractionsPoster
- Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model FeaturesPoster
- Enhancing Image Restoration Transformer via Adaptive Translation EquivariancePoster
- Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense PredictionPoster
- Enhancing Numerical Prediction of MLLMs with Soft LabelingPoster
- Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningPoster
- Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image AlignmentPoster
- Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based SegmentationPoster
- Enhancing Transferability of Targeted Adversarial Examples via Inverse Target Gradient Competition and Spatial Distance StretchingPoster
- Enhancing Transformers Through Conditioned Embedded TokensPoster
- Enhancing Zero-shot Object Counting via Text-guided Local Ranking and Number-evoked Global AttentionPoster
- Enpowering Your Pansharpening Models with Generalizability: Unified Distribution is All You NeedPoster
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMsPoster
- Ensemble Foreground Management for Unsupervised Object DiscoveryPoster
- Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step AlignmentPoster
- Environment-Agnostic Pose: Generating Environment-independent Object Representations for 6D Pose EstimationPoster
- Epipolar Consistent Attention Aggregation Network for Unsupervised Light Field Disparity EstimationPoster
- Epona: Autoregressive Diffusion World Model for Autonomous DrivingPoster
- EquiCaps: Predictor-Free Pose-Aware Pre-Trained Capsule NetworksPoster
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution DetectionPoster
- Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target ConceptsPoster
- Error Recognition in Procedural Videos using Generalized Task GraphPoster
- Estimating 2D Camera Motion with Hybrid Motion BasisPoster
- EvRT-DETR: Latent Space Adaptation of Image Detectors for Event-based VisionPoster
- EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry ImagesPoster
- Evading Data Provenance in Deep Neural NetworksPoster
- Event-Driven Storytelling with Multiple Lifelike Humans in a 3D ScenePoster
- Event-aided Dense and Continuous Point Tracking: Everywhere and AnytimePoster
- Event-based Tiny Object Detection: A Benchmark Dataset and BaselinePoster
- Event-based Visual VibrometryPoster
- Event-boosted Deformable 3D Gaussians for Dynamic Scene ReconstructionPoster
- Event-guided HDR Reconstruction with Diffusion PriorsPoster
- Event-guided Unified Framework for Low-light Video Enhancement, Frame Interpolation, and DeblurringPoster
- EventUPS: Uncalibrated Photometric Stereo Using an Event CameraPoster
- Everything is a Video: Unifying Modalities through Next-Frame PredictionPoster
- Evidential Knowledge DistillationPoster
- EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference AlignmentPoster
- ExCap3D: Expressive 3D Scene Understanding via Object Captioning with Varying DetailPoster
- Explaining Human Preferences via Metrics for Structured 3D ReconstructionPoster
- Exploiting Diffusion Prior for Task-driven Image RestorationPoster
- Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic SegmentationPoster
- Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score PropagationPoster
- ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion PriorsPoster
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image EditingPoster
- Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic SegmentationPoster
- Exploring The Visual Feature Space for Multimodal Neural DecodingPoster
- Exploring View Consistency for Scene-Adaptive Low-Light Light Field Image EnhancementPoster
- Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in RoboticsPoster
- Expressive Talking Human from Single-Image with Imperfect PriorsPoster
- Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
- External Knowledge Injection for CLIP-Based Class-Incremental LearningPoster
- Extrapolated Urban View Synthesis BenchmarkPoster
- F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and RestorationPoster
- FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution DetectionPoster
- FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual RegistersPoster
- FB-Diff: Fourier Basis-guided Diffusion for Temporal Interpolation of 4D Medical ImagingPoster
- FDPT: Federated Discrete Prompt Tuning for Black-Box Visual-Language ModelsPoster
- FE-CLIP: Frequency Enhanced CLIP Model for Zero-Shot Anomaly Detection and SegmentationPoster
- FED-PsyAU: Privacy-Preserving Micro-Expression Recognition via Psychological AU Coordination and Dynamic Facial Motion ModelingPoster
- FEVER-OOD: Free Energy Vulnerability Elimination for Robust Out-of-Distribution DetectionPoster
- FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental UnlearningPoster
- FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image GenerationPoster
- FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal DataPoster
- FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking PortraitPoster
- FLOSS: Free Lunch in Open-vocabulary Semantic SegmentationPoster
- FLSeg: Enhancing Privacy and Robustness in Federated Learning under Heterogeneous Data via Model SegmentationPoster
- FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformancePoster
- FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face RetouchingPoster
- FREE-Merging: Fourier Transform for Efficient Model MergingPoster
- FRET: Feature Redundancy Elimination for Test Time AdaptationPoster
- FROSS: Faster-Than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D ImagesPoster
- FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion DistillationPoster
- FW-Merging: Scaling Model Merging with Frank-Wolfe OptimizationPoster
- Face Retouching with Diffusion Data Generation and Spectral RestorementPoster
- FaceCraft4D: Animated 3D Facial Avatar Generation from a Single ImagePoster
- FaceLift: Learning Generalizable Single Image 3D Face Reconstruction from Synthetic HeadsPoster
- FaceShield: Defending Facial Image against Deepfake ThreatsPoster
- FaceXFormer: A Unified Transformer for Facial AnalysisPoster
- Factorized Learning for Temporally Grounded Video-Language ModelsPoster
- Failure Cases Are Better Learned But Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingPoster
- Fair Generation without Unfair Distortions: Debiasing Text-to-Image Generation with Entanglement-Free AttentionPoster
- FairGen: Enhancing Fairness in Text-to-Image Diffusion Models via Self-Discovering Latent DirectionsPoster
- FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion ModelsPoster
- FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake VideosPoster
- Fast Globally Optimal and Geometrically Consistent 3D Shape MatchingPoster
- Fast Image Super-Resolution via Consistency Rectified FlowPoster
- FastJSMA: Accelerating Jacobian-based Saliency Map Attacks through Gradient DecouplingPoster
- FastPoint: Accelerating 3D Point Cloud Model Inference via Sample Point Distance PredictionPoster
- FastVAR: Linear Visual Autoregressive Modeling via Cached Token PruningPoster
- Faster and Better 3D Splatting via Group TrainingPoster
- Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model AccelerationPoster
- Feature Coding in the Era of Large Models: Dataset, Test Conditions, and BenchmarkPoster
- Feature Extraction and Representation of Pre-training Point Cloud Based on Diffusion ModelsPoster
- Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic SegmentationPoster
- FedAGC: Federated Continual Learning with Asymmetric Gradient CorrectionPoster
- FedDifRC: Unlocking the Potential of Text-to-Image Diffusion Models in Heterogeneous Federated LearningPoster
- FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language ModelsPoster
- FedMeNF: Privacy-Preserving Federated Meta-Learning for Neural FieldsPoster
- FedPall: Prototype-based Adversarial and Collaborative Learning for Federated Learning with Feature DriftPoster
- FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic ManipulationPoster
- FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform QuantizationPoster
- FedXDS: Leveraging Model Attribution Methods to counteract Data Heterogeneity in Federated LearningPoster
- Federated Continual Instruction TuningPoster
- Federated Continuous Category Discovery and LearningPoster
- Federated Domain Generalization with Domain-specific Soft Prompts GenerationPoster
- Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client DataPoster
- Federated Representation Angle LearningPoster
- Feed-Forward SceneDINO for Unsupervised Semantic Scene CompletionPoster
- Few-Shot Image Quality Assessment via Adaptation of Vision-Language ModelsPoster
- Few-Shot Pattern Detection via Template Matching and RegressionPoster
- Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model DeploymentPoster
- FiVE-Bench: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow ModelsPoster
- FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth EstimationPoster
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and ChallengingPoster
- Find Any Part in 3DPoster
- Find a Scapegoat: Poisoning Membership Inference Attack and Defense to Federated LearningPoster
- Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures via Joint Reconstruction and RegistrationPoster
- Fine-Grained Evaluation of Large Vision-Language Models in Autonomous DrivingPoster
- Fine-Tuning Visual Autogressive Models for Subject-Driven GenerationPoster
- Fine-grained Abnormality Prompt Learning for Zero-shot Anomaly DetectionPoster
- Fine-grained Spatiotemporal Grounding on Egocentric VideosPoster
- Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE TrainingPoster
- FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and EditingPoster
- Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric VisionPoster
- Fix-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long TextPoster
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesPoster
- Flash-VStream: Efficient Real-Time Understanding for Long Video StreamsPoster
- FlashDepth: Real-time Streaming Video Depth Estimation at 2K ResolutionPoster
- FlexGen: Flexible Multi-View Generation from Text and Image InputsPoster
- Flexi-FSCIL: Adaptive Knowledge Retention for Breaking the Stability-Plasticity Dilemma in Few-Shot Class-Incremental LearningPoster
- Flow Stochastic Segmentation NetworksPoster
- Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image TokenizationPoster
- Flow-MIL: Constructing Highly-expressive Latent Feature Space For Whole Slide Image Classification Using Normalizing FlowPoster
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical FlowPoster
- FlowChef: Steering of Rectified Flow Models for Controlled GenerationsPoster
- FlowDPS : Flow-Driven Posterior Sampling for Inverse ProblemsPoster
- FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow ModelsPoster
- FlowR: Flowing from Sparse to Dense 3D ReconstructionsPoster
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion BasesPoster
- FlowStyler: Artistic Video Stylization via Transformation Fields TransportsPoster
- FlowTok: Flowing Seamlessly Across Text and Image TokensPoster
- Focal Plane Visual Feature Generation and Matching on a Pixel Processor ArrayPoster
- FonTS: Text Rendering With Typography and Style ControlsPoster
- FontAnimate: High Quality Few-shot Font Generation via Animating Font Transfer ProcessPoster
- ForCenNet: Foreground-Centric Network for Document Image RectificationPoster
- ForeSight: Multi-View Streaming Joint Object Detection and Trajectory ForecastingPoster
- Forecasting Continuous Non-Conservative Dynamical Systems in SO(3)Poster
- Forensic-MoE: Exploring Comprehensive Synthetic Image Detection Traces with Mixture of ExpertsPoster
- Foresight in Motion: Reinforcing Trajectory Prediction with Reward HeuristicsPoster
- ForestFormer3D: A Unified Framework for End-to-End Segmentation of Forest LiDAR 3D Point CloudsPoster
- ForgeLens: Data-Efficient Forgery Focus for Generalizable Forgery Image DetectionPoster
- Forgetting Through Transforming: Enabling Federated Unlearning via Class-Aware Representation TransformationPoster
- FoundIR: Unleashing Million-scale Training Data to Advance Foundation Models for Image RestorationPoster
- FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language ModelsPoster
- FramePainter: Endowing Interactive Image Editing with Video Diffusion PriorsPoster
- Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video GenerationPoster
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single InferencePoster
- Free-running vs Synchronous: Single-Photon Lidar for High-flux 3D ImagingPoster
- Free2Guide: Training-Free Text-to-Video Alignment using Image LVLMPoster
- Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal ConsistencyPoster
- FreeCus: Free Lunch Subject-driven Customization in Diffusion TransformersPoster
- FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise AlignmentPoster
- FreeDance: Towards Harmonic Free-Number Group Dance Generation via a Unified FrameworkPoster
- FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image EditingPoster
- FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion ModelPoster
- FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale FusionPoster
- FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D ReconstructionPoster
- FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection TransformersPoster
- Frequency Domain-Based Diffusion Model for Unpaired Image DehazingPoster
- Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal ForecastingPoster
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image SynthesisPoster
- Frequency-Dynamic Attention Modulation For Dense PredictionPoster
- Frequency-Guided Diffusion for Training-Free Text-Driven Image TranslationPoster
- Frequency-Guided Posterior Sampling for Diffusion-Based Image RestorationPoster
- Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-based Action RecognitionPoster
- From Abyssal Darkness to Blinding Glare: A Benchmark on Extreme Exposure Correction in Real WorldPoster
- From Easy to Hard: Progressive Active Learning Framework for Infrared Small Target Detection with Single Point SupervisionPoster
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image ReasoningPoster
- From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuningPoster
- From Gallery to Wrist: Realistic 3D Bracelet Insertion in VideosPoster
- From Gaze to Movement: Predicting Visual Attention for Autonomous Driving Human-Machine Interaction based on Programmatic Imitation LearningPoster
- From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-TuningPoster
- From Image to Video: An Empirical Study of Diffusion RepresentationsPoster
- From Imitation to Innovation: The Emergence of AI's Unique Artistic Styles and the Challenge of Copyright ProtectionPoster
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial CorrelationsPoster
- From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors via LLM-guided Symbolic ReasoningPoster
- From One to More: Contextual Part Latents for 3D GenerationPoster
- From Panels to Prose: Generating Literary Narratives from ComicsPoster
- From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute TransitionPoster
- From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection TuningPoster
- From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeersPoster
- From Sharp to Blur: Unsupervised Domain Adaptation for 2D Human Pose Estimation Under Extreme Motion Blur Using Event CamerasPoster
- From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward AlignmentPoster
- FuXi-RTM: A Physics-Guided Prediction Framework with Radiative Transfer ModelingPoster
- FullDiT: Video Generative Foundation Models with Multimodal Control via Full AttentionPoster
- Function-centric Bayesian Network for Zero-Shot Object Goal NavigationPoster
- Fuse Before Transfer: Knowledge Fusion for Heterogeneous DistillationPoster
- Fusion Meets Diverse Conditions: A High-diversity Benchmark and Baseline for UAV-based Multimodal Object Detection with Condition CuesPoster
- FusionPhys: A Flexible Framework for Fusing Complementary Sensing Modalities in Remote Physiological MeasurementPoster
- Future-Aware Interaction Network For Motion ForecastingPoster
- Fuzzy Contrastive Decoding to Alleviate Object Hallucination in Large Vision-Language ModelsPoster
- G-DexGrasp: Generalizable Dexterous Grasping Synthesis Via Part-Aware Prior Retrieval and Prior-Assisted GenerationPoster
- G2D: Boosting Multimodal Learning with Gradient-Guided DistillationPoster
- G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction via Evolutionary DiffusionPoster
- G2SF: Geometry-Guided Score Fusion for Multimodal Industrial Anomaly DetectionPoster
- GAP: Gaussianize Any Point Clouds with Text GuidancePoster
- GARF: Learning Generalizable 3D Reassembly for Real-World FracturesPoster
- GAS: Generative Avatar Synthesis from a Single ImagePoster
- GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in InterpretabilityPoster
- GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray DiffusionPoster
- GDKVM: Echocardiography Video Segmentation via Spatiotemporal Key-Value Memory with Gated Delta RulePoster
- GECKO: Gigapixel Vision-Concept Contrastive Pretraining in HistopathologyPoster
- GECO: Geometrically Consistent Embedding with Lightspeed InferencePoster
- GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray DiagnosisPoster
- GENMO: A GENeralist Model for Human MOtionPoster
- GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial TasksPoster
- GEOPARD: Geometric Pretraining for Articulation Prediction in 3D ShapesPoster
- GFPack++: Attention-Driven Gradient Fields for Optimizing 2D Irregular PackingPoster
- GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific AdaptationPoster
- GIViC: Generative Implicit Video CompressionPoster
- GLEAM: Enhanced Transferable Adversarial Attacks for Vision-Language Pre-training Models via Global-Local TransformationsPoster
- GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor ScenePoster
- GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-ExpertsPoster
- GMMamba: Group Masking Mamba for Whole Slide Image ClassificationPoster
- GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal ModelsPoster
- GReg: Geometry-Aware Region Refinement for Sign Language Video GenerationPoster
- GS-ID: Illumination Decomposition on Gaussian Splatting via Adaptive Light Aggregation and Diffusion-Guided Material PriorsPoster
- GS-LIVM: Real-Time Photo-Realistic LiDAR-Inertial-Visual Mapping with Gaussian SplattingPoster
- GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian SplattingPoster
- GSOT3D: Towards Generic 3D Single Object Tracking in the WildPoster
- GSRecon: Efficient Generalizable Gaussian Splatting for Surface Reconstruction from Sparse ViewsPoster
- GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object GenerationPoster
- GT-Loc: Unifying When and Where in Images Through a Joint Embedding SpacePoster
- GT-Mean Loss: A Simple Yet Effective Solution for Brightness Mismatch in Low-Light Image EnhancementPoster
- GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent TrainingPoster
- GUAVA: Generalizable Upper Body 3D Gaussian AvatarPoster
- GVDepth: Zero-Shot Monocular Depth Estimation for Ground Vehicles based on Probabilistic Cue FusionPoster
- GWM: Towards Scalable Gaussian World Models for Robotic ManipulationPoster
- GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo CollectionsPoster
- GaSLight: Gaussian Splats for Spatially-Varying Lighting in HDRPoster
- Gain-MLP: Improving HDR Gain Map Encoding via a Lightweight MLPPoster
- Gait-X: Exploring X modality for Generalized Gait RecognitionPoster
- GameFactory: Creating New Games with Generative Interactive VideosPoster
- GauUpdate: New Object Insertion in 3D Gaussian Fields with Consistent Global IlluminationPoster
- GausSim: Foreseeing Reality by Gaussian Simulator for Elastic ObjectsPoster
- GaussRender: Learning 3D Occupancy with Gaussian RenderingPoster
- Gaussian Splatting with Discretized SDF for Relightable AssetsPoster
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D SynthesisPoster
- Gaussian-based World Model: Gaussian Priors for Voxel-Based Occupancy Prediction and Future Motion PredictionPoster
- GaussianFlowOcc: Sparse and Weakly Supervised Occupancy Estimation using Gaussian Splatting and Temporal FlowPoster
- GaussianOcc: Fully Self-supervised and Efficient 3D Occupancy Estimation with Gaussian SplattingPoster
- GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMsPoster
- GaussianReg: Rapid 2D/3D Registration for Emergency Surgery via Explicit 3D Modeling with Gaussian PrimitivesPoster
- GaussianSpeech: Audio-Driven Personalized 3D Gaussian AvatarsPoster
- GaussianUpdate: Continual 3D Gaussian Splatting Update for Changing EnvironmentsPoster
- GaussianVideo: Efficient Video Representation via Hierarchical Gaussian SplattingPoster
- Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze ScanpathsPoster
- GazeGaussian: High-Fidelity Gaze Redirection with 3D Gaussian SplattingPoster
- Geminio: Language-Guided Gradient Inversion Attacks in Federated LearningPoster
- GenDoP: Auto-regressive Camera Trajectory Generation as a Director of PhotographyPoster
- GenFlow3D: Generative Scene Flow Estimation and Prediction on Point Cloud SequencesPoster
- GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement LearningPoster
- GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric EnhancersPoster
- GenHaze: Pioneering Controllable One-Step Realistic Haze Generation for Real-World DehazingPoster
- GenM3: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion GenerationPoster
- General Compression Framework for Efficient Transformer Object TrackingPoster
- Generalizable Non-Line-of-Sight Imaging with Learnable Physical PriorsPoster
- Generalizable Object Re-Identification via Visual In-Context PromptingPoster
- Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake DetectionPoster
- Generalized Deep Multi-view Clustering via Causal Learning with Partially Aligned Cross-view CorrespondencePoster
ICCV accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.