← All conferences

ICCV 2025 Accepted Papers

The full list of 2,620 papers accepted at ICCV 2025 (IEEE/CVF International Conference on Computer Vision). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Poster: 2,595
  1. "Principal Components" Enable A New Language of ImagesPoster
  2. 2.5 Years in Class: A Multimodal Textbook for Vision-Language PretrainingPoster
  3. 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene UpdatePoster
  4. 2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human VideosPoster
  5. 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language NavigationPoster
  6. 3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage GenerationPoster
  7. 3D Mesh Editing using Masked LRMsPoster
  8. 3D Test-time Adaptation via Graph Spectral Driven Point ShiftPoster
  9. 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object DetectionPoster
  10. 3DGS-LM: Faster Gaussian-Splatting Optimization with Levenberg-MarquardtPoster
  11. 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene UnderstandingPoster
  12. 3DRealCar: An In-the-wild RGB-D Car Dataset with 360-degree ViewsPoster
  13. 3DSRBench: A Comprehensive 3D Spatial Reasoning BenchmarkPoster
  14. 4D Gaussian Splatting SLAMPoster
  15. 4D Visual Pre-training for Robot LearningPoster
  16. 4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object UnderstandingPoster
  17. 4DSegStreamer: Streaming 4D Panoptic Segmentation via Dual ThreadsPoster
  18. 6DOPE-GS: Online 6D Object Pose Estimation using Gaussian SplattingPoster
  19. 7DGS: Unified Spatial-Temporal-Angular Gaussian SplattingPoster
  20. A Conditional Probability Framework for Compositional Zero-shot LearningPoster
  21. A Constrained Optimization Approach for Gaussian Splatting from Coarsely-posed Images and Noisy Lidar Point CloudsPoster
  22. A Differentiable Wave Optics Model for End-to-End Computational Imaging System OptimizationPoster
  23. A Framework for Double-Blind Federated Adaptation of Foundation ModelsPoster
  24. A Good Teacher Adapts Their Knowledge for DistillationPoster
  25. A Hidden Stumbling Block in Generalized Category Discovery: Distracted AttentionPoster
  26. A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors For Visual Place RecognitionPoster
  27. A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D SupervisionPoster
  28. A Linear N-Point Solver for Structure and Motion from Asynchronous TracksPoster
  29. A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty MotionsPoster
  30. A Quality-Guided Mixture of Score-Fusion Experts Framework for Human RecognitionPoster
  31. A Real-world Display Inverse Rendering DatasetPoster
  32. A Recipe for Generating 3D Worlds from a Single ImagePoster
  33. A Simple yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision TasksPoster
  34. A Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with MambaPoster
  35. A Tiny Change, A Giant Leap: Long-Tailed Class-Incremental Learning via Geometric Prototype AlignmentPoster
  36. A Token-level Text Image Foundation Model for Document UnderstandingPoster
  37. A Unified Framework for Motion Reasoning and Generation in Human InteractionPoster
  38. A Unified Framework to BRIDGE Complete and Incomplete Deep Multi-View Clustering under Non-IID Missing PatternsPoster
  39. A Unified Interpretation of Training-Time Out-of-Distribution DetectionPoster
  40. A View-consistent Sampling Method for Regularized Training of Neural Radiance FieldsPoster
  41. A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual SetsPoster
  42. A3GS: Arbitrary Artistic Style into Arbitrary 3D Gaussian SplattingPoster
  43. AAA-Gaussians: Anti-Aliased and Artifact-Free 3D Gaussian RenderingPoster
  44. ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge DistillationPoster
  45. ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-TrainingPoster
  46. AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous DrivingPoster
  47. ADCD-Net: Robust Document Image Forgery Localization via Adaptive DCT Feature and Hierarchical Content DisentanglementPoster
  48. ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing EvaluationPoster
  49. AG2aussian: Anchor-Graph Structured Gaussian Splatting for Instance-Level 3D Scene Understanding and EditingPoster
  50. AGO: Adaptive Grounding for Open World 3D Occupancy PredictionPoster
  51. AHCPTQ: Accurate and Hardware-Compatible Post-Training Quantization for Segment Anything ModelPoster
  52. AIComposer: Any Style and Content Image Composition via Feature IntegrationPoster
  53. AID: Adapting Image2Video Diffusion Models for Instruction-guided Video PredictionPoster
  54. AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language ModelsPoster
  55. AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningPoster
  56. AIM: Amending Inherent Interpretability via Self-Supervised MaskingPoster
  57. AIRA: Activation-Informed Low-Rank Adaptation for Large ModelsPoster
  58. AJAHR: Amputated Joint Aware 3D Human Mesh RecoveryPoster
  59. ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow PredictionsPoster
  60. AM-Adapter: Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-WildPoster
  61. AMD: Adaptive Momentum and Decoupled Contrastive Learning Framework for Robust Long-Tail Trajectory PredictionPoster
  62. AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and SegmentationPoster
  63. AR-1-to-3: Single Image to Consistent 3D Object via Next-View PredictionPoster
  64. AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical ReasoningPoster
  65. ARGUS: Hallucination and Omission Evaluation in Video-LLMsPoster
  66. ARIG: Autoregressive Interactive Head Generation for Real-time ConversationsPoster
  67. ARMO: Autoregressive Rigging for Multi-Category ObjectsPoster
  68. ART: Adaptive Relation Tuning for Generalized Relation PredictionPoster
  69. ASCENT: Annotation-free Self-supervised Contrastive Embeddings for 3D Neuron Tracking in Fluorescence MicroscopyPoster
  70. ASGS: Single-Domain Generalizable Open-Set Object Detection via Adaptive Subgraph SearchingPoster
  71. ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense PredictionPoster
  72. ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingPoster
  73. ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human ModelingPoster
  74. AU-Blendshape for Fine-grained Stylized 3D Facial Expression ManipulationPoster
  75. AURELIA: Test-time Reasoning Distillation in Audio-Visual LLMsPoster
  76. AV-Flow: Transforming Text to Audio-Visual Human-like InteractionsPoster
  77. AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video GenerationPoster
  78. AVAM: a Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question AnsweringPoster
  79. AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMsPoster
  80. AcZeroTS: Active Learning for Zero-shot Tissue Segmentation in Pathology ImagesPoster
  81. Accelerate 3D Object Detection Models via Zero-Shot Attention Key PruningPoster
  82. Accelerating Diffusion Sampling via Exploiting Local Transition CoherencePoster
  83. AccidentalGS: 3D Gaussian Splatting from Accidental Camera MotionPoster
  84. Achieving More with Less: Additive Prompt Tuning for Rehearsal-Free Class-Incremental LearningPoster
  85. Acknowledging Focus Ambiguity in Visual QuestionsPoster
  86. Activation Subspaces for Out-of-Distribution DetectionPoster
  87. Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object DetectionPoster
  88. Active Membership Inference Test (aMINT): Enhancing Model Auditability with Multi-Task Learning.Poster
  89. Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex EnvironmentsPoster
  90. AdaDCP: Learning an Adapter with Discrete Cosine Prior for Clear-to-Adverse Domain GeneralizationPoster
  91. AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous DrivingPoster
  92. AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview DiffusionPoster
  93. Adapt Foundational Segmentation Models with Heterogeneous Searching SpacePoster
  94. Adapting In-Domain Few-Shot Segmentation to New Domains without Source Domain RetrainingPoster
  95. Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak SupervisionPoster
  96. Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part GroundingPoster
  97. Adaptive Caching for Faster Video Generation with Diffusion TransformersPoster
  98. Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time ShiftsPoster
  99. Adaptive Hyper-Graph Convolution Network for Skeleton-based Human Action Recognition with Virtual ConnectionsPoster
  100. Adaptive Learning of High-Value Regions for Semi-Supervised Medical Image SegmentationPoster
  101. Adaptive Prompt Learning via Gaussian Outlier Synthesis for Out-of-distribution DetectionPoster
  102. Adaptive Routing of Text-to-Image Generation Requests Between Large Cloud Model and Light-Weight Edge ModelPoster
  103. AdaptiveAE: An Adaptive Exposure Strategy for HDR Capturing in Dynamic ScenesPoster
  104. Adding Additional Control to One-Step Diffusion with Joint Distribution MatchingPoster
  105. Addressing Representation Collapse in Vector Quantized Models with One Linear LayerPoster
  106. Addressing Text Embedding Leakage in Diffusion-based Image Editing
  107. AdsQA: Towards Advertisement Video UnderstandingPoster
  108. AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?Poster
  109. Advancing Text-to-3D Generation with Linearized Lookahead Variational Score DistillationPoster
  110. Advancing Textual Prompt Learning with Anchored AttributesPoster
  111. Advancing Visual Large Language Model for Multi-granular Versatile PerceptionPoster
  112. Adversarial Attention Perturbations for Large Object Detection TransformersPoster
  113. Adversarial Data Augmentation for Single Domain Generalization via Lyapunov Exponent-Guided OptimizationPoster
  114. Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video SynthesisPoster
  115. Adversarial Exploitation of Data Diversity Improves Visual LocalizationPoster
  116. Adversarial Purification via Super-Resolution and DiffusionPoster
  117. Adversarial Reconstruction Feedback for Robust Fine-grained GeneralizationPoster
  118. Adversarial Robust Memory-Based Continual LearnerPoster
  119. Adversarial Robustness of Discriminative Self-Supervised Learning in VisionPoster
  120. Adversarial Training for Probabilistic RobustnessPoster
  121. AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional RelationsPoster
  122. Aether: Geometric-Aware Unified World ModelingPoster
  123. AffordDexGrasp: Open-set Language-guided Dexterous Grasp with Generalizable-Instructive AffordancePoster
  124. After the Party: Navigating the Mapping From Color to Ambient LightingPoster
  125. Agreement aware and dissimilarity oriented GLOMPoster
  126. AgroBench: Vision-Language Model Benchmark in AgriculturePoster
  127. AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model InferencePoster
  128. Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature RepresentationPoster
  129. AlignDiff: Learning Physically-Grounded Camera Alignment via DiffusionPoster
  130. AlignGuard: Scalable Safety Alignment for Text-to-Image GenerationPoster
  131. Aligning Constraint Generation with Design Intent in Parametric CADPoster
  132. Aligning Effective Tokens with Video Anomaly in Large Language ModelsPoster
  133. Aligning Global Semantics and Local Textures in Generative Video EnhancementPoster
  134. Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text MatchingPoster
  135. Aligning Moments in Time using Video QueriesPoster
  136. Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs ReasoningPoster
  137. All Parts Matter: A Unified Mask-Free Virtual Try-On FrameworkPoster
  138. All in One: Visual-Description-Guided Unified Point Cloud SegmentationPoster
  139. AllGCD: Leveraging All Unlabeled Data for Generalized Category DiscoveryPoster
  140. AllTracker: Efficient Dense Point Tracking at High ResolutionPoster
  141. Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic ApproximationPoster
  142. Allowing Oscillation Quantization: Overcoming Solution Space Limitation in Low Bit-Width QuantizationPoster
  143. Always Skip AttentionPoster
  144. Amodal Depth Anything: Amodal Depth Estimation in the WildPoster
  145. Amodal3R: Amodal 3D Reconstruction from Occluded 2D ImagesPoster
  146. An Efficient Hybrid Vision Transformer for TinyML ApplicationsPoster
  147. An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image RetrievalPoster
  148. An Empirical Study of Autoregressive Pre-training from VideosPoster
  149. An Information-Theoretic Regularizer for Lossy Neural Image CompressionPoster
  150. An Inversion-based Measure of Memorization for Diffusion ModelsPoster
  151. An OpenMind for 3D Medical Vision Self-supervised LearningPoster
  152. Analyzing Finetuning Representation Shift for Multimodal LLMs SteeringPoster
  153. Anchor Token Matching: Implicit Structure Locking for Training-free AR Image EditingPoster
  154. AnimalClue: Recognizing Animals by their TracesPoster
  155. Animate Anyone 2: High-Fidelity Character Image Animation with Environment AffordancePoster
  156. AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh AnimationPoster
  157. AnimeGamer: Infinite Anime Life Simulation with Next Game State PredictionPoster
  158. AnnofreeOD: Detecting All Classes at Low Frame Rates Without Human AnnotationsPoster
  159. Anomaly Detection of Integrated Circuits Package Substrates Using the Large Vision Model SAIC: Dataset Construction, Methodology, and ApplicationPoster
  160. Anti-Tamper Protection for Unauthorized Individual Image GenerationPoster
  161. Any-SSR: How Recursive Least Squares Works in Continual Learning of Large Language ModelPoster
  162. Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing TasksPoster
  163. AnyBimanual: Transferring Unimanual Policy for General Bimanual ManipulationPoster
  164. AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera CalibrationPoster
  165. AnyI2V: Animating Any Conditional Image with Motion ControlPoster
  166. AnyPortal: Zero-Shot Consistent Video Background ReplacementPoster
  167. ArchiSet: Benchmarking Editable and Consistent Single-View 3D Reconstruction of Buildings with Specific Window-to-Wall RatiosPoster
  168. Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsPoster
  169. Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data and Metric PerspectivesPoster
  170. ArgMatch: Adaptive Refinement Gathering for Efficient Dense MatchingPoster
  171. ArgoTweak: Towards Self-Updating HD Maps through Structured PriorsPoster
  172. ArtEditor: Learning Customized Instructional Image Editor from Few-Shot ExamplesPoster
  173. Arti-PG: A Toolbox for Procedurally Synthesizing Large-Scale and Diverse Articulated Objects with Rich AnnotationsPoster
  174. Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene DescriptionPoster
  175. Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question AnsweringPoster
  176. AstroLoc: Robust Space to Ground Image LocalizerPoster
  177. Asynchronous Event Error-Minimizing Noise for Safeguarding Event DatasetPoster
  178. Att-Adapter: A Robust and Precise Domain-Specific Multi-Attributes T2I Diffusion Adapter via Conditional Variational AutoencoderPoster
  179. Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!Poster
  180. Attention to Trajectory: Trajectory-Aware Open-Vocabulary TrackingPoster
  181. Attention to the Burstiness in Visual Prompt Tuning!Poster
  182. Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationPoster
  183. Augmented and Softened Matching for Unsupervised Visible-Infrared Person Re-IdentificationPoster
  184. Augmenting Moment Retrieval: Zero-Dependency Two-Stage LearningPoster
  185. Authentic 4D Driving Simulation with a Video Generation ModelPoster
  186. Auto-Controlled Image Perception in MLLMs via Visual Perception TokensPoster
  187. Auto-Regressive Transformation for Image AlignmentPoster
  188. Auto-Regressively Generating Multi-View Consistent ImagesPoster
  189. Auto-Vocabulary Semantic SegmentationPoster
  190. AutoComPose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsPoster
  191. AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian SplattingPoster
  192. AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial PromptsPoster
  193. AutoScape: Geometry-Consistent Long-Horizon Scene GenerationPoster
  194. Automated Model Evaluation for Object Detection via Prediction Consistency and ReliabilityPoster
  195. Automated Red Teaming for Text-to-Image Models through Feedback-Guided Prompt Iteration with Vision-Language ModelsPoster
  196. Autoregressive Denoising Score Matching is a Good Video Anomaly DetectorPoster
  197. Auxiliary Prompt Tuning of Vision-Language Models for Few-Shot Out-of-Distribution DetectionPoster
  198. Avat3r: Large Animatable Gaussian Reconstruction Model for High-fidelity 3D Head AvatarsPoster
  199. Axis-level Symmetry Detection with Group-Equivariant RepresentationPoster
  200. B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal TokensPoster
  201. BANet: Bilateral Aggregation Network for Mobile Stereo MatchingPoster
  202. BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language ModelsPoster
  203. BATCLIP: Bimodal Online Test-Time Adaptation for CLIPPoster
  204. BUFFER-X: Towards Zero-Shot Point Cloud Registration in Diverse ScenesPoster
  205. BVINet: Unlocking Blind Video Inpainting with Zero AnnotationsPoster
  206. BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant LearningPoster
  207. Back on Track: Bundle Adjustment for Dynamic Scene ReconstructionPoster
  208. Backdoor Attacks on Neural Networks via One-Bit FlipPoster
  209. Backdoor Defense via Enhanced Splitting and Trap IsolationPoster
  210. Backdoor Mitigation by Distance-Driven DetoxificationPoster
  211. Backdooring Self-Supervised Contrastive Learning by Noisy AlignmentPoster
  212. Background Invariance Testing According to Semantic ProximityPoster
  213. BadVideo: Stealthy Backdoor Attack against Text-to-Video GenerationPoster
  214. Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and ReconstructionPoster
  215. Balanced Image Stylization with Style Matching ScorePoster
  216. Balanced Sharpness-Aware Minimization for Imbalanced RegressionPoster
  217. Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot SegmentationPoster
  218. Balancing Task-invariant Interaction and Task-specific Adaptation for Unified Image FusionPoster
  219. Bayesian-Inspired Space-Time SuperpixelsPoster
  220. Benchmarking Burst Super-Resolution for Polarization Images: Noise Dataset and AnalysisPoster
  221. Benchmarking Egocentric Visual-Inertial SLAM at City ScalePoster
  222. Benchmarking Multimodal CoT Reward Model Stepwise by Visual ProgramPoster
  223. Benchmarking Multimodal Large Language Models Against Image CorruptionsPoster
  224. Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-to-3D GenerationPoster
  225. Benefit From Seen: Enhancing Open-Vocabulary Object Detection by Bridging Visual and Textual Co-Occurrence KnowledgePoster
  226. Beyond Blur: A Fluid Perspective on Generative Diffusion ModelsPoster
  227. Beyond Brain Decoding: Visual-Semantic Reconstructions to Mental Creation Extension Based on fMRIPoster
  228. Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line GenerationPoster
  229. Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action RecognitionPoster
  230. Beyond Losses Reweighting: Empowering Multi-Task Learning via the Generalization PerspectivePoster
  231. Beyond Low-Rank Tuning: Model Prior-Guided Rank Allocation for Effective Transfer in Low-Data and Large-Gap Regimes.Poster
  232. Beyond Next-Token: Next-X Prediction for Autoregressive Visual GenerationPoster
  233. Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR RepresentationsPoster
  234. Beyond Perspective: Neural 360-Degree Video CompressionPoster
  235. Beyond RGB: Adaptive Parallel Processing for RAW Object DetectionPoster
  236. Beyond Simple Edits: Composed Video Retrieval with Dense ModificationsPoster
  237. Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object DetectionPoster
  238. Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video DetectionPoster
  239. Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMsPoster
  240. Beyond Training: Dynamic Token Merging for Zero-Shot Video UnderstandingPoster
  241. Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly SearchPoster
  242. Beyond [cls]: Exploring the True Potential of Masked Image Modeling RepresentationsPoster
  243. Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question AnsweringPoster
  244. Beyond the Frame: Generating 360deg Panoramic Videos from Perspective VideosPoster
  245. Beyond the Limits: Overcoming Negative Correlation of Activation-Based Training-Free NASPoster
  246. BezierGS: Dynamic Urban Scene Reconstruction with Bezier Curve Gaussian SplattingPoster
  247. Bi-Level Optimization for Self-Supervised AI-Generated Face DetectionPoster
  248. Bias in Gender Bias Benchmarks: How Spurious Features Distort EvaluationPoster
  249. Bias-Resilient Weakly Supervised Semantic Segmentation Using Normalizing FlowsPoster
  250. Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video RetrievalPoster
  251. Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction DetectionPoster
  252. BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View SynthesisPoster
  253. Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoPoster
  254. Blended Point Cloud Diffusion for Localized Text-guided Shape EditingPoster
  255. Blind Noisy Image Deblurring Using Residual Guidance StrategyPoster
  256. Blind Video Super-Resolution based on Implicit KernelsPoster
  257. Blind2Sound: Self-Supervised Image Denoising without Residual NoisePoster
  258. BlinkTrack: Feature Tracking over 80 FPS via Events and ImagesPoster
  259. BlueNeg: A 35mm Negative Film Dataset for Restoring Channel-Heterogeneous DeteriorationPoster
  260. BokehDiff: Neural Lens Blur with One-Step DiffusionPoster
  261. Bokehlicious: Photorealistic Bokeh Rendering with Controllable AperturesPoster
  262. Bolt3D: Generating 3D Scenes in SecondsPoster
  263. Boost 3D Reconstruction using Diffusion-based Monocular Camera CalibrationPoster
  264. Boosting Adversarial Transferability via Negative Hessian Trace RegularizationPoster
  265. Boosting Adversarial Transferability via Residual Perturbation AttackPoster
  266. Boosting Class Representation via Semantically Related Instances for Robust Long-Tailed Learning with Noisy LabelsPoster
  267. Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and TransferabilityPoster
  268. Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer FeaturesPoster
  269. Boosting MLLM Reasoning with Text-Debiased Hint-GRPOPoster
  270. Boosting Multi-View Indoor 3D Object Detection via Adaptive 3D Volume ConstructionPoster
  271. Boosting Multimodal Learning via Disentangled Gradient LearningPoster
  272. Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-trainingPoster
  273. Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic DataPoster
  274. Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model AdaptationPoster
  275. Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection via Cross-Domain Retrieval AugmentationPoster
  276. Boundary Probing for Input Privacy Protection When Using LMM ServicesPoster
  277. BoxDreamer: Dreaming Box Corners for Generalizable Object Pose EstimationPoster
  278. Breaking Grid Constraints: Dynamic Graph Reconstruction Network for Multi-organ SegmentationPoster
  279. Breaking Rectangular Shackles: Cross-View Object Segmentation for Fine-Grained Object Geo-LocalizationPoster
  280. Breaking the Encoder Barrier for Seamless Video-Language UnderstandingPoster
  281. BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent AlignmentPoster
  282. Bridging 3D Anomaly Localization and Repair via High-Quality Continuous Geometric RepresentationPoster
  283. Bridging Class Imbalance and Partial Labeling via Spectral-Balanced Energy Propagation for Skeleton-based Action RecognitionPoster
  284. Bridging Continuous and Discrete Tokens for Autoregressive Visual GenerationPoster
  285. Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution FrameworkPoster
  286. Bridging Domain Generalization to Multimodal Domain Generalization via Unified RepresentationsPoster
  287. Bridging Local Inductive Bias and Long-Range Dependencies with Pixel-Mamba for End-to-end Whole Slide Image AnalysisPoster
  288. Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging ScenariosPoster
  289. Bridging the Gap between Brain and Machine in Interpreting Visual Semantics: Towards Self-adaptive Brain-to-Text DecodingPoster
  290. Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action RecognitionPoster
  291. Bridging the Sky and Ground: Towards View-Invariant Feature Learning for Aerial-Ground Person Re-IdentificationPoster
  292. Bring Your Rear Cameras for Egocentric 3D Human Pose EstimationPoster
  293. Bringing RNNs Back to Efficient Open-Ended Video UnderstandingPoster
  294. C2MIL: Synchronizing Semantic and Topological Causalities in Multiple Instance Learning for Robust and Interpretable Survival AnalysisPoster
  295. C4D: 4D Made from 3D through Dual CorrespondencesPoster
  296. CA-I2P: Channel-Adaptive Registration Network with Global Optimal SelectionPoster
  297. CA2C: A Prior-Knowledge-Free Approach for Robust Label Noise Learning via Asymmetric Co-learning and Co-trainingPoster
  298. CABLD: Contrast-Agnostic Brain Landmark Detection with Consistency-Based RegularizationPoster
  299. CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task SolversPoster
  300. CAD-Recode: Reverse Engineering CAD Code from Point CloudsPoster
  301. CAFA: a Controllable Automatic Foley ArtistPoster
  302. CAP: Evaluation of Persuasive and Creative Image GenerationPoster
  303. CAPTURE: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object CountingPoster
  304. CARIM: Caption-Based Autonomous Driving Scene Retrieval via Inclusive Text MatchingPoster
  305. CARL: Causality-guided Architecture Representation Learning for an Interpretable Performance PredictorPoster
  306. CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
  307. CAT: A Unified Click-and-Track Framework for Realistic TrackingPoster
  308. CATP-LLM: Empowering Large Language Models for Cost-Aware Tool PlanningPoster
  309. CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View ImagePoster
  310. CAVIS: Context-Aware Video Instance SegmentationPoster
  311. CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in LiteracyPoster
  312. CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian SplattingPoster
  313. CCMNet: Leveraging Calibrated Color Correction Matrices for Cross-Camera Color ConstancyPoster
  314. CE-FAM: Concept-Based Explanation via Fusion of Activation MapsPoster
  315. CF3: Compact and Fast 3D Feature FieldsPoster
  316. CHARM3R: Towards Unseen Camera Height Robust Monocular 3D DetectorPoster
  317. CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE SolversPoster
  318. CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single ImagePoster
  319. CIARD: Cyclic Iterative Adversarial Robustness DistillationPoster
  320. CL-Splats: Continual Learning of Gaussian Splatting with Local OptimizationPoster
  321. CLIP-Adapted Region-to-Text Learning for Generative Open-Vocabulary Semantic SegmentationPoster
  322. CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian SplattingPoster
  323. CLIPSym: Delving into Symmetry Detection with CLIPPoster
  324. CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic SegmentationPoster
  325. CLOT: Closed Loop Optimal Transport for Unsupervised Action SegmentationPoster
  326. CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing ModalitiesPoster
  327. CMB-ML: A Cosmic Microwave Background Dataset for the Oldest Possible Computer Vision TaskPoster
  328. CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD GenerationPoster
  329. CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance ShiftsPoster
  330. CO2-Net: A Physics-Informed Spatio-Temporal Model for Global Surface CO2 ReconstructionPoster
  331. CODA: Repurposing Continuous VAEs for Discrete TokenizationPoster
  332. CODE-CL: Conceptor-Based Gradient Projection for Deep Continual LearningPoster
  333. COIN: Confidence Score-Guided Distillation for Annotation-Free Cell SegmentationPoster
  334. COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound DatasetsPoster
  335. COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language NavigationPoster
  336. COSTARR: Consolidated Open Set Technique with Attenuation for Robust RecognitionPoster
  337. CObL: Toward Zero-Shot Ordinal Layering without User PromptingPoster
  338. CRAM: Large Scale Video Continual Learning with Bootstrapped CompressionPoster
  339. CSD-VAR: Content-Style Decomposition in Visual Autoregressive ModelsPoster
  340. CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath ModelingPoster
  341. CULTURE3D: A Large-Scale and Diverse Dataset of Cultural Landmarks and Terrains for Gaussian-Based Scene RenderingPoster
  342. CVPT: Cross Visual Prompt TuningPoster
  343. CWNet: Causal Wavelet Network for Low-Light Image EnhancementPoster
  344. CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset DistillationPoster
  345. CaliMatch: Adaptive Calibration for Improving Safe Semi-supervised LearningPoster
  346. Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt EnsemblesPoster
  347. CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language ModelPoster
  348. CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion ModelsPoster
  349. Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?Poster
  350. Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge EditingPoster
  351. Can We Achieve Efficient Diffusion Without Self-Attention? Distilling Self-Attention into ConvolutionsPoster
  352. Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D GaussiansPoster
  353. CanFields: Consolidating Diffeomorphic Flows for Non-Rigid 4D Interpolation from Arbitrary-Length SequencesPoster
  354. CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space ModulationPoster
  355. CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language ModelsPoster
  356. CaptionSmiths: Flexibly Controlling Language Pattern in Image CaptioningPoster
  357. Capturing head avatar with hand contacts from a monocular videoPoster
  358. CarGait: Cross-Attention based Re-ranking for Gait recognitionPoster
  359. CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for GuidancePoster
  360. Cassic: Towards Content-Adaptive State-Space Models for Learned Image CompressionPoster
  361. Category-Specific Selective Feature Enhancement for Long-Tailed Multi-Label Image ClassificationPoster
  362. Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot LearningPoster
  363. Causal-Entity Reflected Egocentric Traffic Accident Video SynthesisPoster
  364. Causality-guided Prompt Learning for Vision-language Models via Visual GranulationPoster
  365. Certifiably Optimal Anisotropic Rotation AveragingPoster
  366. CharaConsist: Fine-Grained Consistent Character GenerationPoster
  367. ChartCap: Mitigating Hallucination of Dense Chart CaptioningPoster
  368. ChartPoint: Guiding MLLMs with Grounding Reflection for Chart ReasoningPoster
  369. ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language ModelsPoster
  370. Chimera: Improving Generalist Model with Domain-Specific ExpertsPoster
  371. CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene ReconstructionPoster
  372. CityNav: A Large-Scale Dataset for Real-World Aerial NavigationPoster
  373. ClaraVid: A Holistic Scene Reconstruction Benchmark From Aerial Perspective With Delentropy-Based Complexity ProfilingPoster
  374. Class-Wise Federated Averaging for Efficient PersonalizationPoster
  375. CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge DistillationPoster
  376. ClearSight: Human Vision-Inspired Solutions for Event-Based Motion DeblurringPoster
  377. Client2Vec: Improving Federated Learning by Distribution Shifts Aware Client IndexingPoster
  378. Clink! Chop! Thud! - Learning Object Sounds from Real-World InteractionsPoster
  379. Closed-Loop Transfer for Weakly-supervised Affordance GroundingPoster
  380. Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive InjectionPoster
  381. CoA-VLA: Improving Vision-Language-Action Models via Visual-Text Chain-of-AffordancePoster
  382. CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous DrivingPoster
  383. CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression SegmentationPoster
  384. CoLMDriver: LLM-based Negotiation Benefits Cooperative Autonomous DrivingPoster
  385. CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion ModelsPoster
  386. CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image MatchingPoster
  387. CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred ImagesPoster
  388. CoSMIC: Continual Self-supervised Learning for Multi-Domain Medical Imaging via Conditional Mutual Information MaximizationPoster
  389. CoST: Efficient Collaborative Perception From Unified Spatiotemporal PerspectivePoster
  390. CoStoDet-DDPM: Collaborative Training of Stochastic and Deterministic Models Improves Surgical Workflow Anticipation and RecognitionPoster
  391. CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real VideosPoster
  392. CogCM: Cognition-Inspired Contextual Modeling for Audio-Visual Speech EnhancementPoster
  393. CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMsPoster
  394. Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent DialoguesPoster
  395. Color Matching Using Hypernetwork-Based Kolmogorov-Arnold NetworksPoster
  396. Colors See Colors Ignore: Clothes Changing ReID with Color DisentanglementPoster
  397. CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing GamesPoster
  398. Combinative Matching for Geometric Shape AssemblyPoster
  399. Communication-Efficient Multi-Vehicle Collaborative Semantic Segmentation via Sparse 3D Gaussian SharingPoster
  400. CompCap: Improving Multimodal Large Language Models with Composite CaptionsPoster
  401. CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image GenerationPoster
  402. Competitive Distillation: A Simple Learning Strategy for Improving Visual ClassificationPoster
  403. CompleteMe: Reference-based Human Image CompletionPoster
  404. Compression of 3D Gaussian Splatting with Optimized Feature Planes and Standard Video CodecsPoster
  405. Compression-Aware One-Step Diffusion Model for JPEG Artifact RemovalPoster
  406. ConceptSplit: Decoupled Multi-Concept Personalization of Diffusion Models via Token-wise Adaptation and Attention DisentanglementPoster
  407. Conditional Latent Diffusion Models for Zero-Shot Instance SegmentationPoster
  408. Conditional Visual Autoregressive Modeling for Pathological Image RestorationPoster
  409. ConformalSAM: Unlocking the Potential of Foundational Segmentation Models in Semi-Supervised Semantic Segmentation with Conformal PredictionPoster
  410. Confound from All Sides, Distill with Resilience: Multi-Objective Adversarial Paths to Zero-Shot RobustnessPoster
  411. ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using ConstraintsPoster
  412. Consensus-Driven Active Model SelectionPoster
  413. Consistency Trajectory Matching for One-Step Generative Super-ResolutionPoster
  414. ConsistentCity: Semantic Flow-guided Occupancy DiT for Temporally Consistent Driving Scene SynthesisPoster
  415. ConstStyle: Robust Domain Generalization with Unified Style TransformationPoster
  416. Constraint-Aware Feature Learning for Parametric Point CloudPoster
  417. Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain ReasoningPoster
  418. Contact-Aware Amodal Completion for Human-Object Interaction via Multi-Regional InpaintingPoster
  419. Contact-Aware Refinement of Human Pose Pseudo-Ground Truth via Bioimpedance SensingPoster
  420. Context Guided Transformer Entropy Modeling for Video CompressionPoster
  421. Context-Aware Academic Emotion Dataset and BenchmarkPoster
  422. ContextFace: Generating Facial Expressions from Emotional ContextsPoster
  423. Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic ScenariosPoster
  424. Continual Multiple Instance Learning with Enhanced Localization for Histopathological Whole Slide Image AnalysisPoster
  425. Continual Personalization for Diffusion ModelsPoster
  426. Continuous-Time Human Motion Field from Event CamerasPoster
  427. ContraGS: Codebook-Condensed and Trainable Gaussian Splatting for Fast, Memory-Efficient Reconstruction
  428. Contrastive Flow MatchingPoster
  429. Contrastive Test-Time Composition of Multiple LoRA Models for Image GenerationPoster
  430. Controllable 3D Outdoor Scene Generation via Scene GraphsPoster
  431. Controllable Feature Whitening for Hyperparameter-Free Bias MitigationPoster
  432. Controllable Latent Space Augmentation for Digital PathologyPoster
  433. Controllable Weather Synthesis and Removal with Video Diffusion ModelsPoster
  434. Controllable and Expressive One-Shot Video Head SwappingPoster
  435. Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-ExpertsPoster
  436. Controlling Multimodal LLMs via Reward-guided DecodingPoster
  437. CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential PerceptionPoster
  438. Cooperative Pseudo Labeling for Unsupervised Federated ClassificationPoster
  439. Coordinate-based Speed of Sound Recovery for Aberration-Corrected Photoacoustic Computed TomographyPoster
  440. CopyrightShield: Enhancing Diffusion Model Security Against Copyright Infringement AttacksPoster
  441. CoralSRT: Revisiting Coral Reef Semantic Segmentation by Feature Rectification via Self-supervised GuidancePoster
  442. CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic SegmentationPoster
  443. Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the WildPoster
  444. Correspondence-Free Fast and Robust Spherical Point Pattern RegistrationPoster
  445. Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought ReasoningPoster
  446. CountSE: Soft Exemplar Open-set Object CountingPoster
  447. CounterPC: Counterfactual Feature Realignment for Unsupervised Domain Adaptation on Point CloudsPoster
  448. Counting Stacked ObjectsPoster
  449. Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image AnalysisPoster
  450. CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationPoster
  451. Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLMsPoster
  452. Cross-Architecture Distillation Made Simple with Redundancy SuppressionPoster
  453. Cross-Category Subjectivity Generalization for Style-Adaptive Sketch Re-IDPoster
  454. Cross-Subject Mind Decoding from Inaccurate RepresentationsPoster
  455. Cross-View Isolated Sign Language Recognition via View Synthesis and Feature DisentanglementPoster
  456. Cross-modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and MethodPoster
  457. CryoFastAR: Fast Cryo-EM Ab initio Reconstruction Made EasyPoster
  458. CuMPerLay: Learning Cubical Multiparameter Persistence VectorizationsPoster
  459. CuRe: Cultural Gaps in the Long Tail of Text-to-Image SystemsPoster
  460. Curve-Aware Gaussian Splatting for 3D Parametric Curve ReconstructionPoster
  461. Customizing Domain Adapters for Domain GeneralizationPoster
  462. CutS3D: Cutting Semantics in 3D for 2D Unsupervised Instance SegmentationPoster
  463. Cycle Consistency as Reward: Learning Image-Text Alignment without Human PreferencesPoster
  464. Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object DetectionPoster
  465. CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image TranslationPoster
  466. D-Attn: Decomposed Attention for Large Vision-and-Language ModelPoster
  467. D2ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action RecognitionPoster
  468. D3: Training-Free AI-Generated Video Detection Using Second-Order FeaturesPoster
  469. D3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image DetectionPoster
  470. DAA*: Deep Angular A Star for Image-based Path PlanningPoster
  471. DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference ImagesPoster
  472. DADM: Dual Alignment of Domain and Modality for Face Anti-spoofingPoster
  473. DADet: Safeguarding Image Conditional Diffusion Models against Adversarial and Backdoor Attacks via Diffusion Anomaly DetectionPoster
  474. DALIP: Distribution Alignment-based Language-Image Pre-Training for Domain-Specific DataPoster
  475. DAMap: Distance-aware MapNet for High Quality HD Map ConstructionPoster
  476. DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain LearningPoster
  477. DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene RenderingPoster
  478. DASH: Detection and Assessment of Systematic Hallucinations of VLMsPoster
  479. DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative PerceptionPoster
  480. DAViD: Data-efficient and Accurate Vision Models from Synthetic DataPoster
  481. DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion ModelsPoster
  482. DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent SpacePoster
  483. DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid TokenizerPoster
  484. DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion ModelsPoster
  485. DC-TTA: Divide-and-Conquer Framework for Test-Time Adaptation of Interactive SegmentationPoster
  486. DCHM: Depth-Consistent Human Modeling for Multiview DetectionPoster
  487. DCT-Shield: A Robust Frequency Domain Defense against Malicious Image EditingPoster
  488. DDB: Diffusion Driven Balancing to Address Spurious CorrelationsPoster
  489. DEPTHOR: Depth Enhancement from a Practical Light-Weight dToF Sensor and RGB ImagePoster
  490. DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking HeadsPoster
  491. DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video GenerationPoster
  492. DIA: The Adversarial Exposure of Deterministic Inversion in Diffusion ModelsPoster
  493. DICE: Staleness-Centric Optimizations for Parallel Diffusion MoE InferencePoster
  494. DIH-CLIP: Unleashing the Diversity of Multi-Head Self-Attention for Training-Free Open-Vocabulary Semantic SegmentationPoster
  495. DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative ModelsPoster
  496. DIMO: Diverse 3D Motion Generation for Arbitrary ObjectsPoster
  497. DIP: Unsupervised Dense In-Context Post-training of Visual RepresentationsPoster
  498. DISTA-Net: Dynamic Closely-Spaced Infrared Small Target UnmixingPoster
  499. DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent DiffusionPoster
  500. DIVE: Taming DINO for Subject-Driven Video EditingPoster
  501. DLF: Extreme Image Compression with Dual-generative Latent FusionPoster
  502. DLFR-Gen: Diffusion-based Video Generation with Dynamic Latent Frame RatePoster
  503. DM-EFS: Dynamically Multiplexed Expanded Features Set Form for Robust and Efficient Small Object DetectionPoster
  504. DMQ: Dissecting Outliers of Diffusion Models for Post-Training QuantizationPoster
  505. DMesh++: An Efficient Differentiable Mesh for Complex ShapesPoster
  506. DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse RenderingPoster
  507. DOGR: Towards Versatile Visual Document Grounding and ReferringPoster
  508. DOLLAR: Few-Step Video Generation via Distillation and Latent Reward OptimizationPoster
  509. DONUT: A Decoder-Only Model for Trajectory PredictionPoster
  510. DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose PriorPoster
  511. DRaM-LHM: A Quaternion Framework for Iterative Camera Pose EstimationPoster
  512. DSO: Aligning 3D Generators with Simulation Feedback for Physical SoundnessPoster
  513. DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking TuningPoster
  514. DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary DescriptionsPoster
  515. Dark-ISP: Enhancing RAW Image Processing for Low-Light Object DetectionPoster
  516. Dataset Distillation as Data Compression: A Rate-Utility PerspectivePoster
  517. Dataset Distillation via Vision-Language Category PrototypePoster
  518. Dataset Distillation via the Wasserstein MetricPoster
  519. Dataset Ownership Verification for Pre-trained Masked ModelsPoster
  520. DeFSS: Image-to-Mask Denoising Learning for Few-shot SegmentationPoster
  521. DeGauss: Dynamic-Static Decomposition with Gaussian Splatting for Distractor-free 3D ReconstructionPoster
  522. DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback SynergyPoster
  523. DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity UnderstandingPoster
  524. Debiased Curriculum Adaptation for Safe Transfer Learning in Chest X-ray ClassificationPoster
  525. Debiasing Trace Guidance: Top-down Trace Distillation and Bottom-up Velocity Alignment for Unsupervised Anomaly DetectionPoster
  526. DecAD: Decoupling Anomalies in Latent Space for Multi-Class Unsupervised Anomaly DetectionPoster
  527. Deciphering Cross-Modal Alignment in Large Vision-Language Models via Modality Integration RatePoster
  528. Decoding Correlation-Induced Misalignment in the Stable Diffusion Workflow for Text-to-Image GenerationPoster
  529. Decouple and Track: Benchmarking and Improving Video Diffusion Transformers For Motion TransferPoster
  530. Decouple to Reconstruct: High Quality UHD Restoration via Active Feature Disentanglement and Reversible FusionPoster
  531. Decoupled Diffusion Sparks Adaptive Scene GenerationPoster
  532. Decoupled Multi-Predictor Optimization for Inference-Efficient Model TuningPoster
  533. Deep Adaptive Unfolded Network via Spatial Morphology Stripping and Spectral Filtration for Pan-sharpeningPoster
  534. Deep Incomplete Multi-view Clustering with Distribution Dual-Consistency Recovery GuidancePoster
  535. Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength ImagesPoster
  536. DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement LearningPoster
  537. DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery AnalysisPoster
  538. Deeply Supervised Flow-Based Generative ModelsPoster
  539. Degradation-Modeled Multipath Diffusion for Tunable Metalens PhotographyPoster
  540. Demeter: A Parametric Model of Crop Plant Morphology from the Real WorldPoster
  541. Democratizing High-Fidelity Co-Speech Gesture Video GenerationPoster
  542. Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional TokensPoster
  543. Denoising Token Prediction in Masked Autoregressive ModelsPoster
  544. Dense Policy: Bidirectional Autoregressive Learning of ActionsPoster
  545. Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image GenerationPoster
  546. DepR: Depth Guided Single-view Scene Reconstruction with Instance-level DiffusionPoster
  547. Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationPoster
  548. Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth EstimationPoster
  549. DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationPoster
  550. Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for DermatologyPoster
  551. Describe Anything: Detailed Localized Image and Video CaptioningPoster
  552. Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object RetrievalPoster
  553. Describe, Don't Dictate: Semantic Image Editing with Natural Language IntentPoster
  554. Details Matter for Indoor Open-vocabulary 3D Instance SegmentationPoster
  555. Detect Anything 3D in the WildPoster
  556. Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous CirclePoster
  557. Deterministic Object Pose Confidence Region EstimationPoster
  558. Devil is in the Uniformity: Exploring Diverse Learners within Transformer for Image RestorationPoster
  559. DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-to-Robot HandoverPoster
  560. DexVLG: Dexterous Vision-Language-Grasp Model at ScalePoster
  561. DiGA3D: Coarse-to-Fine Diffusional Propagation of Geometry and Appearance for Versatile 3D InpaintingPoster
  562. DiMPLe - Disentangled Multi-Modal Prompt Learning: Enhancing Out-Of-Distribution Alignment with Invariant and Spurious Feature SeparationPoster
  563. DiSCO-3D : Discovering and Segmenting Sub-Concepts from Open-vocabulary Queries in NeRFPoster
  564. DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationPoster
  565. DiT4SR: Taming Diffusion Transformer for Real-World Image Super-ResolutionPoster
  566. DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion TransformersPoster
  567. DiTaiListener: Controllable High Fidelity Listener Video Generation with DiffusionPoster
  568. Di[M]O: Distilling Masked Diffusion Models into One-step GeneratorPoster
  569. Diagnosing Pretrained Models for Out-of-distribution DetectionPoster
  570. DialNav: Multi-turn Dialog Navigation with a Remote GuidePoster
  571. DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary LookupPoster
  572. Diff2I2P: Differentiable Image-to-Point Cloud Registration with Diffusion PriorPoster
  573. DiffDoctor: Diagnosing Image Diffusion Models Before TreatingPoster
  574. DiffIP: Representation Fingerprints for Robust IP Protection of Diffusion ModelsPoster
  575. DiffRefine: Diffusion-based Proposal Specific Point Cloud Densification for Cross-Domain Object DetectionPoster
  576. DiffSim: Taming Diffusion Models for Evaluating Visual SimilarityPoster
  577. DiffTell: A High-Quality Dataset for Describing Image Manipulation ChangesPoster
  578. DiffVSR: Revealing an Effective Recipe for Taming Robust Video Super-Resolution Against Complex DegradationsPoster
  579. Differentiable Room Acoustic Rendering with Multi-View Vision PriorsPoster
  580. Differential-informed Sample Selection Accelerates Multimodal Contrastive LearningPoster
  581. Differentially Private Fine-Tuning of Diffusion ModelsPoster
  582. DiffuMatch: Category-Agnostic Spectral Diffusion Priors for Robust Non-rigid Shape MatchingPoster
  583. Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion ModelsPoster
  584. Diffusion Curriculum: Synthetic-to-Real Data Curriculum via Image-Guided DiffusionPoster
  585. Diffusion Epistemic Uncertainty with Asymmetric Learning for Diffusion-Generated Image DetectionPoster
  586. Diffusion Guided Adaptive Augmentation for Generalization in Visual Reinforcement LearningPoster
  587. Diffusion Image PriorPoster
  588. Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-ResolutionPoster
  589. Diffusion-Based Extreme High-speed Scenes Reconstruction with the Complementary Vision SensorPoster
  590. Diffusion-Based Imaginative Coordination for Bimanual ManipulationPoster
  591. Diffusion-based 3D Hand Motion Recovery with Intuitive PhysicsPoster
  592. Diffusion-based Source-biased Model for Single Domain Generalized Object DetectionPoster
  593. DimensionX: Create Any 3D and 4D Scenes from a Single Image with Decoupled Video DiffusionPoster
  594. Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene ModelingPoster
  595. Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face RestorationPoster
  596. DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMsPoster
  597. DisCoPatch: Taming Adversarially-driven Batch Statistics for Improved Out-of-Distribution DetectionPoster
  598. DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow DecodingPoster
  599. DisTime: Distribution-based Time Representation for Video Large Language ModelsPoster
  600. Discontinuity-aware Normal Integration for Generic Central Camera ModelsPoster
  601. Discovering Divergent Representations between Text-to-Image ModelsPoster
  602. Discretized Gaussian Representation for Tomographic ReconstructionPoster
  603. DisenQ: Disentangling Q-Former for Activity-BiometricsPoster
  604. Disentangled Clothed Avatar Generation with Layered RepresentationPoster
  605. Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningPoster
  606. Disentangling Instance and Scene Contexts for 3D Semantic Scene CompletionPoster
  607. Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing AccuracyPoster
  608. Dissecting Generalized Category Discovery: Multiplex Consensus under Self-DeconstructionPoster
  609. DistillDrive: End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning ModelPoster
  610. Distilling Diffusion Models to Efficient 3D LiDAR Scene CompletionPoster
  611. Distilling Parallel Gradients for Fast ODE Solvers of Diffusion ModelsPoster
  612. Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyPoster
  613. Diversity-Enhanced Distribution Alignment for Dataset DistillationPoster
  614. Divide-and-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-supervised Continual LearningPoster
  615. Diving into the Fusion of Monocular Priors for Generalized Stereo MatchingPoster
  616. Do It Yourself: Learning Semantic Correspondence from Pseudo-LabelsPoster
  617. DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document UnderstandingPoster
  618. Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?Poster
  619. Domain Generalizable Portrait Style TransferPoster
  620. Domain-aware Category-level Geometry Learning Segmentation for 3D Point CloudsPoster
  621. Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint DetectionPoster
  622. DoppDrive: Doppler-Driven Temporal Aggregation for Improved Radar Object DetectionPoster
  623. Doppler-Aware LiDAR-RADAR Fusion for Weather-Robust 3D DetectionPoster
  624. Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion ModelsPoster
  625. Drawing Developmental Trajectory from Cortical Surface ReconstructionPoster
  626. Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
  627. DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid GuidancePoster
  628. DreamCube: RGB-D Panorama Generation via Multi-plane SynchronizationPoster
  629. DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesPoster
  630. DreamFuse: Adaptive Image Fusion with Diffusion TransformerPoster
  631. DreamLayer: Simultaneous Multi-Layer Generation via Diffusion ModelPoster
  632. DreamRelation: Relation-Centric Video CustomizationPoster
  633. DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image ModelsPoster
  634. DriveArena: A Closed-loop Generative Simulation Platform for Autonomous DrivingPoster
  635. DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous DrivingPoster
  636. DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive TransformersPoster
  637. DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video GenerationPoster
  638. DuCos: Duality Constrained Depth Super-Resolution via Foundation ModelPoster
  639. DuET: Dual Incremental Object Detection via Exemplar-Free Task ArithmeticPoster
  640. Dual Domain Control via Active Learning for Remote Sensing Domain Incremental Object DetectionPoster
  641. Dual Reciprocal Learning of Language-based Human Motion Understanding and GenerationPoster
  642. Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image DiffusionPoster
  643. Dual-Expert Consistency Model for Efficient and High-Quality Video GenerationPoster
  644. Dual-Process Image GenerationPoster
  645. Dual-S3D: Hierarchical Dual-Path Selective SSM-CNN for High-Fidelity Implicit ReconstructionPoster
  646. Dual-Temporal Exemplar Representation Network for Video Semantic SegmentationPoster
  647. Dual-level Prototype Learning for Composite Degraded Image RestorationPoster
  648. DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video CustomizationPoster
  649. DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-based Human Action SegmentationPoster
  650. DuoLoRA : Cycle-consistent and Rank-disentangled Content-Style PersonalizationPoster
  651. DyGS-SLAM: Real-Time Accurate Localization and Gaussian Reconstruction for Dynamic ScenesPoster
  652. DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile ManipulationPoster
  653. DynFaceRestore: Balancing Fidelity and Quality in Diffusion-Guided Blind Face Restoration with Dynamic Blur-Level Mapping and GuidancePoster
  654. DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video UnderstandingPoster
  655. Dynamic Dictionary Learning for Remote Sensing Image SegmentationPoster
  656. Dynamic Group Detection using VLM-augmented Temporal Groupness GraphPoster
  657. Dynamic Multi-Layer Null Space Projection for Vision-Language Continual LearningPoster
  658. Dynamic Multimodal Prototype Learning in Vision-Language ModelsPoster
  659. Dynamic Point Maps: A Versatile Representation for Dynamic 3D ReconstructionPoster
  660. Dynamic Reconstruction of Hand-Object Interaction with Distributed Force-aware Contact RepresentationPoster
  661. Dynamic Typography: Bringing Text to Life via Video Diffusion PriorPoster
  662. Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object DetectionPoster
  663. Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLMPoster
  664. DynamicFace: High-Quality and Consistent Face Swapping for Image and Video using Composable 3D Facial PriorsPoster
  665. DynamicID: Zero-Shot Multi-ID Image Personalization with Flexible Facial EditabilityPoster
  666. E-NeMF: Event-based Neural Motion Field for Novel Space-time View Synthesis of Dynamic ScenesPoster
  667. E-SAM: Training-Free Segment Every Entity ModelPoster
  668. EA-KD: Entropy-based Adaptive Knowledge DistillationPoster
  669. EA-Vit: Efficient Adaptation for Elastic Vision TransformerPoster
  670. EAMamba: Efficient All-Around Vision State Space Model for Image RestorationPoster
  671. EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric FlowPoster
  672. EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image RegistrationPoster
  673. EDM: Efficient Deep Feature MatchingPoster
  674. EDiT: Efficient Diffusion Transformers with Linear Compressed AttentionPoster
  675. EEdit : Rethinking the Spatial and Temporal Redundancy for Efficient Image EditingPoster
  676. EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained ClientsPoster
  677. EMD: Explicit Motion Modeling for High-Quality Street Gaussian SplattingPoster
  678. EMatch: A Unified Framework for Event-based Optical Flow and Stereo MatchingPoster
  679. EMoTive: Event-guided Trajectory Modeling for 3D Motion EstimationPoster
  680. ERNet: Efficient Non-Rigid Registration Network for Point SequencesPoster
  681. ESCNet:Edge-Semantic Collaborative Network for Camouflaged Object DetectionPoster
  682. ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental LearningPoster
  683. ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large ModelsPoster
  684. ETA: Energy-based Test-time Adaptation for Depth CompletionPoster
  685. ETCH: Generalizing Body Fitting to Clothed Humans via Equivariant TightnessPoster
  686. ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and AnsweringPoster
  687. EVDM: Event-based Real-world Video Deblurring with MambaPoster
  688. EVER: Exact Volumetric Ellipsoid Rendering for Real-time View SynthesisPoster
  689. EVEv2: Improved Baselines for Encoder-Free Vision-Language ModelsPoster
  690. EVOLVE: Event-Guided Deformable Feature Transfer and Dual-Memory Refinement for Low-Light Video Object SegmentationPoster
  691. EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionPoster
  692. EYE3:Turn Anything into Naked-eye 3DPoster
  693. Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image EditingPoster
  694. Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingPoster
  695. Easy3D: A Simple Yet Effective Method for 3D Interactive SegmentationPoster
  696. EasyControl: Adding Efficient and Flexible Control for Diffusion TransformerPoster
  697. Edicho: Consistent Image Editing in the WildPoster
  698. Edit360: 2D Image Edits to 3D Assets from Any AnglePoster
  699. EditCLIP: Representation Learning for Image EditingPoster
  700. Effective Training Data Synthesis for Improving MLLM Chart UnderstandingPoster
  701. Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning StrategyPoster
  702. Efficient Autoregressive Shape Generation via Octree-Based Adaptive TokenizationPoster
  703. Efficient Concertormer for Image Deblurring and BeyondPoster
  704. Efficient Event Camera Data Pretraining with Adaptive Prompt FusionPoster
  705. Efficient Fine-Tuning of Large Models via Nested Low-Rank AdaptationPoster
  706. Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation VariationPoster
  707. Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal InteractionsPoster
  708. Efficient Spiking Point Mamba for Point Cloud AnalysisPoster
  709. Efficient Track AnythingPoster
  710. Efficient Unsupervised Shortcut Learning Detection and Mitigation in TransformersPoster
  711. Efficient Visual Place Recognition Through Multimodal Semantic Knowledge IntegrationPoster
  712. EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion ModelsPoster
  713. EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric PerceptionPoster
  714. EgoAgent: A Joint Predictive Agent Model in Egocentric WorldsPoster
  715. EgoM2P: Egocentric Multimodal Multitask Pretraining
  716. EgoMusic-driven Human Dance Motion Estimation with Skeleton MambaPoster
  717. Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language GuidancePoster
  718. Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image DescriptionsPoster
  719. Embodied Navigation with Auxiliary Task of Action Description PredictionPoster
  720. Embodied Representation Alignment with Mirror NeuronsPoster
  721. Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene UnderstandingPoster
  722. EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene UnderstandingPoster
  723. EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile DevicePoster
  724. EmotiCrafter: Text-to-Emotional-Image Generation based on Valence-Arousal ModelPoster
  725. Emulating Self-attention with Convolution for Efficient Image Super-ResolutionPoster
  726. End-to-End Driving with Online Trajectory Evaluation via BEV World ModelPoster
  727. End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationPoster
  728. End-to-End Multi-Modal Diffusion MambaPoster
  729. Engage for All: Making Ordinary Image Descriptions Appealing Again!Poster
  730. Enhanced Event-based Dense Stereo via Cross-Sensor Knowledge DistillationPoster
  731. Enhanced Pansharpening via Quaternion Spatial-Spectral InteractionsPoster
  732. Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model FeaturesPoster
  733. Enhancing Image Restoration Transformer via Adaptive Translation EquivariancePoster
  734. Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense PredictionPoster
  735. Enhancing Numerical Prediction of MLLMs with Soft LabelingPoster
  736. Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningPoster
  737. Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image AlignmentPoster
  738. Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based SegmentationPoster
  739. Enhancing Transferability of Targeted Adversarial Examples via Inverse Target Gradient Competition and Spatial Distance StretchingPoster
  740. Enhancing Transformers Through Conditioned Embedded TokensPoster
  741. Enhancing Zero-shot Object Counting via Text-guided Local Ranking and Number-evoked Global AttentionPoster
  742. Enpowering Your Pansharpening Models with Generalizability: Unified Distribution is All You NeedPoster
  743. Enrich and Detect: Video Temporal Grounding with Multimodal LLMsPoster
  744. Ensemble Foreground Management for Unsupervised Object DiscoveryPoster
  745. Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step AlignmentPoster
  746. Environment-Agnostic Pose: Generating Environment-independent Object Representations for 6D Pose EstimationPoster
  747. Epipolar Consistent Attention Aggregation Network for Unsupervised Light Field Disparity EstimationPoster
  748. Epona: Autoregressive Diffusion World Model for Autonomous DrivingPoster
  749. EquiCaps: Predictor-Free Pose-Aware Pre-Trained Capsule NetworksPoster
  750. Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution DetectionPoster
  751. Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target ConceptsPoster
  752. Error Recognition in Procedural Videos using Generalized Task GraphPoster
  753. Estimating 2D Camera Motion with Hybrid Motion BasisPoster
  754. EvRT-DETR: Latent Space Adaptation of Image Detectors for Event-based VisionPoster
  755. EvaGaussians: Event Stream Assisted Gaussian Splatting from Blurry ImagesPoster
  756. Evading Data Provenance in Deep Neural NetworksPoster
  757. Event-Driven Storytelling with Multiple Lifelike Humans in a 3D ScenePoster
  758. Event-aided Dense and Continuous Point Tracking: Everywhere and AnytimePoster
  759. Event-based Tiny Object Detection: A Benchmark Dataset and BaselinePoster
  760. Event-based Visual VibrometryPoster
  761. Event-boosted Deformable 3D Gaussians for Dynamic Scene ReconstructionPoster
  762. Event-guided HDR Reconstruction with Diffusion PriorsPoster
  763. Event-guided Unified Framework for Low-light Video Enhancement, Frame Interpolation, and DeblurringPoster
  764. EventUPS: Uncalibrated Photometric Stereo Using an Event CameraPoster
  765. Everything is a Video: Unifying Modalities through Next-Frame PredictionPoster
  766. Evidential Knowledge DistillationPoster
  767. EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference AlignmentPoster
  768. ExCap3D: Expressive 3D Scene Understanding via Object Captioning with Varying DetailPoster
  769. Explaining Human Preferences via Metrics for Structured 3D ReconstructionPoster
  770. Exploiting Diffusion Prior for Task-driven Image RestorationPoster
  771. Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic SegmentationPoster
  772. Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score PropagationPoster
  773. ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion PriorsPoster
  774. Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image EditingPoster
  775. Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic SegmentationPoster
  776. Exploring The Visual Feature Space for Multimodal Neural DecodingPoster
  777. Exploring View Consistency for Scene-Adaptive Low-Light Light Field Image EnhancementPoster
  778. Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in RoboticsPoster
  779. Expressive Talking Human from Single-Image with Imperfect PriorsPoster
  780. Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
  781. External Knowledge Injection for CLIP-Based Class-Incremental LearningPoster
  782. Extrapolated Urban View Synthesis BenchmarkPoster
  783. F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and RestorationPoster
  784. FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution DetectionPoster
  785. FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual RegistersPoster
  786. FB-Diff: Fourier Basis-guided Diffusion for Temporal Interpolation of 4D Medical ImagingPoster
  787. FDPT: Federated Discrete Prompt Tuning for Black-Box Visual-Language ModelsPoster
  788. FE-CLIP: Frequency Enhanced CLIP Model for Zero-Shot Anomaly Detection and SegmentationPoster
  789. FED-PsyAU: Privacy-Preserving Micro-Expression Recognition via Psychological AU Coordination and Dynamic Facial Motion ModelingPoster
  790. FEVER-OOD: Free Energy Vulnerability Elimination for Robust Out-of-Distribution DetectionPoster
  791. FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental UnlearningPoster
  792. FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image GenerationPoster
  793. FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal DataPoster
  794. FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking PortraitPoster
  795. FLOSS: Free Lunch in Open-vocabulary Semantic SegmentationPoster
  796. FLSeg: Enhancing Privacy and Robustness in Federated Learning under Heterogeneous Data via Model SegmentationPoster
  797. FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformancePoster
  798. FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face RetouchingPoster
  799. FREE-Merging: Fourier Transform for Efficient Model MergingPoster
  800. FRET: Feature Redundancy Elimination for Test Time AdaptationPoster
  801. FROSS: Faster-Than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D ImagesPoster
  802. FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion DistillationPoster
  803. FW-Merging: Scaling Model Merging with Frank-Wolfe OptimizationPoster
  804. Face Retouching with Diffusion Data Generation and Spectral RestorementPoster
  805. FaceCraft4D: Animated 3D Facial Avatar Generation from a Single ImagePoster
  806. FaceLift: Learning Generalizable Single Image 3D Face Reconstruction from Synthetic HeadsPoster
  807. FaceShield: Defending Facial Image against Deepfake ThreatsPoster
  808. FaceXFormer: A Unified Transformer for Facial AnalysisPoster
  809. Factorized Learning for Temporally Grounded Video-Language ModelsPoster
  810. Failure Cases Are Better Learned But Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingPoster
  811. Fair Generation without Unfair Distortions: Debiasing Text-to-Image Generation with Entanglement-Free AttentionPoster
  812. FairGen: Enhancing Fairness in Text-to-Image Diffusion Models via Self-Discovering Latent DirectionsPoster
  813. FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion ModelsPoster
  814. FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake VideosPoster
  815. Fast Globally Optimal and Geometrically Consistent 3D Shape MatchingPoster
  816. Fast Image Super-Resolution via Consistency Rectified FlowPoster
  817. FastJSMA: Accelerating Jacobian-based Saliency Map Attacks through Gradient DecouplingPoster
  818. FastPoint: Accelerating 3D Point Cloud Model Inference via Sample Point Distance PredictionPoster
  819. FastVAR: Linear Visual Autoregressive Modeling via Cached Token PruningPoster
  820. Faster and Better 3D Splatting via Group TrainingPoster
  821. Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model AccelerationPoster
  822. Feature Coding in the Era of Large Models: Dataset, Test Conditions, and BenchmarkPoster
  823. Feature Extraction and Representation of Pre-training Point Cloud Based on Diffusion ModelsPoster
  824. Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic SegmentationPoster
  825. FedAGC: Federated Continual Learning with Asymmetric Gradient CorrectionPoster
  826. FedDifRC: Unlocking the Potential of Text-to-Image Diffusion Models in Heterogeneous Federated LearningPoster
  827. FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language ModelsPoster
  828. FedMeNF: Privacy-Preserving Federated Meta-Learning for Neural FieldsPoster
  829. FedPall: Prototype-based Adversarial and Collaborative Learning for Federated Learning with Feature DriftPoster
  830. FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic ManipulationPoster
  831. FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform QuantizationPoster
  832. FedXDS: Leveraging Model Attribution Methods to counteract Data Heterogeneity in Federated LearningPoster
  833. Federated Continual Instruction TuningPoster
  834. Federated Continuous Category Discovery and LearningPoster
  835. Federated Domain Generalization with Domain-specific Soft Prompts GenerationPoster
  836. Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client DataPoster
  837. Federated Representation Angle LearningPoster
  838. Feed-Forward SceneDINO for Unsupervised Semantic Scene CompletionPoster
  839. Few-Shot Image Quality Assessment via Adaptation of Vision-Language ModelsPoster
  840. Few-Shot Pattern Detection via Template Matching and RegressionPoster
  841. Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model DeploymentPoster
  842. FiVE-Bench: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow ModelsPoster
  843. FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth EstimationPoster
  844. FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and ChallengingPoster
  845. Find Any Part in 3DPoster
  846. Find a Scapegoat: Poisoning Membership Inference Attack and Defense to Federated LearningPoster
  847. Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures via Joint Reconstruction and RegistrationPoster
  848. Fine-Grained Evaluation of Large Vision-Language Models in Autonomous DrivingPoster
  849. Fine-Tuning Visual Autogressive Models for Subject-Driven GenerationPoster
  850. Fine-grained Abnormality Prompt Learning for Zero-shot Anomaly DetectionPoster
  851. Fine-grained Spatiotemporal Grounding on Egocentric VideosPoster
  852. Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE TrainingPoster
  853. FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and EditingPoster
  854. Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric VisionPoster
  855. Fix-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long TextPoster
  856. FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesPoster
  857. Flash-VStream: Efficient Real-Time Understanding for Long Video StreamsPoster
  858. FlashDepth: Real-time Streaming Video Depth Estimation at 2K ResolutionPoster
  859. FlexGen: Flexible Multi-View Generation from Text and Image InputsPoster
  860. Flexi-FSCIL: Adaptive Knowledge Retention for Breaking the Stability-Plasticity Dilemma in Few-Shot Class-Incremental LearningPoster
  861. Flow Stochastic Segmentation NetworksPoster
  862. Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image TokenizationPoster
  863. Flow-MIL: Constructing Highly-expressive Latent Feature Space For Whole Slide Image Classification Using Normalizing FlowPoster
  864. Flow4Agent: Long-form Video Understanding via Motion Prior from Optical FlowPoster
  865. FlowChef: Steering of Rectified Flow Models for Controlled GenerationsPoster
  866. FlowDPS : Flow-Driven Posterior Sampling for Inverse ProblemsPoster
  867. FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow ModelsPoster
  868. FlowR: Flowing from Sparse to Dense 3D ReconstructionsPoster
  869. FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion BasesPoster
  870. FlowStyler: Artistic Video Stylization via Transformation Fields TransportsPoster
  871. FlowTok: Flowing Seamlessly Across Text and Image TokensPoster
  872. Focal Plane Visual Feature Generation and Matching on a Pixel Processor ArrayPoster
  873. FonTS: Text Rendering With Typography and Style ControlsPoster
  874. FontAnimate: High Quality Few-shot Font Generation via Animating Font Transfer ProcessPoster
  875. ForCenNet: Foreground-Centric Network for Document Image RectificationPoster
  876. ForeSight: Multi-View Streaming Joint Object Detection and Trajectory ForecastingPoster
  877. Forecasting Continuous Non-Conservative Dynamical Systems in SO(3)Poster
  878. Forensic-MoE: Exploring Comprehensive Synthetic Image Detection Traces with Mixture of ExpertsPoster
  879. Foresight in Motion: Reinforcing Trajectory Prediction with Reward HeuristicsPoster
  880. ForestFormer3D: A Unified Framework for End-to-End Segmentation of Forest LiDAR 3D Point CloudsPoster
  881. ForgeLens: Data-Efficient Forgery Focus for Generalizable Forgery Image DetectionPoster
  882. Forgetting Through Transforming: Enabling Federated Unlearning via Class-Aware Representation TransformationPoster
  883. FoundIR: Unleashing Million-scale Training Data to Advance Foundation Models for Image RestorationPoster
  884. FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language ModelsPoster
  885. FramePainter: Endowing Interactive Image Editing with Video Diffusion PriorsPoster
  886. Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video GenerationPoster
  887. Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single InferencePoster
  888. Free-running vs Synchronous: Single-Photon Lidar for High-flux 3D ImagingPoster
  889. Free2Guide: Training-Free Text-to-Video Alignment using Image LVLMPoster
  890. Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal ConsistencyPoster
  891. FreeCus: Free Lunch Subject-driven Customization in Diffusion TransformersPoster
  892. FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise AlignmentPoster
  893. FreeDance: Towards Harmonic Free-Number Group Dance Generation via a Unified FrameworkPoster
  894. FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image EditingPoster
  895. FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion ModelPoster
  896. FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale FusionPoster
  897. FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D ReconstructionPoster
  898. FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection TransformersPoster
  899. Frequency Domain-Based Diffusion Model for Unpaired Image DehazingPoster
  900. Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal ForecastingPoster
  901. Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image SynthesisPoster
  902. Frequency-Dynamic Attention Modulation For Dense PredictionPoster
  903. Frequency-Guided Diffusion for Training-Free Text-Driven Image TranslationPoster
  904. Frequency-Guided Posterior Sampling for Diffusion-Based Image RestorationPoster
  905. Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-based Action RecognitionPoster
  906. From Abyssal Darkness to Blinding Glare: A Benchmark on Extreme Exposure Correction in Real WorldPoster
  907. From Easy to Hard: Progressive Active Learning Framework for Infrared Small Target Detection with Single Point SupervisionPoster
  908. From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image ReasoningPoster
  909. From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuningPoster
  910. From Gallery to Wrist: Realistic 3D Bracelet Insertion in VideosPoster
  911. From Gaze to Movement: Predicting Visual Attention for Autonomous Driving Human-Machine Interaction based on Programmatic Imitation LearningPoster
  912. From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-TuningPoster
  913. From Image to Video: An Empirical Study of Diffusion RepresentationsPoster
  914. From Imitation to Innovation: The Emergence of AI's Unique Artistic Styles and the Challenge of Copyright ProtectionPoster
  915. From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial CorrelationsPoster
  916. From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors via LLM-guided Symbolic ReasoningPoster
  917. From One to More: Contextual Part Latents for 3D GenerationPoster
  918. From Panels to Prose: Generating Literary Narratives from ComicsPoster
  919. From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute TransitionPoster
  920. From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection TuningPoster
  921. From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeersPoster
  922. From Sharp to Blur: Unsupervised Domain Adaptation for 2D Human Pose Estimation Under Extreme Motion Blur Using Event CamerasPoster
  923. From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward AlignmentPoster
  924. FuXi-RTM: A Physics-Guided Prediction Framework with Radiative Transfer ModelingPoster
  925. FullDiT: Video Generative Foundation Models with Multimodal Control via Full AttentionPoster
  926. Function-centric Bayesian Network for Zero-Shot Object Goal NavigationPoster
  927. Fuse Before Transfer: Knowledge Fusion for Heterogeneous DistillationPoster
  928. Fusion Meets Diverse Conditions: A High-diversity Benchmark and Baseline for UAV-based Multimodal Object Detection with Condition CuesPoster
  929. FusionPhys: A Flexible Framework for Fusing Complementary Sensing Modalities in Remote Physiological MeasurementPoster
  930. Future-Aware Interaction Network For Motion ForecastingPoster
  931. Fuzzy Contrastive Decoding to Alleviate Object Hallucination in Large Vision-Language ModelsPoster
  932. G-DexGrasp: Generalizable Dexterous Grasping Synthesis Via Part-Aware Prior Retrieval and Prior-Assisted GenerationPoster
  933. G2D: Boosting Multimodal Learning with Gradient-Guided DistillationPoster
  934. G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction via Evolutionary DiffusionPoster
  935. G2SF: Geometry-Guided Score Fusion for Multimodal Industrial Anomaly DetectionPoster
  936. GAP: Gaussianize Any Point Clouds with Text GuidancePoster
  937. GARF: Learning Generalizable 3D Reassembly for Real-World FracturesPoster
  938. GAS: Generative Avatar Synthesis from a Single ImagePoster
  939. GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in InterpretabilityPoster
  940. GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray DiffusionPoster
  941. GDKVM: Echocardiography Video Segmentation via Spatiotemporal Key-Value Memory with Gated Delta RulePoster
  942. GECKO: Gigapixel Vision-Concept Contrastive Pretraining in HistopathologyPoster
  943. GECO: Geometrically Consistent Embedding with Lightspeed InferencePoster
  944. GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray DiagnosisPoster
  945. GENMO: A GENeralist Model for Human MOtionPoster
  946. GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial TasksPoster
  947. GEOPARD: Geometric Pretraining for Articulation Prediction in 3D ShapesPoster
  948. GFPack++: Attention-Driven Gradient Fields for Optimizing 2D Irregular PackingPoster
  949. GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific AdaptationPoster
  950. GIViC: Generative Implicit Video CompressionPoster
  951. GLEAM: Enhanced Transferable Adversarial Attacks for Vision-Language Pre-training Models via Global-Local TransformationsPoster
  952. GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor ScenePoster
  953. GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-ExpertsPoster
  954. GMMamba: Group Masking Mamba for Whole Slide Image ClassificationPoster
  955. GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal ModelsPoster
  956. GReg: Geometry-Aware Region Refinement for Sign Language Video GenerationPoster
  957. GS-ID: Illumination Decomposition on Gaussian Splatting via Adaptive Light Aggregation and Diffusion-Guided Material PriorsPoster
  958. GS-LIVM: Real-Time Photo-Realistic LiDAR-Inertial-Visual Mapping with Gaussian SplattingPoster
  959. GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian SplattingPoster
  960. GSOT3D: Towards Generic 3D Single Object Tracking in the WildPoster
  961. GSRecon: Efficient Generalizable Gaussian Splatting for Surface Reconstruction from Sparse ViewsPoster
  962. GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object GenerationPoster
  963. GT-Loc: Unifying When and Where in Images Through a Joint Embedding SpacePoster
  964. GT-Mean Loss: A Simple Yet Effective Solution for Brightness Mismatch in Low-Light Image EnhancementPoster
  965. GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent TrainingPoster
  966. GUAVA: Generalizable Upper Body 3D Gaussian AvatarPoster
  967. GVDepth: Zero-Shot Monocular Depth Estimation for Ground Vehicles based on Probabilistic Cue FusionPoster
  968. GWM: Towards Scalable Gaussian World Models for Robotic ManipulationPoster
  969. GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo CollectionsPoster
  970. GaSLight: Gaussian Splats for Spatially-Varying Lighting in HDRPoster
  971. Gain-MLP: Improving HDR Gain Map Encoding via a Lightweight MLPPoster
  972. Gait-X: Exploring X modality for Generalized Gait RecognitionPoster
  973. GameFactory: Creating New Games with Generative Interactive VideosPoster
  974. GauUpdate: New Object Insertion in 3D Gaussian Fields with Consistent Global IlluminationPoster
  975. GausSim: Foreseeing Reality by Gaussian Simulator for Elastic ObjectsPoster
  976. GaussRender: Learning 3D Occupancy with Gaussian RenderingPoster
  977. Gaussian Splatting with Discretized SDF for Relightable AssetsPoster
  978. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D SynthesisPoster
  979. Gaussian-based World Model: Gaussian Priors for Voxel-Based Occupancy Prediction and Future Motion PredictionPoster
  980. GaussianFlowOcc: Sparse and Weakly Supervised Occupancy Estimation using Gaussian Splatting and Temporal FlowPoster
  981. GaussianOcc: Fully Self-supervised and Efficient 3D Occupancy Estimation with Gaussian SplattingPoster
  982. GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMsPoster
  983. GaussianReg: Rapid 2D/3D Registration for Emergency Surgery via Explicit 3D Modeling with Gaussian PrimitivesPoster
  984. GaussianSpeech: Audio-Driven Personalized 3D Gaussian AvatarsPoster
  985. GaussianUpdate: Continual 3D Gaussian Splatting Update for Changing EnvironmentsPoster
  986. GaussianVideo: Efficient Video Representation via Hierarchical Gaussian SplattingPoster
  987. Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze ScanpathsPoster
  988. GazeGaussian: High-Fidelity Gaze Redirection with 3D Gaussian SplattingPoster
  989. Geminio: Language-Guided Gradient Inversion Attacks in Federated LearningPoster
  990. GenDoP: Auto-regressive Camera Trajectory Generation as a Director of PhotographyPoster
  991. GenFlow3D: Generative Scene Flow Estimation and Prediction on Point Cloud SequencesPoster
  992. GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement LearningPoster
  993. GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric EnhancersPoster
  994. GenHaze: Pioneering Controllable One-Step Realistic Haze Generation for Real-World DehazingPoster
  995. GenM3: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion GenerationPoster
  996. General Compression Framework for Efficient Transformer Object TrackingPoster
  997. Generalizable Non-Line-of-Sight Imaging with Learnable Physical PriorsPoster
  998. Generalizable Object Re-Identification via Visual In-Context PromptingPoster
  999. Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake DetectionPoster
  1000. Generalized Deep Multi-view Clustering via Causal Learning with Partially Aligned Cross-view CorrespondencePoster

Looking for submission deadlines instead? See the conference deadline calendar.