← All conferences

ICCV 2025 Accepted Papers

The full list of 2,620 papers accepted at ICCV 2025 (IEEE/CVF International Conference on Computer Vision). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Poster: 2,595
  1. Generalized Few-Shot Point Cloud Segmentation via LLM-Assisted Hyper-Relation MatchingPoster
  2. Generalized Tensor-based Parameter-Efficient Fine-Tuning via Lie Group TransformationsPoster
  3. Generalized and Efficient 2D Gaussian Splatting for Arbitrary-scale Super-ResolutionPoster
  4. Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category DiscoveryPoster
  5. Generate, Transduct, Adapt: Iterative Transduction with VLMsPoster
  6. Generating Multi-Image Synthetic Data for Text-to-Image CustomizationPoster
  7. Generating Physically Stable and Buildable Brick Structures from TextPoster
  8. Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface NetworksPoster
  9. Generative Active Learning for Long-tail Trajectory Prediction via Controllable Diffusion ModelPoster
  10. Generative Adversarial DiffusionPoster
  11. Generative Gaussian Splatting: Generating 3D Scenes with Video Diffusion PriorsPoster
  12. Generative Modeling of Shape-Dependent Self-Contact Human PosesPoster
  13. Generative Video Bi-flowPoster
  14. Generative ZooPoster
  15. Generic Event Boundary Detection via Denoising DiffusionPoster
  16. GenieBlue: Integrating both Linguistic and Multimodal Capabilities for Large Language Models on Mobile DevicesPoster
  17. Geo4D: Leveraging Video Generators for Geometric 4D Scene ReconstructionPoster
  18. GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head AvatarPoster
  19. GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image GenerationPoster
  20. GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View LocalizationPoster
  21. GeoExplorer: Active Geo-localization with Curiosity-Driven ExplorationPoster
  22. GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-based TransformerPoster
  23. GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video DiffusionPoster
  24. GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language FieldsPoster
  25. GeoSplatting: Towards Geometry Guided Gaussian Splatting for Physically-based Inverse RenderingPoster
  26. Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen CategoriesPoster
  27. Geometry DistributionsPoster
  28. GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion PriorsPoster
  29. GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented GenerationPoster
  30. GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal ModelingPoster
  31. GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image GenerationPoster
  32. GlassWizard: Harvesting Diffusion Priors for Glass Surface DetectionPoster
  33. GloPER: Unsupervised Animal Pattern Extraction from Local ReconstructionPoster
  34. Global Motion Corresponder for 3D Point-Based Scene Interpolation under Large MotionPoster
  35. Global Regulation and Excitation via Attention Tuning for Stereo MatchingPoster
  36. Global and Local Entailment Learning for Natural World ImageryPoster
  37. Global-Aware Monocular Semantic Scene Completion with State Space ModelsPoster
  38. Go to Zero: Towards Zero-shot Motion Generation with Million-scale DataPoster
  39. Golden Noise for Diffusion Models: A Learning FrameworkPoster
  40. Gradient Decomposition and Alignment for Incremental Object DetectionPoster
  41. Gradient Extrapolation for Debiased Representation LearningPoster
  42. Gradient Short-Circuit: Efficient Out-of-Distribution Detection via Feature InterventionPoster
  43. Gradient-Reweighted Adversarial Camouflage for Physical Object Detection EvasionPoster
  44. Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept RepresentationsPoster
  45. Graph Domain Adaptation with Dual-branch Encoder and Two-level Alignment for Whole Slide Image-based Survival PredictionPoster
  46. GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language InstructionsPoster
  47. Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-ReferringPoster
  48. GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential GroundingPoster
  49. GroundingSuite: Measuring Complex Multi-Granular Pixel GroundingPoster
  50. Group Inertial Poser: Multi-Person Pose and Global Translation from Sparse Inertial Sensors and Ultra-Wideband Ranging
  51. Group-wise Scaling and Orthogonal Decomposition for Domain-Invariant Feature Extraction in Face Anti-SpoofingPoster
  52. Grouped Speculative Decoding for Autoregressive Image GenerationPoster
  53. Growing a Twig to Accelerate Large Vision-Language ModelsPoster
  54. Guiding Diffusion Models with Adaptive Negative Sampling Without External ResourcesPoster
  55. Guiding Diffusion-Based Articulated Object Generation by Partial Point Cloud Alignment and Physical Plausibility ConstraintsPoster
  56. Guiding Noisy Label Conditional Diffusion Models with Score-based Discriminator CorrectionPoster
  57. H3R: Hybrid Multi-view Correspondence for Generalizable 3D ReconstructionPoster
  58. HADES: Human Avatar with Dynamic Explicit Hair StrandsPoster
  59. HAMSt3R: Human-Aware Multi-view Stereo 3D ReconstructionPoster
  60. HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReIDPoster
  61. HDR Image Generation via Gain Map Decomposed DiffusionPoster
  62. HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationPoster
  63. HERMES: temporal-coHERent long-forM understanding with Episodes and SemanticsPoster
  64. HERO: Human Reaction Generation from VideosPoster
  65. HFD-Teacher: High-Frequency Depth Distillation from Depth Foundation Models for Enhanced Depth CompletionPoster
  66. HIS-GPT: Towards 3D Human-In-Scene Multimodal UnderstandingPoster
  67. HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationPoster
  68. HOMO-Feature: Cross-Arbitrary-Modal Image Matching with Homomorphism of Organized Major OrientationPoster
  69. HORT: Monocular Hand-held Objects Reconstruction with TransformersPoster
  70. HPSv3: Towards Wide-Spectrum Human Preference ScorePoster
  71. HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP ModelsPoster
  72. HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?Poster
  73. HUG: Hierarchical Urban Gaussian Splatting with Block-Based Reconstruction for Large-Scale Aerial ScenesPoster
  74. HUMOTO: A 4D Dataset of Mocap Human Object InteractionsPoster
  75. HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture QuantizationPoster
  76. HVPUNet: Hybrid-Voxel Point-cloud Upsampling NetworkPoster
  77. HairCUP: Hair Compositional Universal Prior for 3D Gaussian AvatarsPoster
  78. Hallucinatory Image Tokens: A Training-free EAZY Approach to Detecting and Mitigating Object Hallucinations in LVLMsPoster
  79. Harmonizing Visual Representations for Unified Multimodal Understanding and GenerationPoster
  80. HarmonySeg: Tubular Structure Segmentation with Deep-Shallow Feature Fusion and Growth-Suppression Balanced LossPoster
  81. Harnessing Input-Adaptive Inference for Efficient VLNPoster
  82. Harnessing Massive Satellite Imagery with Efficient Masked Image ModelingPoster
  83. Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised LearningPoster
  84. Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object DetectionPoster
  85. Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary SegmentationPoster
  86. Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful IllusionsPoster
  87. HazeFlow: Revisit Haze Physical Model as ODE and Non-Homogeneous Haze Generation for Real-World DehazingPoster
  88. HccePose(BF): Predicting Front & Back Surfaces to Construct Ultra-Dense 2D-3D Correspondences for Pose EstimationPoster
  89. Head2Body: Body Pose Generation from Multi-sensory Head-mounted InputsPoster
  90. Heatmap Regression without Soft-Argmax for Facial Landmark DetectionPoster
  91. Heavy Labels Out! Dataset Distillation with Label Space LighteningPoster
  92. Height-Fidelity Dense Global Fusion for Multi-modal 3D Object DetectionPoster
  93. Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language ModelsPoster
  94. Hi-Gaussian: Hierarchical Gaussians under Normalized Spherical Projection for Single-View 3D ReconstructionPoster
  95. Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal BridgingPoster
  96. HiERO: Understanding the Hierarchy of Human Behavior Enhances Reasoning on Egocentric VideosPoster
  97. HiGarment: Cross-modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment ImagePoster
  98. HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal ModelPoster
  99. HiNeuS: High-fidelity Neural Surface Mitigating Low-texture and Reflective AmbiguityPoster
  100. HiP-AD: Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single DecoderPoster
  101. Hierarchical 3D Scene Graphs Construction OutdoorsPoster
  102. Hierarchical Cross-modal Prompt Learning for Vision-Language ModelsPoster
  103. Hierarchical Divide-and-Conquer Grouping for Classification Adaptation of Pre-Trained ModelsPoster
  104. Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal GroundingPoster
  105. Hierarchical Material Recognition from Local AppearancePoster
  106. Hierarchical Variational Test-Time Prompt Generation for Zero-Shot GeneralizationPoster
  107. Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationPoster
  108. Hierarchical-aware Orthogonal Disentanglement Framework for Fine-grained Skeleton-based Action RecognitionPoster
  109. Hierarchy UGP: Hierarchy Unified Gaussian Primitive for Large-Scale Dynamic Scene ReconstructionPoster
  110. Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image RetrievalPoster
  111. High-Precision 3D Measurement of Complex Textured Surfaces Using Multiple Filtering ApproachPoster
  112. High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationPoster
  113. Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image FusionPoster
  114. Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous DrivingPoster
  115. Hipandas: Hyperspectral Image Joint Denoising and Super-Resolution by Image Fusion with the Panchromatic ImagePoster
  116. HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing ImageryPoster
  117. Holistic Tokenizer for Autoregressive Image GenerationPoster
  118. Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model UnlearningPoster
  119. HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion ModelsPoster
  120. HouseTour: A Virtual Real Estate A(I)gentPoster
  121. How Can Objects Help Video-Language Understanding?Poster
  122. How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in An Extensible Escape GamePoster
  123. How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?Poster
  124. How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation ApproachPoster
  125. How To Make Your Cell Tracker Say "I dunno!"Poster
  126. How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor ScenesPoster
  127. Human-Object Interaction from Human-Level InstructionsPoster
  128. Human-in-the-Loop Local Corrections of 3D Scene Layouts via InfillingPoster
  129. HumanOLAT: A Large-Scale Dataset for Full-Body Human Relighting and Novel-View SynthesisPoster
  130. HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion AnomalyPoster
  131. Humans as Checkerboards: Calibrating Camera Motion Scale for World-Coordinate Human Mesh RecoveryPoster
  132. Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated VideosPoster
  133. HumorDB: Can AI understand graphical humor?Poster
  134. HyPiDecoder: Hybrid Pixel Decoder for Efficient Segmentation and DetectionPoster
  135. HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video CodingPoster
  136. Hybrid Layout Control for Diffusion Transformer: Fewer Annotations, Superior AestheticsPoster
  137. Hybrid-TTA: Continual Test-time Adaptation via Dynamic Domain Shift DetectionPoster
  138. Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
  139. Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth EstimationPoster
  140. Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop TrainingPoster
  141. HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image GenerationPoster
  142. Hyper-Depth: Hypergraph-based Multi-Scale Representation Fusion for Monocular Depth EstimationPoster
  143. HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D RegistrationPoster
  144. Hypergraph Clustering Network with Partial Attribute ImputationPoster
  145. I Am Big, You Are Little; I Am Right, You Are WrongPoster
  146. I2-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene ForecastingPoster
  147. I2V3D: Controllable Image-to-video Generation with 3D GuidancePoster
  148. I2VControl: Disentangled and Unified Video Motion Synthesis ControlPoster
  149. IAP: Invisible Adversarial Patch Attack through Perceptibility-Aware Localization and Perturbation OptimizationPoster
  150. ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and EditingPoster
  151. IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using ThemselvesPoster
  152. IDF: Iterative Dynamic Filtering Networks for Generalizable Image DenoisingPoster
  153. IDFace: Face Template Protection for Efficient and Secure IdentificationPoster
  154. IFAdapter: Instance Feature Control for Grounded Text-to-Image GenerationPoster
  155. IGD: Instructional Graphic Design with Multimodal Layer GenerationPoster
  156. IGL-Nav: Incremental 3D Gaussian Localization for Image-goal NavigationPoster
  157. ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhancePoster
  158. IM-LUT: Interpolation Mixing Look-Up Tables for Image Super-ResolutionPoster
  159. IM360: Large-scale Indoor Mapping with 360 CamerasPoster
  160. IMG: Calibrating Diffusion Models via Implicit Multimodal GuidancePoster
  161. IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&APoster
  162. INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in InsurancePoster
  163. INSTINCT: Instance-Level Interaction Architecture for Query-Based Collaborative PerceptionPoster
  164. INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance SamplingPoster
  165. IQA-Adapter: Exploring Knowledge Transfer from Image Quality Assessment to Diffusion-based Generative ModelsPoster
  166. IRASim: A Fine-Grained World Model for Robot ManipulationPoster
  167. IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale BenchmarkPoster
  168. ISP2HRNet: Learning to Reconstruct High Resolution Image from Irregularly Sampled Pixels via Hierarchical Gradient LearningPoster
  169. Identity Preserving 3D Head Stylization with Multiview Score DistillationPoster
  170. Identity-aware Language Gaussian Splatting for Open-vocabulary 3D Semantic SegmentationPoster
  171. Im2Haircut: Single-view Strand-based Hair Reconstruction for Human AvatarsPoster
  172. ImHead: A Large-scale Implicit Morphable Model for Localized Head ModelingPoster
  173. Image Intrinsic Scale Assessment: Bridging the Gap Between Quality and ResolutionPoster
  174. Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred ImagePoster
  175. Image-Guided Shape-from-Template Using Mesh Inextensibility ConstraintsPoster
  176. ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model PersonalizationPoster
  177. ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought ReasoningPoster
  178. Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic SegmentationPoster
  179. Imbalance in Balance: Online Concept Balancing in Generation Models
  180. Implicit Counterfactual Learning for Audio-Visual SegmentationPoster
  181. Importance-Based Token Merging for Efficient Image and Video GenerationPoster
  182. Improved Noise Schedule for Diffusion TrainingPoster
  183. Improving Large Vision and Language Models by Learning from a Panel of PeersPoster
  184. Improving Multimodal Learning via Imbalanced LearningPoster
  185. Improving Noise Efficiency in Privacy-preserving Dataset DistillationPoster
  186. Improving Rectified Flow with Boundary ConditionsPoster
  187. Improving SAM for Camouflaged Object Detection via Dual Stream AdaptersPoster
  188. Incremental Few-Shot Semantic Segmentation via Multi-Level Switchable Visual PromptsPoster
  189. InfGen: A Resolution-Agnostic Paradigm for Scalable Image SynthesisPoster
  190. Inference-Time Diffusion Model DistillationPoster
  191. InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video ModelsPoster
  192. InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score DistillationPoster
  193. InfiniteYou: Flexible Photo Recrafting While Preserving Your IdentityPoster
  194. InfoBridge: Balanced Multimodal Integration through Conditional Dependency ModelingPoster
  195. Information Density Principle for MLLM BenchmarksPoster
  196. Information-Bottleneck Driven Binary Neural Network for Change DetectionPoster
  197. Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional WarpingPoster
  198. InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset ConstructionPoster
  199. InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video GenerationPoster
  200. InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered ScenesPoster
  201. Instance-Level Video Depth in Groups Beyond OcclusionsPoster
  202. Instant GaussianImage: A Generalizable and Self-Adaptive Image Representation via 2D Gaussian SplattingPoster
  203. InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified FlowPoster
  204. InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language ModelsPoster
  205. Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language ModelsPoster
  206. Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMsPoster
  207. Instruction-based Image Editing with Planning, Reasoning, and GenerationPoster
  208. Integrating Biological Knowledge for Robust Microscopy Image Profiling on De Novo Cell LinesPoster
  209. Integrating Task-Specific and Universal Adapters for Pre-Trained Model-based Class-Incremental LearningPoster
  210. Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingPoster
  211. Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive SegmentationPoster
  212. InterGSEdit: Interactive 3D Gaussian Splatting Editing with 3D Geometry-Consistent Attention PriorPoster
  213. InteractAvatar: Modeling Hand-Face Interaction in Photorealistic Avatars with Deformable GaussiansPoster
  214. Interaction-Merged Motion Planning: Effectively Leveraging Diverse Motion Datasets for Robust PlanningPoster
  215. Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object CategoriesPoster
  216. Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language ModelPoster
  217. Interpretable point cloud classification using multiple instance learningPoster
  218. Intervening in Black Box: Concept Bottleneck Model for Enhancing Human Neural Network Mutual UnderstandingPoster
  219. Intra-modal and Cross-modal Synchronization for Audio-visual Deepfake Detection and Temporal LocalizationPoster
  220. Intra-view and Inter-view Correlation Guided Multi-view Novel Class DiscoveryPoster
  221. IntrinsicControlNet: Cross-distribution Image Generation with Real and UnrealPoster
  222. IntroStyle: Training-Free Introspective Style Attribution using Diffusion FeaturesPoster
  223. InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance ModelingPoster
  224. Inverse 3D Microscopy Rendering for Cell Shape Inference with Active MeshPoster
  225. Inverse Image-Based Rendering for Light Field Generation from Single ImagesPoster
  226. Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking DesignPoster
  227. Iris: Breaking GUI Complexity with Adaptive Focus and Self-RefiningPoster
  228. Is CLIP ideal? No. Can we fix it? Yes!Poster
  229. Is Less More? Exploring Token Condensation as Training-free Test-time AdaptationPoster
  230. Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited EntropyPoster
  231. Is Tracking Really More Challenging in First Person Egocentric Vision?Poster
  232. Is Visual in-Context Learning for Compositional Medical Tasks within Reach?Poster
  233. JPEG Processing Neural Operator for Backward-Compatible CodingPoster
  234. JailbreakDiffBench: A Comprehensive Benchmark for Jailbreaking Diffusion ModelsPoster
  235. Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyPoster
  236. Jigsaw++: Imagining Complete Shape Priors for Object ReassemblyPoster
  237. Joint Asymmetric Loss for Learning with Noisy LabelsPoster
  238. Joint Diffusion Models in Continual LearningPoster
  239. Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-level 6D Pose EstimationPoster
  240. Joint Self-Supervised Video Alignment and Action SegmentationPoster
  241. Joint Semantic and Rendering Enhancements in 3D Gaussian Modeling with Anisotropic Local EncodingPoster
  242. JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion TransformersPoster
  243. KOEnsAttack: Towards Efficient Data-Free Black-Box Adversarial Attacks via Knowledge-Orthogonalized Substitute EnsemblesPoster
  244. KV-Edit: Training-Free Image Editing for Precise Background PreservationPoster
  245. Kaleidoscopic Background Attack: Disrupting Pose Estimation with Multi-Fold Radial Symmetry TexturesPoster
  246. Kaputt: A Large-Scale Dataset for Visual Defect DetectionPoster
  247. Keep Your Friends Close, and Your Enemies Farther: Distance-aware Voxel-wise Contrastive Learning for Semi-supervised Multi-organ SegmentationPoster
  248. Kestrel: 3D Multimodal LLM for Part-Aware Grounded DescriptionPoster
  249. Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video ProcessingPoster
  250. KinMo: Kinematic-aware Human Motion Understanding and GenerationPoster
  251. Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIPPoster
  252. Know Your Attention Maps: Class-specific Token Masking for Weakly Supervised Semantic SegmentationPoster
  253. Knowledge Distillation for Learned Image CompressionPoster
  254. Knowledge Distillation with Refined LogitsPoster
  255. Knowledge Transfer from Interaction LearningPoster
  256. LA-MOTR: End-to-End Multi-Object Tracking by Learnable AssociationPoster
  257. LACONIC: A 3D Layout Adapter for Controllable Image CreationPoster
  258. LANGTRAJ: Diffusion Model and Dataset for Language-Conditioned Trajectory SimulationPoster
  259. LATINO-PRO: LAtent consisTency INverse sOlver with PRompt OptimizationPoster
  260. LBM: Latent Bridge Matching for Fast Image-to-Image TranslationPoster
  261. LDIP: Long Distance Information Propagation for Video Super-ResolutionPoster
  262. LDPose: Towards Inclusive Human Pose Estimation for Limb-Deficient Individuals in the WildPoster
  263. LEGION: Learning to Ground and Explain for Synthetic Image DetectionPoster
  264. LEGO-Maker: A Semantic-Driven Algorithm for Text-to-3D GenerationPoster
  265. LGA-Net: Learning Local and Global Affinities for Sparse Scribble based Image ColorizationPoster
  266. LHM: Large Animatable Human Reconstruction Model for Single Image to 3D in SecondsPoster
  267. LIFT: Latent Implicit Functions for Task- and Data-Agnostic EncodingPoster
  268. LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry CompressionPoster
  269. LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region AssistancePoster
  270. LIRA: Reasoning Reconstruction via Multimodal Large Language ModelsPoster
  271. LLM Thought Divergence and Convergence for Dialogue-Based Image Generation ControlPoster
  272. LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object DetectionPoster
  273. LLM-assisted Entropy-based Adaptive Distillation for Unsupervised Fine-grained Visual Representation LearningPoster
  274. LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text MatchingPoster
  275. LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMsPoster
  276. LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D CapabilitiesPoster
  277. LLaVA-CoT: Let Vision Language Models Reason Step-by-StepPoster
  278. LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsPoster
  279. LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsPoster
  280. LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMsPoster
  281. LMM-Det: Make Large Multimodal Models Excel in Object DetectionPoster
  282. LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image EditingPoster
  283. LOMM: Latest Object Memory Management for Temporally Consistent Video Instance SegmentationPoster
  284. LONG3R: Long Sequence Streaming 3D ReconstructionPoster
  285. LOTA: Bit-Planes Guided AI-Generated Image DetectionPoster
  286. LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text PairingPoster
  287. LUDVIG: Learning-Free Uplifting of 2D Visual Features to Gaussian Splatting ScenesPoster
  288. LUSD: Localized Update Score Distillation for Text-Guided Image EditingPoster
  289. LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up TablesPoster
  290. LV-MAE: Learning Long Video Representations through Masked-Embedding AutoencodersPoster
  291. LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsPoster
  292. LVBench: An Extreme Long Video Understanding BenchmarkPoster
  293. LVFace: Progressive Cluster Optimization for Large Vision Models in Face RecognitionPoster
  294. LaCoOT: Layer Collapse through Optimal TransportPoster
  295. LaRender: Training-Free Occlusion Control in Image Generation via Latent RenderingPoster
  296. Laboring on less labors: RPCA Paradigm for Pan-sharpeningPoster
  297. LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature GenerationPoster
  298. LangBridge: Interpreting Image as a Combination of Language EmbeddingsPoster
  299. LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video DiffusionPoster
  300. Language Decoupling with Fine-grained Knowledge Guidance for Referring Multi-object TrackingPoster
  301. Language Driven Occupancy PredictionPoster
  302. Language-Driven Multi-Label Zero-Shot Learning with Semantic GranularityPoster
  303. Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and CompressibilityPoster
  304. Large Multi-modal Models Can Interpret Features in Large Multi-modal ModelsPoster
  305. Large Scene Generation with Cube-Absorb Discrete DiffusionPoster
  306. Large-scale Pre-training for Grounded Video Caption GenerationPoster
  307. Lark: Low-Rank Updates After Knowledge Localization for Few-shot Class-Incremental LearningPoster
  308. Latent Diffusion Models with Masked AutoEncodersPoster
  309. Latent Expression Generation for Referring Image Segmentation and GroundingPoster
  310. Latent Swap Joint Diffusion for 2D Long-Form Latent GenerationPoster
  311. Latent-Reframe: Enabling Camera Control for Video Diffusion Models without TrainingPoster
  312. Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated LearningPoster
  313. LawDIS: Language-Window-based Controllable Dichotomous Image SegmentationPoster
  314. Lay-Your-Scene: Natural Scene Layout Generation with Diffusion TransformersPoster
  315. Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story GenerationPoster
  316. LayerAnimate: Layer-level Control for AnimationPoster
  317. LayerD: Decomposing Raster Graphic Designs into LayersPoster
  318. LayerLock: Non-collapsing Representation Learning with Progressive FreezingPoster
  319. LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion TransformerPoster
  320. LazyMAR: Accelerating Masked Autoregressive Models via Feature CachingPoster
  321. LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityPoster
  322. LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion ModelsPoster
  323. Leaps and Bounds: An Improved Point Cloud Winding Number Formulation for Fast Normal Estimation and Surface ReconstructionPoster
  324. Learn2Synth: Learning Optimal Data Synthesis Using Hypergradients for Brain Image SegmentationPoster
  325. Learnable Feature Patches and Vectors for Boosting Low-light Image Enhancement without External KnowledgePoster
  326. Learnable Fractional Reaction-Diffusion Dynamics for Under-Display ToF Imaging and BeyondPoster
  327. Learnable Logit Adjustment for Imbalanced Semi-Supervised Learning under Class Distribution MismatchPoster
  328. Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report GenerationPoster
  329. Learned Image Compression with Hierarchical Progressive Context ModelingPoster
  330. Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion ModelsPoster
  331. Learning 3D Scene Analogies with Neural Contextual Scene MapsPoster
  332. Learning 4D Embodied World ModelsPoster
  333. Learning A Unified Template for Gait RecognitionPoster
  334. Learning Beyond Still Frames: Scaling Vision-Language Models with VideoPoster
  335. Learning Deblurring Texture Prior from Unpaired Data with Diffusion ModelPoster
  336. Learning Dense Feature Matching via Lifting Single 2D Image to 3D SpacePoster
  337. Learning Efficient and Generalizable Human Representation with Human Gaussian ModelPoster
  338. Learning Few-Step Diffusion Models by Trajectory Distribution MatchingPoster
  339. Learning Hierarchical Line Buffer for Image ProcessingPoster
  340. Learning Implicit Features with Flow-Infused Transformations for Realistic Virtual Try-OnPoster
  341. Learning Interpretable Queries for Explainable Image Classification with Information PursuitPoster
  342. Learning Large Motion Estimation from Intermediate Representations with a High-Resolution Optical Flow Dataset Featuring Long-Range Dynamic MotionPoster
  343. Learning Neural Scene Representation from iToF ImagingPoster
  344. Learning Normal Flow Directly From EventsPoster
  345. Learning Normals of Noisy Points by Local Gradient-Aware Surface FilteringPoster
  346. Learning Null Geodesics for Gravitational Lensing Rendering in General RelativityPoster
  347. Learning Pixel-adaptive Multi-layer Perceptrons for Real-time Image EnhancementPoster
  348. Learning Precise Affordances from Egocentric Videos for Robotic ManipulationPoster
  349. Learning Robust Image Watermarking with Lossless Cover RecoveryPoster
  350. Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsPoster
  351. Learning Separable Fine-Grained Representation via Dendrogram Construction from Coarse Labels for Fine-grained Visual RecognitionPoster
  352. Learning Streaming Video Representation via Multitask TrainingPoster
  353. Learning Visual Hierarchies in Hyperbolic Space for Image RetrievalPoster
  354. Learning Visual Proxy for Compositional Zero-Shot LearningPoster
  355. Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped DisentanglementPoster
  356. Learning an Implicit Physics Model for Image-based Fluid SimulationPoster
  357. Learning on the Go: A Meta-learning Object Navigation ModelPoster
  358. Learning to Generalize without Bias for Open-Vocabulary Action RecognitionPoster
  359. Learning to Inference Adaptively for Multimodal Large Language ModelsPoster
  360. Learning to See Inside Opaque Liquid Containers using Speckle VibrometryPoster
  361. Learning to See in the Extremely DarkPoster
  362. Learning to Unlearn while Retaining: Combating Gradient Conflicts in Machine UnlearningPoster
  363. Less Static, More Private: Towards Transferable Privacy-Preserving Action Recognition by Generative Decoupled LearningPoster
  364. Less is More: Empowering GUI Agent with Context-Aware SimplificationPoster
  365. Less is More: Improving Motion Diffusion Models with Sparse KeyframesPoster
  366. Less-to-More Generalization: Unlocking More Controllability by In-Context GenerationPoster
  367. Leveraging 2D Priors and SDF Guidance for Urban Scene RenderingPoster
  368. Leveraging BEV Paradigm for Ground-to-Aerial Image SynthesisPoster
  369. Leveraging Debiased Cross-modal Attention Maps and Code-based Reasoning for Zero-shot Referring Expression ComprehensionPoster
  370. Leveraging Local Patch Alignment to Seam-cutting for Large Parallax Image StitchingPoster
  371. Leveraging Panoptic Scene Graph for Evaluating Fine-Grained Text-to-Image GenerationPoster
  372. Leveraging Prior Knowledge of Diffusion Model for Person SearchPoster
  373. Leveraging Spatial Invariance to Boost Adversarial TransferabilityPoster
  374. Leveraging the Power of MLLMs for Gloss-Free Sign Language TranslationPoster
  375. LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video DiffusionPoster
  376. LiT: Delving into a Simple Linear Diffusion Transformer for Image GenerationPoster
  377. Liberated-GS: 3D Gaussian Splatting Independent from SfM Point CloudsPoster
  378. Lidar Waveforms are Worth 40x128x33 WordsPoster
  379. Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary ModelingPoster
  380. LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation LearningPoster
  381. LightCity: An Urban Dataset for Outdoor Inverse Rendering and Reconstruction under Multi-illumination ConditionsPoster
  382. LightSwitch: Multi-view Relighting with Material-guided DiffusionPoster
  383. LightsOut: Diffusion-based Outpainting for Enhanced Lens Flare RemovalPoster
  384. Lightweight Gradient-Aware Upscaling of 3D Gaussian Splatting ImagesPoster
  385. Lightweight and Fast Real-time Image Enhancement via Decomposition of the Spatial-aware Lookup TablesPoster
  386. LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette AlignmentPoster
  387. LoRA-FAIR: Federated LoRA Fine-Tuning with Aggregation and Initialization RefinementPoster
  388. LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image GenerationPoster
  389. LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion ModelsPoster
  390. Local Scale Equivariance with Latent Deep Equilibrium CanonicalizerPoster
  391. LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature DecouplingPoster
  392. LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation ModelsPoster
  393. Long Context Tuning for Video GenerationPoster
  394. Long-Context State-Space Video World ModelsPoster
  395. Long-LRM: Long-sequence Large Reconstruction Model for Wide-coverage Gaussian SplatsPoster
  396. Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario GenerationPoster
  397. LongAnimation: Long Animation Generation with Dynamic Global-Local MemoryPoster
  398. LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long VideosPoster
  399. LookOut: Real-World Humanoid Egocentric NavigationPoster
  400. Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification ModelsPoster
  401. Loss Functions for Predictor-based Neural Architecture SearchPoster
  402. Low-Light Image Enhancement Using Event-Based Illumination EstimationPoster
  403. Lumina-Image 2.0: A Unified and Efficient Image Generative FrameworkPoster
  404. Lyra: An Efficient and Speech-Centric Framework for Omni-CognitionPoster
  405. M-Net: MRI Brain Tumor Sequential Segmentation Network via Mesh-CastPoster
  406. M-SpecGene: Generalized Foundation Model for RGBT Multispectral VisionPoster
  407. M2EIT: Multi-Domain Mixture of Experts for Robust Neural Inertial TrackingPoster
  408. M2SFormer: Multi-Spectral and Multi-Scale Attention with Edge-Aware Difficulty Guidance for Image Forgery LocalizationPoster
  409. MA-CIR: A Multimodal Arithmetic Benchmark for Composed Image RetrievalPoster
  410. MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D PerceptionPoster
  411. MATE: Motion-Augmented Temporal Consistency for Event-based Point TrackingPoster
  412. MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual TranslationPoster
  413. MAVias: Mitigate any Visual BiasPoster
  414. MBTI: Masked Blending Transformers with Implicit Positional Encoding for Frame-rate Agnostic Motion EstimationPoster
  415. MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsPoster
  416. MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video UnderstandingPoster
  417. MCID: Multi-aspect Copyright Infringement Detection for Generated ImagesPoster
  418. MDD: A Dataset for Text-and-Music Conditioned Duet Dance GenerationPoster
  419. MDP-Omni: Parameter-free Multimodal Depth Prior-based Sampling for Omnidirectional Stereo MatchingPoster
  420. MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMsPoster
  421. MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic ScenesPoster
  422. MEH: A Multi-Style Dataset and Toolkit for Advancing Egyptian Hieroglyph RecognitionPoster
  423. MEMFOF: High-Resolution Training for Memory-Efficient Multi-Frame Optical Flow EstimationPoster
  424. METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language ModelsPoster
  425. MGSR: 2D/3D Mutual-boosted Gaussian Splatting for High-fidelity Surface Reconstruction under Various Light ConditionsPoster
  426. MGSfM: Multi-Camera Geometry Driven Global Structure-from-MotionPoster
  427. MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video CodingPoster
  428. MIEB: Massive Image Embedding BenchmarkPoster
  429. MINERVA: Evaluating Complex Video ReasoningPoster
  430. MIORe & VAR-MIORe: Benchmarks to Push the Boundaries of RestorationPoster
  431. MM-IFEngine: Towards Multimodal Instruction FollowingPoster
  432. MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsPoster
  433. MMAD: Multi-label Micro-Action Detection in VideosPoster
  434. MMAIF: Multi-task and Multi-degradation All-in-One for Image Fusion with Language GuidancePoster
  435. MMAT-1M: A Large Reasoning Dataset for Multimodal Agent TuningPoster
  436. MMCR: Benchmarking Cross-Source Reasoning in Scientific PapersPoster
  437. MMGeo: Multimodal Compositional Geo-Localization for UAVsPoster
  438. MMOne: Representing Multiple Modalities in One ScenePoster
  439. MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGIPoster
  440. MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder PruningPoster
  441. MOERL: When Mixture-of-Experts Meet Reinforcement Learning for Adverse Weather Image RestorationPoster
  442. MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room EnvironmentsPoster
  443. MOSCATO: Predicting Multiple Object State Change Through ActionsPoster
  444. MOVE: Motion-Guided Few-Shot Video Object SegmentationPoster
  445. MP-HSIR: A Multi-Prompt Framework for Universal Hyperspectral Image RestorationPoster
  446. MPBR: Multimodal Progressive Bidirectional Reasoning for Open-Set Fine-Grained RecognitionPoster
  447. MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationPoster
  448. MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data GenerationPoster
  449. MRGen: Segmentation Data Engine For Underrepresented MRI ModalitiesPoster
  450. MS3D: High-Quality 3D Generation via Multi-Scale Representation ModelingPoster
  451. MSA2: Multi-task Framework with Structure-aware and Style-adaptive Character Representation for Open-set Chinese Text RecognitionPoster
  452. MSQ: Memory-Efficient Bit Sparsification QuantizationPoster
  453. MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video ParsingPoster
  454. MUNBa: Machine Unlearning via Nash BargainingPoster
  455. MUSE-VL: Modeling Unified VLM through Semantic Discrete EncodingPoster
  456. MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic ExpansionPoster
  457. MV-Adapter: Multi-View Consistent Image Generation Made EasyPoster
  458. MVGBench: a Comprehensive Benchmark for Multi-view Generation ModelsPoster
  459. MVQA: Mamba with Unified Sampling for Efficient Video Quality AssessmentPoster
  460. MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance CostPoster
  461. MaGS: Reconstructing and Simulating Dynamic 3D Objects with Mesh-adsorbed Gaussian SplattingPoster
  462. MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language ModelingPoster
  463. MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic DisturbancesPoster
  464. Magic Insert: Style-Aware Drag-and-DropPoster
  465. MagicCity: Geometry-Aware 3D City Generation from Satellite Imagery with Multi-View ConsistencyPoster
  466. MagicColor: Multi-Instance Sketch ColorizationPoster
  467. MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive ControlPoster
  468. MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video ClipsPoster
  469. MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video CustomizationPoster
  470. MagicMirror: ID-Preserved Video Generation in Video Diffusion TransformersPoster
  471. MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidancePoster
  472. Make Me Happier: Evoking Emotions Through Image Diffusion ModelsPoster
  473. Make Your Training Flexible: Towards Deployment-Efficient Video ModelsPoster
  474. MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric SequencePoster
  475. MamV2XCalib: V2X-based Target-less Infrastructure Camera Calibration with State Space ModelPoster
  476. Mamba-3VL: Taming State Space Model for 3D Vision Language LearningPoster
  477. MambaML: Exploring State Space Models for Multi-Label Image ClassificationPoster
  478. Manual-PA: Learning 3D Part Assembly from Instruction DiagramsPoster
  479. Marigold-DC: Zero-Shot Monocular Depth Completion with Guided DiffusionPoster
  480. MaskControl: Spatio-Temporal Control for Masked Motion SynthesisPoster
  481. MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the WildPoster
  482. MaskSAM: Auto-prompt SAM with Mask Classification for Volumetric Medical Image SegmentationPoster
  483. Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and RepresentativenessPoster
  484. MatchDiffusion: Training-free Generation of Match-CutsPoster
  485. MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR DiffusionPoster
  486. MeasureXpert: Automatic Anthropometric Measurement Extraction from Two Unregistered, Partial, Posed, and Dressed Body ScansPoster
  487. Measuring the Impact of Rotation Equivariance on Aerial Object DetectionPoster
  488. MedSegFactory: Text-Guided Generation of Medical Image-Mask PairsPoster
  489. MedVSR: Medical Video Super-Resolution with Cross State-Space PropagationPoster
  490. Medical World ModelPoster
  491. MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionPoster
  492. Membership Inference Attacks with False Discovery Rate ControlPoster
  493. Memory-Efficient 4-bit Preconditioned Stochastic OptimizationPoster
  494. Memory-Efficient Generative Models via Product QuantizationPoster
  495. MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided StylizationPoster
  496. MergeOcc: Bridge the Domain Gap between Different LiDARs for Robust Occupancy PredictionPoster
  497. MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh TokenizationPoster
  498. MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D MeshPoster
  499. MeshMamba: State Space Models for Articulated 3D Mesh Generation and ReconstructionPoster
  500. MeshPad: Interactive Sketch-Conditioned Artist-Reminiscent Mesh Generation and EditingPoster
  501. Met2Net: A Decoupled Two-Stage Spatio-Temporal Forecasting Model for Complex Meteorological SystemsPoster
  502. Meta-Learning Dynamic Center Distance: Hard Sample Mining for Learning with Noisy LabelsPoster
  503. Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned ConceptsPoster
  504. MetaMorph: Multimodal Understanding and Generation via Instruction TuningPoster
  505. MetaScope: Optics-Driven Neural Network for Ultra-Micro Metalens EndoscopyPoster
  506. Metric Convolutions: A Unifying Theory to Adaptive Image ConvolutionsPoster
  507. MiDSummer: Multi-Guidance Diffusion for Controllable Zero-Shot Immersive Gaussian Splatting Scene GenerationPoster
  508. MikuDance: Animating Character Art with Mixed Motion DynamicsPoster
  509. MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnPPoster
  510. Mind the Cost of Scaffold! Benign Clients May Even Become Accomplices of Backdoor AttackPoster
  511. Mind the Gap: Aligning Vision Foundation Models to Image Feature MatchingPoster
  512. Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual LearningPoster
  513. MissRAG: Addressing the Missing Modality Challenge in Multimodal Large Language ModelsPoster
  514. MistSense: Versatile Online Detection of Procedural and Execution MistakesPoster
  515. Mitigating Catastrophic Overfitting in Fast Adversarial Training via Label Information EliminationPoster
  516. Mitigating Geometric Degradation in Fast DownSampling via FastAdapter for Point Cloud SegmentationPoster
  517. Mitigating Object Hallucinations via Sentence-Level Early InterventionPoster
  518. MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization PerspectivePoster
  519. MixA: A Mixed Attention approach with Stable Lightweight Linear Attention to enhance Efficiency of Vision Transformers at the EdgePoster
  520. MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action AnticipationPoster
  521. MixRI: Mixing Features of Reference Images for Novel Object Pose EstimationPoster
  522. Mixed Signals: A Diverse Point Cloud Dataset for Heterogeneous LiDAR V2X CollaborationPoster
  523. Mixture of Experts Guided by Gaussian Splatters Matters: A new Approach to Weakly-Supervised Video Anomaly DetectionPoster
  524. Mixture-of-Scores: Robust Image-Text Data Valuation via Three Lines of CodePoster
  525. MoFRR: Mixture of Diffusion Models for Face Retouching RestorationPoster
  526. MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar ReconstructionPoster
  527. MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile ManipulationPoster
  528. MoMaps: Semantics-Aware Scene Motion Generation with Motion MapsPoster
  529. MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised LearningPoster
  530. Mobile Video DiffusionPoster
  531. MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile DevicesPoster
  532. MobileViCLIP: An Efficient Video-Text Model for Mobile DevicesPoster
  533. ModSkill: Physical Character Skill ModularizationPoster
  534. ModalTune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-task Learning in Digital PathologyPoster
  535. Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion ModelsPoster
  536. Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionPoster
  537. Modeling Saliency Dataset BiasPoster
  538. Moderating the Generalization of Score-based Generative ModelPoster
  539. MolParser: End-to-end Visual Recognition of Molecule Structures in the WildPoster
  540. Moment Quantization for Video Temporal GroundingPoster
  541. Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene ReconstructionPoster
  542. MonSTeR: a Unified Model for Motion, Scene, Text RetrievalPoster
  543. MonoFusion: Sparse-View 4D Reconstruction via Monocular FusionPoster
  544. MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkPoster
  545. MonoMobility: Zero-Shot 3D Mobility Analysis from Monocular VideosPoster
  546. MonoSOWA: Scalable Monocular 3D Object Detector Without Human AnnotationsPoster
  547. Monocular Facial Appearance Capture in the WildPoster
  548. Monocular Semantic Scene Completion via Masked Recurrent NetworksPoster
  549. More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label LearningPoster
  550. MorphoGen: Efficient Unconditional Generation of Long-Range Projection Neuronal Morphology via a Global-to-Local FrameworkPoster
  551. MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining DynamicsPoster
  552. Motal: Unsupervised 3D Object Detection by Modality and Task-specific Knowledge TransferPoster
  553. Motion Synthesis with Sparse and Flexible Keyjoint ControlPoster
  554. Motion-2-to-3: Leveraging 2D Motion Data for 3D Motion GenerationsPoster
  555. MotionAgent: Fine-grained Controllable Video Generation via Motion Field AgentPoster
  556. MotionCtrl: A Real-time Controllable Vision-Language-Motion ModelPoster
  557. MotionDiff: Training-free Zero-shot Interactive Motion Editing via Flow-assisted Multi-view DiffusionPoster
  558. MotionFollower: Editing Video Motion via Score-Guided DiffusionPoster
  559. MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion ParadigmPoster
  560. MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video GenerationPoster
  561. MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent SpacePoster
  562. Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosPoster
  563. Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied NavigationPoster
  564. MuGS: Multi-Baseline Generalizable Gaussian Splatting ReconstructionPoster
  565. Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language ModelsPoster
  566. Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMsPoster
  567. Multi-Modal Few-Shot Temporal Action SegmentationPoster
  568. Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning FrameworkPoster
  569. Multi-Object Sketch Animation by Scene Decomposition and Motion PlanningPoster
  570. Multi-Schema Proximity Network for Composed Image RetrievalPoster
  571. Multi-View 3D Point TrackingPoster
  572. Multi-View Slot Attention Using Paraphrased Texts for Face Anti-SpoofingPoster
  573. Multi-identity Human Image Animation with Structural Video DiffusionPoster
  574. Multi-modal Identity ExtractionPoster
  575. Multi-modal Multi-platform Person Re-Identification: Benchmark and MethodPoster
  576. Multi-modal Segment Anything Model for Camouflaged Scene SegmentationPoster
  577. Multi-scenario Overlapping Text Segmentation with Depth AwarenessPoster
  578. Multi-turn Consistent Image EditingPoster
  579. Multi-view Gaze Target EstimationPoster
  580. MultiADS: Defect-aware Supervision for Multi-type Anomaly Detection and Segmentation in Zero-Shot LearningPoster
  581. MultiModal Action Conditioned Video Simulation
  582. MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language ModelsPoster
  583. Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data GenerationPoster
  584. Multimodal LLM Guided Exploration and Active Mapping using Fisher InformationPoster
  585. Multimodal LLMs as Customized Reward Models for Text-to-Image GenerationPoster
  586. Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference LearningPoster
  587. Multimodal Latent Diffusion Model for Complex Sewing Pattern GenerationPoster
  588. Multimodal Prompt Alignment for Facial Expression RecognitionPoster
  589. Multispectral Demosaicing via Dual CamerasPoster
  590. MultiverSeg: Scalable Interactive Segmentation of Biomedical Imaging Datasets with In-Context GuidancePoster
  591. Music Grounding by Short VideoPoster
  592. Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion ModelingPoster
  593. NAPPure: Adversarial Purification for Robust Image Classification under Non-Additive PerturbationsPoster
  594. NATRA: Noise-Agnostic Framework for Trajectory Prediction with Noisy ObservationsPoster
  595. NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic ReasoningPoster
  596. NETracer: A Topology-Aware Iterative Tracing Approach for Tubular Structure ExtractionPoster
  597. NGD: Neural Gradient Based Deformation for Monocular Garment ReconstructionPoster
  598. Nautilus: Locality-aware Autoencoder for Scalable Mesh GenerationPoster
  599. NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous EnvironmentsPoster
  600. NavQ: Learning a Q-Model for Foresighted Vision-and-Language NavigationPoster
  601. NeRF Is a Valuable Assistant for 3D Gaussian SplattingPoster
  602. NegRefine: Refining Negative Label-Based Zero-Shot OOD DetectionPoster
  603. NeuFrameQ: Neural Frame Fields for Scalable and Generalizable Anisotropic QuadrangulationPoster
  604. NeurOp-Diff: Continuous Remote Sensing Image Super-Resolution via Neural Operator DiffusionPoster
  605. Neural Architecture Search Driven by Locally Guided Diffusion for Personalized Federated LearningPoster
  606. Neural Compression for 3D Geometry SetsPoster
  607. Neural Inverse Rendering for High-Accuracy 3D Measurement of Moving Objects with Fewer Phase-Shifting PatternsPoster
  608. Neural Multi-View Self-Calibrated Photometric Stereo without Photometric Stereo CuesPoster
  609. Neural Shell Texture Splatting: More Details and Fewer PrimitivesPoster
  610. Neural Solver of Dichromatic Reflection Model for Specular Highlight RemovalPoster
  611. NeuralSVG: An Implicit Representation for Text-to-Vector GenerationPoster
  612. Neuromanifold-Regularized KANs for Shape-fair Feature RepresentationsPoster
  613. Neurons: Emulating the Human Visual Cortex Improves Fidelity and Interpretability in fMRI-to-Video ReconstructionPoster
  614. Neuroverse3D: Developing In-Context Learning Universal Model for Neuroimaging in 3DPoster
  615. No More Sibling Rivalry: Debiasing Human-Object Interaction DetectionPoster
  616. No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse ViewsPoster
  617. Noise-Modeled Diffusion Models for Low-Light Spike Image RestorationPoster
  618. Noise2Score3D: Tweedie's Approach for Unsupervised Point Cloud DenoisingPoster
  619. NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and CollaborationPoster
  620. Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology ImagesPoster
  621. NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion PriorsPoster
  622. NormalLoc: Visual Localization on Textureless 3D Models using Surface NormalsPoster
  623. Not All Degradations Are Equal: A Targeted Feature Denoising Framework for Generalizable Image Super-ResolutionPoster
  624. Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static FeaturesPoster
  625. Not Only Vision: Evolve Visual Speech Recognition via Peripheral InformationPoster
  626. Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation ModelsPoster
  627. NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language ModelsPoster
  628. NuiScene: Exploring Efficient Generation of Unbounded Outdoor ScenesPoster
  629. NullSwap: Proactive Identity Cloaking Against Deepfake Face SwappingPoster
  630. O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric ViewsPoster
  631. OCK: Unsupervised Dynamic Video Prediction with Object-Centric KinematicsPoster
  632. OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented GenerationPoster
  633. OCSplats: Observation Completeness Quantification and Label Noise Separation in 3DGSPoster
  634. OD-RASE: Ontology-Driven Risk Assessment and Safety Enhancement for Autonomous DrivingPoster
  635. ODDR: Outlier Detection & Dimension Reduction Based Defense Against Adversarial PatchesPoster
  636. ODP-Bench: Benchmarking Out-of-Distribution Performance PredictionPoster
  637. OMNI-DC: Highly Robust Depth Completion with Multiresolution Depth IntegrationPoster
  638. ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language ModelsPoster
  639. ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action GenerationPoster
  640. OURO: A Self-Bootstrapped Framework for Enhancing Multimodal Scene UnderstandingPoster
  641. OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object DetectionPoster
  642. OV3D-CG: Open-vocabulary 3D Instance Segmentation with Contextual GuidancePoster
  643. OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part DetectionPoster
  644. OVG-HQ: Online Video Grounding with Hybrid-modal QueriesPoster
  645. Oasis: One Image is All You Need for Multimodal Instruction Data SynthesisPoster
  646. Object-centric Video Question Answering with Visual Grounding and ReferringPoster
  647. Object-level Correlation for Few-Shot SegmentationPoster
  648. ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian SplattingPoster
  649. ObjectMate: A Recurrence Prior for Object Insertion and Subject-Driven GenerationPoster
  650. ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric PerspectivesPoster
  651. OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous DrivingPoster
  652. OccluGaussian: Occlusion-Aware Gaussian Splatting for Large Scene Reconstruction and RenderingPoster
  653. Occlusion-robust Stylization for Drawing-based 3D AnimationPoster
  654. Occupancy Learning with Spatiotemporal MemoryPoster
  655. Omegance: A Single Parameter for Various Granularities in Diffusion-Based SynthesisPoster
  656. OminiControl: Minimal and Universal Control for Diffusion TransformerPoster
  657. Omni-scene Perception-oriented Point Cloud Geometry Enhancement for Coordinate QuantizationPoster
  658. OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer ModelsPoster
  659. OmniDiff: A Comprehensive Benchmark for Fine-grained Image Difference CaptioningPoster
  660. OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation ModelsPoster
  661. OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal InpaintingPoster
  662. OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic SegmentationPoster
  663. OmniVTON: Training-Free Universal Virtual Try-OnPoster
  664. On Large Multimodal Models as Open-World Image ClassifiersPoster
  665. On the Complexity-Faithfulness Trade-off of Gradient-Based ExplanationsPoster
  666. On the Generalization of Representation Uncertainty in Earth ObservationPoster
  667. On the Provable Importance of Gradients for Autonomous Language-Assisted Image ClusteringPoster
  668. On the Recovery of Cameras from Fundamental MatricesPoster
  669. On the Robustness Tradeoff in Fine-TuningPoster
  670. On-Device Diffusion Transformer Policy for Efficient Robot ManipulationPoster
  671. One Encoder to Rule them All: Representation Learning for Model-free Visual Reinforcement Learning using Fourier Neural OperatorsPoster
  672. One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution ImagesPoster
  673. One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language ModelsPoster
  674. One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training ModelsPoster
  675. One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt EvolutionPoster
  676. One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object TrajectoryPoster
  677. One-Shot Knowledge Transfer for Scalable Person Re-IdentificationPoster
  678. One-Step Specular Highlight Removal with Adapted Diffusion ModelsPoster
  679. OneGT: One-Shot Geometry-Texture Neural Rendering for Head AvatarsPoster
  680. Online Dense Point Tracking with Streaming MemoryPoster
  681. Online Generic Event Boundary DetectionPoster
  682. Online Language SplattingPoster
  683. Online Reasoning Video Segmentation with Just-in-Time Digital TwinsPoster
  684. Open-Unfairness Adversarial Mitigation for Generalized Deepfake DetectionPoster
  685. Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept CalibrationPoster
  686. Open-World Skill Discovery from Unsegmented Demonstration VideosPoster
  687. Open-ended Hierarchical Streaming Video Understanding with Vision Language ModelsPoster
  688. OpenAnimals: Revisiting Person Re-Identification for Animals Towards Better GeneralizationPoster
  689. OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human AnnotationsPoster
  690. OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing ImagesPoster
  691. OpenSubstance: A High-quality Measured Dataset of Multi-View and -Lighting Images and ShapesPoster
  692. OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal LearningPoster
  693. OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingPoster
  694. Optical Model-Driven Sharpness Mapping for Autofocus in Small Depth-of-Field and Severe Defocus ScenariosPoster
  695. Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information ProcessingPoster
  696. OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic TypographyPoster
  697. Orchid: Image Latent Diffusion for Joint Appearance and Geometry GenerationPoster
  698. OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLMPoster
  699. OuroMamba: A Data-Free Quantization Framework for Vision MambaPoster
  700. Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse RenderingPoster
  701. Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian PointmapsPoster
  702. Outlier-Aware Post-Training Quantization for Image Super-ResolutionPoster
  703. Overcoming Dual Drift for Continual Long-Tailed Visual Question AnsweringPoster
  704. PAN-Crafter: Learning Modality-Consistent Alignment for PAN-SharpeningPoster
  705. PARTE: Part-Guided Texturing for 3D Human Reconstruction from a Single ImagePoster
  706. PASD: A Pixel-Adaptive Swarm Dynamics Approach for Unsupervised Low-Light Image EnhancementPoster
  707. PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic ManipulationPoster
  708. PASTA: Part-Aware Sketch-to-3D Shape Generation with Text-Aligned PriorPoster
  709. PBCAT: Patch-Based Composite Adversarial Training against Physically Realizable Attacks on Object DetectionPoster
  710. PBFG: A New Physically-Based Dataset and Removal of Lens Flares and GlaresPoster
  711. PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-RegularizationsPoster
  712. PEFTDiff: Diffusion-Guided Transferability Estimation for Parameter-Efficient Fine-TuningPoster
  713. PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single ImagePoster
  714. PHATNet: A Physics-guided Haze Transfer Network for Domain-adaptive Real-world Image DehazingPoster
  715. PHD: Personalized 3D Human Body Fitting with Point DiffusionPoster
  716. PINO: Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized GroupsPoster
  717. PLA: Prompt Learning Attack against Text-to-Image Generative ModelsPoster
  718. PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging SparsityPoster
  719. PLAN: Proactive Low-Rank Allocation for Continual LearningPoster
  720. PLMP - Point-Line Minimal Problems for Projective SfMPoster
  721. POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D ReconstructionPoster
  722. PRE-Mamba: A 4D State Space Model for Ultra-High-Frequent Event Camera DerainingPoster
  723. PRIMAL: Physically Reactive and Interactive Motor Model for Avatar LearningPoster
  724. PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding ProjectionPoster
  725. PRM: Photometric Stereo based Large Reconstruction ModelPoster
  726. PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt Relocation
  727. PROGRESSOR: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementPoster
  728. PROL : Rehearsal Free Continual Learning in Streaming Data via Prompt Online LearningPoster
  729. PS-Mamba: Spatial-Temporal Graph Mamba for Pose Sequence RefinementPoster
  730. PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival PredictionPoster
  731. PUMA: Empowering Unified MLLM with Multi-granular Visual GenerationPoster
  732. PUMPS: Skeleton-Agnostic Point-based Universal Motion Pre-Training for Synthesis in Human Motion TasksPoster
  733. PVChat: Personalized Video Chat with One-Shot LearningPoster
  734. PVMamba: Parallelizing Vision Mamba via Dynamic State AggregationPoster
  735. PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and ConsistencyPoster
  736. PanSt3R: Multi-view Consistent Panoptic SegmentationPoster
  737. PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMsPoster
  738. PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama ReconstructionPoster
  739. Parameter-Efficient Adaptation of Geospatial Foundation Models through Embedding DeflectionPoster
  740. Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion ModelsPoster
  741. PartField: Learning 3D Feature Fields for Part Segmentation and BeyondPoster
  742. Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training AccelerationPoster
  743. Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud LocalizationPoster
  744. Passing the Driving Knowledge TestPoster
  745. PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-ResolutionPoster
  746. PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask ConditionsPoster
  747. PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to HistopathologyPoster
  748. PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion ModelPoster
  749. Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative ModelsPoster
  750. Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsPoster
  751. Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion RepresentationPoster
  752. Performing Defocus Deblurring by Modeling its Formation ProcessPoster
  753. PersPose: 3D Human Pose Estimation with Perspective Encoding and Perspective RotationPoster
  754. PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion
  755. PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic DegradationPoster
  756. Personalized Federated Learning under Local SupervisionPoster
  757. Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery SimulationPoster
  758. Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous DistillationPoster
  759. Perspective-Invariant 3D Object DetectionPoster
  760. Perspective-aware 3D Gaussian Inpainting with Multi-view ConsistencyPoster
  761. Ph-GAN: Physics-Inspired GAN for Generating SAR Images Under Limited DataPoster
  762. Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentPoster
  763. Photolithography Overlay Map Generation with Implicit Knowledge Distillation Diffusion TransformerPoster
  764. PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object ModelingPoster
  765. PhysSplat: Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian SplattingPoster
  766. PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from VideosPoster
  767. Physical Degradation Model-Guided Interferometric Hyperspectral Reconstruction with Unfolding TransformerPoster
  768. Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language ModelsPoster
  769. Pi-GPS: Enhancing Geometry Problem Solving by Unleashing the Power of Diagrammatic InformationPoster
  770. Pinco: Position-induced Consistent Adapter for Diffusion Transformer in Foreground-conditioned InpaintingPoster
  771. PixTalk: Controlling Photorealistic Image Processing and Editing with LanguagePoster
  772. PixelStitch: Structure-Preserving Pixel-Wise Bidirectional Warps for Unsupervised Image StitchingPoster
  773. PlaceIt3D: Language-Guided Object Placement in Real 3D ScenesPoster
  774. PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language ModelsPoster
  775. Planar Affine Rectification from Local Change of Scale and OrientationPoster
  776. PlaneRAS: Learning Planar Primitives for 3D Plane RecoveryPoster
  777. Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video CaptioningPoster
  778. Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary SegmentationPoster
  779. PlugMark: A Plug-in Zero-Watermarking Framework for Diffusion ModelsPoster
  780. Point Cloud Self-supervised Learning via 3D to Multi-view Masked LearnerPoster
  781. PointGAC: Geometric-Aware Codebook for Masked Point ModelingPoster
  782. PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface ReconstructionPoster
  783. PolarAnything: Diffusion-based Polarimetric Image SynthesisPoster
  784. Polarimetric Neural Field via Unified Complex-Valued Wave RepresentationPoster
  785. Ponimator: Unfolding Interactive Pose for Versatile Human-human Interaction AnimationPoster
  786. PoseAnchor: Robust Root Position Estimation for 3D Human Pose EstimationPoster
  787. PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D DataPoster
  788. PossLoss: A Reliable and Sensitive Facial Landmark Detection Loss FunctionPoster
  789. Power of Cooperative Supervision: Multiple Teachers Framework for Advanced 3D Semi-Supervised Object DetectionPoster
  790. Preacher: Paper-to-Video Agentic SystemPoster
  791. Precise Action-to-Video Generation Through Visual Action PromptsPoster
  792. Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object UnderstandingPoster
  793. Preserve Anything: Controllable Image Synthesis with Object PreservationPoster
  794. Pretend Benign: A Stealthy Adversarial Attack by Exploiting Vulnerabilities in Cooperative PerceptionPoster
  795. Pretrained Reversible Generation as Unsupervised Visual Representation LearningPoster
  796. PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal ViewPoster
  797. PrimHOI: Compositional Human-Object Interaction via Reusable Primitives
  798. Princeton365: A Diverse Dataset with Accurate Camera PosePoster
  799. Principles of Visual Tokens for Efficient Video UnderstandingPoster
  800. Prior-aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose EstimationPoster
  801. Prior2Former - Evidential Modeling of Mask Transformers for Assumption-Free Open-World Panoptic SegmentationPoster
  802. Privacy-centric Deep Motion Retargeting for Anonymization of Skeleton-Based Motion VisualizationPoster
  803. ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis UsersPoster
  804. ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process JudgesPoster
  805. ProSAM: Enhancing the Robustness of SAM-based Visual Reference Segmentation with Probabilistic PromptsPoster
  806. Proactive Scene Decomposition and ReconstructionPoster
  807. ProbMED: A Probabilistic Framework for Medical Multimodal BindingPoster
  808. ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity RecognitionPoster
  809. Probabilistic Inertial Poser (ProbIP): Uncertainty-aware Human Motion Modeling from Sparse Inertial SensorsPoster
  810. Probabilistic Prototype Calibration of Vision-language Models for Generalized Few-shot Semantic SegmentationPoster
  811. Processing and acquisition traces in visual encoders: What does CLIP know about your camera?Poster
  812. Progressive Artwork Outpainting via Latent Diffusion ModelsPoster
  813. Progressive Distribution Bridging: Unsupervised Adaptation for Large-scale Pre-trained Models via Adaptive Auxiliary DataPoster
  814. Progressive Growing of Video Tokenizers for Temporally Compact Latent SpacesPoster
  815. Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental LearningPoster
  816. Progressive Test Time Energy Adaptation for Medical Image SegmentationPoster
  817. Prompt Guidance and Human Proximal Perception for HOT Prediction with Regional Joint LossPoster
  818. Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLMPoster
  819. Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual InversionPoster
  820. PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware MaskPoster
  821. PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity DiscriminationPoster
  822. Prototype Guided Backdoor Defense via Activation Space Manipulation
  823. Prototype-based Contrastive Learning with Stage-wise Progressive Augmentation for Self-Supervised Fine-Grained LearningPoster
  824. Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video RetrievalPoster
  825. Proxy-Bridged Game Transformer for Interactive Extreme Motion PredictionPoster
  826. Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language ModelsPoster
  827. Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic SegmentationPoster
  828. PseudoMapTrainer: Learning Online Mapping without HD MapsPoster
  829. Punching Bag vs. Punching Person: Motion Transferability in VideosPoster
  830. Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level DynamicsPoster
  831. Purge-Gate: Backpropagation-Free Test-Time Adaptation for Point Clouds Classification via Token purgingPoster
  832. Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene ReconstructionsPoster
  833. Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMsPoster
  834. Q-Norm: Robust Representation Learning via Quality-Adaptive NormalizationPoster
  835. QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingPoster
  836. QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized GenerationPoster
  837. QuEST: Low-bit Diffusion Model Quantization via Efficient Selective FinetuningPoster
  838. Quadratic Gaussian Splatting: High Quality Surface Reconstruction with Second-order Geometric PrimitivesPoster
  839. Quanta Neural Networks: From Photons to PerceptionPoster
  840. Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval via Uncertainty MinimizationPoster
  841. QuickSplat: Fast 3D Surface Reconstruction via Learned Gaussian InitializationPoster
  842. R-LiViT: A LiDAR-Visual-Thermal Dataset Enabling Vulnerable Road User Focused Roadside PerceptionPoster
  843. R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal FormalizationPoster
  844. R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy OptimizationPoster
  845. RA-BUSSeg: Relation-aware Semi-supervised Breast Ultrasound Image Segmentation via Adjacent Propagation and Cross-layer AlignmentPoster
  846. RAGD: Regional-Aware Diffusion Model for Text-to-Image GenerationPoster
  847. RAGDiffusion: Faithful Cloth Generation via External Knowledge AssimilationPoster
  848. RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General GraspingPoster
  849. RALoc: Enhancing Outdoor LiDAR Localization via Rotation AwarenessPoster
  850. RANKCLIP: Ranking-Consistent Language-Image PretrainingPoster
  851. RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot LearningPoster
  852. RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal FusionPoster
  853. REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion LatentsPoster
  854. REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderPoster
  855. REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion TransformersPoster
  856. REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout AlignmentPoster
  857. RESCUE: Crowd Evacuation Simulation via Controlling SDM-United CharactersPoster
  858. RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion PriorsPoster
  859. RI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion PriorsPoster
  860. RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionPoster
  861. RIPE: Reinforcement Learning on Unlabeled Image Pairs for Robust Keypoint ExtractionPoster
  862. RMultiplex200K: Toward Reliable Multimodal Process Supervision for Visual Language Models on TelecommunicationsPoster
  863. ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work ZonesPoster
  864. ROAR: Reducing Inversion Error in Generative Image WatermarkingPoster
  865. ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image GenerationPoster
  866. RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation ModelPoster
  867. RTMap: Real-Time Recursive Mapping with Change Detection and LocalizationPoster
  868. RadGPT: Constructing 3D Image-Text Tumor DatasetsPoster
  869. RadarSplat: Radar Gaussian Splatting for High-Fidelity Data Synthesis and 3D Reconstruction of Autonomous Driving ScenesPoster
  870. Radiant Foam: Real-Time Differentiable Ray TracingPoster
  871. RainbowPrompt: Diversity-Enhanced Prompt-Evolving for Continual LearningPoster
  872. Randomized Autoregressive Visual GenerationPoster
  873. RapVerse: Coherent Vocals and Whole-Body Motion Generation from TextPoster
  874. RareCLIP: Rarity-aware Online Zero-shot Industrial Anomaly DetectionPoster
  875. RayGaussX: Accelerating Gaussian-Based Ray Marching for Real-Time and High-Quality Novel View SynthesisPoster
  876. RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose EstimationPoster
  877. RayZer: A Self-supervised Large View Synthesis ModelPoster
  878. RayletDF: Raylet Distance Fields for Generalizable 3D Surface Reconstruction from Point Clouds or GaussiansPoster
  879. ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous DrivingPoster
  880. ReCamMaster: Camera-Controlled Generative Rendering from A Single VideoPoster
  881. ReCoT: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language ModelsPoster
  882. ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention AdaptationPoster
  883. ReME: A Data-Centric Framework for Training-Free Open-Vocabulary SegmentationPoster
  884. ReMP-AD: Retrieval-enhanced Multi-modal Prompt Fusion for Few-Shot Industrial Visual Anomaly DetectionPoster
  885. RePoseD: Efficient Relative Pose Estimation With Known Depth InformationPoster
  886. ReTracker: Exploring Image Matching for Robust Online Any Point TrackingPoster
  887. Real3D: Towards Scaling Large Reconstruction Models with Real ImagesPoster
  888. RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera ControlPoster
  889. RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video ModelsPoster
  890. Reangle-A-Video: 4D Video Generation as Video-to-Video TranslationPoster
  891. ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question AnsweringPoster
  892. ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco ReconstructionPoster
  893. Recognizing Actions from Robotic View for Natural Human-Robot InteractionPoster
  894. ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene RepresentationPoster
  895. Recover Biological Structure from Sparse-View Diffraction Images with Neural Volumetric PriorPoster
  896. Recovering Parametric Scenes from Very Few Time-of-Flight PixelsPoster
  897. Rectifying Magnitude Neglect in Linear AttentionPoster
  898. Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy RegularizationPoster
  899. RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring ExpressionsPoster
  900. Refer to Any Segmentation Mask Group With Vision-Language PromptsPoster
  901. ReferDINO: Referring Video Object Segmentation with Visual Grounding FoundationsPoster
  902. ReferEverything: Towards Segmenting Everything We Can Speak of in VideosPoster
  903. Reference-based Super-Resolution via Image-based Retrieval-Augmented Generation DiffusionPoster
  904. Referring Expression Comprehension for Small ObjectsPoster
  905. Referring to Any PersonPoster
  906. Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context ReflectionPoster
  907. RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS RegistrationPoster
  908. Region-Level Data Attribution for Text-to-Image Generative ModelsPoster
  909. Region-aware Anchoring Mechanism for Efficient Referring Visual GroundingPoster
  910. Region-based Cluster Discrimination for Visual Representation LearningPoster
  911. Registration beyond Points: General Affine Subspace Alignment via Geodesic Distance on Grassmann ManifoldPoster
  912. Reinforcement Learning-Guided Data Selection via Redundancy AssessmentPoster
  913. Relative Illumination Fields: Learning Medium and Light Independent Underwater ScenesPoster
  914. Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for PrivacyPoster
  915. Removing Cost Volumes from Optical Flow EstimatorsPoster
  916. Removing Out-of-Focus Reflective Flares via Color AlignmentPoster
  917. Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task LearningPoster
  918. Representation Shift: Unifying Token Compression with FlashAttentionPoster
  919. Representing 3D Shapes with 64 Latent Vectors for 3D Diffusion ModelsPoster
  920. Repurposing 2D Diffusion Models with Gaussian Atlas for 3D GenerationPoster
  921. ResGS: Residual Densification of 3D Gaussian for Efficient Detail RecoveryPoster
  922. ResQ: A Novel Framework to Implement Residual Neural Networks on Analog Rydberg Atom Quantum ComputersPoster
  923. ResidualViT for Efficient Temporally Dense Video EncodingPoster
  924. Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task LearningPoster
  925. Resonance: Learning to Predict Social-Aware Pedestrian Trajectories as Co-VibrationsPoster
  926. Rethink Sparse Signals for Pose-guided Text-to-image GenerationPoster
  927. Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction FrameworkPoster
  928. Rethinking Cross-Modal Interaction in Multimodal Diffusion TransformersPoster
  929. Rethinking DPO-style Diffusion Aligning FrameworksPoster
  930. Rethinking Detecting Salient and Camouflaged Objects in Unconstrained ScenesPoster
  931. Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image SynthesisPoster
  932. Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive SettingPoster
  933. Rethinking Key-frame-based Micro-expression Recognition: A Robust and Accurate Framework Against Key-frame ErrorsPoster
  934. Rethinking Layered Graphic Design Generation with a Top-Down ApproachPoster
  935. Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual DisparitiesPoster
  936. Retinex-MEF: Retinex-based Glare Effects Aware Unsupervised Multi-Exposure Image FusionPoster
  937. RetinexMCNet: A Memory Controller Dominated Network for Low-Light Video Enhancement Based on RetinexPoster
  938. Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image SetsPoster
  939. Revelio: Interpreting and leveraging semantic information in diffusion modelsPoster
  940. Reverse Convolution and Its Applications to Image RestorationPoster
  941. Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New InsightsPoster
  942. Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature AlignmentPoster
  943. Revisiting Image Fusion for Multi-Illuminant White-Balance CorrectionPoster
  944. Revisiting Point Cloud Completion: Are We Ready For The Real-World?Poster
  945. Revisiting Pool-based Prompt Learning for Few-shot Class-incremental LearningPoster
  946. RhythmGuassian: Repurposing Generalizable Gaussian Model For Remote Physiological MeasurementPoster
  947. Riemannian-Geometric Fingerprints of Generative ModelsPoster
  948. RnGCam: High-speed video from rolling & global shutter measurementsPoster
  949. RoCo-Sim: Enhancing Roadside Collaborative Perception through Foreground SimulationPoster
  950. RoMo: Robust Motion Segmentation Improves Structure from MotionPoster
  951. RobAVA: A Large-scale Dataset and Baseline Towards Video based Robotic Arm Action UnderstandingPoster
  952. Robin3D: Improving 3D Large Language Model via Robust Instruction TuningPoster
  953. RoboAnnotatorX: A Comprehensive and Universal Annotation Framework for Accurate Understanding of Long-horizon Robot DemonstrationPoster
  954. RoboPearls: Editable Video Simulation for Robot ManipulationPoster
  955. RoboTrom-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionPoster
  956. RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous DrivingPoster
  957. RoboTron-Mani: All-in-One Multimodal Large Model for Robotic ManipulationPoster
  958. RoboTron-Sim: Improving Real-World Driving via Simulated Hard-CasePoster
  959. Robust 3D Object Detection using Probabilistic Point Clouds from Single-Photon LiDARs
  960. Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation SamplingPoster
  961. Robust Adverse Weather Removal via Spectral-based Spatial GroupingPoster
  962. Robust Dataset Condensation using Supervised Contrastive LearningPoster
  963. Robust Low-light Scene Restoration via Illumination TransitionPoster
  964. Robust Machine Unlearning for Quantized Neural Networks via Adaptive Gradient Reweighting with Similar LabelsPoster
  965. Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated PerturbationPoster
  966. Robust Test-Time Adaptation for Single Image Denoising Using Deep Gaussian PriorPoster
  967. Robust Unfolding Network for HDR Imaging with Modulo CamerasPoster
  968. Robust and Efficient 3D Gaussian Splatting for Urban Scene ReconstructionPoster
  969. RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGSPoster
  970. Robustifying Zero-Shot Vision Language Models by Subspaces AlignmentPoster
  971. RogSplat: Robust Gaussian Splatting via Generative PriorsPoster
  972. RomanTex: Decoupling 3D-aware Rotary Positional Embedded Multi-Attention Network for Texture SynthesisPoster
  973. Ross3D: Reconstructive Visual Instruction Tuning with 3D-AwarenessPoster
  974. S2M2: Scalable Stereo Matching Model for Reliable Depth EstimationPoster
  975. S3E: Self-Supervised State Estimation for Radar-Inertial SystemPoster
  976. S3R-GS: Streamlining the Pipeline for Large-Scale Street Scene ReconstructionPoster
  977. S4M: Boosting Semi-Supervised Instance Segmentation with SAMPoster
  978. SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style TransferPoster
  979. SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real WorldPoster
  980. SAC-GNC: SAmple Consensus for adaptive Graduated Non-ConvexityPoster
  981. SAFER: Sharpness Aware layer-selective Finetuning for Enhanced Robustness in vision transformersPoster
  982. SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular VideoPoster
  983. SAGI: Semantically Aligned and Uncertainty Guided AI Image InpaintingPoster
  984. SALAD -- Semantics-Aware Logical Anomaly DetectionPoster
  985. SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model FailuresPoster
  986. SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory TreePoster
  987. SAM4D: Segment Anything in Camera and LiDAR StreamsPoster
  988. SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of ExpertsPoster
  989. SAMO: A Lightweight Sharpness-Aware Approach for Multi-Task Optimization with Joint Global-Local PerturbationPoster
  990. SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action RecognitionPoster
  991. SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical ImagesPoster
  992. SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationPoster
  993. SAS: Segment Any 3D Scene with Integrated 2D PriorsPoster
  994. SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse AutoencodersPoster
  995. SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement LearningPoster
  996. SC-Lane: Slope-aware and Consistent Road Height Estimation Framework for 3D Lane DetectionPoster
  997. SCAN: Bootstrapping Contrastive Pre-training for Data EfficiencyPoster
  998. SCFlow: Implicitly Learning Style and Content Disentanglement with Flow ModelsPoster
  999. SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance SegmentationPoster
  1000. SD2Actor: Continuous State Decomposition via Diffusion Embeddings for Robotic ManipulationPoster

Looking for submission deadlines instead? See the conference deadline calendar.