← Search

Guangming Lu

45 accepted papers

2026

Beyond Heuristics: Learnable Density Control for 3D Gaussian Splatting

ICML 2026poster

While 3D Gaussian Splatting (3DGS) has demonstrated impressive real-time rendering performance, its efficacy remains constrained by a reliance on heuristic density control. Despite numerous refinements to these handcrafted rules, such methods inherently lack the flexibility to adapt to diverse scene…

Cited by 0SourceScholar
2026

Coverage ≠ Exposure: Auditable Control of Same-Support Tail Failures under Multimodal Missingness

ICML 2026poster

Real-world multimodal systems inevitably face partial observability due to sensor dropout and degradation. Standard robustness methods can improve average performance, but they often remain unreliable in rare, adverse long-tail conditions. Under a locked same-support contract, we uncover a same-supp…

Cited by 0SourceScholar
2026

DiffTrans: Differentiable Geometry-Materials Decomposition for Reconstructing Transparent Objects

ICLR 2026poster

Reconstructing transparent objects from a set of multi-view images is a challenging task due to the complicated nature and indeterminate behavior of light propagation. Typical methods are primarily tailored to specific scenarios, such as objects following a uniform topology, exhibiting ideal transpa…

Cited by 0SourceScholar
2026

High-Fidelity Virtual Try-On beyond Paired Data Scarcity via Diffusion-based Cycle-Consistent Learning

CVPR 2026

Diffusion-based virtual try-on methods rely on vast high-quality garment-person pairs, which are scarce in practice due to the high cost of data collection and preprocessing, limiting their performance in real-world scenarios.To overcome this bottleneck, we propose Cycle-Consistent Virtual Try-On (C

Cited by 0SourceScholar
2026

PointRePar : SpatioTemporal Point Relation Parsing for Robust Category-Unified 3D Tracking

ICLR 2026poster

3D single object tracking (SOT) remains a highly challenging task due to the inherent crux in learning representations from point clouds to effectively capture both spatial shape features and temporal motion features. Most existing methods employ a category-specific optimization paradigm, training t…

Cited by 0SourceScholar
2026

Prompt Tuning for CLIP on the Pretrained Manifold

ICML 2026poster

Prompt tuning introduces learnable prompt vectors that adapt pretrained vision-language models to downstream tasks in a parameter-efficient manner. However, under limited supervision, prompt tuning alters pretrained representations and drives downstream features away from the pretrained manifold tow…

Cited by 0SourceScholar
2026

Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity

ICLR 2026poster

Vision-language models (VLMs) face significant computational inefficiencies caused by excessive generation of visual tokens. While prior work shows that a large fraction of visual tokens are redundant, existing compression methods struggle to balance \textit{importance preservation} and \textit{info…

Cited by 0SourcecodeScholar
2026

ReMoE: Region-Mixture Experts for Adversarially-Robust Vision Transformers

CVPR 2026

Vision Transformers (ViTs) achieve state-of-the-art performance on a wide range of vision tasks, yet they remain highly vulnerable to adversarial perturbations due to the lack of explicit region-level semantic modeling. Adversarial perturbations are typically local and spatially structured, whereas

Cited by 0SourcecodeScholar
2026

Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity

CVPR 2026

Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authentic to real-world photography remains challenging. This is partly due to biases in existing evaluation paradigms: human ratings and preference-trained metrics often

Cited by 0SourcecodeScholar
2025

A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers

NeurIPS 2025poster

Poison-only Clean-label Backdoor Attacks (PCBAs) aim to covertly inject attacker-desired behavior into DNNs by merely poisoning the dataset without changing the labels. To effectively implant a backdoor, multiple triggers are proposed for various attack requirements of Attack Success Rate (ASR) and…

Cited by 0SourceScholar
2025

ALRMR-GEC: Adjusting Learning Rate Based on Memory Rate to Optimize the Edit Scorer for Grammatical Error Correction

AAAI 2025technical

Edit-based approaches for Grammatical Error Correction (GEC) have attracted volume attention due to their outstanding explanations of the correction process and rapid inference. Through exploring the characteristics of the generalized and specific knowledge learning for GEC, we discover that efficie…

2025

Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases

IJCAI 2025

Existing Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, t

2025

D2ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition

ICCV 2025poster

Adapting pre-trained image models to video modality has proven to be an effective strategy for robust few-shot action recognition. In this work, we explore the potential of adapter tuning in image-to-video model adaptation and propose a novel video adapter tuning framework, called Disentangled-and-D…

2025

EditInfinity: Image Editing with Binary-Quantized Generative Models

NeurIPS 2025poster

Adapting pretrained diffusion-based generative models for text-driven image editing with negligible tuning overhead has demonstrated remarkable potential. A classical adaptation paradigm, as followed by these methods, first infers the generative trajectory inversely for a given source image by image…

Cited by 0SourcecodeScholar
2025

Efficient and Separate Authentication Image Steganography Network

ICML 2025spotlight

Image steganography hides multiple images for multiple recipients into a single cover image. All secret images are usually revealed without authentication, which reduces security among multiple recipients. It is elegant to design an authentication mechanism for isolated reception. We explore such me…

2025

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

ICCV 2025poster

Recent advances in point cloud perception have demonstrated remarkable progress in scene understanding through vision-language alignment leveraging large language models (LLMs). However, existing methods may still encounter challenges in handling complex instructions that require accurate spatial re…

Cited by 0SourcePDFScholar
2025

Learning Compatible Multi-Prize Subnetworks for Asymmetric Retrieval

CVPR 2025poster

Asymmetric retrieval is a typical scenario in real-world retrieval systems, where compatible models of varying capacities are deployed on platforms with different resource configurations. Existing methods generally train pre-defined networks or subnetworks with capacities specifically designed for p…

2025

SSHR: More Secure Generative Steganography with High-Quality Revealed Secret Images

ICML 2025poster

Image steganography ensures secure information transmission and storage by concealing secret messages within images. Recently, the diffusion model has been incorporated into the generative image steganography task, with text prompts being employed to guide the entire process. However, existing metho…

Cited by 0SourcePDFScholar
2025

Towards Robust Visual Question Answering via Prompt-Driven Geometric Harmonization

AAAI 2025technical

Visual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations pos…

Cited by 0SourcePDFScholar
2024

CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive Learning

AAAI 2024technical

Dental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intrica…

2024

Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification

ICASSP 2024accepted

The feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across d…

Cited by 0SourceScholar
2024

Domain-Rectifying Adapter for Cross-Domain Few-Shot Segmentation

CVPR 2024poster

Few-shot semantic segmentation (FSS) has achieved great success on segmenting objects of novel classes supported by only a few annotated samples. However existing FSS methods often underperform in the presence of domain shifts especially when encountering new domain styles that are unseen during tra…

2024

Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing

IJCAI 2024poster

Cross-modal hashing has garnered considerable research interest due to its rapid retrieval and low storage costs. However, the majority of existing methods suffer from the limitations of context loss and information redundancy, particularly in simulated textual environments enriched with manually an…

Cited by 3SourcePDFScholar
2024

Robust 3D Tracking with Quality-Aware Shape Completion

AAAI 2024technical

3D single object tracking remains a challenging problem due to the sparsity and incompleteness of the point clouds. Existing algorithms attempt to address the challenges in two strategies. The first strategy is to learn dense geometric features based on the captured sparse point cloud. Nevertheless,…

Cited by 6SourcePDFScholar
2024

SA²VP: Spatially Aligned-and-Adapted Visual Prompt

AAAI 2024technical

As a prominent parameter-efficient fine-tuning technique in NLP, prompt tuning is being explored its potential in computer vision. Typical methods for visual prompt tuning follow the sequential modeling paradigm stemming from NLP, which represents an input image as a flattened sequence of token embe…

2024

TAROT: A Hierarchical Framework with Multitask co-pretraining on Semi-Structured Data Towards Effective Person-Job fit

ICASSP 2024accepted

Person-job fit is an essential part of online recruitment platforms in serving various downstream applications like Job Search and Candidate Recommendation. Recently, pretrained large language models have further enhanced the effectiveness by leveraging richer textual information in user profiles an…

Cited by 0SourceScholar
2024

UniVoxel: Fast Inverse Rendering by Unified Voxelization of Scene Representation

ECCV 2024poster

"Typical inverse rendering methods focus on learning implicit neural scene representations by modeling the geometry, materials and illumination separately, which entails significant computations for optimization. In this work we design a Unified Voxelization framework for explicit learning of scene…

2024

WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-only Supervised Text Spotting

ECCV 2024poster

"Transcription-only Supervised Text Spotting aims to learn text spotters relying only on transcriptions but no text boundaries for supervision, thus eliminating expensive boundary annotation. The crux of this task lies in locating each transcription in scene text images without location annotations.…

2023

Hierarchical Contrastive Learning for Pattern-Generalizable Image Corruption Detection

ICCV 2023accepted

Effective image restoration with large-size corruptions, such as blind image inpainting, entails precise detection of corruption region masks which remains extremely challenging due to diverse shapes and patterns of corruptions. In this work, we present a novel method for automatic corruption detect…

2022

DHWP: Learning High-Quality Short Hash Codes Via Weight Pruning

ICASSP 2022accepted

Hashing is widely used in large-scale image retrieval because of its efficiency in storage and computation. Although longer hash codes can lead to higher search accuracy, the retrieval cost increases linearly with the increase of the number of hash bits. Most deep hashing methods suffer from the pro…

Cited by 0SourceScholar
2022

Few-Shot Object Detection by Knowledge Distillation Using Bag-of-Visual-Words Representations

ECCV 2022poster

"While fine-tuning based methods for few-shot object detection have achieved remarkable progress, a crucial challenge that has not been addressed well is the potential class-specific overfitting on base classes and sample-specific overfitting on novel classes. In this work we design a novel knowledg…

Cited by 18SourcePDFScholar
2022

Improved Deep Unsupervised Hashing with Fine-grained Semantic Similarity Mining for Multi-Label Image Retrieval

IJCAI 2022poster

In this paper, we study deep unsupervised hashing, a critical problem for approximate nearest neighbor research. Most recent methods solve this problem by semantic similarity reconstruction for guiding hashing network learning or contrastive learning of hash codes. However, in multi-label scenarios,…

Cited by 16SourcePDFScholar
2022

Learning Modal-Invariant and Temporal-Memory for Video-Based Visible-Infrared Person Re-Identification

CVPR 2022poster

Thanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the "probe-to-gallery", almost all existing RGB-IR based cro…

Cited by 64PDFcodeScholar
2022

Multi-faceted Distillation of Base-Novel Commonality for Few-Shot Object Detection

ECCV 2022poster

"Most of existing methods for few-shot object detection follow the fine-tuning paradigm, which potentially assumes that the class-agnostic generalizable knowledge can be learned and transferred implicitly from base classes with abundant samples to novel classes with limited samples via such a two-st…

2022

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

EMNLP 2022main

Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a…

2021

A Bipolar Myoelectric Sensor-Enabled Human-Machine Interface Based On Spinal Module Activations

ICRA 2021poster

The surface electromyography (sEMG) signal-based human-machine interface (HMI) has been widely used for various scenarios of physical human-robot interaction. However, current HMIs based on bipolar myoelectric sensors are hindered by the limitations of global sEMG features, which are prone to variab…

Cited by 3SourceScholar
2021

Bidirectional Hierarchical Attention Networks based on Document-level Context for Emotion Cause Extraction

EMNLP 2021finding

Emotion cause extraction (ECE) aims to extract the causes behind the certain emotion in text. Some works related to the ECE task have been published and attracted lots of attention in recent years. However, these methods neglect two major issues: 1) pay few attentions to the effect of document-level…

2021

Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion Recognition

ICASSP 2021accepted

Many studies on automatic speech emotion recognition (SER) have been devoted to extracting meaningful emotional features for generating emotion-relevant representations. However, they generally ignore the complementary learning of static and dynamic features, leading to limited performances. In this…

Cited by 0SourceScholar
2021

Prototype-Supervised Adversarial Network for Targeted Attack of Deep Hashing

CVPR 2021poster

Due to its powerful capability of representation learning and high-efficiency computation, deep hashing has made significant progress in large-scale image retrieval. However, deep hashing networks are vulnerable to adversarial examples, which is a practical secure problem but seldom studied in hashi…

Cited by 62PDFcodeScholar