← Search

Hongtao Lu

31 accepted papers

2026

Dejavu: Towards Experience Feedback Learning for Embodied Intelligence

CVPR 2026

Embodied agents face a fundamental limitation: once deployed in real-world environments, they cannot easily acquire new knowledge to improve task performance. In this paper, we propose Dejavu, a general post-deployment learning framework that augments a frozen Vision-Language-Action (VLA) policy wit

Cited by 0SourcecodeScholar
2026

MVP-Nav: Multi-layer Value Map Planner Navigator

RSS 2026poster

Zero-Shot Object Goal Navigation (ZSON) is an important task for robots. While Multimodal Large Language Models (MLLMs) have empowered robots with significant semantic reasoning capabilities, current RGB-only navigation methods still struggle to align high-level discrete logic with low-level continu…

Cited by 0SourceScholar
2026

Recovering Hidden Reward in Diffusion-Based Policies

ICML 2026poster

This paper introduces EnergyFlow, a framework that unifies generative action modeling with inverse reinforcement learning by parameterizing a scalar energy function whose gradient is the denoising field. We establish that under maximum-entropy optimality, the score function learned via denoising sco…

Cited by 0SourceScholar
2025

BECAME: Bayesian Continual Learning with Adaptive Model Merging

ICML 2025poster

Continual Learning (CL) strives to learn incrementally across tasks while mitigating catastrophic forgetting. A key challenge in CL is balancing stability (retaining prior knowledge) and plasticity (learning new tasks). While representative gradient projection methods ensure stability, they often li…

2025

Discretized Gaussian Representation for Tomographic Reconstruction

ICCV 2025poster

Computed Tomography (CT) enables detailed cross-sectional imaging but continues to face challenges in balancing reconstruction quality and computational efficiency. While deep learning-based methods have significantly improved image quality and noise reduction, they typically require large-scale tra…

2025

Few-shot Implicit Function Generation via Equivariance

CVPR 2025highlight

Implicit Neural Representations (INRs) have emerged as a powerful framework for representing continuous signals. However, generating diverse INR weights remains challenging due to limited training data. We introduce Few-shot Implicit Function Generation, a new problem setup that aims to generate div…

2025

How Does Topology Bias Distort Message Passing in Graph Recommender? A Dirichlet Energy Perspective

NeurIPS 2025poster

Graph-based recommender systems have achieved remarkable effectiveness by modeling high-order interactions between users and items. However, such approaches are significantly undermined by popularity bias, which distorts the interaction graph’s structure—referred to as topology bias. This leads to o…

Cited by 0SourcecodeScholar
2024

Changenet: Multi-Temporal Asymmetric Change Detection Dataset

ICASSP 2024accepted

Change Detection (CD) has been attracting extensive interests with the availability of bi-temporal datasets. However, due to the huge cost of multi-temporal images acquisition and labeling, existing change detection datasets are small in quantity, short in temporal, and low in practicability. Theref…

Cited by 0SourceScholar
2024

Discrete Latent Perspective Learning for Segmentation and Detection

ICML 2024spotlight

In this paper, we address the challenge of Perspective-Invariant Learning in machine learning and computer vision, which involves enabling a network to understand images from varying perspectives to achieve consistent semantic interpretation. While standard approaches rely on the labor-intensive col…

Cited by 9SourcePDFScholar
2024

FedHCA2: Towards Hetero-Client Federated Multi-Task Learning

CVPR 2024poster

Federated Learning (FL) enables joint training across distributed clients using their local data privately. Federated Multi-Task Learning (FMTL) builds on FL to handle multiple tasks assuming model congruity that identical model architecture is deployed in each client. To relax this assumption and t…

2024

Task Indicating Transformer for Task-Conditional Dense Predictions

ICASSP 2024accepted

The task-conditional model is a distinctive stream for efficient multi-task learning. Existing works encounter a critical limitation in learning task-agnostic and task-specific representations, primarily due to shortcomings in global context modeling arising from CNN-based architectures, as well as…

Cited by 0SourceScholar
2024

UNIDEAL: Curriculum Knowledge Distillation Federated Learning

ICASSP 2024accepted

Federated Learning (FL) has emerged as a promising approach to enable collaborative learning among multiple clients while preserving data privacy. However, cross-domain FL tasks, where clients possess data from different domains or distributions, remain a challenging problem due to the inherent hete…

Cited by 0SourceScholar
2024

YOLO-Med : Multi-Task Interaction Network for Biomedical Images

ICASSP 2024accepted

Object detection and semantic segmentation are pivotal components in biomedical image analysis. Current single-task networks exhibit promising outcomes in both detection and segmentation tasks. Multi-task networks have gained prominence due to their capability to simultaneously tackle segmentation a…

Cited by 0SourceScholar
2023

A Simple Framework for Text-Supervised Semantic Segmentation

CVPR 2023poster

Text-supervised semantic segmentation is a novel research topic that allows semantic segments to emerge with image-text contrasting. However, pioneering methods could be subject to specifically designed network architectures. This paper shows that a vanilla contrastive language-image pre-training (C…

2023

Generating Human Motion From Textual Descriptions With Discrete Representations

CVPR 2023poster

In this work, we investigate a simple and must-known conditional generative framework based on Vector Quantised-Variational AutoEncoder (VQ-VAE) and Generative Pre-trained Transformer (GPT) for human motion generation from textural descriptions. We show that a simple CNN-based VQ-VAE with commonly u…

2023

Guided Patch-Grouping Wavelet Transformer with Spatial Congruence for Ultra-High Resolution Segmentation

IJCAI 2023poster

Most existing ultra-high resolution (UHR) segmentation methods always struggle in the dilemma of balancing memory cost and local characterization accuracy, which are both taken into account in our proposed Guided Patch-Grouping Wavelet Transformer (GPWFormer) that achieves impressive performances. I…

Cited by 17SourcePDFScholar
2023

Spammer Detection on Short Video Applications: A new Challenge and Baselines

ICASSP 2023accepted

Users can interact with the advertisements and share their impressions through the review system on short video applications. However, spammers may post false or malicious comments to mislead normal users due to profit-driven reasons, damaging the community’s positive atmosphere. In this paper, we i…

Cited by 0SourceScholar
2023

Stuart: Individualized Classroom Observation of Students with Automatic Behavior Recognition And Tracking

ICASSP 2023accepted

Each student matters, but it is hardly for instructors to observe all the students during the courses and provide helps to the needed ones immediately. In this paper, we present StuArt, a novel automatic system designed for the individualized classroom observation, which empowers instructors to conc…

Cited by 0SourceScholar
2023

Ultra-High Resolution Segmentation With Ultra-Rich Context: A Novel Benchmark

CVPR 2023poster

With the increasing interest and rapid development of methods for Ultra-High Resolution (UHR) segmentation, a large-scale benchmark covering a wide range of scenes with full fine-grained dense annotations is urgently needed to facilitate the field. To this end, the URUR dataset is introduced, in the…

2022

Structural and Statistical Texture Knowledge Distillation for Semantic Segmentation

CVPR 2022poster

Existing knowledge distillation works for semantic segmentation mainly focus on transfering high-level contextual knowledge from teacher to student. However, low-level texture knowledge is also of vital importance for characterizing the local structural pattern and global statistical property, such…

Cited by 77PDFScholar
2021

Laplacian Regularized Tensor Low-Rank Minimization for Hyperspectral Snapshot Compressive Imaging

ICASSP 2021accepted

Snapshot Compressive Imaging (SCI) systems, including hyperspectral compressive imaging and video compressive imaging, are designed to depict high-dimensional signals with limited data by mapping multiple images into one. One key module of SCI systems is a high quality reconstruction algorithm for o…

Cited by 0SourceScholar
2020

Learning Texture Transformer Network for Image Super-Resolution

CVPR 2020poster

We study on image super-resolution (SR), which aims to recover realistic textures from a low-resolution (LR) image. Recent progress has been made by taking high-resolution images as references (Ref), so that relevant textures can be transferred to LR images. However, existing SR approaches neglect t…

Cited by 1107PDFcodeScholar
2019

Attribute-Driven Feature Disentangling and Temporal Aggregation for Video Person Re-Identification

CVPR 2019poster

Video-based person re-identification plays an important role in surveillance video analysis, expanding image-based methods by learning features of multiple frames. Most existing methods fuse features by temporal average-pooling, without exploring the different frame weights caused by various viewpoi…

Cited by 184PDFScholar
2018

An Adversarial Approach to Hard Triplet Generation

ECCV 2018poster

While deep neural networks have demonstrated competitive results for many visual recognition and image retrieval tasks, the major challenge lies in distinguishing similar images from different categories (i.e., hard negative examples) while clustering images with large variations from the same categ…

Cited by 121SourcePDFScholar
2018

Anisotropic Total Variation Regularized Low-Rank Tensor Completion Based On Tensor Nuclear Norm for Color Image Inpainting

ICASSP 2018accepted

In this paper, we propose a novel low-rank tensor completion (LRTC) model under the circulant algebra for color image inpainting, which simultaneously preserves the low-rank structures of images, and also explore the local smooth and piecewise priors of the images in the spatial domain. First, color…

Cited by 0SourceScholar