← Search

Sheng Zhong

33 accepted papers

2026

Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning

ICML 2026poster

Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concerns regarding privacy and storage. Thus, generative replay has emerged as a viable alternative, synthesizing old data usin…

Cited by 0SourceScholar
2026

Diverse Human Driving Vehicle Simulation in Background Traffic for Autonomous Driving Tests

AAAI 2026technical

Realistic background traffic is critical to the simulation platforms for autonomous driving (AD) testing. Given that most vehicles in reality are driven by human beings, introducing human driving (HD) vehicles to the background traffic is necessary to be able to discover more problems of the tested

Cited by 0SourcePDFScholar
2026

Learnability-Driven Knowledge Assimilation for Class-Incremental Semantic Segmentation

ICML 2026poster

Class-incremental semantic segmentation learns new classes while retaining old ones without access to past data. Although existing methods alleviate catastrophic forgetting on old classes, new-class performance remains limited. We identify the key bottleneck arises from low-margin regions, where the…

Cited by 0SourceScholar
2026

MSCD-GS: Motion-Separated Cooperative Deblurring Dynamic Reconstruction via Gaussian Splatting

CVPR 2026

Although 4D reconstruction based on Gaussian Splatting has achieved many impressive results, reconstructing real-world images captured by a casual monocular camera remains a significant challenge. In dynamic scenes, as the camera and objects move during the exposure time, these input images inevitab

Cited by 0SourceScholar
2026

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

ICML 2026poster

Despite the rapid progress of Vision-Language-Action (VLA) models, the prevailing paradigm of predicting discrete waypoints remains fundamentally misaligned with the intrinsic continuity of physical motion. This discretization imposes rigid sampling rates, lacks high-order differentiability, and int…

Cited by 0SourceScholar
2026

Real-Time Motion Segmentation with Event-Based Normal Flow

ICRA 2026poster

Event-based cameras are bio-inspired sensors with pixels that independently and asynchronously respond to brightness changes at microsecond resolution, offering the potential to handle visual tasks in challenging scenarios. However, due to the sparse information content in individual events, directl…

2026

TriFusion-IDS: A Multimodal Graph-Tabular-Text Contrastive Framework for Cross-Dataset Intrusion Detection

AAAI 2026technical

Traditional Intrusion Detection Systems (IDS) are typically trained in specific network environments, and their performance often degrades significantly when deployed in new environments with different attack categories. To address this challenge, we propose and define the task of cross-dataset intr

Cited by 0SourcePDFScholar
2025

DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes

AAAI 2025technical

Vision-centric autonomous driving systems require diverse data for robust training and evaluation, which can be augmented by manipulating object positions and appearances within existing scene captures. While recent advancements in diffusion models have shown promise in video editing, their applicat…

2025

High-dimension Prototype is a Better Incremental Object Detection Learner

ICLR 2025poster

Incremental object detection (IOD), surpassing simple classification, requires the simultaneous overcoming of catastrophic forgetting in both recognition and localization tasks, primarily due to the significantly higher feature space complexity. Integrating Knowledge Distillation (KD) would mitigate…

Cited by 0SourcePDFScholar
2025

Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task Learning

AAAI 2025technical

Generative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often…

Cited by 0SourcePDFScholar
2025

MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM

IROS 2025

Recent advancements in 3D Gaussian Splatting (3DGS) have made a significant impact on rendering and reconstruction techniques. Current research predominantly focuses on improving rendering performance and reconstruction quality using high-performance desktop GPUs, largely overlooking applications fo

Cited by 3SourceScholar
2025

ObfusLM: Privacy-preserving Language Model Service against Embedding Inversion Attacks

ACL 2025long

As the rapid expansion of Machine Learning as a Service (MLaaS) for language models, concerns over the privacy of client inputs during inference or fine-tuning have correspondingly escalated. Recently, solutions have been proposed to safeguard client privacy by obfuscation techniques. However, the s…

2025

SAP: Privacy-Preserving Fine-Tuning on Language Models with Split-and-Privatize Framework

IJCAI 2025

Pre-trained Language Models (PLM) have enabled a cost-effective approach to handling various downstream applications via Parameter-Efficient-Fine-Tuning (PEFT) techniques. In this context, service providers have introduced a popular fine-tuning-based product service known as Model-as-a-Service (MaaS

Cited by 0SourcePDFScholar
2025

The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations

EMNLP 2025

Estimating the difficulty of input questions as perceived by large language models (LLMs) is essential for accurate performance evaluation and adaptive inference. Existing methods typically rely on repeated response sampling, auxiliary models, or fine-tuning the target model itself, which may incur

Cited by 0SourcePDFScholar
2024

Make Lossy Compression Meaningful for Low-Light Images

AAAI 2024technical

Low-light images frequently occur due to unavoidable environmental influences or technical limitations, such as insufficient lighting or limited exposure time. To achieve better visibility for visual perception, low-light image enhancement is usually adopted. Besides, lossy image compression is vita…

2024

Molecule Design by Latent Prompt Transformer

NeurIPS 2024spotlight

This work explores the challenging problem of molecule design by framing it as a conditional generative modeling task, where target biological properties or desired chemical constraints serve as conditioning variables. We propose the Latent Prompt Transformer (LPT), a novel generative model comprisi…

Cited by 2SourcePDFScholar
2024

MuSR: Multi-Scale 3D Scenes Reconstruction based on Monocular Video

ICASSP 2024accepted

Three-dimensional (3D) scene reconstruction, particularly from monocular videos, is a significant challenge in large-scale scenarios due to difficulty handling varying object sizes and high computational resource needs. This paper introduces MuSR, a novel multi-scale reconstruction method addressing…

Cited by 0SourceScholar
2024

SNIDA: Unlocking Few-Shot Object Detection with Non-linear Semantic Decoupling Augmentation

CVPR 2024poster

Once only a few-shot annotated samples are available the performance of learning-based object detection would be heavily dropped. Many few-shot object detection (FSOD) methods have been proposed to tackle this issue by adopting image-level augmentations in linear manners. Nevertheless those handcraf…

Cited by 9SourcePDFScholar
2024

VIDAR: Data Quality Improvement for Monocular 3D Reconstruction through In-situ Visual Interaction

ICRA 2024poster

3D reconstruction based on monocular videos has attracted wide attention, and existing reconstruction methods usually work in a reconstruction-after-scanning manner. However, these methods suffer from insufficient data collection problems due to the lack of effective guidance for users during the sc…

Cited by 2SourceScholar
2023

CHSEL: Producing Diverse Plausible Pose Estimates from Contact and Free Space Data

RSS 2023poster

This paper proposes a novel method for estimating the set of plausible poses of a rigid object from a set of points with volumetric information, such as whether each point is in free space or on the surface of the object. In particular, we study how pose can be estimated from force and tactile data…

2022

Category-Aware Transformer Network for Better Human-Object Interaction Detection

CVPR 2022poster

Human-Object Interactions (HOI) detection, which aims to localize a human and a relevant object while recognizing their interaction, is crucial for understanding a still image. Recently, tranformer-based models have significantly advanced the progress of HOI detection. However, the capability of the…

Cited by 46PDFScholar
2022

Improving Human-Object Interaction Detection via Phrase Learning and Label Composition

AAAI 2022technical

Human-Object Interaction (HOI) detection is a fundamental task in high-level human-centric scene understanding. We propose PhraseHOI, containing a HOI branch and a novel phrase branch, to leverage language prior and improve relation expression. Specifically, the phrase branch is supervised by semant…

Cited by 44SourcePDFScholar
2022

Soft Tracking Using Contacts for Cluttered Objects to Perform Blind Object Retrieval

RA-L 2022

Retrieving an object from cluttered spaces such as cupboards, refrigerators, or bins requires tracking objects with limited or no visual sensing. In these scenarios, contact feedback is necessary to estimate the pose of the objects, yet the objects are movable while their shapes and number may be un

Cited by 16SourcecodeScholar
2021

TAMPC: A Controller for Escaping Traps in Novel Environments

RA-L 2021

We propose an approach to online model adaptation and control in the challenging case of hybrid and discontinuous dynamics where actions may lead to difficult-to-escape “trap” states, under a given controller. We first learn dynamics for a system without traps from a randomly collected training set

Cited by 8SourcecodeScholar
2019

Learning Robust Facial Landmark Detection via Hierarchical Structured Ensemble

ICCV 2019poster

Heatmap regression-based models have significantly advanced the progress of facial landmark detection. However, the lack of structural constraints always generates inaccurate heatmaps resulting in poor landmark detection performance. While hierarchical structure modeling methods have been proposed t…

Cited by 80PDFScholar
2017

Hyper-Laplacian Regularized Unidirectional Low-Rank Tensor Recovery for Multispectral Image Denoising

CVPR 2017poster

Recent low-rank based matrix/tensor recovery methods have been widely explored in multispectral images (MSI) denoising. These methods, however, ignore the difference of the intrinsic structure correlation along spatial sparsity, spectral correlation and non-local self-similarity mode. In this paper,…

Cited by 220PDFScholar