← Search

Cheng Shi

22 accepted papers

2025

A Method for Removing Reflections from Water Surface Images Based on Pre-trained Image Restoration

ICASSP 2025accepted

Reflections on the water surface hinder the extraction of valuable information from water surface images. To remove reflections from water surface images, we construct a synthetic dataset and propose a multi-task network for water surface reflection detection and removal. Specifically, we first use…

Cited by 0SourceScholar
2025

Discovering Compositional Hallucinations in LVLMs

NeurIPS 2025poster

Large language models (LLMs) and vision-language models (LVLMs) have driven the paradigm shift towards general-purpose foundation models. However, both of them are prone to hallucinations, which compromise their factual accuracy and reliability. While existing research primarily focuses on isolated…

Cited by 0SourceScholar
2025

Joint Graph Rewiring and Feature Denoising via Spectral Resonance

ICLR 2025oral

When learning from graph data, the graph and the node features both give noisy information about the node labels. In this paper we propose an algorithm to **j**ointly **d**enoise the features and **r**ewire the graph (JDR), which improves the performance of downstream node classification graph neura…

2025

Rethinking Query-based Transformer for Continual Image Segmentation

CVPR 2025poster

Class-incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-in objectness of query-based transformers, which mitigates catastrophic forgetting of mask proposals, current methods of…

2025

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

ICCV 2025poster

Temporal sentence grounding aims to identify exact moments in a video that correspond to a given textual query, typically addressed with detection transformer (DETR) solutions. However, we find that typical strategies designed to enhance DETR do not improve, and may even degrade, its performance in…

Cited by 0SourcePDFScholar
2024

Part2Object: Hierarchical Unsupervised 3D Instance Segmentation

ECCV 2024poster

"Unsupervised 3D instance segmentation aims to segment objects from a 3D point cloud without any annotations. Existing methods face the challenge of either too loose or too tight clustering, leading to under-segmentation or over-segmentation. To address this issue, we propose Part2Object, hierarchic…

2024

The Devil is in the Object Boundary: Towards Annotation-free Instance Segmentation using Foundation Models

ICLR 2024poster

Foundation models, pre-trained on a large amount of data have demonstrated impressive zero-shot capabilities in various downstream tasks. However, in object detection and instance segmentation, two fundamental computer vision tasks heavily reliant on extensive human annotations, foundation models su…

2023

Contrastive Grouping With Transformer for Referring Image Segmentation

CVPR 2023poster

Referring image segmentation aims to segment the target referent in an image conditioning on a natural language expression. Existing one-stage methods employ per-pixel classification frameworks, which attempt straightforwardly to align vision and language at the pixel level, thus failing to capture…

2023

Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator

NeurIPS 2023poster

Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt. This study focuses on zero-shot text-to-video generation considering the data- and cost-efficient. To generate a s…

2022

Spatial and Visual Perspective-Taking via View Rotation and Relation Reasoning for Embodied Reference Understanding

ECCV 2022poster

"Embodied Reference Understanding studies the reference understanding in an embodied fashion, where a receiver requires to locate a target object referred to by both language and gesture of the sender in a shared physical environment. Its main challenge lies in how to make the receiver with the egoc…

2020

Feature Augmented Memory with Global Attention Network for VideoQA

IJCAI 2020poster

Recently, Recurrent Neural Network (RNN) based methods and Self-Attention (SA) based methods have achieved promising performance in Video Question Answering (VideoQA). Despite the success of these works, RNN-based methods tend to forget the global semantic contents due to the inherent drawbacks of t…

Cited by 0SourcePDFScholar
2020

Robust Reinforcement Learning via Adversarial training with Langevin Dynamics

NeurIPS 2020poster

We introduce a \emph{sampling} perspective to tackle the challenging task of training robust Reinforcement Learning (RL) agents. Leveraging the powerful Stochastic Gradient Langevin Dynamics, we present a novel, scalable two-player RL algorithm, which is a sampling variant of the two-player policy g…

Cited by 73SourcePDFScholar