← Search

Sibei Yang

48 accepted papers

2026

Chart Deep Research in LVLMs via Parallel Relative Policy Optimization

ICLR 2026poster

With the rapid advancement of data science, charts have evolved from simple numerical presentation tools to essential instruments for insight discovery and decision-making support. However, current chart data intelligence exhibits significant limitations in deep research capabilities, with existing…

Cited by 0SourceScholar
2026

RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation

ICLR 2026poster

In this paper, we propose a 3D asset-referenced diffusion model for image generation, exploring how to integrate 3D assets into image diffusion models. Existing reference-based image generation methods leverage large-scale pretrained diffusion models and demonstrate strong capability in generating d…

Cited by 0SourcecodeScholar
2025

Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement

CVPR 2025poster

Generalized Category Discovery (GCD) aims to recognize unlabeled images from known and novel classes by distinguishing novel classes from known ones, while also transferring knowledge from another set of labeled images with known classes. Existing GCD methods rely on self-supervised vision transform…

Cited by 0SourcePDFScholar
2025

Auto-Search and Refinement: An Automated Framework for Gender Bias Mitigation in Large Language Models

NeurIPS 2025poster

Pre-training large language models (LLMs) on vast text corpora enhances natural language processing capabilities but risks encoding social biases, particularly gender bias. While parameter-modification methods like fine-tuning mitigate bias, they are resource-intensive, unsuitable for closed-source…

Cited by 0SourceScholar
2025

CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMs

ICLR 2025poster

In this paper, we present a 3D visual grounding method called CityAnchor for localizing an urban object in a city-scale point cloud. Recent developments in multiview reconstruction enable us to reconstruct city-scale point clouds but how to conduct visual grounding on such a large-scale urban point…

Cited by 0SourcePDFScholar
2025

DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing

ACL 2025finding

Large Language Models (LLMs) are widely applied in decision making, but their deployment is threatened by jailbreak attacks, where adversarial users manipulate model behavior to bypass safety measures. Existing defense mechanisms, such as safety fine-tuning and model editing, either require extensiv…

2025

Discovering Compositional Hallucinations in LVLMs

NeurIPS 2025poster

Large language models (LLMs) and vision-language models (LVLMs) have driven the paradigm shift towards general-purpose foundation models. However, both of them are prone to hallucinations, which compromise their factual accuracy and reliability. While existing research primarily focuses on isolated…

Cited by 0SourceScholar
2025

Discovering Influential Neuron Path in Vision Transformers

ICLR 2025poster

Vision Transformer models exhibit immense power yet remain opaque to human understanding, posing challenges and risks for practical applications. While prior research has attempted to demystify these models through input attribution and neuron role analysis, there's been a notable gap in considerin…

2025

Don’t Say No: Jailbreaking LLM by Suppressing Refusal

ACL 2025finding

Ensuring the safety alignment of Large Language Models (LLMs) is critical for generating responses consistent with human values. However, LLMs remain vulnerable to jailbreaking attacks, where carefully crafted prompts manipulate them into producing toxic content. One category of such attacks reformu…

2025

Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats

NeurIPS 2025poster

Despite their impressive performance across a wide range of tasks, Large Vision-Language Models (LVLMs) remain prone to hallucination. In this study, we propose a comprehensive intervention framework aligned with the transformer’s causal architecture in LVLMs, integrating the effects of different in…

Cited by 0SourceScholar
2025

MVTokenFlow: High-quality 4D Content Generation using Multiview Token Flow

ICLR 2025poster

In this paper, we present MVTokenFlow for high-quality 4D content creation from monocular videos. Recent advancements in generative models such as video diffusion models and multiview diffusion models enable us to create videos or 3D models. However, extending these generative models for dynamic 4D…

2025

No More Sibling Rivalry: Debiasing Human-Object Interaction Detection

ICCV 2025poster

Detection transformers have been applied to human-object interaction (HOI) detection, enhancing the localization and recognition of human-action-object triplets in images. Despite remarkable progress, this study identifies a critical issue--"Toxic Siblings" bias--which hinders the interaction decode…

Cited by 0SourcePDFScholar
2025

Rethinking Query-based Transformer for Continual Image Segmentation

CVPR 2025poster

Class-incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-in objectness of query-based transformers, which mitigates catastrophic forgetting of mask proposals, current methods of…

2025

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

CVPR 2025poster

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or explicit instruction strictly corresponds to a specific affordance…

Cited by 3SourcePDFScholar
2025

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

ICCV 2025poster

Temporal sentence grounding aims to identify exact moments in a video that correspond to a given textual query, typically addressed with detection transformer (DETR) solutions. However, we find that typical strategies designed to enhance DETR do not improve, and may even degrade, its performance in…

Cited by 0SourcePDFScholar
2025

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

ICCV 2025poster

Recent advancements in language-grounded autonomous driving have been significantly promoted by the sophisticated cognition and reasoning capabilities of large language models (LLMs). However, current LLM-based approaches encounter critical challenges: (1) Failure analysis reveals that frequent coll…

2025

VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction

CVPR 2025poster

Virtual Try-On (VTON) is a transformative technology in e-commerce and fashion design, enabling realistic digital visualization of clothing on individuals. In this work, we propose VTON 360, a novel 3D VTON method that addresses the open challenge of achieving high-fidelity VTON that supports any-vi…

Cited by 1SourcePDFScholar
2025

Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context

ICCV 2025poster

Large Vision-Language Models (LVLMs) have made significant progress in recent years but are also prone to hallucination issues. They exhibit more hallucinations in longer, free-form responses, often attributed to accumulated uncertainties. In this paper, we ask: Does increased hallucination result s…

2024

OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers

CVPR 2024poster

We have recently seen tremendous progress in realistic text-to-motion generation. Yet the existing methods often fail or produce implausible motions with unseen text inputs which limits the applications. In this paper we present OMG a novel framework which enables compelling motion generation from z…

2024

RealDex: Towards Human-like Grasping for Robotic Dexterous Hand

IJCAI 2024poster

In this paper, we introduce RealDex, a pioneering dataset capturing authentic dexterous hand grasping motions infused with human behavioral patterns, enriched by multi-view and multimodal visual data. Utilizing a teleoperation system, we seamlessly synchronize human-robot hand poses in real time. Th…

2024

The Devil is in the Object Boundary: Towards Annotation-free Instance Segmentation using Foundation Models

ICLR 2024poster

Foundation models, pre-trained on a large amount of data have demonstrated impressive zero-shot capabilities in various downstream tasks. However, in object detection and instance segmentation, two fundamental computer vision tasks heavily reliant on extensive human annotations, foundation models su…

2024

WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language

ECCV 2024poster

"We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a novel method, dubbed WildRefer, for this task by fully utilizing the rich appe…

2023

CCQ: Cross-Class Query Network for Partially Labeled Organ Segmentation

AAAI 2023technical

Learning multi-organ segmentation from multiple partially-labeled datasets attracts increasing attention. It can be a promising solution for the scarcity of large-scale, fully labeled 3D medical image segmentation datasets. However, existing algorithms of multi-organ segmentation on partially-labele…

2023

Contrastive Grouping With Transformer for Referring Image Segmentation

CVPR 2023poster

Referring image segmentation aims to segment the target referent in an image conditioning on a natural language expression. Existing one-stage methods employ per-pixel classification frameworks, which attempt straightforwardly to align vision and language at the pixel level, thus failing to capture…

2023

DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

NeurIPS 2023poster

A long-standing goal of AI systems is to perform complex multimodal reasoning like humans. Recently, large language models (LLMs) have made remarkable strides in such multi-step reasoning on the language modality solely by leveraging the chain of thought (CoT) to mimic human thinking. However, the t…

Cited by 100SourcePDFScholar
2023

Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator

NeurIPS 2023poster

Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt. This study focuses on zero-shot text-to-video generation considering the data- and cost-efficient. To generate a s…

2023

Grounded Image Text Matching with Mismatched Relation Reasoning

ICCV 2023poster

This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image,…

Cited by 8PDFcodeScholar
2022

Spatial and Visual Perspective-Taking via View Rotation and Relation Reasoning for Embodied Reference Understanding

ECCV 2022poster

"Embodied Reference Understanding studies the reference understanding in an embodied fashion, where a receiver requires to locate a target object referred to by both language and gesture of the sender in a shared physical environment. Its main challenge lies in how to make the receiver with the egoc…

2021

Bottom-Up Shift and Reasoning for Referring Image Segmentation

CVPR 2021poster

Referring image segmentation aims to segment the referent that is the corresponding object or stuff referred by a natural language expression in an image. Its main challenge lies in how to effectively and efficiently differentiate between the referent and other objects of the same category as the re…

Cited by 101PDFcodeScholar
2021

Preservational Learning Improves Self-Supervised Medical Image Models by Reconstructing Diverse Contexts

ICCV 2021poster

Preserving maximal information is the basic principle of designing self-supervised learning methodologies. To reach this goal, contrastive learning adopts an implicit way which is contrasting image pairs. However, we believe it is not fully optimal to simply use the contrastive estimation for preser…

Cited by 116PDFcodeScholar
2018

Multi-Evidence Filtering and Fusion for Multi-Label Classification, Object Detection and Semantic Segmentation Based on Weakly Supervised Learning

CVPR 2018poster

Supervised object detection and semantic segmentation require object or even pixel level annotations. When there exist image level labels only, it is challenging for weakly supervised algorithms to achieve accurate predictions. The accuracy achieved by top weakly supervised algorithms is still signi…

Cited by 251SourcePDFScholar