← Search

Shijie Zhou

22 accepted papers

2026

Can Simple Denoising Improve Uniform State Diffusion Models?

ICML 2026poster

Recent Uniform-state Diffusion Models (USDMs), initialized from a uniform prior, offer the promise of fast text generation due to their inherent self-correction ability compared to masked diffusion models. However, they still rely on complex loss formulations with additional computational overhead, …

Cited by 0SourceScholar
2026

Learning Adaptive and Expandable Mixture Model for Continual Learning

AAAI 2026technical

Continuous learning constitutes a fundamental capability of artificial intelligence systems, enabling them to incrementally assimilate novel information without succumbing to catastrophic forgetting. Recent research has leveraged Pre-Trained Models (PTMs) to enhance continual learning efficacy. Neve

Cited by 0SourcePDFScholar
2026

RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation

CVPR 2026

Recent advances in Vision-Language-Action (VLA) models have enabled robots to execute increasingly complex tasks. However, VLA models trained through imitation learning struggle to operate reliably in dynamic environments and often fail under Out-of-Distribution (OOD) conditions. To address this iss

Cited by 0SourceScholar
2026

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

CVPR 2026

Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their inability to capture fine-grained 3D geometry and spatial relationships. While recent efforts have introduced multi-view ge

Cited by 0SourceScholar
2026

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

CVPR 2026

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has sparked interest in extending these models to 3D scenes, with the goal of human-like visual-spatial intelligence. However, achieving deep spatial understanding comparable to human capabilities remains challenging fo

Cited by 0SourcecodeScholar
2025

4K4DGen: Panoramic 4D Generation at 4K Resolution

ICLR 2025spotlight

The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques either focus solely on dynamic objects or perform outpainting from a single perspe…

2025

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

COLING 2025main

Large multimodal models still struggle with text-rich images because of inadequate training data. Self-Instruct provides an annotation-free way for generating instruction data, but its quality is poor, as multimodal alignment remains a hurdle even for the largest models. In this work, we propose LLa…

2025

Dynamic Siamese Expansion Framework for Improving Robustness in Online Continual Learning

NeurIPS 2025poster

Continual learning requires the model to continually capture novel information without forgetting prior knowledge. Nonetheless, existing studies predominantly address the catastrophic forgetting, often neglecting enhancements in model robustness. Consequently, these methodologies fall short in real-…

Cited by 0SourceScholar
2025

Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields

CVPR 2025poster

Recent advancements in 2D and multimodal models have achieved remarkable success by leveraging large-scale training on extensive datasets. However, extending these achievements to enable free-form interactions and high-level semantic operations with complex 3D/4D scenes remains challenging. This dif…

Cited by 1SourcePDFScholar
2025

Learning Expandable and Adaptable Representations for Continual Learning

NeurIPS 2025poster

Extant studies predominantly address catastrophic forgetting within a simplified continual learning paradigm, typically confined to a singular data domain. Conversely, real-world applications frequently encompass multiple, evolving data domains, wherein models often struggle to retain many critical…

Cited by 0SourceScholar
2025

Learning Multi-Source and Robust Representations for Continual Learning

NeurIPS 2025poster

Plasticity and stability denote the ability to assimilate new tasks while preserving previously acquired knowledge, representing two important concepts in continual learning. Recent research addresses stability by leveraging pre-trained models to provide informative representations, yet the efficacy…

Cited by 0SourcecodeScholar
2025

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

ICCV 2025poster

We introduce LLaVA-Reward, an efficient reward model designed to automatically evaluate text-to-image (T2I) generations across multiple perspectives, leveraging pretrained multimodal large language models (MLLMs). Existing MLLM-based approaches require instruction-following data for supervised fine-…

2025

Reinforcement Learning for Large Language Models via Group Preference Reward Shaping

EMNLP 2025

Large Language Models (LLMs) require alignment via reinforcement learning (RL) to effectively perform task-specific objectives, such as human preference alignment and enhanced reasoning. While Proximal Policy Optimization (PPO) is widely adopted, its computational overhead, stemming from additional

Cited by 0SourcePDFScholar
2025

Simple Distillation for One-Step Diffusion Models

NeurIPS 2025poster

Diffusion models have established themselves as leading techniques for image generation. However, their reliance on an iterative denoising process results in slow sampling speeds, which limits their applicability to interactive and creative applications. An approach to overcoming this limitation inv…

Cited by 0SourceScholar
2025

SwarmAgentic: Towards Fully Automated Agentic System Generation via Swarm Intelligence

EMNLP 2025

The rapid progress of Large Language Models has advanced agentic systems in decision-making, coordination, and task execution. Yet, existing agentic system generation frameworks lack full autonomy, missing from-scratch agent generation, self-optimizing agent functionality, and collaboration, limitin

Cited by 0SourcePDFScholar
2025

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

ICCV 2025poster

Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts--abili…

Cited by 0SourcePDFScholar
2025

X-Dyna: Expressive Dynamic Human Image Animation

CVPR 2025highlight

We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the subject and the surrounding environment. Building on prior appro…

2024

Enhancing Tracking Robustness with Auxiliary Adversarial Defense Networks

ECCV 2024poster

"Adversarial attacks in visual object tracking have significantly degraded the performance of advanced trackers by introducing imperceptible perturbations into images. However, there is still a lack of research on designing adversarial defense methods for object tracking. To address these issues, we…

Cited by 1SourcePDFScholar
2024

Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields

CVPR 2024highlight

3D scene representations have gained immense popularity in recent years. Methods that use Neural Radiance fields are versatile for traditional tasks such as novel view synthesis. In recent times some work has emerged that aims to extend the functionality of NeRF beyond view synthesis for semanticall…

2024

Large Spatial Model: End-to-end Unposed Images to Semantic 3D

NeurIPS 2024poster

Reconstructing and understanding 3D structures from a limited number of images is a classical problem in computer vision. Traditional approaches typically decompose this task into multiple subtasks, involving several stages of complex mappings between different data representations. For example, den…

2023

ALTO: Alternating Latent Topologies for Implicit 3D Reconstruction

CVPR 2023poster

This work introduces alternating latent topologies (ALTO) for high-fidelity reconstruction of implicit 3D surfaces from noisy point clouds. Previous work identifies that the spatial arrangement of latent encodings is important to recover detail. One school of thought is to encode a latent vector for…

Cited by 35SourcePDFScholar
2021

Topic-Aware Dialogue Generation with Two-Hop Based Graph Attention

ICASSP 2021accepted

Generating on-topic responses and understanding the background information of context are both significant for dialogue generation. However, few works simultaneously concentrate on these two issues. For this purpose, we propose an open-domain topic-aware dialogue generation model via joint learning.…

Cited by 0SourceScholar