← Search

Jianan Wang

37 accepted papers

2026

Context and Diversity Matter: The Emergence of In-Context Learning in World Models

ICLR 2026poster

The capability of predicting environmental dynamics underpins both biological neural systems and general embodied AI in adapting to their surroundings. Yet prevailing approaches rest on static world models that falter when confronted with novel or rare configurations. We investigate in-context learn…

Cited by 0SourceScholar
2026

DSPv2: Improved Dense Policy for Effective and Generalizable Whole-Body Mobile Manipulation

ICRA 2026poster

Learning whole-body mobile manipulation via imitation is essential for generalizing robotic skills to diverse environments and complex tasks. However, this goal is hindered by significant challenges, particularly in effectively processing complex observation, achieving robust generalization, and gen…

2026

MetaphorVU: Towards Metaphorical Video Understanding

ICML 2026spotlight

Metaphorical videos are prevalent across various real-world scenarios to convey complex ideas, and understanding them typically requires high-order cognitive capabilities. The lack of systematic studies on metaphorical video understanding not only constrains the real-world applicability of MLLMs but…

Cited by 0SourceScholar
2026

Refacade: Editing Object with Given Reference Texture

CVPR 2026

Recent advances in diffusion models have brought remarkable progress in image and video editing, yet some tasks remain underexplored. In this paper, we extend Object Retexture into video domain, which transfers local textures from a reference object to a target object in images or videos. To perform

Cited by 0SourcecodeScholar
2026

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

ICML 2026poster

With the current surge in spatial reasoning, researchers have made significant progress in understanding indoor scenes, but still struggle with more diverse applications. This paper aims to advance all-scale spatial reasoning by tackling two key challenges: 1) the heavy reliance on indoor 3D scans a…

Cited by 0SourcecodeScholar
2026

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

ICML 2026poster

It is infeasible to encompass all possible disturbances within the training dataset. This raises a critical question regarding the robustness of Vision-Language-Action (VLA) models when encountering unseen real-world visual disturbances, particularly under imperfect visual conditions. In this work, …

Cited by 0SourceScholar
2026

Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer

ICLR 2026poster

Text-guided color editing in images and videos is a fundamental yet unsolved problem, requiring fine-grained manipulation of color attributes, including albedo, light source color, and ambient lighting, while preserving physical consistency in geometry, material properties, and light-matter interact…

Cited by 0SourceScholar
2025

BeSimulator: A Large Language Model Powered Text-based Behavior Simulator

EMNLP 2025

Traditional robot simulators focus on physical process modeling and realistic rendering, often suffering from high computational costs, inefficiencies, and limited adaptability. To handle this issue, we concentrate on behavior simulation in robotics to analyze and validate the logic behind robot beh

2025

CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility

AAAI 2025technical

Video inpainting is a crucial task with diverse applications, including fine-grained video editing, video recovery, and video dewatermarking. However, most existing video inpainting methods primarily focus on visual content completion while neglecting text information. There are only a limited numbe…

2025

Code-BT: A Code-Driven Approach to Behavior Tree Generation for Robot Tasks Planning with Large Language Models

IJCAI 2025

Behavior trees(BTs) provide a systematic and structured control architecture extensively employed in game AI and robotic behavior control, owing to their modularity, reactivity, and reusability. Nonetheless, manual BTs design requires significant expertise and becomes inefficient as task complexity

Cited by 0SourcePDFScholar
2025

ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models

CoRL 2025poster

Learning real-world robotic manipulation is challenging, particularly when limited demonstrations are available. Existing methods for few-shot manipulation often rely on simulation-augmented data or pre-built modules like grasping and pose estimation, which struggle with sim-to-real gaps and lack ex…

Cited by 0SourceScholar
2025

DreamCube: RGB-D Panorama Generation via Multi-plane Synchronization

ICCV 2025poster

3D panorama synthesis is a promising yet challenging task that demands high-quality and diverse visual appearance and geometry of the generated omnidirectional content. Existing methods leverage rich image priors from pre-trained 2D foundation models to circumvent the scarcity of 3D panoramic data,…

Cited by 0SourcePDFScholar
2025

GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D Reconstruction

CVPR 2025poster

Garments are common in daily life and are important for embodied intelligence community. Current category-level garments pose tracking works focus on predicting point-wise canonical correspondence and learning a shape deformation in point cloud sequences. In this paper, motivated by the 2D warping s…

Cited by 0SourcePDFScholar
2025

HeGTa: Leveraging Heterogeneous Graph-enhanced Large Language Models for Few-shot Complex Table Understanding

AAAI 2025technical

Table Understanding (TU) has achieved promising advancements, but it faces the challenges of the scarcity of manually labeled tables and the presence of complex table structures. To address these challenges, we propose HeGTa, a heterogeneous graph (HG)-enhanced large language model (LLM) designed fo…

Cited by 2SourcePDFScholar
2025

MiniMax-Remover: Taming Bad Noise Helps Video Object Removal

NeurIPS 2025poster

Recent advances in video diffusion models have driven rapid progress in video editing techniques. However, video object removal, a critical subtask of video editing, remains challenging due to issues such as hallucinated objects and visual artifacts. Furthermore, existing methods often rely on compu…

Cited by 0SourcecodeScholar
2025

R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render Strategy

AAAI 2025technical

Human life is filled with articulated objects. Previous works for estimating the pose of category-level articulated objects rely on costly 3D point clouds or RGB-D images. In this paper, our goal is to estimate category-level articulation poses from a single RGB image, where we propose R2-Art, a nov…

Cited by 0SourcePDFScholar
2025

Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model

ICLR 2025poster

Room layout estimation from multiple-perspective images is poorly investigated due to the complexities that emerge from multi-view geometry, which requires muti-step solutions such as camera intrinsic and extrinsic estimation, image matching, and triangulation. However, in 3D reconstruction, the adv…

2024

DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation

CVPR 2024poster

We propose DiffSHEG a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation. While previous works focused on co-speech gesture or expression generation individually the joint generation of synchronized expressions and gestures remains barely explored. To address th…

Cited by 37SourcePDFScholar
2024

DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation

ICLR 2024poster

Text-to-image diffusion models pre-trained on billions of image-text pairs have recently enabled 3D content creation by optimizing a randomly initialized differentiable 3D representation with score distillation. However, the optimization process suffers slow convergence and the resultant 3D models o…

Cited by 22SourcePDFScholar
2024

MRMLREC: A Two-Stage Approach for Addressing Data Sparsity in MOOC Video Recommendation (Student Abstract)

AAAI 2024technical

With the abundance of learning resources available on massive open online courses (MOOCs) platforms, the issue of interactive data sparsity has emerged as a significant challenge.This paper introduces MRMLREC, an efficient MOOC video recommendation which consists of two main stages: multi-relational…

Cited by 4SourcePDFScholar
2024

Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts

ICLR 2024poster

Recent text-to-3D generation methods achieve impressive 3D content creation capacity thanks to the advances in image diffusion models and optimizing strategies. However, current methods struggle to generate correct 3D content for a complex prompt in semantics, i.e., a prompt describing multiple inte…

Cited by 42SourcePDFScholar
2024

Rethinking 3D Convolution in $\ell_p$-norm Space

NeurIPS 2024spotlight

Convolution is a fundamental operation in the 3D backbone. However, under certain conditions, the feature extraction ability of traditional convolution methods may be weakened. In this paper, we introduce a new convolution method based on $\ell_p$-norm. For theoretical support, we prove the univer…

Cited by 9SourcePDFScholar
2024

TOSS: High-quality Text-guided Novel View Synthesis from a Single Image

ICLR 2024poster

In this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image. While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the cha…

Cited by 18SourcePDFScholar
2024

Text-Video Completion Networks With Motion Compensation And Attention Aggregation

ICASSP 2024accepted

The purpose of video inpainting is to fill a specified area with reasonable content. However, in the case of multiple targets and complex textures, current methods struggle to distinguish between feature information of the targets, leading to confusing or fuzzy inpainting results. In this paper, we…

Cited by 0SourceScholar
2023

DisCo-CLIP: A Distributed Contrastive Loss for Memory Efficient CLIP Training

CVPR 2023highlight

We propose DisCo-CLIP, a distributed memory-efficient CLIP training approach, to reduce the memory consumption of contrastive loss when training contrastive learning models. Our approach decomposes the contrastive loss and its gradient computation into two parts, one to calculate the intra-GPU gradi…

2023

DreamWaltz: Make a Scene with Complex 3D Animatable Avatars

NeurIPS 2023poster

We present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects, creating high-quality and animatable 3D avatars remains chall…

2023

Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes

CVPR 2023poster

Humans have long been recorded in a variety of forms since antiquity. For example, sculptures and paintings were the primary media for depicting human beings before the invention of cameras. However, most current human-centric computer vision tasks like human pose estimation and human image generati…

2023

HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation

ICCV 2023oral

Controllable human image generation (HIG) has attracted significant attention from academia and industry for its numerous real-life applications. State-of-the-art solutions, such as ControlNet and T2I-Adapter, introduce an additional learnable branch on top of the frozen pre-trained stable diffusion…

Cited by 90PDFcodeScholar
2023

LipsFormer: Introducing Lipschitz Continuity to Vision Transformers

ICLR 2023poster

We present a Lipschitz continuous Transformer, called LipsFormer, to pursue training stability both theoretically and empirically for Transformer-based models. In contrast to previous practical tricks that address training instability by learning rate warmup, layer normalization, attention formulati…

2023

Multi-View MOOC Quality Evaluation via Information-Aware Graph Representation Learning

AAAI 2023technical

In this paper, we study the problem of MOOC quality evaluation that is essential for improving the course materials, promoting students' learning efficiency, and benefiting user services. While achieving promising performances, current works still suffer from the complicated interactions and relati…

Cited by 5SourcePDFScholar
2023

TabPrompt: Graph-based Pre-training and Prompting for Few-shot Table Understanding

EMNLP 2023long findings

Table Understanding (TU) is a crucial aspect of information extraction that enables machines to comprehend the semantics behind tabular data. However, existing methods of TU cannot deal with the scarcity of labeled tabular data. In addition, these methods primarily focus on the textual content withi…

Cited by 0SourceScholar
2022

ManiTrans: Entity-Level Text-Guided Image Manipulation via Token-Wise Semantic Alignment and Generation

CVPR 2022oral

Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical application. In this work, we study a novel task on text-guided image manipulation on the entity level in the real world. Th…

Cited by 19PDFcodeScholar
2021

Gated Linear Networks

AAAI 2021technical

This paper presents a new family of backpropagation-free neural architectures, Gated Linear Networks (GLNs). What distinguishes GLNs from contemporary neural networks is the distributed and local nature of their credit assignment mechanism; each neuron directly predicts the target, forgoing the abil…

Cited by 48SourcePDFScholar
2021

Implicit Sentiment Analysis with Event-centered Text Representation

EMNLP 2021main

Implicit sentiment analysis, aiming at detecting the sentiment of a sentence without sentiment words, has become an attractive research topic in recent years. In this paper, we focus on event-centric implicit sentiment analysis that utilizes the sentiment-aware event contained in a sentence to infer…

2020

A Combinatorial Perspective on Transfer Learning

NeurIPS 2020poster

Human intelligence is characterized not only by the capacity to learn complex skills, but the ability to rapidly adapt and acquire new skills within an ever-changing environment. In this work we study how the learning of modular solutions can allow for effective generalization to both unseen and pot…

2020

Online Learning in Contextual Bandits using Gated Linear Networks

NeurIPS 2020poster

We introduce a new and completely online contextual bandit algorithm called Gated Linear Contextual Bandits (GLCB). This algorithm is based on Gated Linear Networks (GLNs), a recently introduced deep learning architecture with properties well-suited to the online setting. Leveraging data-dependent g…

2019

Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds

NeurIPS 2019spotlight

We propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cl…