← Search

Bin Lin

22 accepted papers

2026

360Explorer: Exploring 4D Controllable World in Panoramic Videos

AAAI 2026technical

We present 360Explorer, a novel approach for generating 4D controllable panoramic videos conditioned on user-provided 3D instructions for exploring and manipulating dynamic worlds. Compared to existing perspective-based methods struggle to address spatial consistency during camera rotation in place,

Cited by 1SourcePDFScholar
2026

GIR-Bench: Versatile Benchmark for Generating Images with Reasoning

ICLR 2026poster

Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous reasoning-centric benchmark to systematically evaluate the align…

Cited by 0SourcecodeScholar
2026

JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization

CVPR 2026

Agent-based editing models have substantially advanced interactive experiences, processing quality, and creative flexibility. However, two critical challenges persist: (1) instruction hallucination--text-only chain-of-thought (CoT) reasoning cannot fully prevent factual errors due to inherent inform

Cited by 0SourcecodeScholar
2026

Look-Back: Implicit Visual Re-focusing in MLLM Reasoning

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in multimodal reasoning. However, they often excessively rely on textual information during the later stages of inference, neglecting the crucial integration of visual input. Current methods typically address this by explicit

Cited by 0SourcePDFScholar
2026

NeuroMamba: A Universal Spatiotemporal Module for Robust Perception in Degraded Sensory Streams

ICML 2026poster

In open-world intelligent systems, processing continuous sensory streams disrupted by heterogeneous degradation sources presents a fundamental challenge: reconciling the inherent tension between observational completeness and reconstruction fidelity. Methods that prioritize completeness by bridging …

Cited by 0SourceScholar
2026

Next Patch Prediction for AutoRegressive Visual Generation

AAAI 2026technical

Autoregressive models, built based on the Next Token Prediction (NTP) paradigm, show great potential in developing a unified framework that integrates both language and vision tasks. Pioneering works introduce NTP to autoregressive visual generation tasks. In this work, we rethink the NTP for autore

Cited by 0SourcePDFScholar
2026

Twins: Learn to Predict Unified Representations with Focal Loss

ICML 2026poster

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations—semantic features (e.g., ViT) for understa…

Cited by 0SourceScholar
2026

VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference

CVPR 2026

Long-context video understanding and generation pose a significant computational challenge for Transformer-based video models due to the quadratic complexity of self-attention. While existing sparse attention methods employ coarse-grained patterns to improve efficiency, they typically incur redundan

Cited by 0SourcecodeScholar
2026

WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

ICML 2026poster

Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and evaluation standards predominantly focus on image realism and shallow text-image alignment, lacking a comprehensive assessment of complex semantic understanding and…

Cited by 0SourceScholar
2026

WOW-Seg: A Word-free Open World Segmentation Model

ICLR 2026poster

Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open set of object categories encountered in the real world. However, traditional closed-set segmentation approaches struggle to adapt to complex open…

Cited by 0SourceScholar
2025

Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle

AAAI 2025technical

Recent 3D large reconstruction models typically employ a two-stage process, including first generate multi-view images by a multi-view diffusion model, and then utilize a feed-forward model to reconstruct images to 3D content. However, multi-view diffusion models often produce low-quality and incons…

Cited by 18SourcePDFScholar
2025

DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

ICCV 2025poster

In this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and user-friendly manner. Concretely, baseline methods relying on only…

2025

ImgEdit: A Unified Image Editing Dataset and Benchmark

NeurIPS 2025poster

Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quality data and insufficient benchmarks. To overcome these limitations, we introduce…

Cited by 0SourcecodeScholar
2025

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

NeurIPS 2025poster

Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose **OpenS2V-Nexus**, consisting of (i) **OpenS2V‑Eval**, a fine‑grained benc…

Cited by 0SourceScholar
2025

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

CVPR 2025poster

Video Variational Autoencoder (VAE) encodes videos into a low-dimensional latent space, becoming a key component of most Latent Video Diffusion Models (LVDMs) to reduce model training costs. However, as the resolution and duration of generated videos increase, the encoding cost of Video VAEs becomes…

2024

A Buddy Temporal-Spatial Calibration Method for Airborne Sensors in Multi-UAV Systems

RA-L 2024

The temporal and spatial relationships among various onboard sensors are crucial for ensuring the accuracy of visual observations carried out by the multiple unmanned aerial vehicles (multi-UAV) system. To leverage the integration within the multi-UAV system, we propose a buddy calibration method fo

Cited by 5SourceScholar
2024

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

ICLR 2024poster

The video-language (VL) pretraining has achieved remarkable improvement in multiple downstream tasks. However, the current VL pretraining framework is hard to extend to multiple modalities (N modalities, N ≥ 3) beyond vision and language. We thus propose LanguageBind, taking the language as the bind…

2024

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

NeurIPS 2024poster

We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions. The series comprises: 1) ShareGPT4Video, 40K GPT4V annotated dense captions of videos wit…

Cited by 156SourcePDFScholar
2024

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

EMNLP 2024main

Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding. Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models. However, due to the lack of unified tok…

2023

Robust Target Interception Strategy for a USV With Experimental Validation

RA-L 2023

This letter addresses the problem of designing an interception strategy for an underactuated uncrewed surface vessel (USV) in the presence of uncertain external disturbances and unknown internal parameters i.e., linear and nonlinear damping coefficients, vehicle mass, etc. The interception strategy

Cited by 11SourceScholar
2022

Flexible Collision-free Platooning Method for Unmanned Surface Vehicle with Experimental Validations

IROS 2022poster

This paper addresses the flexible formation problem for unmanned surface vehicles in the presence of obstacles. Building upon the leader-follower formation scheme, a hybrid line-of-sight based flexible platooning method is proposed for follower vehicle to keep tracking the leader ship. A fusion arti…

Cited by 3SourceScholar