← Search

zhibin wang

27 accepted papers

2026

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

CVPR 2026

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception through irreversible information disposal or inhibit long-range temporal modeling via rigid, predefined sparse patterns. This

Cited by 0SourceScholar
2026

Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning

AAAI 2026technical

Video large language models have demonstrated remarkable capabilities in video understanding tasks. However, the redundancy of video tokens introduces significant computational overhead during inference, limiting their practical deployment. Many compression algorithms are proposed to prioritize reta

Cited by 0SourcePDFScholar
2026

GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding

CVPR 2026

Video Large Language Models (VLMs) have achieved remarkable success in video understanding, but the significant computational cost from processing dense frames severely limits their practical application. Existing methods alleviate this by selecting keyframes, but their greedy decision-making, combi

Cited by 0SourceScholar
2026

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs

AAAI 2026technical

Video Large Language Models (VLLMs) excel in video understanding, but their excessive visual tokens pose a significant computational challenge for real-world applications. Current methods aim to enhance inference efficiency by visual token pruning. However, they do not consider the dynamic character

Cited by 0SourcePDFScholar
2026

ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis

ICLR 2026poster

While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasoning remains underdeveloped. This gap stems primarily from a critical data bottleneck: existing datasets lack the challengi…

Cited by 0SourcecodeScholar
2026

SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation

CVPR 2026

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to fine-grained spatial relationships, often producing images that appear

Cited by 0SourcecodeScholar
2026

StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression

AAAI 2026technical

Video Large Language Models (Video-LLMs) have demonstrated significant potential in the areas of video captioning, search, and summarization. However, current Video-LLMs still face challenges with long real-world videos. Recent methods have introduced a retrieval mechanism that retrieves query-relev

Cited by 0SourcePDFScholar
2025

CR2PQ: Continuous Relative Rotary Positional Query for Dense Visual Representation Learning

ICLR 2025poster

Dense visual contrastive learning (DRL) shows promise for learning localized information in dense prediction tasks, but struggles with establishing pixel/patch correspondence across different views (cross-contrasting). Existing methods primarily rely on self-contrasting the same view with variations…

Cited by 0SourcePDFScholar
2024

Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based Approach

ICLR 2024poster

Image outpainting aims to generate the content of an input sub-image beyond its original boundaries. It is an important task in content generation yet remains an open problem for generative models. This paper pushes the technical frontier of image outpainting in two directions that have not been res…

2024

Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

ACL 2024findings

The remarkable multimodal capabilities demonstrated by OpenAI’s GPT-4 have sparked significant interest in the development of multimodal Large Language Models (LLMs). A primary research objective of such models is to align visual and textual modalities effectively while comprehending human instructi…

2024

I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing

NeurIPS 2024poster

Significant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in this field is the establishment of a comprehensive evaluation benchmark for accurately assessing editing results and prov…

2024

MeshXL: Neural Coordinate Field for Generative 3D Foundation Models

NeurIPS 2024poster

The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately,…

2024

Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models

CVPR 2024poster

This paper presents Paint3D a novel coarse-to-fine generative framework that is capable of producing high-resolution lighting-less and diverse 2K UV texture maps for untextured 3D meshes conditioned on text or image inputs. The key challenge addressed is generating high-quality textures without embe…

2023

D2Q-DETR: Decoupling and Dynamic Queries for Oriented Object Detection with Transformers

ICASSP 2023accepted

Despite the promising results, existing oriented object detection methods usually involve heuristically designed rules, e.g., RRoI generation, rotated NMS. In this paper, we propose an end-to-end framework for oriented object detection, which simplifies the model pipeline and obtains superior perfor…

Cited by 0SourceScholar
2023

Efficient Mask Correction for Click-Based Interactive Image Segmentation

CVPR 2023poster

The goal of click-based interactive image segmentation is to extract target masks with the input of positive/negative clicks. Every time a new click is placed, existing methods run the whole segmentation network to obtain a corrected mask, which is inefficient since several clicks may be needed to r…

2023

Foundation Model Drives Weakly Incremental Learning for Semantic Segmentation

CVPR 2023poster

Modern incremental learning for semantic segmentation methods usually learn new categories based on dense annotations. Although achieve promising results, pixel-by-pixel labeling is costly and time-consuming. Weakly incremental learning for semantic segmentation (WILSS) is a novel and attractive tas…

Cited by 15SourcePDFScholar
2023

Frequency Domain Disentanglement for Arbitrary Neural Style Transfer

AAAI 2023technical

Arbitrary neural style transfer has been a popular research topic due to its rich application scenarios. Effective disentanglement of content and style is the critical factor for synthesizing an image with arbitrary style. The existing methods focus on disentangling feature representations of conten…

Cited by 5SourcePDFScholar
2023

K2NN: Self-Supervised Learning with Hierarchical Nearest Neighbors for Remote Sensing

ICASSP 2023accepted

Self-supervised learning aims to learn applicable pre-trained models from massive unlabeled data. Besides image-level pretext tasks, many recent pixel-level studies have been pro-posed to learn dense information in each image. However, most of those methods focus on obtaining pair of matched patches…

Cited by 0SourceScholar
2023

LMSeg: Language-guided Multi-dataset Segmentation

ICLR 2023poster

It’s a meaningful and attractive topic to build a general and inclusive segmentation model that can recognize more categories in various scenarios. A straightforward way is to combine the existing fragmented segmentation datasets and train a multi-dataset network. However, there are two major issues…

Cited by 20SourcePDFScholar
2023

Patch-level Contrastive Learning via Positional Query for Visual Pre-training

ICML 2023poster

Dense contrastive learning (DCL) has been recently explored for learning localized information for dense prediction tasks (e.g., detection and segmentation). It still suffers the difficulty of mining pixels/patches correspondence between two views. A simple way is inputting the same view twice and a…

2023

Point-Teaching: Weakly Semi-supervised Object Detection with Point Annotations

AAAI 2023technical

Point annotations are considerably more time-efficient than bounding box annotations. However, how to use cheap point annotations to boost the performance of semi-supervised object detection is still an open question. In this work, we present Point-Teaching, a weakly- and semi-supervised object dete…

2023

Robust Geometry-Preserving Depth Estimation Using Differentiable Rendering

ICCV 2023poster

In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate for mix-dataset training, enhancing generalization across dive…

Cited by 6PDFScholar
2023

SwinRDM: Integrate SwinRNN with Diffusion Model towards High-Resolution and High-Quality Weather Forecasting

AAAI 2023technical

Data-driven medium-range weather forecasting has attracted much attention in recent years. However, the forecasting accuracy at high resolution is unsatisfactory currently. Pursuing high-resolution and high-quality weather forecasting, we develop a data-driven model SwinRDM which integrates an impro…

Cited by 55SourcePDFScholar
2022

Poseur: Direct Human Pose Regression with Transformers

ECCV 2022poster

"We propose a direct, regression-based approach to 2D human pose estimation from single images. We formulate the problem as a sequence prediction task, which we solve using a Transformer network. This network directly learns a regression mapping from images to the keypoint coordinates, without resor…

2021

A Simple Baseline for Semi-Supervised Semantic Segmentation With Strong Data Augmentation

ICCV 2021poster

Recently, significant progress has been made on semantic segmentation. However, the success of supervised semantic segmentation typically relies on a large amount of labeled data, which is time-consuming and costly to obtain. Inspired by the success of semi-supervised learning methods in image class…

Cited by 150PDFcodeScholar
2021

Instant-Teaching: An End-to-End Semi-Supervised Object Detection Framework

CVPR 2021poster

Supervised learning based object detection frameworks demand plenty of laborious manual annotations, which may not be practical in real applications. Semi-supervised object detection (SSOD) can effectively leverage unlabeled data to improve the model performance, which is of great significance for t…

Cited by 254PDFScholar