← Search

Bing-Kun Bao

17 accepted papers

2026

Predict Before You Explore: Predictive Planning with Specialized Memory for Embodied Question Answering

CVPR 2026

Embodied Question Answering (EQA) requires agents to navigate 3D environments, accumulate visual evidence, and reason over partial observations to answer questions. However, current agents struggle to maintain coherent, long-horizon behavior: planning remains reactive, causing inconsistent actions,

Cited by 0SourcecodeScholar
2026

ProcessMaker: A Generalized Process Visualization Framework with Adaptive Sequence Steps on Diffusion Transformers

CVPR 2026

Procedural sequence generation aims to create intermediate images through multi-step processes, which is applied in industrial design, educational tutorial, and book illustration. However, existing methods often focus on a specific domain or initialize several expert networks for different domains,

Cited by 0SourcecodeScholar
2026

SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation

CVPR 2026

Customized video generation aims to produce videos that faithfully preserve the subject's appearance from reference images while maintaining temporally consistent motion from reference videos. Existing methods struggle to ensure both subject appearance similarity and motion pattern consistency due t

Cited by 0SourcecodeScholar
2025

BeFA: A General Behavior-driven Feature Adapter for Multimedia Recommendation

AAAI 2025technical

Multimedia recommender systems focus on utilizing behavioral information and content information to model user preferences. Typically, it employs pre-trained feature encoders to extract content features, then fuses them with behavioral features. However, pre-trained feature encoders often extract fe…

2025

InstantPainting: Expanding GANs for Efficient Text-Conditioned Image Generation Platform

AAAI 2025technical

Text-conditioned image generation enables cross-modal comprehension. Recent emergence of many platforms have found applications in diverse domains like assisted designing and video gaming. However, there still exist challenges in existing platforms due to their expensive training and time-consuming…

Cited by 0SourcePDFScholar
2025

Leveraging Group Classification with Descending Soft Labeling for Deep Imbalanced Regression

AAAI 2025technical

Deep imbalanced regression (DIR), where the target values have a highly skewed distribution and are also continuous, is an intriguing yet under-explored problem in machine learning. While recent works have already shown that incorporating various classification-based regularizers can produce enha…

2025

Mind Individual Information! Principal Graph Learning for Multimedia Recommendation

AAAI 2025technical

Graph Neural Network (GNN)-based methods have recently emerged as effective approaches for multimedia recommendation. Typically, these methods employ message passing on the user-item interaction graph, and model user preferences by exploiting co-occurrence patterns. Despite their effectiveness, we a…

2025

SCVBench: A Benchmark with Multi-turn Dialogues for Story-Centric Video Understanding

IJCAI 2025

Video understanding seeks to enable machines to interpret visual content across three levels: action, event, and story. Existing models are limited in their ability to perform high-level long-term story understanding, due to (1) the oversimplified treatment of temporal information and (2) the traini

2023

DE-net: Dynamic Text-Guided Image Editing Adversarial Networks

AAAI 2023technical

Text-guided image editing models have shown remarkable results. However, there remain two problems. First, they employ fixed manipulation modules for various editing requirements (e.g., color changing, texture changing, content adding and removing), which results in over-editing or insufficient edit…

2023

GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis

CVPR 2023poster

Synthesizing high-fidelity complex images from text is challenging. Based on large pretraining, the autoregressive and diffusion models can synthesize photo-realistic images. Although these large models have shown notable progress, there remain three flaws. 1) These models require tremendous trainin…

2023

UTM: A Unified Multiple Object Tracking Model With Identity-Aware Feature Enhancement

CVPR 2023poster

Recently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate robust trajectories, where identity association is independent of other components. H…

Cited by 66SourcePDFScholar
2022

DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis

CVPR 2022oral

Synthesizing high-quality realistic images from text descriptions is a challenging task. Existing text-to-image Generative Adversarial Networks generally employ a stacked architecture as the backbone yet still remain three flaws. First, the stacked architecture introduces the entanglements between g…

Cited by 349PDFcodeScholar