← Search

Boyang Wang

16 accepted papers

2026

DiffPlace: Street View Generation Via Place-Controllable Diffusion Model Enhancing Place Recognition

ICRA 2026poster

Generative models have advanced significantly in realistic image synthesis, with diffusion models excelling in quality and stability. Recent multi-view diffusion models improve 3D-aware street view generation, but they struggle to produce place-aware and background-consistent urban scenes from text,…

2026

U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding

ICLR 2026poster

Ultrasound is a widely-used imaging modality critical to global healthcare, yet its interpretation remains challenging due to its varying image quality on operators, noises, and anatomical structures. Although large vision-language models (LVLMs) have demonstrated impressive multimodal capabilities…

Cited by 0SourceScholar
2025

Component-Aware Unsupervised Logical Anomaly Generation for Industrial Anomaly Detection

ICRA 2025

Anomaly detection is critical in industrial manufacturing for ensuring product quality and improving efficiency in automated processes. The scarcity of anomalous samples limits traditional detection methods, making anomaly generation essential for expanding the data repository. However, recent gener

Cited by 2SourceScholar
2025

Frame In-N-Out: Unbounded Controllable Image-to-Video Generation

NeurIPS 2025poster

Controllability, temporal coherence, and detail synthesis remain the most critical challenges in video generation. In this paper, we focus on a commonly used yet underexplored cinematic technique known as Frame In and Frame Out. Specifically, starting from image-to-video generation, users can contro…

Cited by 0SourceScholar
2025

H-PCC: Point Cloud Compression With Hybrid Mode Selection and Content Adaptive Down-Sampling

RA-L 2025

LiDAR sensors are integral to autonomous driving and augmented reality applications, providing essential depth information. However, managing the substantial volume of LiDAR point cloud data is crucial for practical application, necessitating efficient compression algorithms. Similar to other data c

Cited by 5SourceScholar
2025

McEval: Massively Multilingual Code Evaluation

ICLR 2025poster

Code large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks.…

2025

Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments

IROS 2025

Anomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects i

Cited by 0SourcecodeScholar
2025

SIR-DIFF: Sparse Image Sets Restoration with Multi-View Diffusion Model

CVPR 2025poster

The computer vision community has developed numerous techniques for digitally restoring true scene information from single-view degraded photographs, an important yet extremely ill-posed task. In this work, we tackle image restoration from a different perspective by jointly denoising multiple photog…

2025

SNS-Bench: Defining, Building, and Assessing Capabilities of Large Language Models in Social Networking Services

ICML 2025poster

With the rapid advancement of Social Networking Services (SNS), the need for intelligent and efficient interaction within diverse platforms has become more crucial. Large Language Models (LLMs) play an important role in SNS as they possess the potential to revolutionize user experience, content gene…

Cited by 0SourcePDFScholar
2025

This&That: Language-Gesture Controlled Video Generation for Robot Planning

ICRA 2025

Clear, interpretable instructions are invaluable for complex tasks, helping to clarify goals and anticipate necessary steps. In this work, we propose a robot learning framework for communicating, planning, and executing a wide range of tasks, dubbed This&That. This&That solves general tasks by lever

Cited by 41SourcecodeScholar
2024

APISR: Anime Production Inspired Real-World Anime Super-Resolution

CVPR 2024poster

While real-world anime super-resolution (SR) has gained increasing attention in the SR community existing methods still adopt techniques from the photorealistic domain. In this paper we analyze the anime production workflow and rethink how to use characteristics of it for the sake of the real-world…

2024

Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution

ECCV 2024poster

"Pre-trained diffusion models utilized for image generation encapsulate a substantial reservoir of a priori knowledge pertaining to intricate textures. Harnessing the potential of leveraging this a priori knowledge in the context of image super-resolution presents a compelling avenue. Nonetheless, p…

Cited by 3SourcePDFScholar
2024

FD-UAD: Unsupervised Anomaly Detection Platform Based on Defect Autonomous Imaging and Enhancement

IJCAI 2024poster

In industrial quality control, detecting defects is essential. However, manual checks and machine vision encounter challenges in complex conditions, as defects vary among products made of different materials and shapes. We create FD-UAD, Unsupervised Anomaly Detection Platform Based on Defect Autono…

Cited by 0SourcePDFScholar
2024

LogFormer: A Pre-train and Tuning Pipeline for Log Anomaly Detection

AAAI 2024technical

Log anomaly detection is a key component in the field of artificial intelligence for IT operations (AIOps). Considering log data of variant domains, retraining the whole network for unknown domains is inefficient in real industrial scenarios. However, previous deep models merely focused on extractin…

2023

Isolating Trajectory Tracking From Motion Control: A Model Predictive Control and Robust Control Framework for Unmanned Ground Vehicles

RA-L 2023

This letter studies the trajectory tracking and motion control problems of unmanned ground vehicles (UGVs). A model predictive control and robust control (MPC-RC) framework for UGVs is proposed to improve tracking accuracy, yaw stability and robustness in a modular fashion without introducing comple

Cited by 28SourceScholar
2023

M2C: Towards Automatic Multimodal Manga Complement

EMNLP 2023short findings

Multimodal manga analysis focuses on enhancing manga understanding with visual and textual features, which has attracted considerable attention from both natural language processing and computer vision communities. Currently, most comics are hand-drawn and prone to problems such as missing pages, te…

Cited by 0SourcecodeScholar