← Search

Yaqian Zhao

10 accepted papers

2025

DeMAC: Enhancing Multi-Agent Coordination with Dynamic DAG and Manager-Player Feedback

EMNLP 2025

Multi-agent systems (MAS) powered by large language models (LLMs) have shown potential in tackling multifaceted problems through advanced understanding and reasoning. However, they struggle to adapt to evolving task dependencies and to handle uncertainties, such as shifting priorities or unpredictab

Cited by 0SourcePDFScholar
2025

DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation

ICCV 2025poster

Spatio-temporal consistency is a critical topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily foc…

2025

Improving Height Prediction for Vision-Based Roadside 3D Object Detection

ICASSP 2025accepted

Roadside vision-based 3D object detection is vital in many applications, such as autonomous driving. The mainstream methods enhance the accuracy of distance estimation by converting predicted height distribution into depth distribution. However, predicting object’s height in roadside perception is c…

Cited by 0SourceScholar
2025

MMEditor: Multimodal Prompt-Driven 3D Gaussian Splatting Editing

ICASSP 2025accepted

We propose a multimodal 3D scene editing framework MMEditor to create or modify objects within an extant 3D Gaussian Splatting (3DGS) according to text and image prompts. MMEditor employs a multimodal image editing module to iteratively optimize 3D Gaussians in editing regions for delicate and multi…

Cited by 0SourceScholar
2024

Glance, Focus and Refinement Network for Remote Sensing Change Detection

ICASSP 2024accepted

Existing change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively…

Cited by 0SourceScholar
2024

Image Content Generation with Causal Reasoning

AAAI 2024technical

The emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed the reasoning potential reflected in the generated textual content. However, this current ability for causal reasoning…

2024

Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline

NeurIPS 2024poster

Existing video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad appl…

2024

Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change Detection

ICASSP 2024accepted

Existing deep learning based remote sensing change detection (RSCD) methods only rely on binary ground-truth to guide the network learning while neglecting the useful semantic guidance. As a result, the network can be readily misled by irrelevant category changes, leading to degraded performance and…

Cited by 15SourceScholar
2023

Enhancing Network by Reinforcement Learning and Neural Confined Local Search

IJCAI 2023poster

It has been found that many real networks, such as power grids and the Internet, are non-robust, i.e., attacking a small set of nodes would cause the paralysis of the entire network. Thus, the Network Enhancement Problem~(NEP), i.e., improving the robustness of a given network by modifying its struc…

Cited by 2SourcePDFScholar
2023

Group-Wise Co-Salient Object Detection with Siamese Transformers Via Brownian Distance Covariance Matching

ICASSP 2023accepted

Co-salient object detection (CoSOD) aims to discover and segment foreground targets in a group of images with the same semantic category. Existing mainstream approaches often employ convolutional neural networks (CNNs) to learn the semantic-invariant features from a group of images. Despite demonstr…

Cited by 0SourceScholar