← Search

Xinyi Zhang

20 accepted papers

2026

FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning

ICLR 2026poster

Search has emerged as core infrastructure for LLM-based agents and is widely viewed as critical on the path toward more general intelligence. Finance is a particularly demanding proving ground: analysts routinely conduct complex, multi-step searches over time-sensitive, domain-specific data, making…

Cited by 0SourcecodeScholar
2025

AKI360: Enabling Highly Interactive 360-degree Video Streaming by Adaptive Keyframe Interval

ICASSP 2025accepted

360-degree video is a panoramic video technology designed to offer audience an immersive visual experience. In Motion Constrained Tile Set (MCTS)-based streaming schemes, the server only updates Field-Of-View (FOV) coordinates when codec generating keyframes, thus the keyframe interval significantly…

Cited by 0SourceScholar
2025

An Information Criterion for Controlled Disentanglement of Multimodal Data

ICLR 2025poster

Multimodal representation learning seeks to relate and decompose information inherent in multiple modalities. By disentangling modality-specific information from information that is shared across modalities, we can improve interpretability and robustness and enable downstream tasks such as the gener…

2025

Coarse-to-Fine 3D Part Assembly via Semantic Super-Parts and Symmetry-Aware Pose Estimation

NeurIPS 2025poster

We propose a novel two-stage framework, Coarse-to-Fine Part Assembly (CFPA), for 3D shape assembly from basic parts. Effective part assembly demands precise local geometric reasoning for accurate component assembly, as well as global structural understanding to ensure semantic coherence and plausibl…

Cited by 0SourceScholar
2025

Exploiting Continuous Motion Clues for Vision-Based Occupancy Prediction

AAAI 2025technical

Occupancy networks aim to reconstruct the surroundings with occupied semantic voxels. However, frequent object occlusions often occur in dynamic real-world scenarios, which cannot be captured by independent frames. Most existing occupancy networks generate results without explicitly considering past…

2025

First-order State Space Model for Lightweight Image Super-resolution

ICASSP 2025accepted

State space models (SSMs), particularly Mamba, have shown promise in NLP tasks and are increasingly applied to vision tasks. However, most Mamba-based vision models focus on network architecture and scan paths, with little attention to the SSM module. In order to explore the potential of SSMs, we mo…

Cited by 0SourceScholar
2025

Joint Knowledge Editing for Information Enrichment and Probability Promotion

AAAI 2025technical

Knowledge stored in large language models requires timely updates to reflect the dynamic nature of real-world information. To update the knowledge, most knowledge editing methods focus on the low layers, since recent probes into the knowledge recall process reveal that the answer information is enri…

2025

No Loss, No Gain: Gated Refinement and Adaptive Compression for Prompt Optimization

NeurIPS 2025poster

Prompt engineering is crucial for leveraging the full potential of large language models (LLMs). While automatic prompt optimization offers a scalable alternative to costly manual design, generating effective prompts remains challenging. Existing methods often struggle to stably generate improved pr…

Cited by 0SourcecodeScholar
2025

Pose Magic: Efficient and Temporally Consistent Human Pose Estimation with a Hybrid Mamba-GCN Network

AAAI 2025technical

Current state-of-the-art (SOTA) methods in 3D Human Pose Estimation (HPE) are primarily based on Transformers. However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work, we leverage recen…

Cited by 3SourcePDFScholar
2025

Training-free Generation of Temporally Consistent Rewards from VLMs

ICCV 2025poster

Recent advances in vision-language models (VLMs) have significantly improved performance in embodied tasks such as goal decomposition and visual comprehension. However, providing accurate rewards for robotic manipulation without fine-tuning VLMs remains challenging due to the absence of domain-speci…

2024

Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs

NeurIPS 2024poster

Diffusion-based image super-resolution (SR) models have attracted substantial interest due to their powerful image restoration capabilities. However, prevailing diffusion models often struggle to strike an optimal balance between efficiency and performance. Typically, they either neglect to exploit…

2023

Learning Efficient Policies for Picking Entangled Wire Harnesses: An Approach to Industrial Bin Picking

RA-L 2023

Wire harnesses are essential connecting components in manufacturing industry but are challenging to be automated in industrial tasks such as bin picking. They are long, flexible and tend to get entangled when randomly placed in a bin. This makes it difficult for the robot to grasp a single one in de

Cited by 27SourcecodeScholar
2023

Learning to Dexterously Pick or Separate Tangled-Prone Objects for Industrial Bin Picking

RA-L 2023

Industrial bin picking for tangled-prone objects requires the robot to either pick up untangled objects or perform separation manipulation when the bin contains no isolated objects. The robot must be able to flexibly perform appropriate actions based on the current observation. It is challenging due

Cited by 5SourceScholar
2023

Unsupervised Surface Anomaly Detection with Diffusion Probabilistic Model

ICCV 2023poster

Unsupervised surface anomaly detection aims at discovering and localizing anomalous patterns using only anomaly-free training samples. Reconstruction-based models are among the most popular and successful methods, which rely on the assumption that anomaly regions are more difficult to reconstruct. H…

Cited by 77PDFScholar
2022

Blind Face Restoration via Integrating Face Shape and Generative Priors

CVPR 2022poster

Blind face restoration, which aims to reconstruct high-quality images from low-quality inputs, can benefit many applications. Although existing generative-based methods achieve significant progress in producing high-quality images, they often fail to restore natural face shapes and high-fidelity fac…

Cited by 48PDFcodeScholar
2021

Learning To Restore Hazy Video: A New Real-World Dataset and a New Method

CVPR 2021poster

Most of the existing deep learning-based dehazing methods are trained and evaluated on the image dehazing datasets, where the dehazed images are generated by only exploiting the information from the corresponding hazy ones. On the other hand, the video dehazing algorithms, which can acquire more sat…

Cited by 103PDFScholar
2020

Multi-Scale Boosted Dehazing Network With Dense Feature Fusion

CVPR 2020poster

In this paper, we propose a Multi-Scale Boosted Dehazing Network with Dense Feature Fusion based on the U-Net architecture. The proposed method is designed based on two principles, boosting and error feedback, and we show that they are suitable for the dehazing problem. By incorporating the Strength…

Cited by 1034PDFcodeScholar