← Search

Xinyu Sun

7 accepted papers

2026

Virtual Community: An Open World for Humans, Robots, and Society

ICLR 2026poster

The rapid progress of AI and robotics may profoundly transform society, as humans and robots begin to coexist in shared communities, bringing both opportunities and challenges. To explore this future, we present Virtual Community—an open-world platform for humans, robots, and society—built on a univ…

Cited by 0SourcecodeScholar
2025

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

CVPR 2025poster

Research on 3D Vision-Language Models (3D-VLMs) is gaining increasing attention, which is crucial for developing embodied AI within 3D scenes, such as visual navigation and embodied question answering. Due to the high density of visual features, especially in large 3D scenes, accurately locating tas…

2025

Language-Guided Hybrid Representation Learning for Visual Grounding on Remote Sensing Images

IJCAI 2025

Visual grounding (VG) refers to detecting the specific objects in images based on linguistic expressions, and it has profound significance in the advanced interpretation of natural images. In remote sensing image interpretation, visual grounding is limited by characteristics such as the complex scen

Cited by 0SourcePDFScholar
2025

Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference Learning

ICCV 2025poster

The image signal processing (ISP) pipeline is responsible for converting the RAW images collected from the sensor into high-quality RGB images. It contains a series of image processing modules and associated ISP hyperparameters. Recent learning-based approaches aim to automate ISP hyperparameter opt…

Cited by 0SourcePDFScholar
2024

RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal Processing

AAAI 2024technical

Hardware image signal processing (ISP), aiming at converting RAW inputs to RGB images, consists of a series of processing blocks, each with multiple parameters. Traditionally, ISP parameters are manually tuned in isolation by imaging experts according to application-specific quality and performance…

Cited by 4SourcePDFScholar
2023

FGPrompt: Fine-grained Goal Prompting for Image-goal Navigation

NeurIPS 2023poster

Learning to navigate to an image-specified goal is an important but challenging task for autonomous systems like household robots. The agent is required to well understand and reason the location of the navigation goal from a picture shot in the goal position. Existing methods try to solve this prob…

Cited by 13SourcePDFScholar
2023

Masked Motion Encoding for Self-Supervised Video Representation Learning

CVPR 2023poster

How to learn discriminative video representation from unlabeled videos is challenging but crucial for video analysis. The latest attempts seek to learn a representation model by predicting the appearance contents in the masked regions. However, simply masking and recovering appearance contents may n…