← Search

Siqi Li

21 accepted papers

2026

Benchmarking and Enhancing VLM for Compressed Image Understanding

ICML 2026poster

With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs has become increasingly important. Existing VLMs predominantly digest and understand high-bitrate compressed images, while their ability to interpret l…

Cited by 0SourceScholar
2026

Hyper-PCN: Hypergraph-Based Point Cloud Completion via High-Order Correlation Modeling

CVPR 2026

Point cloud completion is an important yet challenging problem in 3D computer vision, which aims to reconstruct complete and dense 3D shapes from partial point clouds. Although transformer-based and geometry-based approaches have made significant progress, they often struggle to capture the complex,

Cited by 0SourcecodeScholar
2025

Chameleon: Fast-Slow Neuro-Symbolic Lane Topology Extraction

ICRA 2025

Lane topology extraction involves detecting lanes and traffic elements and determining their relationships, a key perception task for mapless autonomous driving. This task requires complex reasoning, such as determining whether it is possible to turn left into a specific lane. To address this challe

Cited by 12SourcecodeScholar
2025

ERetinex: Event Camera Meets Retinex Theory for Low-Light Image Enhancement

ICRA 2025

Low-light image enhancement aims to restore the under-exposure image captured in dark scenarios. Under such scenarios, traditional frame-based cameras may fail to capture the structure and color information due to the exposure time limitation. Event cameras are bio-inspired vision sensors that respo

Cited by 5SourcecodeScholar
2025

GraphI2P: Image-to-Point Cloud Registration with Exploring Pattern of Correspondence via Graph Learning

CVPR 2025poster

Although the fusion of images and LiDAR point clouds is crucial to many applications in computer vision, the relative poses of cameras and LiDAR scanners are often unknown. The general registration pipeline first establishes correspondences and then performs pose estimation based on the generated ma…

Cited by 0SourcePDFScholar
2025

Hyper-Depth: Hypergraph-based Multi-Scale Representation Fusion for Monocular Depth Estimation

ICCV 2025poster

Monocular depth estimation (MDE) is a fundamental problem in computer vision with wide-ranging applications in various downstream tasks. While multi-scale features are perceptually critical for MDE, existing transformer-based methods have yet to leverage them explicitly. To address this limitation,…

Cited by 0SourcePDFScholar
2025

Learning Symmetric Legged Locomotion via State Distribution Symmetrization

IROS 2025

Morphological symmetry is a fundamental characteristic of legged animals and robots. Most existing Deep Reinforcement Learning approaches for legged locomotion neglect to exploit this inherent symmetry, often producing unnatural and suboptimal behaviors such as dominant legs or non-periodic gaits. T

Cited by 0SourceScholar
2025

Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEs

CVPR 2025poster

Place recognition (PR) aims at retrieving the query place from a database and plays a crucial role in various applications, including navigation, autonomous driving, and augmented reality. While previous multi-modal PR works have mainly focused on the same-view scenario in which ground-view descript…

Cited by 0SourcePDFScholar
2025

SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMs

NeurIPS 2025poster

Recent multimodal large language models (MLLMs) marry modality-specific vision or audio encoders with a shared text decoder. While the encoder is compute- intensive but memory-light, the decoder is the opposite, yet state-of-the-art serving stacks still time-multiplex these complementary kernels, id…

Cited by 0SourcecodeScholar
2025

UAVScenes: A Multi-Modal Dataset for UAVs

ICCV 2025poster

Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward localization and 3D reconstruction tasks, or only support ma…

2025

UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition

ICASSP 2025accepted

Recent advancements in scaling up models have significantly improved performance in Automatic Speech Recognition (ASR) tasks. However, training large ASR models from scratch remains costly. To address this issue, we introduce UME, a novel method that efficiently Upcycles pretrained dense ASR checkpo…

Cited by 0SourceScholar
2024

MaxQ: Multi-Axis Query for N:M Sparsity Network

CVPR 2024poster

N:M sparsity has received increasing attention due to its remarkable performance and latency trade-off compared with structured and unstructured sparsity. However existing N:M sparsity methods do not differentiate the relative importance of weights among blocks and leave important weights underappre…

2024

Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration Approach

EMNLP 2024main

Direct speech translation (ST) models often struggle with rare words. Incorrect translation of these words can have severe consequences, impacting translation quality and user trust. While rare word translation is inherently challenging for neural models due to sparse learning signals, real-world sc…

2024

OvSW: Overcoming Silent Weights for Accurate Binary Neural Networks

ECCV 2024poster

"Binary Neural Networks (BNNs) have been proven to be highly effective for deploying deep neural networks on mobile and embedded platforms. Most existing works focus on minimizing quantization errors, improving representation ability, or designing gradient approximations to alleviate gradient mismat…

2024

PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments

IROS 2024poster

Robotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are limited in their adaptability across different object categorie…

Cited by 4SourceScholar
2024

Structured Optimal Brain Pruning for Large Language Models

EMNLP 2024main

The massive parameters and computational demands hinder the widespread application of Large Language Models (LLMs). Network pruning provides a practical solution to this problem. However, existing pruning works for LLMs mainly focus on unstructured pruning or necessitate post-pruning fine-tuning. Th…

Cited by 1SourcePDFScholar
2023

SUBP: Soft Uniform Block Pruning for 1$\times$N Sparse CNNs Multithreading Acceleration

NeurIPS 2023poster

The study of sparsity in Convolutional Neural Networks (CNNs) has become widespread to compress and accelerate models in environments with limited resources. By constraining N consecutive weights along the output channel to be group-wise non-zero, the recent network with 1$\times$N sparsity has rece…

2023

UFO2: A Unified Pre-Training Framework for Online and Offline Speech Recognition

ICASSP 2023accepted

In this paper, we propose a Unified pre-training Framework for Online and Offline (UFO2) Automatic Speech Recognition (ASR), which 1) simplifies the two separate training workflows for online and offline modes into one process, and 2) improves the Word Error Rate (WER) performance with limited utter…

Cited by 0SourceScholar
2021

Event Stream Super-Resolution via Spatiotemporal Constraint Learning

ICCV 2021poster

Event cameras are bio-inspired sensors that respond to brightness changes asynchronously and output in the form of event streams instead of frame-based images. They own outstanding advantages compared with traditional cameras: higher temporal resolution, higher dynamic range, and lower power consump…

Cited by 21PDFScholar