← Search

Lai Wei

21 accepted papers

2026

Artificial Hippocampus Networks for Efficient Long-Context Modeling

ICML 2026poster

Long-sequence modeling faces a fundamental trade-off between the efficiency of compressive fixed-size memory in RNN-like models and the fidelity of lossless growing memory in attention-based Transformers. Inspired by the Multi-Store Model in cognitive science, we introduce a memory framework of arti…

Cited by 0SourceScholar
2026

SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation

AAAI 2026technical

This work proposes a grammar-based chunking strategy that segments input streams into semantically complete units by parsing dependency relations (e.g., noun phrase boundaries, verb-object structures) and punctuation features. The method ensures chunk coherence and minimizes semantic fragmentation.

Cited by 0SourcePDFScholar
2026

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

ICML 2026poster

Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-with-Images" methods alleviate this by iteratively zooming into regions of interes…

Cited by 0SourceScholar
2025

AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator

COLING 2025main

Artificial intelligence has significantly revolutionized healthcare, particularly through large language models (LLMs) that demonstrate superior performance in static medical question answering benchmarks. However, evaluating the potential of LLMs for real-world clinical applications remains challen…

2025

ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning

EMNLP 2025

Large Reasoning Models (LRMs) perform strongly in complex reasoning tasks via Chain-of-Thought (CoT) prompting, but often suffer from verbose outputs, increasing computational overhead. Existing fine-tuning-based compression methods either operate post-hoc pruning, risking disruption to reasoning co

Cited by 0SourcePDFScholar
2025

DSFormer: Deformable Pointformer for 3D Salient Object Detection

ICASSP 2025accepted

Due to the irregularity of 3D point clouds, it is extremely challenging to detect the most salient objects from them and segment contours. DSFormer is proposed for 3D salient object detection, addressing challenges like small objects, multiple objects and complex backgrounds. It employs an encoder-d…

Cited by 0SourceScholar
2025

Ensuring Force Safety in Vision-Guided Robotic Manipulation via Implicit Tactile Calibration

CoRL 2025poster

In unstructured environments, robotic manipulation tasks involving objects with constrained motion trajectories—such as door opening—often experience discrepancies between the robot's vision-guided end-effector trajectory and the object's constrained motion path. Such discrepancies generate uninten…

Cited by 0SourceScholar
2025

First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training

NeurIPS 2025poster

Improving Multi-modal Large Language Models (MLLMs) in the post-training stage typically relies on supervised fine-tuning (SFT) or reinforcement learning (RL), which require expensive and manually annotated multi-modal data--an ultimately unsustainable resource. This limitation has motivated a growi…

Cited by 0SourcecodeScholar
2024

Adaptive Online Experimental Design for Causal Discovery

ICML 2024spotlight

Causal discovery aims to uncover cause-and-effect relationships encoded in causal graphs by leveraging observational, interventional data, or their combination. The majority of existing causal discovery methods are developed assuming infinite interventional data. We focus on interventional data effi…

Cited by 1SourcePDFScholar
2024

Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models

NeurIPS 2024poster

Large Language Models (LLMs) have transformed natural language processing and extended their powerful capabilities to multi-modal domains. As LLMs continue to advance, it is crucial to develop diverse and appropriate metrics for their evaluation. In this paper, we introduce a novel rank-based metric…

2024

Discriminatively Fuzzy Multi-View K-means Clustering with Local Structure Preserving

AAAI 2024technical

Multi-view K-means clustering successfully generalizes K-means from single-view to multi-view, and obtains excellent clustering performance. In every view, it makes each data point close to the center of the corresponding cluster. However, multi-view K-means only considers the compactness of each cl…

Cited by 4SourcePDFScholar
2024

Unidirectional Brain-Computer Interface: Artificial Neural Network Encoding Natural Images to FMRI Response in the Visual Cortex

ICASSP 2024accepted

While significant advancements in artificial intelligence (AI) have catalyzed progress across various domains, its full potential in understanding visual perception remains underexplored. We propose an artificial neural network dubbed VISION, an acronym for "Visual Interface System for Imaging Outpu…

Cited by 0SourceScholar
2023

Adaptive Graph Convolutional Subspace Clustering

CVPR 2023poster

Spectral-type subspace clustering algorithms have shown excellent performance in many subspace clustering applications. The existing spectral-type subspace clustering algorithms either focus on designing constraints for the reconstruction coefficient matrix or feature extraction methods for finding…

2023

Approximate Allocation Matching for Structural Causal Bandits with Unobserved Confounders

NeurIPS 2023poster

Structural causal bandit provides a framework for online decision-making problems when causal information is available. It models the stochastic environment with a structural causal model (SCM) that governs the causal relations between random variables. In each round, an agent applies an interventio…

2023

SCoDA: Domain Adaptive Shape Completion for Real Scans

CVPR 2023poster

3D shape completion from point clouds is a challenging task, especially from scans of real-world objects. Considering the paucity of 3D shape ground truths for real scans, existing works mainly focus on benchmarking this task on synthetic data, e.g. 3D computer-aided design models. However, the doma…

2021

Multi-Robot Gaussian Process Estimation and Coverage: Deterministic Sequencing Algorithm and Regret Analysis

ICRA 2021poster

We study the problem of multi-robot coverage over an unknown, nonuniform sensory field. Modeling the sensory field as a realization of a Gaussian Process and using Bayesian techniques, we devise a policy which aims to balance the tradeoff between learning the sensory function and covering the enviro…

Cited by 18SourceScholar
2020

Expedited Multi-Target Search with Guaranteed Performance via Multi-fidelity Gaussian Processes

IROS 2020poster

We consider a scenario in which an autonomous vehicle equipped with a downward facing camera operates in a 3D environment and is tasked with searching for an unknown number of stationary targets on the 2D floor of the environment. The key challenge is to minimize the search time while ensuring a hig…

Cited by 8SourceScholar