← Search

Wei Luo

29 accepted papers

2026

4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

CVPR 2026

World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct realistic, dynamic, and physically consistent 3D/4D worlds from images, videos, or text. These models not only need to prod

Cited by 0SourcecodeScholar
2026

Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

AAAI 2026technical

We propose Anomagic, a zero-shot anomaly generation method that produces semantically coherent anomalies without requiring any exemplar anomalies. By unifying both visual and textual cues through a crossmodal prompt encoding scheme, Anomagic leverages rich contextual information to steer an inpaint

Cited by 0SourcePDFScholar
2026

Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory

ICML 2026poster

Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this failure remains poorly understood. In this work, we reveal that hallucinations are strongly associated with a human-like attention distraction phenomenon, where …

Cited by 0SourceScholar
2026

DeepWriter: A Multi-Agent Collaboration Framework for Information-rich Ultra-long Book Writing

AAAI 2026technical

Long-form books are among the most information-rich and structurally complex forms of written content, often exceeding 100,000 words. While recent methods have enabled basic long-text generation, they remain limited in two key aspects: the inability to generate ultra-long content at book scale, and

Cited by 0SourcePDFScholar
2026

Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking Across Datasets, Models, and Generated Content

IJCAI 2026

Large language models (LLMs) are substantial investments and increasingly deployed in high-stakes domains, making it critical to protect LLM-related assets and to trace their provenance.Identity technologies such as fingerprinting and watermarking address these needs by enabling ownership verificati

Cited by 0Scholar
2026

Parameter-, Memory-, Time-Efficient Multi-Task Dense Vision Adaptation

AAAI 2026technical

While adapting pretrained vision models to downstream dense prediction tasks is widely used, current methods often overlook adaptation efficiency, especially in the context of multi-task learning (MTL). Although parameter-efficient fine-tuning (PEFT) methods can enhance parameter efficiency, broader

Cited by 0SourcePDFScholar
2026

ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer Acceleration

AAAI 2026technical

Diffusion Transformers (DiTs) have achieved state-of-the-art performance in generative modeling, yet their high computational cost hinders real-time deployment. While feature caching offers a promising training-free acceleration solution by exploiting temporal redundancy, existing methods suffer fro

Cited by 0SourcePDFScholar
2026

Robot Deformable Object Manipulation Via NMPC-Generated Demonstrations in Deep Reinforcement Learning (I)

ICRA 2026poster

In this work, we conducted research on deformable object manipulation by robots based on demonstration-enhanced reinforcement learning (RL). We present FADERL (FuzzyAugmented Demonstration-Embedded Reinforcement Learning),a novel framework for robotic manipulation of deformable objects that signific…

Cited by 0Scholar
2026

TDSS: Task Dynamic-Synergistic Skill Adaptation for Boosting Efficient and Scalable Multi-Task Learning in Dense Visual Prediction

AAAI 2026technical

The transfer of knowledge from large-scale pre-trained models to diverse downstream tasks has achieved remarkable success. Beyond the traditional full fine-tuning paradigm, Parameter-Efficient Fine-Tuning (PEFT) has emerged as a more efficient model adaptation approach. However, applying existing PE

Cited by 0SourcePDFScholar
2026

Tool-Grasp: A 6-DoF Functional Grasping Framework for General-Purpose Hand Tools

ICRA 2026poster

Detecting functional grasp poses for tool operation is critical for robots in complex real-world tasks, yet existing methods lack this capability. Key challenges are: 1) Scarce realworld datasets with fine-grained functional labels and task-valid grasp annotations, as their construction requires dom…

Cited by 0Scholar
2025

Automated Detection of Pre-training Text in Black-box LLMs

IJCAI 2025

Detecting whether a given text is a member in the pre-training data of Large Language Models (LLMs) is crucial for ensuring data privacy and copyright protection. Most existing methods rely on the LLM's hidden information (e.g., model parameters or token probabilities), making them ineffective in th

2025

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing

AAAI 2025technical

Large multimodal language models (MLLMs) have revolutionized natural language processing and visual understanding, but often contain outdated or inaccurate information. Current multimodal knowledge editing evaluations are limited in scope and potentially biased, focusing on narrow tasks and failing…

2025

Exploring Intrinsic Normal Prototypes within a Single Image for Universal Anomaly Detection

CVPR 2025poster

Anomaly detection (AD) is essential for industrial inspection, yet existing methods typically rely on "comparing" test images to normal references from a training set. However, variations in appearance and positioning often complicate the alignment of these references with the test image, limiting d…

2025

Improve Speech Translation Through Text Rewrite

COLING 2025industry

Despite recent progress in Speech Translation (ST) research, the challenges posed by inherent speech phenomena that distinguish transcribed speech from written text are not well addressed. The informal and erroneous nature of spontaneous speech is inadequately represented in the typical parallel tex…

2025

Improving Robustness of Post-hoc Calibration Against Common Corruptions By Learnable Augmentation

ICASSP 2025accepted

Various research has addressed the overconfidence problem, and we focus on improving the robustness of post-hoc calibration (e.g., temperature scaling, TS) when the test set shifts from the training set by image corruption. TS is greatly affected by the validation set, which previous work has propos…

Cited by 0SourceScholar
2025

LLM-Friendly Knowledge Representation for Customer Support

COLING 2025industry

We propose a practical approach by integrating Large Language Models (LLMs) with a framework designed to navigate the complexities of Airbnb customer support operations. In this paper, our methodology employs a novel reformatting technique, the Intent, Context, and Action (ICA) format, which transfo…

Cited by 2SourcePDFScholar
2025

Test-Time Learning for Large Language Models

ICML 2025poster

While Large Language Models (LLMs) have exhibited remarkable emergent capabilities through extensive pre-training, they still face critical limitations in generalizing to specialized domains and handling diverse linguistic variations, known as distribution shifts. In this paper, we propose a Test-T…

Cited by 0SourcePDFScholar
2025

Uncovering Argumentative Flow: A Question-Focus Discourse Structuring Framework

EMNLP 2025

Understanding the underlying argumentative flow in analytic argumentative writing is essential for discourse comprehension, especially in complex argumentative discourse such as think-tank commentary. However, existing structure modeling approaches often rely on surface-level topic segmentation, fai

Cited by 0SourcePDFScholar
2024

Beyond Linguistic Cues: Fine-grained Conversational Emotion Recognition via Belief-Desire Modelling

COLING 2024main

Emotion recognition in conversation (ERC) is essential for dialogue systems to identify the emotions expressed by speakers. Although previous studies have made significant progress, accurate recognition and interpretation of similar fine-grained emotion properly accounting for individual variability…

Cited by 2SourcePDFScholar
2024

Divergence-Guided Simultaneous Speech Translation

AAAI 2024technical

To achieve high-quality translation with low latency, a Simultaneous Speech Translation (SimulST) system relies on a policy module to decide whether to translate immediately or wait for additional streaming input, along with a translation model capable of effectively handling partial speech input. P…

2024

IBD-PSC: Input-level Backdoor Detection via Parameter-oriented Scaling Consistency

ICML 2024poster

Deep neural networks (DNNs) are vulnerable to backdoor attacks, where adversaries can maliciously trigger model misclassifications by implanting a hidden backdoor during model training. This paper proposes a simple yet effective input-level backdoor detection (dubbed IBD-PSC) as a `firewall' to filt…

2024

MARIO: MAth Reasoning with code Interpreter Output - A Reproducible Pipeline

ACL 2024findings

Large language models (LLMs) have significantly improved in understanding natural language but still lack in mathematical reasoning, a hurdle on the path to true artificial general intelligence. The training of large language models, based on next-token prediction, struggles to capture the precise n…

2023

A Canonicalization-Enhanced Known Fact-Aware Framework For Open Knowledge Graph Link Prediction

IJCAI 2023poster

Open knowledge graph (OpenKG) link prediction aims to predict missing factual triples in the form of (head noun phrase, relation phrase, tail noun phrase). Since triples are not canonicalized, previous methods either focus on canonicalizing noun phrases (NPs) to reduce graph sparsity, or utilize tex…

2023

Adaptive Policy with Wait-k Model for Simultaneous Translation

EMNLP 2023long main

Simultaneous machine translation (SiMT) requires a robust read/write policy in conjunction with a high-quality translation model. Traditional methods rely on either a fixed wait-k policy coupled with a standalone wait-k translation model, or an adaptive policy jointly trained with the translation m…

Cited by 0SourceScholar
2023

Better Simultaneous Translation with Monotonic Knowledge Distillation

ACL 2023long

Simultaneous machine translation (SiMT) presents a unique challenge as it requires generating target tokens before the source sentence is fully consumed. This can lead to the hallucination problem, where target tokens are generated without support from the source sentence. The prefix-to-prefix train…

2023

Do We Need an Encoder-Decoder to Model Dynamical Systems on Networks?

IJCAI 2023poster

As deep learning gains popularity in modelling dynamical systems, we expose an underappreciated misunderstanding relevant to modelling dynamics on networks. Strongly influenced by graph neural networks, latent vertex embeddings are naturally adopted in many neural dynamical network models. However,…

2022

Discrete Cross-Modal Alignment Enables Zero-Shot Speech Translation

EMNLP 2022main

End-to-end Speech Translation (ST) aims at translating the source language speech into target language text without generating the intermediate transcriptions. However, the training of end-to-end methods relies on parallel ST data, which are difficult and expensive to obtain. Fortunately, the superv…

2019

Cross-X Learning for Fine-Grained Visual Categorization

ICCV 2019poster

Recognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features ar…

Cited by 241PDFcodeScholar