← Search

Weizhe Huang

4 accepted papers

2026

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference

ICML 2026poster

DeepSeek-OCR leverages visual–text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural information. Moreover, current token pruning methods for conventional vision–language models (VLMs) fail to preserve textua…

Cited by 0SourceScholar
2025

Logic-of-Thought: Injecting Logic into Contexts for Full Reasoning in Large Language Models

NAACL 2025long

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks but their performance in complex logical reasoning tasks remains unsatisfactory. Although some prompting methods, such as Chain-of-Thought, can improve the reasoning ability of LLMs to some extent, they suffe…

Cited by 9SourcePDFScholar
2025

S2-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency

NAACL 2025long

Large language models (LLMs) have demonstrated remarkable capabilities across various natural language processing (NLP) scenarios, but they still face challenges when handling complex arithmetic and logical reasoning tasks. While Chain-Of-Thought (CoT) reasoning, self-consistency (SC) and self-corre…

Cited by 1SourcePDFScholar
2023

A Bounded Ability Estimation for Computerized Adaptive Testing

NeurIPS 2023poster

Computerized adaptive testing (CAT), as a tool that can efficiently measure student's ability, has been widely used in various standardized tests (e.g., GMAT and GRE). The adaptivity of CAT refers to the selection of the most informative questions for each student, reducing test length. Existing CAT…