← Search

Cheng Chang

14 accepted papers

2026

Diffusion-Enhanced Tree Planning for Autonomous Driving

RA-L 2026

In highly interactive urban driving, decision making is often naturally multi-stage, and decisions at different stages can lead to different reactions from surrounding vehicles. This calls for stage-wise evaluation and selection. Tree-based planning naturally supports multi-stage search and evaluati

Cited by 0SourceScholar
2026

MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual Knowledge

ICML 2026poster

As real-world knowledge continues to evolve, the parametric knowledge acquired by multimodal models during pretraining becomes increasingly difficult to remain consistent with real-world knowledge. Existing research on multimodal knowledge updating focuses only on learning previously unknown knowled…

Cited by 0SourceScholar
2026

WARC-Bench: Web Archive based Benchmark for GUI Subtask Executions

ICLR 2026poster

Training web agents to navigate complex, real-world websites requires them to master subtasks—short-horizon interactions on multiple UI components (e.g., choosing the correct date in a date picker, or scrolling in a container to extract information). We introduce WARC-Bench (Web Archive Benchmark),…

Cited by 0SourceScholar
2025

DUSTED: Dual-Attention Enhanced Spatial Transcriptomics Denoiser

AAAI 2025technical

Spatially Resolved Transcriptomics (SRT) has become an indispensable tool in various fields, including tumor microenvironment identification, neurobiology, and the study of complex tissue architecture. However, the accuracy of these insights is often compromised by noise in spatial transcriptomics…

2025

Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

ICLR 2025oral

Multimodal large language models (MLLMs) are transforming the capabilities of graphical user interface (GUI) agents, facilitating their transition from controlled simulations to complex, real-world applications across various platforms. However, the effectiveness of these agents hinges on the robust…

2024

Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models

COLING 2024main

Recent advancements in Chain-of-Thought prompting have facilitated significant breakthroughs for Large Language Models (LLMs) in complex reasoning tasks. Current research enhances the reasoning performance of LLMs by sampling multiple reasoning chains and ensembling based on the answer frequency. Ho…

2024

ContraNovo: A Contrastive Learning Approach to Enhance De Novo Peptide Sequencing

AAAI 2024technical

De novo peptide sequencing from mass spectrometry (MS) data is a critical task in proteomics research. Traditional de novo algorithms have encountered a bottleneck in accuracy due to the inherent complexity of proteomics data. While deep learning-based methods have shown progress, they reduce the pr…

2024

Mixing Left and Right-Hand Driving Data in a Hierarchical Framework With LLM Generation

RA-L 2024

Data-driven trajectory prediction is critical in autonomous vehicles, which requires high-quality data. However, discussions about the compatibility of data collected from different countries remain limited, with a typical issue being the different driving rules in various countries. Therefore, we p

Cited by 5SourceScholar
2024

Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles

EMNLP 2024main

Recent works leverage LLMs to roleplay realistic social scenarios, aiding novices in practicing their social skills. However, simulating sensitive interactions, such as in the domain of mental health, is challenging. Privacy concerns restrict data access, and collecting expert feedback, although vit…

2023

Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication

EMNLP 2023long main

Large Language Models (LLMs) have recently made significant strides in complex reasoning tasks through the Chain-of-Thought technique. Despite this progress, their reasoning is often constrained by their intrinsic understanding, lacking external insights. To address this, we propose Exchange-of-Thou…

Cited by 0SourcecodeScholar
2019

Guided Similarity Separation for Image Retrieval

NeurIPS 2019oral

Despite recent progress in computer vision, image retrieval remains a challenging open problem. Numerous variations such as view angle, lighting and occlusion make it difficult to design models that are both robust and efficient. Many leading methods traverse the nearest neighbor graph to exploit hi…

2019

L2 Learners' Emotion Production in Video Dubbing Practices

ICASSP 2019accepted

Video dubbing is a new type of language learning practice. Because of the fun it brings into learning, video dubbing mobile applications have become quite popular. During video dubbing, learners not only mimic characters' pronunciations but also other voicing characteristics, e.g., emotions. In this…

Cited by 4SourceScholar
2018

Policy Adaptation for Deep Reinforcement Learning-Based Dialogue Management

ICASSP 2018accepted

Policy optimization is the core part of statistical dialogue management. Deep reinforcement learning has been successfully used for dialogue policy optimization for a static pre-defined domain. However, when the domain changes dynamically, e.g. a new previously unseen concept (or slot) which can be…

Cited by 0SourceScholar