← Search

Ruoxi Chen

6 accepted papers

2026

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

ICLR 2026poster

Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field. However, existing benchmarks fail to differentiate question difficulty, limiting their ability…

Cited by 0SourcecodeScholar
2025

Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment

ICLR 2025spotlight

Many real-world user queries (e.g. *"How do to make egg fried rice?"*) could benefit from systems capable of generating responses with both textual steps with accompanying images, similar to a cookbook. Models designed to generate interleaved text and images face challenges in ensuring consistency w…

Cited by 8SourcePDFScholar
2024

An Online Rcm Adjusting System for Robot-Assisted Retinal Surgeries

IROS 2024poster

In robot-assisted retinal surgery, a Remote Center of Motion (Rcm) allows the surgical instrument to rotate around a distal fixed point without any lateral translations. The Rcm point should be perfectly aligned inside the trocar. Otherwise, unexpected tool translations at the expected remote center…

Cited by 0SourceScholar
2024

CatchBackdoor: Backdoor Detection via Critical Trojan Neural Path Fuzzing

ECCV 2024poster

"The success of deep neural networks (DNNs) in real-world applications has benefited from abundant pre-trained models. However, the backdoored pre-trained models can pose a significant trojan threat to the deployment of downstream DNNs. Numerous backdoor detection methods have been proposed but are…

Cited by 2SourcePDFScholar
2024

EditShield: Protecting Unauthorized Image Editing by Instruction-guided Diffusion Models

ECCV 2024poster

"Text-to-image diffusion models have emerged as an evolutionary for producing creative content in image synthesis. Based on the impressive generation abilities of these models, instruction-guided diffusion models can edit images with simple instructions and input images. While they empower users to…

Cited by 12SourcePDFScholar
2024

MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark

ICML 2024oral

Multimodal Large Language Models (MLLMs) have gained significant attention recently, showing remarkable potential in artificial general intelligence. However, assessing the utility of MLLMs presents considerable challenges, primarily due to the absence multimodal benchmarks that align with human pre…