← Search

Zeyu Lu

10 accepted papers

2025

ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems

CVPR 2025poster

Much previous AI research has focused on developing monolithic models to maximize their intelligence, with the primary goal of enhancing performance on specific tasks. In contrast, this work attempts to study using LLM-based agents to design collaborative AI systems autonomously. To explore this pro…

2025

Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots

NAACL 2025findings

Multi-modal Large Language Models have shown remarkable progress in visual contexts, yet their ability to convert visual figures into executable code remains underexplored. To address this, we introduce Plot2Code, a comprehensive benchmark designed to assess MLLMs’ visual coding capabilities. Plot2C…

2025

ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model

CVPR 2025poster

The scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion to…

Cited by 6SourcePDFScholar
2024

FiT: Flexible Vision Transformer for Diffusion Model

ICML 2024spotlight

In the context of this reality, existing diffusion models, such as Diffusion Transformers, often face challenges when processing image resolutions outside of their trained domain. To overcome this limitation, we present the Flexible Vision Transformer (FiT), a transformer architecture specifically d…

2024

LLaMA Pro: Progressive LLaMA with Block Expansion

ACL 2024long

Humans generally acquire new skills without compromising the old; however, the opposite holds for Large Language Models (LLMs), e.g., from LLaMA to CodeLLaMA. To this end, we propose a new post-pretraining method for LLMs with an expansion of Transformer blocks. We tune the expanded blocks using onl…

2023

$\pi$-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation

ICML 2023poster

Foundation models have achieved great advances in multi-task learning with a unified interface of unimodal and multimodal tasks. However, the potential of such multi-task learners has not been exploited during transfer learning. In this work, we present a universal parameter-efficient transfer learn…

2023

Interaction Control for Tool Manipulation on Deformable Objects Using Tactile Feedback

RA-L 2023

The human sense of touch enables us to perform delicate tasks on deformable objects and/or in a vision-denied environment. To achieve similar desirable interactions for robots, such as administering a swab test, tactile information sensed beyond the tool-in-hand is crucial for contact state estimati

Cited by 9SourceScholar
2023

Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images

NeurIPS 2023poster

Photos serve as a way for humans to record what they experience in their daily lives, and they are often regarded as trustworthy sources of information. However, there is a growing concern that the advancement of artificial intelligence (AI) technology may produce fake photos, which can create confu…

2022

GTac-Gripper: A Reconfigurable Under-Actuated Four-Fingered Robotic Gripper With Tactile Sensing

RA-L 2022

Humans can use different grasping poses and forces for everyday objects of different shapes and sizes. Grasping and manipulating everyday objects have been longstanding challenges in robotics. Performing multiple grasping configurations is difficult for robotic end-effectors with limited degrees of

Cited by 24SourceScholar
2020

A Deep Learning Based End-to-End Locomotion Mode Detection Method for Lower Limb Wearable Robot Control

IROS 2020poster

To function effectively in real-world environments, powered wearable robots such as exoskeletons and robotic prostheses must recognize the user's motion intent by detecting the user's locomotion modes such as walking, stair ascent and descent or ramp ascent and descent. Traditionally, intent detecti…

Cited by 19SourceScholar