← Search

Zhen Zeng

20 accepted papers

2025

A Multi-Style Chinese Characters Writing Intelligent Tool Based on Small-scale Training Data

AAAI 2025technical

Chinese characters are a unique blend of language and art, featuring diverse artistic styles. Mastering these styles requires extensive practice and limits public participation. To encourage broader participation, we developed a real-time, interactive tool that supports multiple Chinese character ar…

Cited by 0SourcePDFScholar
2025

AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations

ACL 2025long

State-of-the-art multimodal web agents, powered by Multimodal Large Language Models (MLLMs), can autonomously execute many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs). Current strategies for building web agents rely on (i) the generalizability of u…

Cited by 0SourcePDFScholar
2025

LAW: Legal Agentic Workflows for Custody and Fund Services Contracts

COLING 2025industry

Legal contracts in the custody and fund services domain govern critical aspects such as key provider responsibilities, fee schedules, and indemnification rights. However, it is challenging for an off-the-shelf Large Language Model (LLM) to ingest these contracts due to the lengthy unstructured strea…

2025

LETS-C: Leveraging Text Embedding for Time Series Classification

ACL 2025long

Recent advancements in language modeling have shown promising results when applied to time series data. In particular, fine-tuning pre-trained large language models (LLMs) for time series classification tasks has achieved state-of-the-art (SOTA) performance on standard benchmarks. However, these LLM…

Cited by 0SourcePDFScholar
2025

Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models

ICCV 2025poster

Existing knowledge editing works for MultiModal Large Language Models primarily focus on text-oriented, coarse-grained scenarios, where modifying textual content alone is sufficient. As a result, they fail to capture the unique challenges of multimodal editing, particularly when visual information i…

2024

Boosting Neural Cognitive Diagnosis with Student’s Affective State Modeling

AAAI 2024technical

Cognitive Diagnosis Modeling aims to infer students' proficiency level on knowledge concepts from their response logs. Existing methods typically model students’ response processes as the interaction between students and exercises or concepts based on hand-crafted or deeply-learned interaction funct…

2024

Evaluating Large Language Models on Time Series Feature Understanding: A Comprehensive Taxonomy and Benchmark

EMNLP 2024main

Large Language Models (LLMs) offer the potential for automatic time series analysis and reporting, which is a critical task across many domains, spanning healthcare, finance, climate, energy, and many more. In this paper, we propose a framework for rigorously evaluating the capabilities of LLMs on t…

Cited by 8SourcePDFScholar
2023

Self-Supervised Graph Learning for Long-Tailed Cognitive Diagnosis

AAAI 2023technical

Cognitive diagnosis is a fundamental yet critical research task in the field of intelligent education, which aims to discover the proficiency level of different students on specific knowledge concepts. Despite the effectiveness of existing efforts, previous methods always considered the mastery leve…

2022

Composable Causality in Semantic Robot Programming

ICRA 2022poster

Assembly tasks are challenging for robot manipulation because the robot must reason over the composed effects of actions and execute multi-objective behaviors. Robots typically use predefined priorities provided by users to determine how to compose controller behaviors, but we want the robot to auto…

Cited by 2SourceScholar
2022

Elephants Don't Pack Groceries: Robot Task Planning for Low Entropy Belief States

RA-L 2022

Recent advances in computational perception have significantly improved the ability of autonomous robots to perform state estimation with low entropy. Such advances motivate a reconsideration of robot decision-making under uncertainty. Current approaches to solving sequential decision-making problem

Cited by 7SourceScholar
2021

LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation

ICASSP 2021accepted

In this paper, we propose a novel conditional convolution network, named location-variable convolution, to model the dependencies of the waveform sequence. Different from the use of unified convolution kernels in WaveNet to capture the dependencies of arbitrary waveform, the location-variable convol…

Cited by 0SourceScholar
2021

Probabilistic Inference in Planning for Partially Observable Long Horizon Problems

IROS 2021poster

For autonomous service robots to successfully perform long horizon tasks in the real world, they must act intelligently in partially observable environments. Most Task and Motion Planning approaches assume full observability of their state space, making them ineffective in stochastic and partially o…

Cited by 11SourceScholar
2021

Semantic Linking Maps for Active Visual Object Search (Extended Abstract)

IJCAI 2021poster

We aim for mobile robots to function in a variety of common human environments, which requires them to efficiently search previously unseen target objects. We can exploit background knowledge about common spatial relations between landmark objects and target objects to narrow down search space. In t…

Cited by 0SourcePDFScholar
2020

Aligntts: Efficient Feed-Forward Text-to-Speech System Without Explicit Alignment

ICASSP 2020accepted

Targeting at both high efficiency and performance, we propose AlignTTS to predict the mel-spectrum in parallel. AlignTTS is based on a Feed-Forward Transformer which generates mel-spectrum from a sequence of characters, and the duration of each character is determined by a duration predictor. Instea…

Cited by 0SourceScholar
2020

GraphTTS: Graph-to-Sequence Modelling in Neural Text-to-Speech

ICASSP 2020accepted

This paper leverages the graph-to-sequence method in neural text-to-speech (GraphTTS), which maps the graph embedding of the input sequence to spectrograms. The graphical inputs consist of node and edge representations constructed from input texts. The encoding of these graphical inputs incorporates…

Cited by 0SourceScholar
2018

Semantic Mapping with Simultaneous Object Detection and Localization

IROS 2018poster

We present a filtering-based method for semantic mapping to simultaneously detect objects and localize their 6 degree-of-freedom pose. For our method, called Contextual Temporal Mapping (or CT-Map), we represent the semantic map as a belief over object classes and poses across an observed scene. Inf…

Cited by 39SourceScholar
2018

Semantic Robot Programming for Goal-Directed Manipulation in Cluttered Scenes

ICRA 2018poster

We present the Semantic Robot Programming (SRP) paradigm as a convergence of robot programming by demonstration and semantic mapping. In SRP, a user can directly program a robot manipulator by demonstrating a snapshot of their intended goal scene in workspace. The robot then parses this goal as a sc…

Cited by 55SourceScholar