← Search

Yi Zheng

17 accepted papers

2026

Decomposition of Concept-Level Rules in Visual Scenes

ICLR 2026poster

Human cognition is compositional, and one can parse a visual scene into independent concepts and the corresponding concept-changing rules. By contrast, many vision-language systems process images holistically, with limited support for explicit decomposition. And previous methods of decomposing conce…

Cited by 0SourceScholar
2026

Discovering Decoupled Functional Modules in Large Language Models

AAAI 2026technical

Understanding the internal functional organization of Large Language Models (LLMs) is crucial for improving their trustworthiness and performance. However, how LLMs organize different functions into modules remains highly unexplored. To bridge this gap, we formulate a function module discovery probl

Cited by 0SourcePDFScholar
2026

Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer

ICLR 2026oral

The probabilistic diffusion model (DM), generating content by inferencing through a recursive chain structure, has emerged as a powerful framework for visual generation. After pre-training on enormous data, the model needs to be properly aligned to meet requirements for downstream applications. How…

Cited by 0SourcecodeScholar
2026

LogART: Pushing the Limit of Efficient Logarithmic Post-Training Quantization

ICLR 2026poster

Efficient deployment of deep neural networks increasingly relies on Post-Training Quantization (PTQ). Logarithmic PTQ, in particular, promises multiplier-free hardware efficiency, but its performance is often limited by the nonlinear and symmetric quantization grid and standard rounding-to-nearest (…

Cited by 0SourcecodeScholar
2026

Representation-Aware Modularity: Efficient Cross-Task Generalization for LLMs

IJCAI 2026

Cross-task generalization (CTG) enables large language models (LLMs) to handle unseen tasks proficiently, enhancing their adaptability in real-world scenarios. However, existing methods relying on per-token dynamic routing to multiple trained LoRA adapters face high computational and GPU memory cost

Cited by 0Scholar
2025

FISTAPruner: Layer-wise Post-training Pruning for Large Language Models

EMNLP 2025

Pruning is a critical strategy for compressing trained large language models (LLMs), aiming at substantial memory conservation and computational acceleration without compromising performance. However, existing pruning methods typically necessitate inefficient retraining for billion-scale LLMs or rel

Cited by 0SourcePDFScholar
2024

EFTNAS: Searching for Efficient Language Models in First-Order Weight-Reordered Super-Networks

COLING 2024main

Transformer-based models have demonstrated outstanding performance in natural language processing (NLP) tasks and many other domains, e.g., computer vision. Depending on the size of these models, which have grown exponentially in the past few years, machine learning practitioners might be restricted…

2024

LoNAS: Elastic Low-Rank Adapters for Efficient Large Language Models

COLING 2024main

Large Language Models (LLMs) continue to grow, reaching hundreds of billions of parameters and making it challenging for Deep Learning practitioners with resource-constrained systems to use them, e.g., fine-tuning these models for a downstream task of their interest. Adapters, such as low-rank adapt…

2023

AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing

ACL 2023findings

Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP). In this work, we present AraMUS, the largest Arabic PLM with 11B parameters trained on 529GB of high-quality Arabic textual data. AraMUS ac…

2023

Learning-Based Distortion Compensation for a Hybrid Simulator of Space Docking

RA-L 2023

By effectively utilizing the fidelity of a physical simulation and the flexibility of a numerical simulation, the hybrid simulation is applicable to test the complicated docking contact process of various kinds of spacecraft. However, the hybrid simulation of space docking often has a divergence or

Cited by 5SourceScholar
2023

Video Captioning via Relation-Aware Graph Learning

ICASSP 2023accepted

Recent neural models for video captioning usually employed an encoder-decoder framework. However, most approaches either neglected the spatial and temporal interactions between objects in a video or implicitly modelled the interactions, resulting in less desired performance. In this paper, we propos…

Cited by 0SourceScholar
2022

Delving Deep into Regularity: A Simple but Effective Method for Chinese Named Entity Recognition

NAACL 2022findings

Recent years have witnessed the improving performance of Chinese Named Entity Recognition (NER) from proposing new frameworks or incorporating word lexicons. However, the inner composition of entity mentions in character-level Chinese NER has been rarely studied. Actually, most mentions of regular t…

Cited by 67SourcePDFScholar
2022

ReCo: Reliable Causal Chain Reasoning via Structural Causal Recurrent Neural Networks

EMNLP 2022main

Causal chain reasoning (CCR) is an essential ability for many decision-making AI systems, which requires the model to build reliable causal chains by connecting causal pairs. However, CCR suffers from two main transitive problems: threshold effect and scene drift. In other words, the causal pairs to…

2022

Recognition and Prediction of Surgical Gestures and Trajectories Using Transformer Models in Robot-Assisted Surgery

IROS 2022poster

Surgical activity recognition and prediction can help provide important context in many Robot-Assisted Surgery (RAS) applications, for example, surgical progress monitoring and estimation, surgical skill evaluation, and shared control strategies during teleoperation. Transformer models were first de…

Cited by 20SourceScholar
2021

Online Pseudo Label Generation by Hierarchical Cluster Dynamics for Adaptive Person Re-Identification

ICCV 2021poster

Adaptive person re-identification (adaptive ReID) targets at transferring learned knowledge from the labeled source domain to the unlabeled target domain. Pseudo-label-based methods that alternatively generate pseudo labels and optimize the training model have demonstrated great effectiveness in thi…

Cited by 119PDFScholar
2021

Scene Synthesis via Uncertainty-Driven Attribute Synchronization

ICCV 2021poster

Developing deep neural networks to generate 3D scenes is a fundamental problem in neural synthesis with immediate applications in architectural CAD, computer graphics, as well as in generating virtual robot training environments. This task is challenging because 3D scenes exhibit diverse patterns, r…

Cited by 39PDFcodeScholar
2017

No Spurious Local Minima in Nonconvex Low Rank Problems: A Unified Geometric Analysis

ICML 2017poster

In this paper we develop a new framework that captures the common landscape underlying the common non-convex low-rank matrix problems including matrix sensing, matrix completion and robust PCA. In particular, we show for all above problems (including asymmetric cases): 1) all local minima are also g…

Cited by 554SourcePDFScholar