← Search

Yizhi Wang

19 accepted papers

2026

BIOARC: Discovering Optimal Neural Architectures for Biological Foundation Models

ICML 2026poster

Foundation models have revolutionized AI, yet biological applications often repurpose general architectures without accounting for the intrinsic structural and functional properties of distinct modalities, such as genomic and proteomic sequences. Consequently, these architectures lack the inductive …

Cited by 0SourceScholar
2026

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

ICLR 2026poster

We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference s…

Cited by 0SourcecodeScholar
2026

Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation

AAAI 2026technical

Layer pruning is a viable technique for compressing large language models while achieving acceleration proportional to the pruning ratio. In this work, we identify that removing any layer induces a magnitude gap in hidden states, and demonstrate that a simple compensation operation leads to superior

Cited by 0SourcePDFScholar
2026

VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization

CVPR 2026

Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired data of simple editing operations, which fundamentally limits t

Cited by 0SourcecodeScholar
2025

From Characters to Subwords: Modeling Unit Conversion for Low-resource Speech Recognition

ICASSP 2025accepted

Multilingual automatic speech recognition (ASR) models greatly facilitate recognizing low-resource languages by sharing representations across similar languages. However, the commonly adopted modeling units, e.g., character-level modeling, lack language-specific information, resulting in a susceptib…

Cited by 0SourceScholar
2025

GALA: Geometry-Aware Local Adaptive Grids for Detailed 3D Generation

ICLR 2025poster

We propose GALA, a novel representation of 3D shapes that (i) excels at capturing and reproducing complex geometry and surface details, (ii) is computationally efficient, and (iii) lends itself to 3D generative modelling with modern, diffusion-based schemes. The key idea of GALA is to exploit both t…

2025

Hierarchical Decision-Making for Autonomous Navigation: Integrating Deep Reinforcement Learning and Fuzzy Logic in Four-Wheel Independent Steering and Driving Systems

IROS 2025

This paper presents a hierarchical decision-making framework for autonomous navigation in four-wheel independent steering and driving (4WISD) systems. The proposed approach integrates deep reinforcement learning (DRL) for high-level navigation with fuzzy logic for low-level control to ensure both ta

Cited by 2SourceScholar
2025

Partial Label Causal Representation Learning for Instance-Dependent Supervision and Domain Generalization

AAAI 2025technical

Partial label learning (PLL) addresses situations where each training example is associated with a set of candidate labels, among which only one corresponds to the true class label. As the candidate labels often come from crowdsourced workers, their generation is inherently dependent on the features…

Cited by 0SourcePDFScholar
2024

HIQ: One-Shot Network Quantization for Histopathological Image Classification

ICASSP 2024accepted

To deploy neural networks on clinical edge devices, quantization is the most commonly used method to compress the models, which requires a calibration set of hundreds of real images. However, due to privacy concerns, the scarcity of private histopathological images hinders the application of quantiz…

Cited by 0SourceScholar
2024

Slice3D: Multi-Slice Occlusion-Revealing Single View 3D Reconstruction

CVPR 2024poster

We introduce multi-slice reasoning a new notion for single-view 3D reconstruction which challenges the current and prevailing belief that multi-view synthesis is the most natural conduit between single-view and 3D. Our key observation is that object slicing is a more direct and hence more advantageo…

Cited by 7SourcePDFScholar
2024

SweepNet: Unsupervised Learning Shape Abstraction via Neural Sweepers

ECCV 2024poster

"Shape abstraction is an important task for simplifying complex geometric structures while retaining essential features. Sweep surfaces, commonly found in human-made objects, aid in this process by effectively capturing and representing object geometry, thereby facilitating abstraction. In this pape…

Cited by 0SourcePDFScholar
2023

ARO-Net: Learning Implicit Fields From Anchored Radial Observations

CVPR 2023poster

We introduce anchored radial observations (ARO), a novel shape encoding for learning implicit field representation of 3D shapes that is category-agnostic and generalizable amid significant shape variations. The main idea behind our work is to reason about shapes through partial observations from a s…

2023

DS-Fusion: Artistic Typography via Discriminated and Stylized Diffusion

ICCV 2023poster

We introduce a novel method to automatically generate an artistic typography by stylizing one or more letter fonts to visually convey the semantics of an input word, while ensuring that the output remains readable. To address an assortment of challenges with our task at hand including conflicting go…

Cited by 25PDFcodeScholar
2023

DeepVecFont-v2: Exploiting Transformers To Synthesize Vector Fonts With Higher Quality

CVPR 2023poster

Vector font synthesis is a challenging and ongoing problem in the fields of Computer Vision and Computer Graphics. The recently-proposed DeepVecFont achieved state-of-the-art performance by exploiting information of both the image and sequence modalities of vector fonts. However, it has limited capa…

2023

TexQ: Zero-shot Network Quantization with Texture Feature Distribution Calibration

NeurIPS 2023poster

Quantization is an effective way to compress neural networks. By reducing the bit width of the parameters, the processing efficiency of neural network models at edge devices can be notably improved. Most conventional quantization methods utilize real datasets to optimize quantization parameters and…

2022

Aesthetic Text Logo Synthesis via Content-Aware Layout Inferring

CVPR 2022poster

Text logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, few attention has been paid to this task which needs to take many factors (e.g., fonts, linguistics, topics, etc.) into cons…

Cited by 33PDFcodeScholar
2022

BILCO: An Efficient Algorithm for Joint Alignment of Time Series

NeurIPS 2022accept

Multiple time series data occur in many real applications and the alignment among them is usually a fundamental step of data analysis. Frequently, these multiple time series are inter-dependent, which provides extra information for the alignment task and this information cannot be well utilized in t…

2019

muSSP: Efficient Min-cost Flow Algorithm for Multi-object Tracking

NeurIPS 2019poster

Min-cost flow has been a widely used paradigm for solving data association problems in multi-object tracking (MOT). However, most existing methods of solving min-cost flow problems in MOT are either direct adoption or slight modifications of generic min-cost flow algorithms, yielding sub-optimal com…

2016

Graphical Time Warping for Joint Alignment of Multiple Curves

NeurIPS 2016poster

Dynamic time warping (DTW) is a fundamental technique in time series analysis for comparing one curve to another using a flexible time-warping function. However, it was designed to compare a single pair of curves. In many applications, such as in metabolomics and image series analysis, alignment is…

Cited by 16SourcePDFScholar