← Search

Zixun Sun

6 accepted papers

2026

Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token Selection

CVPR 2026

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet processing long visual token sequences remains computationally expensive. Existing approaches mitigate this cost by reducing image tokens, either by discarding them after the visual encod

Cited by 0SourcecodeScholar
2025

AdaV: Adaptive Text-visual Redirection for Vision-Language Models

ACL 2025finding

The success of Vision-Language Models (VLMs) often relies on high-resolution schemes that preserve image details, while these approaches also generate an excess of visual tokens, leading to a substantial decrease in model efficiency. A typical VLM includes a visual encoder, a text encoder, and an LL…

2024

Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation

ICML 2024poster

Large-scale pretrained models have proven immensely valuable in handling data-intensive modalities like text and image. However, fine-tuning these models for certain specialized modalities, such as protein sequence and cosmic ray, poses challenges due to the significant modality discrepancy and scar…

Cited by 4SourcePDFScholar
2024

RE-SORT: Removing Spurious Correlation in Multilevel Interaction for CTR Prediction

UAI 2024poster

Click-through rate (CTR) prediction is a critical task in recommendation systems, serving as the ultimate filtering step to sort items for a user. Most recent cutting-edge methods primarily focus on investigating complex implicit and explicit feature interactions; however, these methods neglect the…

2024

Weight Diffusion for Future: Learn to Generalize in Non-Stationary Environments

NeurIPS 2024poster

Enabling deep models to generalize in non-stationary environments is vital for real-world machine learning, as data distributions are often found to continually change. Recently, evolving domain generalization (EDG) has emerged to tackle the domain generalization in a time-varying system, where the…

Cited by 0SourcePDFScholar
2021

Discovering Interpretable Latent Space Directions of GANs Beyond Binary Attributes

CVPR 2021poster

Generative adversarial networks (GANs) learn to map noise latent vectors to high-fidelity image outputs. It is found that the input latent space shows semantic correlations with the output image space. Recent works aim to interpret the latent space and discover meaningful directions that correspond…

Cited by 64PDFcodeScholar