← Search

Feihong He

5 accepted papers

2026

Efficient Bilevel Optimization for CKA-Guided MoE Upcycling

ICML 2026poster

Upcycling, a strategy that initializes Mixture-of-Experts (MoE) by replicating pre-trained feed-forward or MoE networks to expand model capacity, has become a popular method in continual learning due to its effectiveness in mitigating catastrophic forgetting. However, existing paradigms rely on indi…

Cited by 0SourceScholar
2026

Plasticity Activation via Polar Operator: A Plug-in Method for Balancing Stability and Plasticity

ICML 2026poster

Continual learning (CL) seeks models that acquire new knowledge while avoiding catastrophic forgetting. However, many methods that mitigate forgetting constrain parameter updates and thereby reduce model plasticity. We revisit the singular value spectrum of gradients in representative CL methods and…

Cited by 0SourceScholar
2025

ExAct: A Video-Language Benchmark for Expert Action Analysis

NeurIPS 2025poster

We present ExAct, a new video-language benchmark for expert-level understanding of skilled physical human activities. Our new benchmark contains 3,521 expert-curated video question-answer pairs spanning 11 physical activities in 6 domains: Sports, Bike Repair, Cooking, Health, Music, and Dance. ExAc…

Cited by 0SourcecodeScholar
2024

CartoonDiff: Training-free Cartoon Image Generation with Diffusion Transformer Models

ICASSP 2024accepted

Image cartoonization has attracted significant interest in the field of image generation. However, most of the existing image cartoonization techniques require re-training models using images of cartoon style. In this paper, we present CartoonDiff, a novel training-free sampling approach which gener…

Cited by 0SourceScholar
2024

Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

ACL 2024long

Multimodal Large Language Models (MLLMs) have demonstrated proficiency in handling a variety of visual-language tasks. However, current MLLM benchmarks are predominantly designed to evaluate reasoning based on static information about a single image, and the ability of modern MLLMs to extrapolate fr…