← Search

Hefeng Wu

12 accepted papers

2026

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

CVPR 2026

Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks.However, real-world robotic tasks often involve long-horizon, multi-step problem-solving and require generalization for continual skill acquisition, extending beyond single actions or

Cited by 0SourceScholar
2026

Massive Editing for Large Language Models Based on Dynamic Weight Generation

ICLR 2026poster

Knowledge Editing (KE) is a field that studies how to modify some knowledge in Large Language Models (LLMs) at a low cost (compared to pre-training). Currently, performing large-scale edits on LLMs while ensuring the Reliability, Generality, and Locality metrics of the edits remain a challenge. This…

Cited by 0SourceScholar
2025

MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models

ACL 2025long

Long Context Understanding (LCU) is a critical area for exploration in current large language models (LLMs). However, due to the inherently lengthy nature of long-text data, existing LCU benchmarks for LLMs often result in prohibitively high evaluation costs, like testing time and inference expenses…

2025

RoboPearls: Editable Video Simulation for Robot Manipulation

ICCV 2025poster

The development of generalist robot manipulation policies has seen significant progress, driven by large-scale demonstration data across diverse environments. However, the high cost and inefficiency of collecting real-world demonstrations hinder the scalability of data acquisition. While existing si…

Cited by 0SourcePDFScholar
2025

Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention

NeurIPS 2025poster

Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for understanding egocentric human behavior. However, achieving such segmentation rob…

Cited by 0SourceScholar
2025

RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

EMNLP 2025

Routing large language models (LLMs) is a new paradigm that uses a router to recommend the best LLM from a pool of candidates for a given input. In this paper, our comprehensive analysis with more than 8,500 LLMs reveals a novel model-level scaling up phenomenon in Routing LLMs, i.e., a capable rout

2022

Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels

AAAI 2022technical

Training the multi-label image recognition models with partial labels, in which merely some labels are known while others are unknown for each image, is a considerably challenging and practical task. To address this task, current algorithms mainly depend on pre-training classification or similarity…

2022

Structured Semantic Transfer for Multi-Label Recognition with Partial Labels

AAAI 2022technical

Multi-label image recognition is a fundamental yet practical task because real-world images inherently possess multiple semantic labels. However, it is difficult to collect large-scale multi-label annotations due to the complexity of both the input images and output label spaces. To reduce the annot…

2021

AU-Expression Knowledge Constrained Representation Learning for Facial Expression Recognition

ICRA 2021poster

Recognizing human emotion/expressions automatically is quite an expected ability for intelligent robotics, as it can promote better communication and cooperation with humans. Current deep-learning-based algorithms may achieve impressive performance in some lab-controlled environments, but they alway…

Cited by 28SourcecodeScholar
2021

Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting

CVPR 2021poster

Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover potential pedestrians in unconstrained scenarios. In this work, we f…

Cited by 165PDFcodeScholar
2019

ADCrowdNet: An Attention-Injective Deformable Convolutional Network for Crowd Understanding

CVPR 2019poster

We propose an attention-injective deformable convolutional network called ADCrowdNet for crowd understanding that can address the accuracy degradation problem of highly congested noisy scenes. ADCrowdNet contains two concatenated networks. An attention-aware network called Attention Map Generator (A…

Cited by 355PDFScholar
2019

Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition

ICCV 2019poster

Recognizing multiple labels of images is a practical and challenging task, and significant progress has been made by searching semantic-aware regions and modeling label dependency. However, current methods cannot locate the semantic regions accurately due to the lack of part-level supervision or sem…

Cited by 384PDFcodeScholar