← Search

Fei Ma

32 accepted papers

2026

A Principle-Driven Adaptive Policy for Group Cognitive Stimulation Dialogue for Elderly with Cognitive Impairment

AAAI 2026technical

Cognitive impairment is becoming a major public health challenge. Cognitive Stimulation Therapy (CST) is an effective intervention for cognitive impairment, but traditional methods are difficult to scale, and existing digital systems struggle with group dialogues and cognitive stimulation principles

Cited by 0SourcePDFScholar
2026

Cross Domain Test Time Scaling: Scale Knowledge and Reasoning on Cross Domains

IJCAI 2026

Test-time scaling (TTS) has demonstrated remarkable potential in enhancing the reasoning capabilities of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs). However, its application has primarily been limited to domains such as mathematics and programming, owing to their reasoning

Cited by 0Scholar
2026

D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies

AAAI 2026technical

Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial General Intelligence. While most existing datasets and benchmarks for training and evaluating GUI agents are static and id

Cited by 5SourcePDFScholar
2026

EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?

ICLR 2026poster

Descriptive Multimodal Emotion Recognition (DMER) has garnered increasing research attention. Unlike traditional discriminative paradigms that rely on predefined emotion taxonomies, DMER aims to describe human emotional state using free-form natural language, enabling finer-grained and more interpre…

Cited by 0SourcecodeScholar
2026

Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability

ICLR 2026poster

Temporal Video Grounding (TVG) aims to localize video segments corresponding to a given textual query, which often describes human actions. However, we observe that current methods, usually optimizing for high temporal Intersection-over-Union (IoU), frequently struggle to accurately recognize or und…

Cited by 0SourceScholar
2026

LottieGPT: Tokenizing Vector Animation for Autoregressive Generation

CVPR 2026

Despite rapid progress in video generation, existing models are incapable of producing vector animation, a dominant and highly expressive form of multimedia on the Internet. Vector animations offer resolution-independence, compactness, semantic structure, and editable parametric motion representatio

Cited by 0SourceScholar
2026

Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning

ICLR 2026poster

Vision token pruning has proven to be an effective acceleration technique for the Efficient Vision Language Model (VLM). However, existing pruning methods demonstrate excellent performance preservation in visual question answering (VQA) and suffer substantial degradation on visual grounding (VG) tas…

Cited by 0SourcecodeScholar
2026

SCNS: Continual Personalization of Diffusion Models via Submodular Concept Neuron Selection

ICML 2026poster

Custom diffusion models (CDMs) have demonstrated impressive success in visual personalization tasks by enabling the generation of user-specific concepts. However, existing CDMs typically assume that personalized concepts are static and rely on costly model merging or sequential updates that are pron…

Cited by 0SourceScholar
2026

Scalable Event Cloud Network for Event-based Classification

ICML 2026oral

Event cameras are biologically inspired sensors garnering significant attention from both industry and academia. Mainstream methods favor frame and voxel representations, which reach a satisfactory performance while introducing time-consuming transformations, bulky models, and sacrificing fine-grain…

Cited by 0SourceScholar
2026

T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding

CVPR 2026

Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dynamics. Existing Vision-LMMs typically perceive temporal dynamics via positional encoding, text-based timestamps, or visu

Cited by 0SourceScholar
2026

ViSA-Gait: Leveraging Vision Foundation Models for Semantic Anchored Gait Recognition

IJCAI 2026

Gait recognition has achieved remarkable success in constrained environments, yet its performance often degrades significantly in cross-domain and cross-vertical-view scenarios. This is primarily due to the fact that domain-specific silhouette geometry causes models to overfit to extrinsic geometric

Cited by 0Scholar
2025

Active Multimodal Distillation for Few-shot Action Recognition

IJCAI 2025

Owing to its rapid progress and broad application prospects, few-shot action recognition has attracted considerable interest. However, current methods are predominantly based on limited single-modal data, which does not fully exploit the potential of multimodal information. This paper presents a nov

Cited by 0SourcePDFScholar
2025

DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language Models

ICASSP 2025accepted

Inderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing,…

Cited by 0SourceScholar
2025

GaussianPU: Color Point Cloud Upsampling via 3D Gaussian Splatting

IROS 2025

Dense colored point clouds enhance visual perception and are of significant value in various robotic applications. However, existing learning-based point cloud upsampling methods are constrained by computational resources and batch processing strategies, which often require subdividing point clouds

Cited by 1SourceScholar
2025

Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction

IJCAI 2025

Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive obje

2025

Observation-Graph Interaction and Key-Detail Guidance for Vision and Language Navigation

IROS 2025

Vision and Language Navigation (VLN) requires an agent to navigate through environments following natural language instructions. However, existing methods often struggle with effectively integrating visual observations and instruction details during navigation, leading to suboptimal path planning an

Cited by 2SourceScholar
2025

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

AAAI 2025technical

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and identity-consistent talking heads from just a few minutes of tr…

Cited by 4SourcePDFScholar
2025

ReMask-Animate: Refined Character Image Animation Using Mask-Guided Adapters

AAAI 2025technical

Pose-controlled human video generation is of significant interest and finds extensive applications in areas such as automated advertising and content creation on social media platforms. While existing methods employing pose sequences and reference images for human image animation have exhibited nota…

Cited by 0SourcePDFScholar
2025

RoMa: A Robust Model Watermarking Scheme for Protecting IP in Diffusion Models

NeurIPS 2025poster

Preserving intellectual property (IP) within a pre-trained diffusion model is critical for protecting the model's copyright and preventing unauthorized model deployment. In this regard, model watermarking is a common practice for IP protection that embeds traceable information within models and allo…

Cited by 0SourcecodeScholar
2025

Subgraph Invariant Learning Towards Large-Scale Graph Node Classification

AAAI 2025technical

Graph Neural Networks (GNNs) have shown efficacy in graph node classification, but face computational challenges on large-scale graphs. Although existing graph reduction methods address these issues, they still require high computational resources and fail to prioritize robust performance on out-of-…

2025

Universal Visuo-Tactile Video Understanding for Embodied Interaction

NeurIPS 2025poster

Tactile perception is essential for embodied agents to understand the physical attributes of objects that cannot be determined through visual inspection alone. While existing methods have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate…

Cited by 0SourceScholar
2025

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism

ACL 2025long

Large Vision-Language Models (LVLMs) have shown exceptional performance in multimodal tasks, but their effectiveness in complex visual reasoning is still constrained, especially when employing Chain-of-Thought prompting techniques. In this paper, we propose VReST, a novel training-free approach that…

2025

VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweening

IJCAI 2025

We propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar

Cited by 0SourcePDFScholar
2025

VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models

COLING 2025main

Visual Language Models (VLMs) have rapidly progressed with the recent success of large language models. However, there have been few attempts to incorporate efficient linear Recurrent Neural Networks (RNNs) architectures into VLMs. In this study, we introduce VisualRWKV, the first application of a l…

2024

A Language-Driven Navigation Strategy Integrating Semantic Maps and Large Language Models

IROS 2024poster

Accurate perception of semantic and spatial information is crucial for robots performing language-driven navigation tasks. Existing approaches utilize visual-language models to extract semantic information from the environment and construct maps. However, constrained by the generalization and accura…

Cited by 0SourceScholar
2024

An Active Noise Control System Based On Soundfield Interpolation Using A Physics-Informed Neural Network

ICASSP 2024accepted

Conventional multiple-point active noise control (ANC) systems require placing error microphones within the region of interest (ROI), inconveniencing users. This paper designs a feasible monitoring microphone arrangement placed outside the ROI, providing a user with more freedom of movement. The sou…

Cited by 0SourceScholar
2024

Image Augmentation with Controlled Diffusion for Weakly-Supervised Semantic Segmentation

ICASSP 2024accepted

Weakly-supervised semantic segmentation (WSSS), which aims to train segmentation models solely using image-level labels, has achieved significant attention. Existing methods primarily focus on generating high-quality pseudo labels using available images and their image-level labels. However, the qua…

Cited by 0SourceScholar
2023

Spherical Sector Harmonics Based Soundfield Radial Extrapolation And Robustness Analysis

ICASSP 2023accepted

The development of spherical sector harmonics benefits the sound- field decomposition and analysis over a spherical sector region. However, research on the soundfield radial extrapolation from one spherical sector region to another concentric sector region with a different radius is still insufficie…

Cited by 0SourceScholar
2021

Pairwise Half-graph Discrimination: A Simple Graph-level Self-supervised Strategy for Pre-training Graph Neural Networks

IJCAI 2021poster

Self-supervised learning has gradually emerged as a powerful technique for graph representation learning. However, transferable, generalizable, and robust representation learning on graph data still remains a challenge for pre-training graph neural networks. In this paper, we propose a simple and ef…

Cited by 21SourcePDFScholar
2021

Semi-Supervised Multimodal Image Translation for Missing Modality Imputation

ICASSP 2021accepted

Missing data is a common problem in multimodal and multi-view learning. It raises a critical challenge for most multimodal algorithms, which are unable to deal with incomplete datasets. Rather than discarding entries with missing modalities, this paper aims to reconstruct the complete image-based mu…

Cited by 0SourceScholar
2018

Reference Signal Generation for Broadband ANC Systems in Reverberant Rooms

ICASSP 2018accepted

One major issue of implementing broadband active noise control systems in reverberant rooms is the lack of reference signals. In this work, by exploiting the spatial sound field characteristics, a time-domain sound field separation method is developed to generate the reference signal for broadband a…

Cited by 0SourceScholar