← Search

Feng Jiang

51 accepted papers

2026

A Semi-Active Occupational Shoulder Exoskeleton for Overhead Work With Free Mode and Personalized Assistive Torque

RA-L 2026

Current passive or semi-active shoulder exoskeletons for overhead work provide fixed assistive torque for all participants and tasks, which lacks adaptability. In addition, due to the need to store energy at low elevation angles, they may increase physical demand on the user when assistance is not r

Cited by 0SourceScholar
2026

A Semi-Active Occupational Shoulder Exoskeleton for Overhead Work with Free Mode and Personalized Assistive Torque

ICRA 2026poster

Current passive or semi-active shoulder exoskeletons for overhead work provide fixed assistive torque for all participants and tasks, which lacks adaptability. In addition, due to the need to store energy at low elevation angles, they may increase physical demand on the user when assistance is not r…

Cited by 0SourceScholar
2026

BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaboration

CVPR 2026

Many everyday objects are difficult to directly grasp (e.g., a flat iPad) or manipulate functionally (e.g., opening the cap of a pen lying on a desk). Such tasks require sequential, asymmetric coordination between two arms, where one arm performs preparatory manipulation that enables the other's goa

Cited by 0SourceScholar
2026

CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical Generation

AAAI 2026technical

Theme detection is a fundamental task in user-centric dialogue systems, aiming to identify the latent topic of each utterance without relying on predefined schemas. Unlike intent induction, which operates within fixed label spaces, theme detection requires cross-dialogue consistency and alignment wi

Cited by 0SourcePDFScholar
2026

GRAM-DTI: Adaptive Multimodal Representation Learning for Drug–Target Interaction Prediction

ICLR 2026poster

Drug target interaction (DTI) prediction is a cornerstone of computational drug discovery, enabling rational design, repurposing, and mechanistic insights. While deep learning has advanced DTI modeling, existing approaches primarily rely on SMILES–protein pairs and fail to exploit the rich multimoda…

Cited by 0SourceScholar
2026

Hyperbolic Gramian Volumes for Multimodal Alignment

CVPR 2026

Multimodal contrastive learning typically relies on pairwise similarities for alignment, but recent work has shown that Gramian volumes can capture higher-order correlations across modalities. However, Euclidean Gramian volumes suffer from volume collapse under L2 normalization, concentrating near u

Cited by 0SourceScholar
2026

IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models

ICML 2026poster

Generative Reward Models (GRMs) have demonstrated strong performance in reward modeling, due to their interpretability and potential for refinement through reinforcement learning (RL). However, widely used pairwise GRMs create a computational bottleneck in reinforcement learning from human feedback …

Cited by 0SourceScholar
2026

Learning from Guidelines: Structured Prompt Optimization for Expert Annotation Tasks

AAAI 2026technical

Deep learning has significantly advanced numerous fields by training on extensive annotated datasets. However, this data-driven paradigm faces limitations such as limited adaptability and high annotation costs, particularly when precise adherence to detailed, domain-specific guidelines is required i

Cited by 0SourcePDFScholar
2026

PHOTONS: Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views

AAAI 2026technical

We present PHOTONS (Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views), a real-time framework for novel view synthesis without requiring camera calibration. Our method reconstructs consistent 3D Gaussian point clouds and synthesizes 2K photo-realistic novel vie

Cited by 0SourcePDFScholar
2026

SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation

ICLR 2026poster

Large language models (LLMs) are increasingly applied in scientific research, offering new capabilities for knowledge discovery and reasoning. In single-cell biology, however, evaluation practices for both general and specialized LLMs remain inadequate: existing benchmarks are fragmented across task…

Cited by 0SourcecodeScholar
2026

Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent

CVPR 2026

Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven Image Clustering Agent, the first universal framework that bridges these gaps through textual guidelines. To incorporate complex guidelines without

Cited by 0SourceScholar
2025

A Ranking Scheme for Trust Region Multi-agent Reinforcement Learning

ICASSP 2025accepted

In multi-agent reinforcement learning (MARL), trust region (TR) methods are widely used because they effectively mitigate the nonstationarity of multi-agent systems and facilitate collaboration among diverse agent types. Based on the multi-agent advantage decomposition lemma, TR methods adopt a sequ…

Cited by 0SourceScholar
2025

Aligning Language Models Using Follow-up Likelihood as Reward Signal

AAAI 2025technical

In natural human-to-human conversations, participants often receive feedback signals from one another based on their follow-up reactions. These reactions can include verbal responses, facial expressions, changes in emotional state, and other non-verbal cues. Similarly, in human-machine interactions,…

2025

EGENN: An Efficient Graph-Enhanced Neural Network for Multivariate Time Series Forecasting

ICASSP 2025accepted

Graph Neural Network (GNN) has been widely applied in multivariate time series forecasting due to its excellent relationship modeling capabilities. However, current methods still face limitations in computational efficiency or time series expression capabilities. To address these issues, we propose…

Cited by 0SourceScholar
2025

FreqLLM: Frequency-Aware Large Language Models for Time Series Forecasting

IJCAI 2025

Large Language Models (LLMs) have recently shown promise in Time Series Forecasting (TSF) by effectively capturing intricate time-domain dependencies. However, our preliminary experiments reveal that standard LLM-based approaches often fail to capture global correlations, limiting predictive perform

2025

GoBERT: Gene Ontology Graph Informed BERT for Universal Gene Function Prediction

AAAI 2025technical

Exploring the functions of genes and gene products is crucial to a wide range of fields, including medical research, evolutionary biology, and environmental science. However, discovering new functions largely relies on expensive and exhaustive wet lab experiments. Existing methods of automatic funct…

Cited by 0SourcePDFScholar
2025

Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles

ACL 2025long

User simulators are crucial for replicating human interactions with dialogue systems, supporting both collaborative training and automatic evaluation, especially for large language models (LLMs). However, current role-playing methods face challenges such as a lack of utterance-level authenticity and…

2025

MRANet: An Encoder-Decoder Network with Multi-Scale Residual Atrous-Spatial Pyramid Pooling for Seismic Phase Picking

ICASSP 2025accepted

Seismic phase picking is one of the critical challenges in seismic data processing. With the advancement of deep learning, numerous neural network architectures have been employed to explore the correlations between seismic waveforms and the underlying information. However, existing methods predomin…

Cited by 0SourceScholar
2025

PPTP: Performance-Guided Physiological Signal-Based Trust Prediction in Sequential Human-Robot Collaboration

RA-L 2025

Trust prediction is a key issue in human-robot collaboration, especially in construction scenarios where maintaining appropriate trust calibration is critical for safety and efficiency. This paper introduces the Performance-guided Physiological signal-based Trust Prediction (PPTP), a novel framework

Cited by 0SourceScholar
2025

TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence

NeurIPS 2025spotlight

Molecular property prediction aims to learn representations that map chemical structures to functional properties. While multimodal learning has emerged as a powerful paradigm to learn molecular representations, prior works have largely overlooked textual and taxonomic information of molecules for r…

Cited by 0SourcecodeScholar
2025

Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models

NAACL 2025long

Data selection for fine-tuning large language models (LLMs) aims to choose a high-quality subset from existing datasets, allowing the trained model to outperform baselines trained on the full dataset. However, the expanding body of research lacks a clear, unified framework, and the variability in ex…

Cited by 5SourcePDFScholar
2025

Zero-Shot Composed Image Retrieval via Dual-Stream Instruction-Aware Distillation

ICCV 2025poster

Composed Image Retrieval (CIR) targets the retrieval of images conditioned on a reference image and a textual modification, but constructing labeled triplets (reference image, textual modification, target image) is inherently challenging. Existing Zero-Shot CIR (ZS-CIR) approaches often rely on well…

Cited by 0SourcePDFScholar
2024

Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark

COLING 2024main

Topic segmentation and outline generation strive to divide a document into coherent topic sections and generate corresponding subheadings, unveiling the discourse topic structure of a document. Compared with sentence-level topic structure, the paragraph-level topic structure can quickly grasp and un…

2024

CMB: A Comprehensive Medical Benchmark in Chinese

NAACL 2024long

Large Language Models (LLMs) provide a possibility to make a great breakthrough in medicine. The establishment of a standardized medical benchmark becomes a fundamental cornerstone to measure progression. However, medical environments in different regions have their local characteristics, e.g., the…

2024

Causal Subgraphs and Information Bottlenecks: Redefining OOD Robustness in Graph Neural Networks

ECCV 2024poster

"Graph Neural Networks (GNNs) are increasingly popular in processing graph-structured data, yet they face significant challenges when training and testing distributions diverge, common in real-world scenarios. This divergence often leads to substantial performance drops in GNN models. To address thi…

Cited by 0SourcePDFScholar
2024

Deep Ad-hoc Sub-Team Partition Learning for Multi-Agent Air Combat Cooperation

IROS 2024poster

In the future, unmanned autonomous air combat will encounter large-scale confrontation scenarios, where agents must consider complex time-varying relationships among aircraft when making decisions. Previous works have already introduced Multi-Agent Reinforcement Learning (MARL) into air combat and s…

Cited by 2SourceScholar
2024

Humans or LLMs as the Judge? A Study on Judgement Bias

EMNLP 2024main

Adopting human and large language models (LLM) as judges (*a.k.a* human- and LLM-as-a-judge) for evaluating the performance of LLMs has recently gained attention. Nonetheless, this approach concurrently introduces potential biases from human and LLMs, questioning the reliability of the evaluation re…

2024

PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User Simulator

ACL 2024long

The unparalleled performance of closed-sourced ChatGPT has sparked efforts towards its democratization, with notable strides made by leveraging real user and ChatGPT dialogues, as evidenced by Vicuna. However, due to challenges in gathering dialogues involving human participation, current endeavors…

Cited by 5SourcePDFScholar
2024

TS-Align: A Teacher-Student Collaborative Framework for Scalable Iterative Finetuning of Large Language Models

EMNLP 2024finding

Mainstream approaches to aligning large language models (LLMs) heavily rely on human preference data, particularly when models require periodic updates. The standard process for iterative alignment of LLMs involves collecting new human feedback for each update. However, the data collection process i…

2024

Uncovering the Potential of ChatGPT for Discourse Analysis in Dialogue: An Empirical Study

COLING 2024main

Large language models, like ChatGPT, have shown remarkable capability in many downstream tasks, yet their ability to understand discourse structures of dialogues remains less explored, where it requires higher level capabilities of understanding and reasoning. In this paper, we aim to systematically…

2023

Aprogressive Image Dehazing Framework with inter and Intra Contrastive Learning

ICASSP 2023accepted

Image dehazing, aims to estimate latent haze-free images from hazy images, suffering from a lot of lost information. Existing contrastive learning methods tend to utilize hazefree images as positive samples without consideration of negative samples. Even if negative samples are employed, the connect…

Cited by 0SourceScholar
2023

EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text Detection

ICASSP 2023accepted

Text detection in natural scenarios, has made significant progress with the deep learning architecture. Towards arbitrary-shaped text detection, fracture detection is the major concern due to the lack of semantic relationship within an instance in existing methods. To circumvent this dilemma, we pro…

Cited by 0SourceScholar
2023

Factual Relation Discrimination for Factuality-oriented Abstractive Summarization

EMNLP 2023long findings

Most neural abstractive summarization models are capable of producing high-quality summaries. However, they still frequently contain factual errors. Existing factuality-oriented abstractive summarization models only consider the integration of factual information and ignore the causes of factual err…

Cited by 0SourceScholar
2023

Hierarchical Interactive Reconstruction Network for Video Compressive Sensing

ICASSP 2023accepted

Deep network-based image and video Compressive Sensing (CS) has attracted increasing attentions in recent years. However, in the existing deep network-based CS methods, a simple stacked convolutional network is usually adopted, which not only weakens the perception of rich contextual prior knowledge…

Cited by 0SourceScholar
2023

HuatuoGPT, Towards Taming Language Model to Be a Doctor

EMNLP 2023long findings

In this paper, we present HuatuoGPT, a Large Language Model (LLM) for medical consultation. The core recipe of HuatuoGPT is to leverage both distilled data from **ChatGPT** and real-world data from **doctors** in the supervised fine-tuning stage. This is not only because purely using **ChatGPT**-di…

Cited by 0SourcecodeScholar
2023

Improving Dialogue Discourse Parsing via Reply-to Structures of Addressee Recognition

EMNLP 2023long main

Dialogue discourse parsing aims to reflect the relation-based structure of dialogue by establishing discourse links according to discourse relations. To alleviate data sparsity, previous studies have adopted multitasking approaches to jointly learn dialogue discourse parsing with related tasks (e.g.…

Cited by 0SourcecodeScholar
2023

Multi-to-Single Knowledge Distillation for Point Cloud Semantic Segmentation

ICRA 2023poster

3D point cloud semantic segmentation is one of the fundamental tasks for environmental understanding. Although significant progress has been made in recent years, the performance of classes with few examples or few points is still far from satisfactory. In this paper, we propose a novel multi-to-sin…

Cited by 6SourcecodeScholar
2022

Key Mention Pairs Guided Document-Level Relation Extraction

COLING 2022main

Document-level Relation Extraction (DocRE) aims at extracting relations between entities in a given document. Since different mention pairs may express different relations or even no relation, it is crucial to identify key mention pairs responsible for the entity-level relation labels. However, most…

2021

A Bipolar Myoelectric Sensor-Enabled Human-Machine Interface Based On Spinal Module Activations

ICRA 2021poster

The surface electromyography (sEMG) signal-based human-machine interface (HMI) has been widely used for various scenarios of physical human-robot interaction. However, current HMIs based on bipolar myoelectric sensors are hindered by the limitations of global sEMG features, which are prone to variab…

Cited by 3SourceScholar
2021

Hierarchical Macro Discourse Parsing Based on Topic Segmentation

AAAI 2021technical

Hierarchically constructing micro (i.e., intra-sentence or inter-sentence) discourse structure trees using explicit boundaries (e.g., sentence and paragraph boundaries) has been proved to be an effective strategy. However, it is difficult to apply this strategy to document-level macro (i.e., inter-p…

2021

Not Just Classification: Recognizing Implicit Discourse Relation on Joint Modeling of Classification and Generation

EMNLP 2021main

Implicit discourse relation recognition (IDRR) is a critical task in discourse analysis. Previous studies only regard it as a classification task and lack an in-depth understanding of the semantics of different relations. Therefore, we first view IDRR as a generation task and further propose a metho…

2021

Residual Relaxation for Multi-view Representation Learning

NeurIPS 2021poster

Multi-view methods learn representations by aligning multiple views of the same image and their performance largely depends on the choice of data augmentation. In this paper, we notice that some other useful augmentations, such as image rotation, are harmful for multi-view methods because they cause…

Cited by 40SourcePDFScholar
2020

Chinese Paragraph-level Discourse Parsing with Global Backward and Local Reverse Reading

COLING 2020main

Discourse structure tree construction is the fundamental task of discourse parsing and most previous work focused on English. Due to the cultural and linguistic differences, existing successful methods on English discourse parsing cannot be transformed into Chinese directly, especially in paragraph…

Cited by 8SourcePDFScholar
2020

Classify and Explain: An Interpretable Convolutional Neural Network For Lung Cancer Diagnosis

ICASSP 2020accepted

The deep network-based computer-aided diagnosis systems have encountered many difficulties in practical applications because of its "black box" feature. The crux of the problem is that these models should be explainable - the model should provide doctors rationales that can explain the diagnosis. In…

Cited by 0SourceScholar
2020

Multi-Stage Residual Hiding for Image-Into-Audio Steganography

ICASSP 2020accepted

The widespread application of audio communication technologies has speeded up audio data flowing across the Internet, which made it a popular carrier for covert communication. In this paper, we present a cross-modal steganography method for hiding image content into audio carriers while preserving t…

Cited by 0SourceScholar
2019

Scalable Convolutional Neural Network for Image Compressed Sensing

CVPR 2019poster

Recently, deep learning based image Compressed Sensing (CS) methods have been proposed and demonstrated superior reconstruction quality with low computational complexity. However, the existing deep learning based image CS methods need to train different models for different sampling ratios, which in…

Cited by 195PDFcodeScholar
2018

An Efficient Deep Convolutional Laplacian Pyramid Architecture for Cs Reconstruction At Low Sampling Ratios

ICASSP 2018accepted

The compressed sensing (CS) has been successfully applied to image compression in the past few years as most image signals are sparse in a certain domain. Several CS reconstruction models have been proposed and obtained superior performance. However, these methods suffer from blocking artifacts or r…

Cited by 0SourceScholar