← Search

Jingjing Wang

38 accepted papers

2026

$\textit{S}$-SPPO: Semantic-Calibrated Self-Play Preference Optimization

ICML 2026poster

Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley-Terry instantiation of DPO is limited in modeling common departures from transitivity in human preferences. To address this, recent work has introd…

Cited by 0SourceScholar
2026

LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection

IJCAI 2026

Human experts often struggle to select the best option from a large set of items with multiple competing objectives, a process bottlenecked by the difficulty of formalizing complex, implicit preferences. To address this, we introduce \textbf{LISTEN} (\textbf{L}LM-based \textbf{I}terative \textbf{S}e

Cited by 0Scholar
2026

Learning Fingerprints for Medical Time Series with Redundancy-Constrained Information Maximization

ICML 2026poster

Learning meaningful representations from medical time series (MedTS), such as ECG or EEG signals, is a critical challenge. These signals are often high-dimensional, variable-length, and rife with noise. Existing self-supervised approaches, such as Masked Autoencoders (MAEs), are highly effective for…

Cited by 0SourceScholar
2026

Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios

AAAI 2026technical

In the literature, existing human-centric emotional motion generation methods primarily focus on boosting performance within a single scale-fixed dataset, largely neglecting the flexible and scale-increasing motion scenarios (e.g., sports, dance), whereas effectively learning these newly emerging sc

Cited by 0SourcePDFScholar
2025

Compliance Control with Dynamic and Self-Sensing Hydraulic Artificial Muscles for Wearable Assistive Devices

ICRA 2025

While wearable robots that utilize intrinsically soft materials for actuation offer enhanced safety and biological compatibility, the challenges of sensing and control significantly affect their performance. The control problem in such systems is inherently complex, and the inclusion of 'softness' i

Cited by 1SourceScholar
2025

Denoising Diffusion Models are Good General Gaze Feature Learners

IJCAI 2025

Since the collection of labeled gaze data is laborious and time-consuming, methods which can learn generalizable features by leveraging large-scale available unlabeled data are desirable. In recent years, we have witnessed the tremendous capabilities of diffusion models in generating images as well

Cited by 0SourcePDFScholar
2025

DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation

ICCV 2025poster

Spatio-temporal consistency is a critical topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily foc…

2025

Gaze Label Alignment: Alleviating Domain Shift for Gaze Estimation

AAAI 2025technical

Gaze estimation methods encounter significant performance deterioration when being evaluated across different domains, because of the domain gap between the testing and training data. Existing methods try to solve this issue by reducing the deviation of data distribution, however, they ignore the ex…

Cited by 1SourcePDFScholar
2025

KD-RIEKF: Kinodynamic Right-Invariant EKF for Legged Robot State Estimation

IROS 2025

We present KD-RIEKF, a novel state estimation framework that incorporates kinodynamic constraints into the Right-Invariant Extended Kalman Filter (RIEKF). Our framework integrates generalized momentum-based contact estimation, centroidal dynamics, and a noise-adaptive module, improving state estimat

Cited by 0SourceScholar
2025

LightCity: An Urban Dataset for Outdoor Inverse Rendering and Reconstruction under Multi-illumination Conditions

ICCV 2025poster

We propose an outdoor scene dataset and propose a series of benchmarks based on it.Inverse rendering in urban scenes is pivotal for applications like autonomous driving and digital twins, yet it faces significant challenges due to complex illumination conditions, including multi-illumination and ind…

Cited by 0SourcePDFScholar
2025

SQLFixAgent: Towards Semantic-Accurate Text-to-SQL Parsing via Consistency-Enhanced Multi-Agent Collaboration

AAAI 2025technical

While fine-tuned large language models (LLMs) excel in generating grammatically valid SQL in Text-to-SQL parsing, they often struggle to ensure semantic accuracy in queries, leading to user confusion and diminished system usability. To tackle this challenge, we introduce SQLFixAgent, a new consisten…

2025

Table-Critic: A Multi-Agent Framework for Collaborative Criticism and Refinement in Table Reasoning

ACL 2025long

Despite the remarkable capabilities of large language models (LLMs) in various reasoning tasks, they still struggle with table reasoning tasks, particularly in maintaining consistency throughout multi-step reasoning processes. While existing approaches have explored various decomposition strategies,…

2024

Arbitrary-Scale Point Cloud Upsampling by Voxel-Based Network with Latent Geometric-Consistent Learning

AAAI 2024technical

Recently, arbitrary-scale point cloud upsampling mechanism became increasingly popular due to its efficiency and convenience for practical applications. To achieve this, most previous approaches formulate it as a problem of surface approximation and employ point-based networks to learn surface repre…

2024

CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic Model

AAAI 2024technical

Gaze estimation methods often experience significant performance degradation when evaluated across different domains, due to the domain gap between the testing and training data. Existing methods try to address this issue using various domain generalization approaches, but with little success becaus…

Cited by 13SourcePDFScholar
2024

ChatASU: Evoking LLM’s Reflexion to Truly Understand Aspect Sentiment in Dialogues

COLING 2024main

Aspect Sentiment Understanding (ASU) in interactive scenarios (e.g., Question-Answering and Dialogue) has attracted ever-more interest in recent years and achieved important progresses. However, existing studies on interactive ASU largely ignore the coreference issue for opinion targets (i.e., aspec…

2024

How to Understand “Support”? An Implicit-enhanced Causal Inference Approach for Weakly-supervised Phrase Grounding

COLING 2024main

Weakly-supervised Phrase Grounding (WPG) is an emerging task of inferring the fine-grained phrase-region matching, while merely leveraging the coarse-grained sentence-image pairs for training. However, existing studies on WPG largely ignore the implicit phrase-region matching relations, which are cr…

2024

LG-Gaze: Learning Geometry-aware Continuous Prompts for Language-Guided Gaze Estimation

ECCV 2024poster

"The ability of gaze estimation models to generalize is often significantly hindered by various factors unrelated to gaze, especially when the training dataset is limited. Current strategies aim to address this challenge through different domain generalization techniques, yet they have had limited s…

Cited by 5SourcePDFScholar
2024

NERF-GAZE: A Head-Eye Redirection Parametric Model for Gaze Estimation

ICASSP 2024accepted

Gaze estimation is a fundamental aspect of many visual tasks. However, the high cost of acquiring gaze datasets with 3D annotations hinders the optimization and application of gaze estimation models. In this work, we propose a novel Head-Eye redirection parametric model based on Neural Radiance Fiel…

Cited by 0SourceScholar
2024

Suitable is the Best: Task-Oriented Knowledge Fusion in Vulnerability Detection

NeurIPS 2024poster

Deep learning technologies have demonstrated remarkable performance in vulnerability detection. Existing works primarily adopt a uniform and consistent feature learning pattern across the entire target set. While designed for general-purpose detection tasks, they lack sensitivity towards target code…

Cited by 0SourcePDFScholar
2024

TopicDiff: A Topic-enriched Diffusion Approach for Multimodal Conversational Emotion Detection

COLING 2024main

Multimodal Conversational Emotion (MCE) detection, generally spanning across the acoustic, vision and language modalities, has attracted increasing interest in the multimedia community. Previous studies predominantly focus on learning contextual information in conversations with only a few consideri…

2023

Brant: Foundation Model for Intracranial Neural Signal

NeurIPS 2023poster

We propose a foundation model named Brant for modeling intracranial recordings, which learns powerful representations of intracranial neural signals by pre-training, providing a large-scale, off-the-shelf model for medicine. Brant is the largest model in the field of brain signals and is pre-trained…

2023

Learning Expressive And Generalizable Motion Features For Face Forgery Detection

ICASSP 2023accepted

Previous face forgery detection methods mainly focus on appearance features, which may be easily attacked by sophisticated manipulation. Considering the majority of current face manipulation methods generate fake faces based on a single frame, which do not take frame consistency and coordination int…

Cited by 0SourceScholar
2023

MDR-MFI:Multi-Branch Decoupled Regression and Multi-Scale Feature Interaction for Partial-to-Partial Cloud Registration

ICASSP 2023accepted

Point cloud registration is a fundamental task in the 3D vision field. Many previous works adopt the regression model to estimate the transformation parameters. However, these methods couple the estimation of rotation and translation via a single regression branch, which suffers from the mutual inte…

Cited by 0SourceScholar
2023

Rethinking the Approximation Error in 3D Surface Fitting for Point Cloud Normal Estimation

CVPR 2023poster

Most existing approaches for point cloud normal estimation aim to locally fit a geometric surface and calculate the normal from the fitted surface. Recently, learning-based methods have adopted a routine of predicting point-wise weights to solve the weighted least-squares surface fitting problem. De…

2023

Single Domain Dynamic Generalization for Iris Presentation Attack Detection

ICASSP 2023accepted

Iris presentation attack detection (PAD) has achieved great success under intra-domain settings but easily degrades on unseen domains. Conventional domain generalization methods mitigate the gap by learning domain-invariant features. However, they ignore the discriminative information in the domain-…

Cited by 0SourceScholar
2023

Source-free Depth for Object Pop-out

ICCV 2023poster

Depth cues are known to be useful for visual perception. However, direct measurement of depth is often impracticable. Fortunately, though, modern learning-based methods offer promising depth maps by inference in the wild. In this work, we adapt such depth inference models for object segmentation usi…

Cited by 70PDFcodeScholar
2022

FBNet: Feedback Network for Point Cloud Completion

ECCV 2022poster

"The rapid development of point cloud learning has driven point cloud completion into a new era. However, the information flows of most existing completion methods are solely feedforward, and high-level information is rarely reused to improve low-level feature learning. To this end, we propose a nov…

2022

Few-Shot One-Class Domain Adaptation Based On Frequency For Iris Presentation Attack Detection

ICASSP 2022accepted

Iris presentation attack detection (PAD) has achieved remarkable success to ensure the reliability and security of iris recognition systems. Most existing methods exploit discriminative features in the spatial domain and report outstanding performance under intra-dataset settings. However, the degra…

Cited by 0SourceScholar
2022

Learning Multiple Explainable and Generalizable Cues for Face Anti-Spoofing

ICASSP 2022accepted

Although previous CNN based face anti-spoofing methods have achieved promising performance under intra-dataset testing, they suffer from poor generalization under cross-dataset testing. The main reason is that they learn the network with only binary supervision, which may learn arbitrary cues overfi…

Cited by 0SourceScholar
2022

Unimodal-Concentrated Loss: Fully Adaptive Label Distribution Learning for Ordinal Regression

CVPR 2022poster

Learning from a label distribution has achieved promising results on ordinal regression tasks such as facial age and head pose estimation wherein, the concept of adaptive label distribution learning (ALDL) has drawn lots of attention recently for its superiority in theory. However, compared with the…

Cited by 29PDFScholar
2021

Multi-Level Adaptive Region of Interest and Graph Learning for Facial Action Unit Recognition

ICASSP 2021accepted

In facial action unit (AU) recognition tasks, regional feature learning and AU relation modeling are two effective aspects which are worth exploring. However, the limited representation capacity of regional features makes it difficult for relation models to embed AU relationship knowledge. In this p…

Cited by 0SourceScholar
2021

Self-Domain Adaptation for Face Anti-Spoofing

AAAI 2021technical

Although current face anti-spoofing methods achieve promising results under intra-dataset testing, they suffer from poor generalization to unseen attacks. Most existing works adopt domain adaptation (DA) or domain generalization (DG) techniques to address this problem. However, the target domain is…

Cited by 117SourcePDFScholar
2020

Multimodal Topic-Enriched Auxiliary Learning for Depression Detection

COLING 2020main

From the perspective of health psychology, human beings with long-term and sustained negativity are highly possible to be diagnosed with depression. Inspired by this, we argue that the global topic information derived from user-generated contents (e.g., texts and images) is crucial to boost the perf…