← Search

Yun Liu

67 accepted papers

2026

LLMS ON TRIAL: Evaluating Judicial Fairness For Large Language Models

ICLR 2026poster

Large Language Models (LLMs) are increasingly used in high-stakes fields, such as law, where their decisions can directly impact people's lives. When LLMs act as judges, the ability to fairly resolve judicial issues is necessary to ensure their trustworthiness. Based on theories of judicial fairness…

Cited by 0SourcecodeScholar
2026

Nighttime Flare Removal via Wavelet-Guided and Gated-Enhanced Spatial-Frequency Fusion Network

AAAI 2026technical

Nighttime flares, caused by complex scattering and reflections from artificial light sources, significantly degrade image quality and hinder downstream visual tasks. Existing deflare networks usually struggle to jointly capture and fuse latent spatial and frequency features. In this paper, we propos

Cited by 0SourcePDFScholar
2026

Self-Refine Learning in LLM Multi-Agent Systems for Legal Norm Cognition and Compliance

IJCAI 2026

As large language models (LLMs) increasingly serve as autonomous agents in social simulations, ensuring their ability to understand and comply with legal norms is essential. Yet, current LLM agents frequently exhibit reward hacking (RH) behaviors by optimizing metrics at the expense of norm adherenc

Cited by 0Scholar
2026

Training Dynamics-Aware Multi-Factor Curriculum Learning for Target Speaker Extraction

ICASSP 2026oral

Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due to different interacting factors. Previous curriculum learning approaches for TSE typically address these factors separa…

Cited by 0SourcePDFScholar
2026

UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking

CVPR 2026

Generating lifelike conversational avatars requires modeling not just isolated speakers, but the dynamic, reciprocal interaction of speaking and listening.However, modeling the listener is exceptionally challenging: direct audio-driven training fails, producing stiff, static listening motions. This

Cited by 0SourcecodeScholar
2026

Unleashing Humanoid Reaching Potential Via Real-World-Ready Skill Space

ICRA 2026poster

Humans possess a large reachable space in the 3D world, enabling interactions with objects at varying heights and distances. However, realizing such large-space reaching on humanoids is a complex whole-body control (WBC) problem. Learning from scratch often leads to optimization difficulty and poor …

2026

Unleashing Humanoid Reaching Potential via Real-World-Ready Skill Space

RA-L 2026

Humans possess a large reachable space in the 3D world, enabling interactions with objects at varying heights and distances. However, realizing such large-space reaching on humanoids is a complex whole-body control (WBC) problem. Learning from scratch often leads to optimization difficulty and poor

Cited by 23SourcecodeScholar
2026

Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided Alignment

AAAI 2026technical

Unsupervised visual anomaly detection from multi-view images presents a significant challenge: distinguishing genuine defects from benign appearance variations caused by viewpoint changes. Existing methods, often designed for single-view inputs, treat multiple views as a disconnected set of images,

Cited by 0SourcePDFScholar
2026

VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

CVPR 2026

Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and communicate spatial layouts through language. However, existing methods largely rely on shallow text-point cloud correspond

Cited by 0SourcecodeScholar
2026

WiTTA-Bench: Benchmarking Test-Time Adaptation for WiFi Sensing

CVPR 2026

WiFi sensing offers passive and privacy-preserving perception that complements vision-based sensing, but its performance degrades sharply under domain shifts caused by changes in environment, subjects, or hardware. This challenge is exacerbated in real-world deployments where source data are unavail

Cited by 0SourcecodeScholar
2025

CORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangement

CVPR 2025poster

Understanding how humans cooperatively rearrange household objects is critical for VR/AR and human-robot interaction. However, in-depth studies on modeling these behaviors are under-researched due to the lack of relevant datasets. We fill this gap by presenting CORE4D, a novel large-scale 4D human-o…

2025

Dehaze-RetinexGAN: Real-World Image Dehazing via Retinex-based Generative Adversarial Network

AAAI 2025technical

Deep learning based dehazing networks trained on paired synthetic data have shown impressive performance, but they struggle with significant degradation in generalization ability on real-world hazy scenes. In this paper, we propose Dehaze-RetinexGAN, a lightweight Retinex-based Generative Adversari…

Cited by 0SourcePDFScholar
2025

Exploiting Temporal State Space Sharing for Video Semantic Segmentation

CVPR 2025poster

Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we i…

2025

Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model

CVPR 2025poster

Generalized few-shot 3D point cloud segmentation (GFS-PCS) adapts models to new classes with few support samples while retaining base class segmentation. Existing GFS-PCS methods enhance prototypes via interacting with support or query features but remain limited by sparse knowledge from few-shot sa…

2025

J&H: Evaluating the Robustness of Large Language Models Under Knowledge-Injection Attacks in Legal Domain

AAAI 2025technical

As the scale and capabilities of Large Language Models (LLMs) increase, their applications in knowledge-intensive fields such as legal domain have garnered widespread attention. However, it remains doubtful whether these LLMs make judgments based on domain knowledge for reasoning. If LLMs base their…

2025

JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoning

EMNLP 2025

In recent years, Large Language Models (LLMs) have been widely applied to legal tasks. To enhance their understanding of legal texts and improve reasoning accuracy, a promising approach is to incorporate legal theories. One of the most widely adopted theories is the Four-Element Theory (FET), which

2025

ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping

CVPR 2025highlight

In this paper, we introduce ManiVideo, a novel method for generating consistent and temporally coherent bimanual hand-object manipulation videos from given motion sequences of hands and objects. The core idea of ManiVideo is the construction of a multi-layer occlusion (MLO) representation that learn…

Cited by 1SourcePDFScholar
2025

Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation

ICLR 2025spotlight

Few-shot 3D point cloud segmentation (FS-PCS) aims at generalizing models to segment novel categories with minimal annotated support samples. While existing FS-PCS methods have shown promise, they primarily focus on unimodal point cloud inputs, overlooking the potential benefits of leveraging multim…

2025

RADAR: Benchmarking Language Models on Imperfect Tabular Data

NeurIPS 2025poster

Language models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness—the ability to recognize, reason over, and appropriately handle data artifacts such as missing values, outliers, and logical inconsistencies—remains underexplored. These artifacts…

Cited by 0SourcecodeScholar
2025

Runtime Analysis of Evolutionary NAS for Multiclass Classification

ICML 2025poster

Evolutionary neural architecture search (ENAS) is a key part of evolutionary machine learning, which commonly utilizes evolutionary algorithms (EAs) to automatically design high-performing deep neural architectures. During past years, various ENAS methods have been proposed with exceptional performa…

Cited by 0SourcePDFScholar
2025

Scaling Wearable Foundation Models

ICLR 2025poster

Wearable sensors have become ubiquitous thanks to a variety of health tracking features. The resulting continuous and longitudinal measurements from everyday life generate large volumes of data. However, making sense of these observations for scientific and actionable insights is non-trivial. Inspir…

Cited by 6SourcePDFScholar
2025

SensorLM: Learning the Language of Wearable Sensors

NeurIPS 2025poster

We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descr…

Cited by 0SourcecodeScholar
2025

SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback Optimization

CVPR 2025poster

Snowfall presents significant challenges for visual data processing, necessitating specialized desnowing algorithms. However, existing models often fail to generalize effectively due to their heavy reliance on synthetic datasets. Furthermore, current real-world snowfall datasets are limited in scale…

Cited by 0SourcePDFScholar
2025

SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis

ICCV 2025poster

Synthesizing realistic human-object interaction motions is a critical problem in VR/AR and human animation. Unlike the commonly studied scenarios involving a single human or hand interacting with one object, we address a more generic multi-body setting with arbitrary numbers of humans, hands, and ob…

Cited by 0SourcePDFScholar
2025

Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts

NeurIPS 2025poster

Single-source Domain Generalized Object Detection (SDGOD), as a cutting-edge research topic in computer vision, aims to enhance model generalization capability in unseen target domains through single-source domain training. Current mainstream approaches attempt to mitigate domain discrepancies via d…

Cited by 0SourceScholar
2024

Enhancing Generalizable 6D Pose Tracking of an In-Hand Object With Tactile Sensing

RA-L 2024

When manipulating an object to accomplish complex tasks, humans rely on both vision and touch to keep track of the object's 6D pose. However, most existing object pose tracking systems in robotics rely exclusively on visual signals, which hinder a robot's ability to manipulate objects effectively. T

Cited by 25SourcecodeScholar
2024

Frame-By-Frame Motion Retargeting With Self-Collision Avoidance From Diverse Human Demonstrations

RA-L 2024

Human-robot motion retargeting is a complex nonlinear problem, due to heterogeneous kinematic configuration between human and robot. Recent efforts aim to tackle the generalizability of motion retargeting on diverse robots, yet challenges persist in handling unseen human motions with varying scales

Cited by 4SourceScholar
2024

LEEC for Judicial Fairness: A Legal Element Extraction Dataset with Extensive Extra-Legal Labels

IJCAI 2024poster

An extensive label system is pivotal to facilitate judicial fairness and social justice. Prior empirical research and our interview with legal professionals underscore the importance of extra-legal factors in criminal trials. To help identify sentencing biases and facilitate downstream applications,…

2024

Mobile-Seed: Joint Semantic Segmentation and Boundary Detection for Mobile Robots

RA-L 2024

Precise and rapid delineation of sharp boundaries and robust semantics is essential for numerous downstream robotic tasks, such as robot grasping and manipulation, real-time semantic mapping, and online sensor calibration performed on edge computing units. Although boundary detection and semantic se

Cited by 26SourcecodeScholar
2024

Rethinking Few-shot 3D Point Cloud Semantic Segmentation

CVPR 2024poster

This paper revisits few-shot 3D point cloud semantic segmentation (FS-PCS) with a focus on two significant issues in the state-of-the-art: foreground leakage and sparse point distribution. The former arises from non-uniform point sampling allowing models to distinguish the density disparities betwee…

2024

STARD: A Chinese Statute Retrieval Dataset Derived from Real-life Queries by Non-professionals

EMNLP 2024finding

Statute retrieval aims to find relevant statutory articles for specific queries. This process is the basis of a wide range of legal applications such as legal advice, automated judicial decisions, legal document drafting, etc. Existing statute retrieval benchmarks emphasize formal and professional q…

2024

TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding

CVPR 2024poster

Humans commonly work with multiple objects in daily life and can intuitively transfer manipulation skills to novel objects by understanding object functional regularities. However existing technical approaches for analyzing and synthesizing hand-object manipulation are mostly limited to handling a s…

2023

CAMS: CAnonicalized Manipulation Spaces for Category-Level Functional Hand-Object Manipulation Synthesis

CVPR 2023poster

In this work, we focus on a novel task of category-level functional hand-object manipulation synthesis covering both rigid and articulated object categories. Given an object geometry, an initial human hand pose as well as a sparse control sequence of object poses, our goal is to generate a physicall…

2023

DEHRFormer: Real-Time Transformer for Depth Estimation and Haze Removal from Varicolored Haze Scenes

ICASSP 2023accepted

Varicolored haze caused by chromatic casts poses haze removal and depth estimation challenges. Recent learning-based depth estimation methods are mainly targeted at dehazing first and estimating depth subsequently from haze-free scenes. This way, the inner connections between colored haze and scene…

Cited by 0SourceScholar
2023

Feature Modulation Transformer: Cross-Refinement of Global Representation via High-Frequency Prior for Image Super-Resolution

ICCV 2023poster

Transformer-based methods have exhibited remarkable potential in single image super-resolution (SISR) by effectively extracting long-range dependencies. However, most of the current research in this area has prioritized the design of transformer blocks to capture global information, while overlookin…

Cited by 79PDFcodeScholar
2023

Indiscernible Object Counting in Underwater Scenes

CVPR 2023poster

Recently, indiscernible scene understanding has attracted a lot of attention in the vision community. We further advance the frontier of this field by systematically studying a new challenge named indiscernible object counting (IOC), the goal of which is to count objects that are blended with respec…

2023

Joint Noise Reduction and Listening Enhancement for Full-End Speech Enhancement

ICASSP 2023accepted

Speech enhancement (SE) methods mainly focus on recovering clean speech from noisy input. In real-world speech communication, however, noises often exist in not only speaker but also listener environments. Although SE methods can suppress the noise contained in the speaker’s voice, they cannot deal…

Cited by 0SourceScholar
2023

MSP-Former: Multi-Scale Projection Transformer for Single Image Desnowing

ICASSP 2023accepted

Snow removal causes challenges due to its characteristic of complex degradations. To this end, targeted treatment of multi-scale snow degradations is critical for the network to learn effective snow removal. In order to handle the diverse scenes, we propose a multi-scale projection transformer (MSP-…

Cited by 0SourceScholar
2023

UniDexGrasp++: Improving Dexterous Grasping Policy Learning via Geometry-Aware Curriculum and Iterative Generalist-Specialist Learning

ICCV 2023oral

We propose a novel, object-agnostic method for learning a universal policy for dexterous object grasping from realistic point cloud observations and proprioceptive information under a table-top setting, namely UniDexGrasp++. To address the challenge of learning the vision-based policy across thousan…

Cited by 88PDFScholar
2023

Unsupervised Legal Evidence Retrieval via Contrastive Learning with Approximate Aggregated Positive

AAAI 2023technical

Verifying the facts alleged by the prosecutors before the trial requires the judges to retrieve evidence within the massive materials accompanied. Existing Legal AI applications often assume the facts are already determined and fail to notice the difficulty of reconstructing them. To build a practic…

2022

Coarse-To-Fine Feature Mining for Video Semantic Segmentation

CVPR 2022poster

The contextual information plays a core role in semantic segmentation. As for video semantic segmentation, the contexts include static contexts and motional contexts, corresponding to static content and moving content in a video clip, respectively. The static contexts are well exploited in image sem…

Cited by 75PDFcodeScholar
2022

HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction

CVPR 2022poster

We present HOI4D, a large-scale 4D egocentric dataset with rich annotations, to catalyze the research of category-level human-object interaction. HOI4D consists of 2.4M RGB-D egocentric video frames over 4000 sequences collected by 9 participants interacting with 800 different object instances from…

Cited by 183PDFcodeScholar
2022

LEVEN: A Large-Scale Chinese Legal Event Detection Dataset

ACL 2022findings

Recognizing facts is the most fundamental step in making judgments, hence detecting events in the legal documents is important to legal case analysis tasks. However, existing Legal Event Detection (LED) datasets only concern incomprehensive event types and have limited annotated data, which restrict…

2022

Mining Relations among Cross-Frame Affinities for Video Semantic Segmentation

ECCV 2022poster

"The essence of video semantic segmentation (VSS) is how to leverage temporal information for prediction. Previous efforts are mainly devoted to developing new techniques to calculate the cross-frame affinities such as optical flow and attention. Instead, this paper contributes from a different angl…

2022

Multi-Scale Interaction for Real-Time LiDAR Data Segmentation on an Embedded Platform

RA-L 2022

Real-time semantic segmentation of LiDAR data is crucial for autonomously driving vehicles and robots, which are usually equipped with an embedded platform and have limited computational resources. Approaches that operate directly on the point cloud use complex spatial aggregation operations, which

Cited by 101SourcecodeScholar
2022

Perceiving and Modeling Density for Image Dehazing

ECCV 2022poster

"In the real world, the degradation of images taken under haze can be quite complex, where the spatial distribution of haze varies from image to image. Recent methods adopt deep neural networks to recover clean scenes from hazy images directly. However, due to the generic design of network architect…

2022

The PCG-AIID System for L3DAS22 Challenge: MIMO and MISO Convolutional Recurrent Network for Multi Channel Speech Enhancement and Speech Recognition

ICASSP 2022accepted

This paper described the PCG-AIID system for L3DAS22 challenge in Task 1: 3D speech enhancement in office reverberant environment. We proposed a two-stage framework to address multi-channel speech denoising and dereverberation. In the first stage, a multiple input and multiple out-put (MIMO) network…

Cited by 0SourceScholar
2022

Uformer: A Unet Based Dilated Complex & Real Dual-Path Conformer Network for Simultaneous Speech Enhancement and Dereverberation

ICASSP 2022accepted

Complex spectrum and magnitude are considered as two major features of speech enhancement and dereverberation. Traditional approaches always treat these two features separately, ignoring their underlying relationship. In this paper, we propose Uformer, a Unet based dilated complex & real dual-path c…

Cited by 0SourceScholar
2022

Zero Pixel Directional Boundary by Vector Transform

ICLR 2022poster

Boundaries or contours are among the primary visual cues used by human and computer vision systems. One of the key problems in boundary detection is the loss formulation, which typically leads to class imbalance and, as a consequence, to thick boundaries which require non-differential post-processin…

2021

DOTS: Decoupling Operation and Topology in Differentiable Architecture Search

CVPR 2021poster

Differentiable Architecture Search (DARTS) has attracted extensive attention due to its efficiency in searching for cell structures. DARTS mainly focuses on the operation search and derives the cell topology from the operation weights. However, the operation weights can not indicate the importance o…

Cited by 72PDFcodeScholar
2021

Densely Connected Multi-Stage Model with Channel Wise Subband Feature for Real-Time Speech Enhancement

ICASSP 2021accepted

Research on single channel speech enhancement (SE) has a long tradition, but two main practical problems still remain unsolved. Firstly, it’s hard to balance between enhancement quality and computational efficiency, and low-latency always brings loss of quality. Secondly, enhancement in specific sce…

Cited by 0SourceScholar
2021

Matching Distributions between Model and Data: Cross-domain Knowledge Distillation for Unsupervised Domain Adaptation

ACL 2021long

Unsupervised Domain Adaptation (UDA) aims to transfer the knowledge of source domain to the unlabeled target domain. Existing methods typically require to learn to adapt the target model by exploiting the source data and sharing the network architecture across domains. However, this pipeline makes t…

Cited by 22SourcePDFScholar
2021

MiniSeg: An Extremely Minimum Network for Efficient COVID-19 Segmentation

AAAI 2021technical

The rapid spread of the new pandemic, i.e., COVID-19, has severely threatened global health. Deep-learning-based computer-aided screening, e.g., COVID-19 infected CT area segmentation, has attracted much attention. However, the publicly available COVID-19 training data are limited, easily causing ov…

2019

Multi-Level Context Ultra-Aggregation for Stereo Matching

CVPR 2019poster

Exploiting multi-level context information to cost volume can improve the performance of learning-based stereo matching methods. In recent years, 3-D Convolution Neural Networks (3-D CNNs) show the advantages in regularizing cost volume but are limited by unary features learning in matching cost com…

Cited by 137PDFScholar
2019

Scoot: A Perceptual Metric for Facial Sketches

ICCV 2019poster

While it is trivial for humans to quickly assess the perceptual similarity between two images, the underlying mechanism are thought to be quite complex. Despite this, the most widely adopted perceptual metrics today, such as SSIM and FSIM, are simple, shallow functions, and fail to consider many fac…

Cited by 57PDFScholar
2018

Crowd Counting With Deep Negative Correlation Learning

CVPR 2018poster

Deep convolutional networks (ConvNets) have achieved unprecedented performances on many computer vision tasks. However, their adaptations to crowd counting on single images are still in their infancy and suffer from severe over-fitting. Here we propose a new learning strategy to produce generalizabl…

2018

Structured Skip List: A Compact Data Structure for 3D Reconstruction

IROS 2018poster

The model produced by 3D reconstruction algorithm is usually represented by voxels. The management of these voxels is usually divided into two categories: ordered and unordered methods. The ordered method holds too many empty voxels to maintain data order which leads to a low storage efficiency. On…

Cited by 3SourceScholar
2017

Structure-Measure: A New Way to Evaluate Foreground Maps

ICCV 2017spotlight

Foreground map evaluation is crucial for gauging the progress of object segmentation algorithms, in particular in the filed of salient object detection where the purpose is to accurately detect and segment the most salient object in a scene. Several widely-used measures such as Area Under the Curve…

Cited by 1925PDFcodeScholar
2015

Time-reversal space-time codes in asynchronous two-way double-antenna relay networks

ICASSP 2015accepted

We consider an asynchronous two-way relay network, in which two double-antenna relays assist in the communication between two single-antenna terminals through analog network coding. The asynchronous transmission between relays and terminals causes symbol misalignments and results in diversity loss i…

Cited by 0SourceScholar