← Search

Wu Liu

42 accepted papers

2026

CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict

ICML 2026poster

The intelligent verdict is essential for handling voluminous demands of E-commerce dispute. Unlike the legal dispute, it necessitates identifying pivotal clues from redundant multimodal evidence chains, relying on informal transaction rules for dispute verdicts. The complex ``clues-dispute" causal l…

Cited by 0SourceScholar
2026

DiasR: Dual-Modal Identity-Anchored Sparse Routing for Efficient Multi-Subject Video Generation

ICML 2026poster

Personalized multi-subject video generation is a promising direction within the field of controllable video generation; however, existing methods face challenges in maintaining cross-frame identity consistency and incur high computational overhead. To address these issues, we propose DiasR, an effic…

Cited by 0SourceScholar
2026

GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents

AAAI 2026technical

Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacking the ability to adaptively determine when, whether, and how to observe the int

Cited by 0SourcePDFScholar
2026

HyperGait: Unleashing the Power of Parsing for Gait Recognition in the Wild via Hypergraph

CVPR 2026

In recent years, the gait parsing sequence has become increasingly popular due to its higher information entropy than the binary silhouette and the keypoint-based skeleton. However, existing parsing-based gait recognition methods have not fully explored the complex, non-linear relationships between

Cited by 0SourceScholar
2026

In-Context Generation with Regional Constraints for Instructional Video Editing

ICML 2026poster

The In-context generation paradigm has demonstrated strong power in instructional image editing for better synthesis quality. Nevertheless, shaping such in-context learning for instructional video editing is not trivial. Without specifying editing regions, the results can suffer from the issue of in…

Cited by 0SourceScholar
2026

Multi-level Causal LLM-based Text-to-Motion Generation with Human Alignment

CVPR 2026

Although progress has been made in LLM-based text-driven motion generation, it still has the limitations of generating fine-grained and semantically consistent motions. These limitations stem from: 1) fine-grained motion quantization errors; 2) mismatches between causal reasoning language and non-ca

Cited by 0SourceScholar
2026

Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing

CVPR 2026

Recent advancements in diffusion-based image editing pose a significant threat to the authenticity of digital visual content. Traditional embedding-based watermarking methods often introduce perceptible perturbations to maintain robustness, inevitably compromising visual fidelity. Meanwhile, existin

Cited by 0SourcecodeScholar
2026

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

ICML 2026poster

Reliable watermarking of panoramic imagery is fundamentally challenged by arbitrary 3D rotations. As panoramas are defined on the sphere, they naturally transform under the action of $SO(3)$, rendering conventional planar representations and augmentation-based robustness strategies inadequate and de…

Cited by 0SourceScholar
2025

A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image Segmentation

CVPR 2025poster

The limited data annotations have made semi-supervised learning (SSL) increasingly popular in medical image analysis. However, the use of pseudo labels in SSL degrades the performance of decoders that heavily rely on high-accuracy annotations. This issue is particularly pronounced in class-imbalance…

2025

ACEBench: A Comprehensive Evaluation of LLM Tool Usage

EMNLP 2025

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex problems. However, existing benchmarks for evaluating LLMs’ tool usage face several limitations: (1) limited evaluation

Cited by 0SourcePDFScholar
2025

HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation

CVPR 2025poster

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale videos with accurate captions for HOI. To address this issue,…

2025

MotionPro: A Precise Motion Controller for Image-to-Video Generation

CVPR 2025poster

Animating images with interactive motion control has garnered popularity for image-to-video (I2V) generation. Modern approaches typically rely on large Gaussian kernels to extend motion trajectories as condition without explicitly defining movement region, leading to coarse motion control and failin…

2025

PlugMark: A Plug-in Zero-Watermarking Framework for Diffusion Models

ICCV 2025poster

Diffusion models have significantly advanced the field of image synthesis, making the protection of their intellectual property (IP) a critical concern. Existing IP protection methods primarily focus on embedding watermarks into generated images by altering the structure of the diffusion process. Ho…

Cited by 0SourcePDFScholar
2025

RoboNotonecta: A Backswimmer-inspired Swimming Miniature Robot with Efficient Low-power Propulsion and Agile Aquatic Maneuverability

IROS 2025

In this letter, we present the design, manufacturing, and performance test of a centimeter-scale two-legged swimming robot, RoboNotonecta, inspired by the efficient and agile locomotion of an aquatic beetle backswimmer (Notonectid). The robot utilizes a crank-slider and slider-rocker paddling mechan

Cited by 0SourceScholar
2024

A Fish-Like Underwater Miniature Robot Capable of High-Speed and Controllable Locomotion

RA-L 2024

This letter presents the design, fabrication, and performance test of a fish-like underwater miniature robot directly driven by two piezoelectric actuators. Considering the different medium conditions and the ever-present challenge of waterproofing in underwater motion, the miniaturization of underw

Cited by 12SourceScholar
2024

A Multi-Modal Tailless Flapping-Wing Robot Capable of Flying, Crawling, Self-Righting and Horizontal Take-Off

RA-L 2024

Multi-modal flapping robots demonstrate potential capabilities to accomplish assigned missions in confined indoor and outdoor environments. In this letter, inspired by the multimodal movement of insects, we report a multimodal tailless flapping wing robot capable of flying, crawling, self-righting a

Cited by 27SourceScholar
2024

HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse Poses

CVPR 2024poster

We present HumanNeRF-SE a simple yet effective method that synthesizes diverse novel pose images with simple input. Previous HumanNeRF works require a large number of optimizable parameters to fit the human images. Instead we reload these approaches by combining explicit and implicit human represent…

Cited by 4SourcePDFScholar
2024

Norma: A Noise Robust Memory-Augmented Framework for Whole Slide Image Classification

ECCV 2024poster

"In recent years, the Whole Slide Image (WSI) classification task has achieved great advancement due to the success of Multiple Instance Learning (MIL). However, the MIL-based studies usually consider instances within each bag as unordered, potentially resulting in the missing of local and global co…

2023

Learning To Segment Every Referring Object Point by Point

CVPR 2023poster

Referring Expression Segmentation (RES) can facilitate pixel-level semantic alignment between vision and language. Most of the existing RES approaches require massive pixel-level annotations, which are expensive and exhaustive. In this paper, we propose a new partially supervised training paradigm f…

2023

RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments

NeurIPS 2023poster

Intention-oriented object detection aims to detect desired objects based on specific intentions or requirements. For instance, when we desire to "lie down and rest", we instinctively seek out a suitable option such as a "bed" or a "sofa" that can fulfill our needs. Previous work in this area is limi…

Cited by 14SourcePDFScholar
2023

TRACE: 5D Temporal Regression of Avatars With Dynamic Cameras in 3D Environments

CVPR 2023poster

Although the estimation of 3D human pose and shape (HPS) is rapidly progressing, current methods still cannot reliably estimate moving humans in global coordinates, which is critical for many applications. This is particularly challenging when the camera is also moving, entangling human and camera m…

2022

CAViT: Contextual Alignment Vision Transformer for Video Object Re-identification

ECCV 2022poster

"Video object re-identification (reID) aims at re-identifying the same object under non-overlapping cameras by matching the video tracklets with cropped video frames. The key point is how to make full use of spatio-temporal interactions to extract more accurate representation. However, there are dil…

2022

Gait Recognition in the Wild With Dense 3D Representations and a Benchmark

CVPR 2022poster

Existing studies for gait recognition are dominated by 2D representations like the silhouette or skeleton of the human body in constrained scenes. However, humans live and walk in the unconstrained 3D space, so projecting the 3D human body onto the 2D plane will discard a lot of crucial information…

Cited by 192PDFcodeScholar
2022

Genre-Conditioned Long-Term 3D Dance Generation Driven by Music

ICASSP 2022accepted

Dancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre.…

Cited by 0SourceScholar
2022

Learning Monocular Mesh Recovery of Multiple Body Parts Via Synthesis

ICASSP 2022accepted

In this paper, we focus on simultaneously recovering the 3D mesh of multiple body parts from a single RGB image. One of the main challenges is that available datasets with full-body 3D annotations are very limited. This results in poor generalization ability of existing learning-based methods. Exist…

Cited by 0SourceScholar
2022

Putting People in Their Place: Monocular Regression of 3D People in Depth

CVPR 2022poster

Given an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene co…

Cited by 180PDFcodeScholar
2022

SiRi: A Simple Selective Retraining Mechanism for Transformer-Based Visual Grounding

ECCV 2022poster

"In this paper, we investigate how to achieve better referring visual grounding with modern vision-language transformers, and propose a simple yet powerful Selective Retraining (SiRi) mechanism. Particularly, SiRi conveys a significant principle to the research of visual grounding, i.e, a better ini…

2021

Explainable Person Re-Identification With Attribute-Guided Metric Distillation

ICCV 2021poster

Despite the great progress of person re-identification (ReID) with the adoption of Convolutional Neural Networks, current ReID models are opaque and only outputs a scalar distance between two persons. There are few methods providing users semantically understandable explanations for why two persons…

Cited by 57PDFcodeScholar
2021

Group-aware Label Transfer for Domain Adaptive Person Re-identification

CVPR 2021poster

Unsupervised Domain Adaptive (UDA) person re-identification (ReID) aims at adapting the model trained on a labeled source-domain dataset to a target-domain dataset without any further annotations. Most successful UDA-ReID approaches combine clustering-based pseudo-label prediction with representatio…

Cited by 231PDFcodeScholar
2021

Learning Instance-Level Spatial-Temporal Patterns for Person Re-Identification

ICCV 2021poster

Person re-identification (Re-ID) aims to match pedestrians under dis-joint cameras. Most Re-ID methods formulate it as visual representation learning and image search, and its accuracy is consequently affected greatly by the search space. Spatial-temporal information has been proven to be efficient…

Cited by 30PDFcodeScholar
2021

Monocular, One-Stage, Regression of Multiple 3D People

ICCV 2021poster

This paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a On…

Cited by 327PDFcodeScholar
2021

Neural Architecture Search for Joint Human Parsing and Pose Estimation

ICCV 2021poster

Human parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process w…

Cited by 26PDFcodeScholar
2020

Content-Consistent Matching for Domain Adaptive Semantic Segmentation

ECCV 2020poster

This paper considers the adaptation of semantic segmentation from the synthetic source domain to the real target domain. Different from most previous explorations that often aim at developing adversarial-based domain alignment solutions, we tackle this challenging task from a new perspective, mph{i.…

2020

When Pedestrian Detection Meets Nighttime Surveillance: A New Benchmark

IJCAI 2020poster

Pedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising per…

2019

Foreground-Aware Pyramid Reconstruction for Alignment-Free Occluded Person Re-Identification

ICCV 2019poster

Re-identifying a person across multiple disjoint camera views is important for intelligent video surveillance, smart retailing and many other applications. However, existing person re-identification methods are challenged by the ubiquitous occlusion over persons and suffer performance degradation. T…

Cited by 252PDFScholar
2019

Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled Representation

ICCV 2019poster

We describe an end-to-end method for recovering 3D human body mesh from single images and monocular videos. Different from the existing methods try to obtain all the complex 3D pose, shape, and camera parameters from one coupling feature, we propose a skeleton-disentangling based framework, which di…

Cited by 214PDFcodeScholar
2019

Social Relation Recognition From Videos via Multi-Scale Spatial-Temporal Reasoning

CVPR 2019poster

Discovering social relations, e.g., kinship, friendship, etc., from visual contents can make machines better interpret the behaviors and emotions of human beings. Existing studies mainly focus on recognizing social relations from still images while neglecting another important media--video. On one h…

Cited by 94PDFScholar
2018

Joint License Plate Super-Resolution and Recognition in One Multi-Task Gan Framework

ICASSP 2018accepted

License plate recognition (LPR) plays an important role in intelligent transport systems. The existed LPR systems are mostly based on hand-crafted methods for detection, segmentation, and recognition, which cannot accurately recognize the license plate in unconstrained surveillance environments. In…

Cited by 0SourceScholar
2015

Multi-Task Deep Visual-Semantic Embedding for Video Thumbnail Selection

CVPR 2015poster

Given the tremendous growth of online videos, video thumbnail, as the common visualization form of video content, is becoming increasingly important to influence user's browsing and searching experience. However, conventional methods for video thumbnail selection often fail to produce satisfying res…

Cited by 285SourcePDFScholar