← Search

Yu Jiang

22 accepted papers

2026

HumanPro: Single-view 3D Clothed Human Reconstruction with Progressive Normal Guidance

AAAI 2026technical

Reconstructing fine-grained geometry of clothed human from single-view image is a challenging task, particularly in accurately recovering complex shapes and generating clothes details. To address these limitations, we propose a novel approach named HumanPro, which estimates high-quality human normal

Cited by 0SourcePDFScholar
2026

HyperSign: Hierarchical Hypergraph-based Co-occurrence Modeling for Sign Language Recognition and Translation

AAAI 2026technical

Effectively capturing co-occurrence signals, such as hand shapes, facial expressions, and body postures, is critical for semantic understanding in sign language recognition (SLR) and translation (SLT). Although skeleton data offer greater efficiency and robustness than RGB inputs, existing methods t

Cited by 0SourcePDFScholar
2026

Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation Models

CVPR 2026

With the emergence of 3D foundation models, there is growing interest in fine-tuning them for downstream tasks, where LoRA is the dominant fine-tuning paradigm. As 3D datasets exhibit distinct variations in texture, geometry, camera motion, and lighting, there are interesting fundamental questions:

Cited by 0SourceScholar
2026

ProbeLLM: Automating Principled Diagnosis of LLM Failures

ICML 2026poster

Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has been enabled by dynamic test generation, existing approaches often discover isolated failure cases, lack principled con…

Cited by 0SourceScholar
2025

ArtGS: 3D Gaussian Splatting for Interactive Visual-Physical Modeling and Manipulation of Articulated Objects

IROS 2025

Articulated object manipulation remains a critical challenge in robotics due to the complex kinematic constraints and the limited physical reasoning of existing methods. In this work, we introduce ArtGS, a novel framework that extends 3D Gaussian Splatting (3DGS) by integrating visual-physical model

Cited by 8SourceScholar
2025

Augmenting Short Enrollment Speech via Synthesis for Target Speaker Extraction

ICASSP 2025accepted

A high-quality enrollment speech is crucial to target speaker extraction (TSE), since it provides essential cues for identifying the target speaker in the mixture. However, real applications usually only permit a short enrollment speech, e.g. a wakeup word for a mobile device, that provides limited…

Cited by 0SourceScholar
2025

Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge

ICASSP 2025accepted

In this paper, we present the system developed by our team, CCATTS, for the LIMMITS’25 challenge, focusing on few-shot and zero-shot TTS. We adopt a two-stage TTS strategy. In track 1, we fine-tune the pre-trained ZMM-TTS model and successfully achieve multilingual low-latency TTS. In track 2, we pr…

Cited by 0SourceScholar
2025

ERetinex: Event Camera Meets Retinex Theory for Low-Light Image Enhancement

ICRA 2025

Low-light image enhancement aims to restore the under-exposure image captured in dark scenarios. Under such scenarios, traditional frame-based cameras may fail to capture the structure and color information due to the exposure time limitation. Event cameras are bio-inspired vision sensors that respo

Cited by 5SourcecodeScholar
2025

Joint 3D Point Cloud Segmentation Using Real-Sim Loop: From Panels to Trees and Branches

ICRA 2025

Modern orchards are planted in structured rows with distinct panel divisions to improve management. Accurate and efficient joint segmentation of point cloud from Panel to Tree and Branch (P2TB) is essential for robotic operations. However, most current segmentation methods focus on single-instance s

Cited by 2SourceScholar
2025

iG-6DoF: Model-free 6DoF Pose Estimation for Unseen Object via Iterative 3D Gaussian Splatting

CVPR 2025poster

Traditional methods in pose estimation often rely on precise 3D models or additional data such as depth and normals, limiting their generalization, especially when objects undergo large translations or rotations. We propose iG-6DoF, a novel model-free 6D pose estimation method iterative 3D Gaussian…

Cited by 0SourcePDFScholar
2024

(Real2Sim)−1: 3D Branch Point Cloud Completion for Robotic Pruning in Apple Orchards

IROS 2024poster

Robotic branch pruning, a rapidly growing field addressing labor shortages in agriculture, requires detailed perception of branch geometry and topology. However, point clouds obtained in agricultural settings often lack completeness, limiting pruning accuracy. This work addressed point cloud quality…

Cited by 1SourceScholar
2024

At Which Training Stage Does Code Data Help LLMs Reasoning?

ICLR 2024spotlight

Large Language models (LLMs) have exhibited remarkable reasoning capabilities and become the foundation of language technologies. Inspired by the great success of code data in training LLMs, we naturally wonder at which training stage introducing code data can really help LLMs reasoning. To this end…

2023

Adaptive Constraint Partition Based Optimization Framework for Large-Scale Integer Linear Programming (Student Abstract)

AAAI 2023technical

Integer programming problems (IPs) are challenging to be solved efficiently due to the NP-hardness, especially for large-scale IPs. To solve this type of IPs, Large neighborhood search (LNS) uses an initial feasible solution and iteratively improves it by searching a large neighborhood around the cu…

Cited by 5SourcePDFScholar
2023

GNN&GBDT-Guided Fast Optimizing Framework for Large-scale Integer Programming

ICML 2023poster

The latest two-stage optimization framework based on graph neural network (GNN) and large neighborhood search (LNS) is the most popular framework in solving large-scale integer programs (IPs). However, the framework can not effectively use the embedding spatial information in GNN and still highly re…

Cited by 17SourcePDFScholar
2023

Self-Paced Learning Based Graph Convolutional Neural Network for Mixed Integer Programming (Student Abstract)

AAAI 2023technical

Graph convolutional neural network (GCN) based methods have achieved noticeable performance in solving mixed integer programming problems (MIPs). However, the generalization of existing work is limited due to the problem structure. This paper proposes a self-paced learning (SPL) based GCN network (S…

Cited by 3SourcePDFScholar
2023

Stacking-Based Attention Temporal Convolutional Network for Action Segmentation

ICASSP 2023accepted

Action segmentation plays an important role in video understanding, which is implemented by frame-wise action classification. Recent works on action segmentation capture long-term dependencies by increasing temporal convolution layers in Temporal Convolution Networks (TCNs). However, high layers in…

Cited by 0SourceScholar
2023

Vision-Based Vineyard Navigation Solution with Automatic Annotation

IROS 2023poster

Autonomous navigation is crucial for achieving the full automation of agricultural research and production management using agricultural robots. In this paper, we present a vision-based autonomous navigation approach for agriculture robots in trellised cropping systems, which stands out for its rema…

Cited by 4SourceScholar
2022

An Adaptive Approach to Whole-Body Balance Control of Wheel-Bipedal Robot Ollie

IROS 2022poster

The wheel-bipedal robot has the advantages of both wheeled robots and legged robots, but as a cost, it is more challenging to perform flexible movements in various surroundings while keeping it balanced. The inaccurate dynamics of the robot makes the balance problem even more intractable. To solve t…

Cited by 28SourceScholar
2022

Near Real-Time Vineyard Downy Mildew Detection and Severity Estimation

IROS 2022poster

The global grape and wine industry has been considerably impacted by diseases such as downy mildew (DM). Agricultural robots have demonstrated great potential to accurately and rapidly map DM infection for precision applications. Although the robots can autonomously acquire high-resolution images in…

Cited by 6SourceScholar
2021

Event Stream Super-Resolution via Spatiotemporal Constraint Learning

ICCV 2021poster

Event cameras are bio-inspired sensors that respond to brightness changes asynchronously and output in the form of event streams instead of frame-based images. They own outstanding advantages compared with traditional cameras: higher temporal resolution, higher dynamic range, and lower power consump…

Cited by 21PDFScholar
2020

Joint Policy Search for Multi-agent Collaboration with Imperfect Information

NeurIPS 2020poster

To learn good joint policies for multi-agent collaboration with incomplete information remains a fundamental challenge. While for two-player zero-sum games, coordinate-ascent approaches (optimizing one agent's policy at a time, e.g., self-play) work with guarantees, in multi-agent cooperative settin…

2019

Towards VQA Models That Can Read

CVPR 2019poster

Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new "TextVQA" dataset to…

Cited by 1328PDFcodeScholar