← Search

Xiangyang Zhu

24 accepted papers

2026

Generalizable Video Quality Assessment via Weak-to-Strong Learning

CVPR 2026

Video quality assessment (VQA) seeks to predict the perceptual quality of a video in alignment with human visual perception, serving as a fundamental tool for quantifying quality degradation across video processing workflows. The dominant VQA paradigm relies on supervised training with human-labeled

Cited by 0SourcecodeScholar
2026

Image Quality Assessment for Embodied AI

ICLR 2026poster

Embodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application. Traditionally, Image Quality Assessment (IQA) methods are applied to predict human preferences for distorted images; however, there…

Cited by 0SourcecodeScholar
2026

MedOmni-45°: A Safety–Performance Benchmark for Reasoning-Oriented LLMs in Medicine

AAAI 2026technical

With the rapid integration of large language models (LLMs) into medical decision-support aids, ensuring reliability in reasoning steps—not just final answers—is increasingly critical. Two key safety dimensions are Chain-of-Thought (CoT) faithfulness, which assesses alignment of the model’s reasoning

Cited by 0SourcePDFScholar
2026

SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models

AAAI 2026technical

In the rapidly evolving landscape of Multimodal Large Language Models (MLLMs), the safety concerns of their outputs have earned significant attention. Although numerous datasets have been proposed, they may become outdated with MLLM advancements and are susceptible to data contamination issues. To a

Cited by 0SourcePDFScholar
2026

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

ICML 2026poster

The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchmarks often suffer from limited risk coverage and a reliance on subjective evaluation. To address thess problems, we intr…

Cited by 0SourceScholar
2026

VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning

AAAI 2026technical

Video quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainab

Cited by 0SourcePDFScholar
2025

Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

ICCV 2025poster

We introduce Lumina-Image 2.0, an advanced text-to-image (T2I) model that surpasses previous state-of-the-art methods across multiple benchmarks. Lumina-Image 2.0 is characterized by two key features: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image to…

2024

Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT

NeurIPS 2024poster

Lumina-T2X is a nascent family of Flow-based Large Diffusion Transformers (Flag-DiT) that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters chall…

2024

Multidirectional Bending Soft Pneumatic Actuator With Fishbone-Like Strain-Limiting Layer for Dexterous Manipulation

RA-L 2024

Soft pneumatic actuators (SPAs), due to their compliance and adaptiveness, are promising solutions for manipulation. However, most SPAs have only simple motion modes and cannot perform the compound motion required for complex manipulation. In this letter, we propose a parallel-chamber actuator capab

Cited by 17SourceScholar
2024

No Time to Train: Empowering Non-Parametric Networks for Few-shot 3D Scene Segmentation

CVPR 2024highlight

To reduce the reliance on large-scale datasets recent works in 3D segmentation resort to few-shot learning. Current 3D few-shot segmentation methods first pre-train models on 'seen' classes and then evaluate their generalization performance on 'unseen' classes. However the prior pre-training stage n…

2023

Efficient Multi-View Inverse Rendering Using a Hybrid Differentiable Rendering Method

IJCAI 2023poster

Recovering the shape and appearance of real-world objects from natural 2D images is a long-standing and challenging inverse rendering problem. In this paper, we introduce a novel hybrid differentiable rendering method to efficiently reconstruct the 3D geometry and reflectance of a scene from multi-v…

2023

Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior Refinement

ICCV 2023poster

The popularity of Contrastive Language-Image Pre-training (CLIP) has propelled its application to diverse downstream vision tasks. To improve its capacity on downstream tasks, few-shot learning has become a widely-adopted technique. However, existing methods either exhibit limited performance or suf…

Cited by 89PDFcodeScholar
2023

PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning

ICCV 2023poster

Large-scale pre-trained models have shown promising open-world performance for both vision and language tasks. However, their transferred capacity on 3D point clouds is still limited and only constrained to the classification task. In this paper, we first collaborate CLIP and GPT to be a unified 3D…

Cited by 241PDFcodeScholar
2022

Obstacle Avoidance of Resilient UAV Swarm Formation with Active Sensing System in the Dense Environment

IROS 2022poster

This paper proposes a perception-shared and swarm trajectory global optimal (STGO) algorithm fused UAVs formation motion planning framework aided by an active sensing system. First, the point cloud received by each UAV is fit by the gaussian mixture model (GMM) and transmitted in the swarm. Resampli…

Cited by 21SourceScholar
2022

Triply Periodic Channels Enable Soft Pneumatic Linear Actuator With Single Material and Scalability

RA-L 2022

In this letter, we propose a new class of soft pneumatic linear actuators with single-material, uniaxial deformation, high energy density, and scalability, purely based on periodic curved air channels. The shape of channels is implicitly parameterized by modified triply periodic minimal surfaces (mT

Cited by 12SourceScholar
2021

Computationally Efficient Trajectory Planning for High Speed Obstacle Avoidance of a Quadrotor With Active Sensing

RA-L 2021

Quadrotor with active sensing was proposed recently to overcome the view field limitation and achieved an excellent perception ability in obstacle avoidance tasks. To realize high-speed flights of this quadrotor in unknown and cluttered environments, a computationally efficient trajectory planner is

Cited by 21SourceScholar
2019

Buckling-induced Shape Morphing using Dielectric Elastomer Actuators Patterned with Spatially-varying Electrodes

IROS 2019poster

Shape morphing is at the core of future research, which shows promise for wide applications ranging from reconfigurable electronics to soft material robots. In this paper, we present a novel buckling-induced mechanism for shape morphing using dielectric elastomer actuators (DEAs), by bonding the pla…

Cited by 7SourceScholar
2019

Model-Based Estimation of the Gravity-Loaded Shape and Scene Depth for a Slim 3-Actuator Continuum Robot with Monocular Visual Feedback

ICRA 2019poster

Fruitful developments on continuum robots have been witnessed in recent years due to their movements and manipulation capabilities in confined spaces. Due to the nature that a continuum robot has an infinite number of DoFs (Degrees of Freedom), majority of the existing systems deployed abundant actu…

Cited by 12SourceScholar
2018

Continuum Manipulator with Redundant Backbones and Constrained Bending Curvature for Continuously Variable Stiffness

IROS 2018poster

Snake-like manipulators can navigate and perform manipulation in confined spaces. Their recent implementations in surgical robots attracted a lot of attentions. These slender manipulators usually possess either a hyper-redundant articulated vertebrate structure or a continuum one. Primary design con…

Cited by 22SourceScholar
2015

A single-actuator prosthetic hand using a continuum differential mechanism

ICRA 2015poster

Substantial progresses have been made in building versatile anthropomorphic prosthetic hands in the past two decades using emerging technologies. However the trade-offs between functionality, reliability, affordability, appearance, etc. have not been fully settled. Many existing designs, particularl…

Cited by 57SourceScholar