← Search

Ziyang Song

9 accepted papers

2026

Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation

CVPR 2026

Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their application to medical imaging remains challenging due to the high variability and complexity of medical data. Specifically, medical images often exhibit

Cited by 0SourcecodeScholar
2025

FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity

CVPR 2025poster

In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing works often fail to learn complex physical motions at boundarie…

2024

UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models

EMNLP 2024main

Sequential decision-making refers to algorithms that take into account the dynamics of the environment, where early decisions affect subsequent decisions. With large language models (LLMs) demonstrating powerful capabilities between tasks, we can’t help but ask: Can Current LLMs Effectively Make Seq…

Cited by 2SourcePDFScholar
2023

ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation

ICCV 2023poster

We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion TransFormer (ActFormer) under a GAN training scheme, equippe…

Cited by 74PDFScholar
2023

NVFi: Neural Velocity Fields for 3D Physics Learning from Dynamic Videos

NeurIPS 2023poster

In this paper, we aim to model 3D scene dynamics from multi-view videos. Unlike the majority of existing works which usually focus on the common task of novel view synthesis within the training time period, we propose to simultaneously learn the geometry, appearance, and physical velocity of 3D scen…

2021

Multi-Scale Matching Networks for Semantic Correspondence

ICCV 2021poster

Deep features have been proven powerful in building accurate dense semantic correspondences in various previous works. However, the multi-scale and pyramidal hierarchy of convolutional neural networks has not been well studied to learn discriminative pixel-level features for semantic correspondence.…

Cited by 50PDFcodeScholar