← Search

Ziyi Chen

29 accepted papers

2026

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

ICLR 2026poster

Vision-Language Models (VLMs) achieve strong results on multimodal tasks such as visual question answering, yet they can still fail even when the correct visual evidence is present. In this work, we systematically investigate whether these failures arise from not perceiving the evidence or from not…

Cited by 0SourceScholar
2026

SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation

CVPR 2026

Embodied navigation that adheres to social norms remains an open research challenge. Our SocialNav is a foundational model for socially-aware navigation with a hierarchical "brain-action" architecture, capable of understanding high-level social norms and generating low-level, socially compliant traj

Cited by 0SourcecodeScholar
2026

UniVBench: Towards Unified Evaluation for Video Foundation Models

CVPR 2026

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation benchmarks remain fragmented and limited in scope, as they each

Cited by 0SourcecodeScholar
2025

Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement

CVPR 2025poster

Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, many rely on strong visual controls, such as pose images or TPS keypoint movements,…

2025

Serialization based Point Cloud Oversegmentation

ICCV 2025poster

Point cloud oversegmentation, as a fundamental preprocessing step for 3D understanding, is a challenging task due to its spatial proximity and semantic similarity requirements. Most existing works struggle to efficiently group semantically consistent points into superpoints while maintaining spatial…

2025

Towards Optimal Multi-draft Speculative Decoding

ICLR 2025poster

Large Language Models (LLMs) have become an indispensable part of natural language processing tasks. However, autoregressive sampling has become an efficiency bottleneck. Multi-Draft Speculative Decoding (MDSD) is a recent approach where, when generating each token, a small draft model generates mul…

Cited by 2SourcePDFScholar
2025

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

ICCV 2025poster

Vision-Language Models (VLMs) have shown promising capabilities in handling various multimodal tasks, yet they struggle in long-context scenarios, particularly tasks involving videos, high-resolution images, or lengthy image-text documents. In our work, we first conduct an empirical analysis of VLMs…

2024

Cascade Speculative Drafting for Even Faster LLM Inference

NeurIPS 2024poster

Introduced to enhance the efficiency of large language model (LLM) inference, speculative decoding operates by having a smaller model generate a draft. A larger target model then reviews this draft to align with its output, and any acceptance by the target model results in a reduction of the number…

2024

Co-speech Gesture Video Generation with 3D Human Meshes

ECCV 2024poster

"Co-speech gesture video generation is an enabling technique for many digital human applications. Substantial progress has been made in creating high-quality talking head videos. However, existing hand gesture video generation methods are primarily limited by the widely adopted 2D skeleton-based ges…

Cited by 1SourcePDFScholar
2024

Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier Domain

AAAI 2024technical

RAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB image…

2024

Generated and Pseudo Content guided Prototype Refinement for Few-shot Point Cloud Segmentation

NeurIPS 2024spotlight

Few-shot 3D point cloud semantic segmentation aims to segment query point clouds with only a few annotated support point clouds. Existing prototype-based methods learn prototypes from the 3D support set to guide the segmentation of query point clouds. However, they encounter the challenge of low pro…

Cited by 1SourcePDFScholar
2024

NC-SDF: Enhancing Indoor Scene Reconstruction Using Neural SDFs with View-Dependent Normal Compensation

CVPR 2024poster

State-of-the-art neural implicit surface representations have achieved impressive results in indoor scene reconstruction by incorporating monocular geometric priors as additional supervision. However we have observed that multi-view inconsistency between such priors poses a challenge for high-qualit…

Cited by 2SourcePDFScholar
2023

Generalized-Smooth Nonconvex Optimization is As Efficient As Smooth Nonconvex Optimization

ICML 2023poster

Various optimal gradient-based algorithms have been developed for smooth nonconvex optimization. However, many nonconvex machine learning problems do not belong to the class of smooth functions and therefore the existing algorithms are sub-optimal. Instead, these problems have been shown to satisfy…

2022

Data sampling affects the complexity of online SGD over dependent data

UAI 2022poster

Conventional machine learning applications typically assume that data samples are independently and identically distributed (i.i.d.). However, practical scenarios often involve a data-generating process that produces highly dependent data samples, which are known to heavily bias the stochastic optim…

Cited by 4SourcePDFScholar
2022

Finding Correlated Equilibrium of Constrained Markov Game: A Primal-Dual Approach

NeurIPS 2022accept

Constrained Markov game is a fundamental problem that covers many applications, where multiple players compete with each other under behavioral constraints. The existing literature has proved the existence of Nash equilibrium for constrained Markov games, which turns out to be PPAD-complete and cann…

Cited by 12SourcePDFScholar
2022

Sample Efficient Stochastic Policy Extragradient Algorithm for Zero-Sum Markov Game

ICLR 2022poster

Two-player zero-sum Markov game is a fundamental problem in reinforcement learning and game theory. Although many algorithms have been proposed for solving zero-sum Markov games in the existing literature, many of them either require a full knowledge of the environment or are not sample-efficient. I…

Cited by 21SourcePDFScholar
2022

Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time Analysis

ICML 2022spotlight

Actor-critic (AC) algorithms have been widely used in decentralized multi-agent systems to learn the optimal joint control policy. However, existing decentralized AC algorithms either need to share agents’ sensitive information or lack communication-efficiency. In this work, we develop decentralized…

Cited by 36SourcePDFScholar
2022

Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis

EMNLP 2022finding

Multimodal speech emotion recognition (SER) and sentiment analysis (SA) are important techniques for human-computer interaction. Most existing multimodal approaches utilize either shallow cross-modal fusion of pretrained features, or deep cross-modal fusion with raw features. Recently, attempts have…

2022

Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment

ICASSP 2022accepted

Automatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous efforts typically only model one aspect (e.g., accuracy) at one granularity (e.g.,…

Cited by 0SourceScholar
2021

Greedy-GQ with Variance Reduction: Finite-time Analysis and Improved Complexity

ICLR 2021poster

Greedy-GQ is a value-based reinforcement learning (RL) algorithm for optimal control. Recently, the finite-time analysis of Greedy-GQ has been developed under linear function approximation and Markovian sampling, and the algorithm is shown to achieve an $\epsilon$-stationary point with a sample comp…

Cited by 19SourcePDFScholar
2021

Proximal Gradient Descent-Ascent: Variable Convergence under KŁ Geometry

ICLR 2021poster

The gradient descent-ascent (GDA) algorithm has been widely applied to solve minimax optimization problems. In order to achieve convergent policy parameters for minimax optimization, it is important that GDA generates convergent variable sequences rather than convergent sequences of function value o…

Cited by 37SourcePDFScholar
2021

The Thinkit System for Icassp2021 M2voc Challenge

ICASSP 2021accepted

In this paper, we introduce the low resource text-to-speech system from the ThinkIT team submitted to Multi-Speaker Multi-Style Voice Cloning Challenge (M2VoC). The challenge has two tasks: few-shot track1 provides 100 samples for each person and one-shot track2 offers 5 samples only. Each track con…

Cited by 0SourceScholar