← Search

Qi Yang

24 accepted papers

2026

CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering

CVPR 2026

Knowledge-based visual question answering (KB-VQA) demonstrates significant potential for handling knowledge-intensive tasks. However, conflicts arise between static parametric knowledge in vision language models (VLMs) and dynamically retrieved information due to the static model knowledge from pre

Cited by 0SourcecodeScholar
2026

On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional Study

ICLR 2026poster

Counterfactual reasoning has emerged as a crucial technique for generalizing the reasoning capabilities of large language models (LLMs). By generating and analyzing counterfactual scenarios, researchers can assess the adaptability and reliability of model decision-making. Although prior work has sho…

Cited by 0SourceScholar
2026

R-4B: Incentivizing General-Purpose Auto-Thinking in MLLMs via Bi-Mode Annealing and Reinforce Learning

CVPR 2026

Multimodal Large Language Models (MLLMs) with explicit step-by-step reasoning have achieved strong performance on complex tasks. However, such reasoning is unnecessary for many simple queries and introduces substantial computational overhead. To address this inefficiency, we present R-4B, an auto-th

Cited by 0SourcecodeScholar
2026

RAP: Fast Feedforward Rendering-Free Attribute-Guided Primitive Importance Score Prediction for Efficient 3D Gaussian Splatting Processing

CVPR 2026

3D Gaussian Splatting (3DGS) has emerged as a leading technology for high-quality 3D scene reconstruction. However, the iterative refinement and densification process leads to the generation of a large number of primitives, each contributing to the reconstruction to a substantially different extent.

Cited by 0SourcecodeScholar
2025

A Hierarchical Compression Technique for 3D Gaussian Splatting Compression

ICASSP 2025accepted

3D Gaussian Splatting (GS) demonstrates excellent rendering quality and generation speed in novel view synthesis. However, substantial data size poses challenges for storage and transmission, making 3D GS compression an essential technology. Current 3D GS compression research primarily focuses on de…

Cited by 0SourceScholar
2025

A Quality-Aware Sampling Framework for Efficient 3D Point Cloud Transmission

ICASSP 2025accepted

The large volume of data from the point cloud brings significant demands on network bandwidth. However, the current transmission framework only considers using lossy compression to control the size of data, while ignoring visually redundant information due to the setting of rendering devices. Based…

Cited by 0SourceScholar
2025

ADC-GS: Anchor-Driven Deformable and Compressed Gaussian Splatting for Dynamic Scene Reconstruction

IJCAI 2025

Existing 4D Gaussian Splatting methods rely on per-Gaussian deformation from a canonical space to target frames, which overlooks redundancy among adjacent Gaussian primitives and result in suboptimal performance. To address this limitation, we propose Anchor-Driven Deformable and Compressed Gaussian

2025

Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-to-3D Generation

ICCV 2025poster

Text-to-3D generation has achieved remarkable progress in recent years, yet evaluating these methods remains challenging for two reasons: i) existing benchmarks lack fine-grained evaluation on different prompt categories and evaluation dimensions; ii) previous evaluation metrics only focus on a sing…

Cited by 0SourcePDFScholar
2025

CushionCatch: A Compliant Catching Mechanism for Mobile Manipulators via Combined Optimization and Learning

IROS 2025

Catching flying objects with a cushioning process is a skill commonly performed by humans, yet it remains a significant challenge for robots. In this paper, we present a framework that combines optimization and learning to achieve compliant catching on mobile manipulators (CCMM). First, we propose a

Cited by 1SourceScholar
2025

HybridGS: High-Efficiency Gaussian Splatting Data Compression using Dual-Channel Sparse Representation and Point Cloud Encoder

ICML 2025poster

Most existing 3D Gaussian Splatting (3DGS) compression schemes focus on producing compact 3DGS representation via implicit data embedding. They have long encoding and decoding times and highly customized data format, making it difficult for widespread deployment. This paper presents a new 3DGS compr…

2025

KD-RIEKF: Kinodynamic Right-Invariant EKF for Legged Robot State Estimation

IROS 2025

We present KD-RIEKF, a novel state estimation framework that incorporates kinodynamic constraints into the Right-Invariant Extended Kalman Filter (RIEKF). Our framework integrates generalized momentum-based contact estimation, centroidal dynamics, and a noise-adaptive module, improving state estimat

Cited by 0SourceScholar
2025

Keypoint-Aware RAG for Robotic Manipulation: In-Context Constraint Learning via Large-Scale Retrieval

IROS 2025

Recent advances in robotic manipulation leverage foundation models pre-trained on internet-scale data, where keypoint-based representations have shown promising results in spatial reasoning. However, existing approaches primarily focus on zero-shot generalization or human-collected demonstrations, w

Cited by 0SourcecodeScholar
2025

Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

NeurIPS 2025poster

The task of Knowlegde-Based Visual Question Answering (KB-VQA) requires the model to understand visual features and retrieve external knowledge. Retrieval-Augmented Generation (RAG) have been employed to address this problem through knowledge base querying. However, existing work demonstrate two lim…

Cited by 0SourceScholar
2025

LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry Compression

ICCV 2025poster

Existing AI-based point cloud compression methods struggle with dependence on specific training data distributions, which limits their real-world deployment. Implicit Neural Representation (INR) methods solve the above problem by encoding overfitted network parameters to the bitstream, resulting in…

Cited by 0SourcePDFScholar
2025

Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger

ICML 2025spotlight

Recent advancements in Large Vision Language Models (LVLMs) have significantly improved performance in Visual Question Answering (VQA) tasks through multimodal Retrieval-Augmented Generation (RAG). However, existing methods still face challenges, such as the scarcity of knowledge with reasoning exam…

Cited by 0SourcePDFScholar
2025

Sarcasm-R1: Enhancing Sarcasm Detection through Focused Reasoning

EMNLP 2025

Sarcasm detection is a crucial yet challenging task in natural language processing. Existing methods primarily rely on supervised learning or prompt engineering, which often struggle to capture the complex reasoning process required for effective sarcasm detection. This paper proposes a novel approa

2024

Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment

CVPR 2024poster

No-reference point cloud quality assessment (NR-PCQA) aims to automatically evaluate the perceptual quality of distorted point clouds without available reference which have achieved tremendous improvements due to the utilization of deep neural networks. However learning-based NR-PCQA methods suffer…

Cited by 17SourcePDFScholar
2024

Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation

CVPR 2024highlight

Recently an audio-visual segmentation (AVS) task has been introduced aiming to group pixels with sounding objects within a given video. This task necessitates a first-ever audio-driven pixel-level understanding of the scene posing significant challenges. In this paper we propose an innovative audio-…

2024

Exploring the Capability of Multimodal LLMs with Yonkoma Manga: The YManga Dataset and Its Challenging Tasks

EMNLP 2024finding

Yonkoma Manga, characterized by its four-panel structure, presents unique challenges due to its rich contextual information and strong sequential features. To address the limitations of current multimodal large language models (MLLMs) in understanding this type of data, we create a novel dataset nam…

2024

SJTU-TMQA: A Quality Assessment Database for Static Mesh with Texture Map

ICASSP 2024accepted

In recent years, static meshes with texture maps have become one of the most prevalent digital representations of 3D shapes in various applications, such as animation, gaming, medical imaging, and cultural heritage applications. However, little research has been done on the quality assessment of tex…

Cited by 0SourceScholar
2024

Text Fluoroscopy: Detecting LLM-Generated Text through Intrinsic Features

EMNLP 2024main

Large language models (LLMs) have revolutionized the domain of natural language processing because of their excellent performance on various tasks. Despite their impressive capabilities, LLMs also have the potential to generate texts that pose risks of misuse. Consequently, detecting LLM-generated t…

Cited by 3SourcePDFScholar
2023

A Small-Scale Untethered Tensegrity Robot With High Velocity and Multi Locomotion Modes

RA-L 2023

Tensegrity mobile robots are well-appraised for their high stiffness-to-mass ratio and superior structural compliance. However, traditional untethered tensegrity mobile robots usually have low velocity due to the large actuation force required by coupling effects among stiff struts and soft cables.

Cited by 15SourceScholar
2023

Orca: A Few-shot Benchmark for Chinese Conversational Machine Reading Comprehension

EMNLP 2023long findings

The conversational machine reading comprehension (CMRC) task aims to answer questions in conversations, which has been a hot research topic in recent years because of its wide applications. However, existing CMRC benchmarks in which each conversation is assigned a static passage are inconsistent wit…

Cited by 0SourcecodeScholar
2022

No-Reference Point Cloud Quality Assessment via Domain Adaptation

CVPR 2022poster

We present a novel no-reference quality assessment metric, the image transferred point cloud quality assessment (IT-PCQA), for 3D point clouds. For quality assessment, deep neural network (DNN) has shown compelling performance on no-reference metric design. However, the most challenging issue for no…

Cited by 101PDFcodeScholar