← Search

chi zhang

197 accepted papers

2026

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

ICLR 2026poster

Vision-Language-Action (VLA) models pre-trained on large, diverse datasets show remarkable potential for general-purpose robotic manipulation. However, a primary bottleneck remains in adapting these models to downstream tasks, especially when the robot's embodiment or the task itself differs from th…

Cited by 0SourcecodeScholar
2026

BrainJanus: A Foundation Model for Unified Understanding and Generation across Brain, Vision, and Language

ICML 2026poster

Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding and decoding as isolated tasks, relying heavily on unimodal alignment and extern…

Cited by 0SourceScholar
2026

CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free Fusion

CVPR 2026

Personalized image generation requires effectively balancing content fidelity with stylistic consistency when synthesizing images based on text and reference examples. Low-Rank Adaptation (LoRA) offers an efficient personalization approach, with potential for precise control through combining LoRA w

Cited by 0SourcecodeScholar
2026

Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation

CVPR 2026

Text-to-Image (T2I) generation has achieved remarkable progress in recent years. Meanwhile, reinforcement learning methods, particularly those based on Group Relative Policy Optimization (GRPO), have attracted widespread attention and been successfully applied to T2I tasks. However, the uniform samp

Cited by 0SourcecodeScholar
2026

DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use

ICML 2026poster

Recent work increasingly synthesizes agentic tasks for post-training tool-using LLMs, yet robust generalization under shifts in tasks and toolsets remains an open challenge. We trace this brittleness to insufficient diversity in synthesized training tasks. Scaling diversity is difficult because trai…

Cited by 0SourceScholar
2026

DiffQ: UNIFIED PARAMETER INITIALIZATION FOR VARIATIONAL QUANTUM ALGORITHMS VIA DIFFUSION MODELS

ICASSP 2026oral

Variational Quantum Algorithms (VQAs) are widely used in the noisy intermediate-scale quantum (NISQ) era, but their trainability and performance depend critically on initialization parameters that shape the optimization landscape. Existing machine learning-based initializers achieve state-of-the-art…

Cited by 0SourcePDFScholar
2026

EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual Alignment

CVPR 2026

Recent Vision-Language-Action (VLA) models map visual-textual inputs to robotic actions via end-to-end architectures, yet this approach entangles visual understanding with task-specific actions. This leads to an exhaustive collection of full operational sequences and parameter redundancy across task

Cited by 0SourceScholar
2026

Evaluating, Synthesizing, and Enhancing for Customer Support Conversation

AAAI 2026technical

Effective customer support requires not only accurate problem-solving but also structured and empathetic communication aligned with professional standards. However, existing dialogue datasets often lack strategic guidance, and real-world service data is difficult to access and annotate. To address t

Cited by 0SourcePDFScholar
2026

Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven Consistency

AAAI 2026technical

Super-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applicatio

Cited by 0SourcePDFScholar
2026

FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning

ICLR 2026poster

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for enhancing the reasoning capabilities of large language models (LLMs). In this context, models explore reasoning trajectories and exploit rollouts with correct answers as positive signals for policy optimiz…

Cited by 0SourceScholar
2026

FINMCP-BENCH: BENCHMARKING LLM AGENTS FOR REAL-WORLD FINANCIAL TOOL USE UNDER THE MODEL CONTEXT PROTOCOL

ICASSP 2026poster

This paper introduces \textbf{FinMCP-Bench}, a novel benchmark for evaluating large language models (LLMs) in solving real-world financial problems through tool invocation of financial model context protocols. FinMCP-Bench contains 613 samples spanning 10 main scenarios and 33 sub-scenarios, featuri…

Cited by 0SourcePDFScholar
2026

FW-VTON: FLATTENING-AND-WARPING FOR PERSON-TO-PERSON VIRTUAL TRY-ON

ICASSP 2026poster

Traditional virtual try-on methods primarily focus on the garment-to-person try-on task, which requires flat garment representations. In contrast, this paper introduces a novel approach to the person-to-person try-on task. Unlike the garment-to-person try-on task, the person-to-person task only invo…

Cited by 0SourcePDFScholar
2026

Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration

CVPR 2026

Diffusion models have achieved impressive generative quality across modalities like 2D images, videos, and 3D shapes, but their inference remains computationally expensive due to the iterative denoising process. While recent caching-based methods effectively reuse redundant computations to speed up

Cited by 0SourceScholar
2026

Few-Step Diffusion Sampling Through Instance-Aware Discretizations

CVPR 2026

Diffusion and flow matching models generate high-fidelity data by simulating paths defined by Ordinary or Stochastic Differential Equations (ODEs/SDEs), starting from a tractable prior distribution. The probability flow ODE formulation enables the use of advanced numerical solvers to accelerate samp

Cited by 0SourceScholar
2026

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models

IJCAI 2026

Process Reward Models (PRMs) supervise intermediate reasoning steps in large language models (LLMs), but existing PRMs are mainly trained on general-domain data and struggle with the structured, symbolic, and fact-sensitive nature of financial reasoning. Financial tasks require not only correct fina

Cited by 0Scholar
2026

FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing

CVPR 2026

Text-driven video editing aims to modify video content based on natural language instructions. While recent training-free methods have leveraged pretrained diffusion models, they often rely on an inversion-editing paradigm. This paradigm maps the video to a latent space before editing. However, the

Cited by 0SourcecodeScholar
2026

Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context

CVPR 2026

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory for consistency, or iterative 3D reconstruction and inpainting, which accumulate errors during inference due to incorrect

Cited by 0SourceScholar
2026

HAMLET: Hyperadaptive Agent-based Modeling for Live Embodied Theatrics

ICLR 2026poster

Creating an immersive and interactive theatrical experience is a long-term goal in the field of interactive narrative. The emergence of large language model (LLM) is providing a new path to achieve this goal. However, existing LLM-based drama generation methods often result in agents that lack initi…

Cited by 0SourcecodeScholar
2026

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

ICLR 2026poster

The pursuit of human-like conversational agents has long been guided by the Turing test. For modern speech-to-speech (S2S) systems, a critical yet unanswered question is whether they can converse like humans. To tackle this, we conduct the first Turing test for S2S systems, collecting 2,968 human ju…

Cited by 0SourceScholar
2026

IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation

CVPR 2026

Dataset Distillation aims to synthesize compact datasets that can approximate the training efficacy of large-scale real datasets, offering an efficient solution to the increasing computational demands of modern deep learning. Recently, diffusion-based dataset distillation methods have shown great pr

Cited by 0SourcecodeScholar
2026

Improving Diffusion Generalization with Weak-to-Strong Segmented Guidance

CVPR 2026

Diffusion models generate synthetic images through an iterative refinement process. However, the misalignment between the simulation-free objective and the iterative process often causes accumulated gradient error along the sampling trajectory, which leads to unsatisfactory results and a failure to

Cited by 0SourcecodeScholar
2026

Informative Subgraph Extraction with Deep Reinforcement Learning for Drug-Drug Interaction Prediction

AAAI 2026technical

Drug-drug interaction (DDI) prediction is pivotal for drug safety and clinical decision-making. Recently, subgraph-based methods utilizing knowledge graphs (KGs) and domain information have achieved promising results by extracting informative subgraphs for DDI prediction. However, existing subgraph

Cited by 0SourcePDFScholar
2026

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning

AAAI 2026technical

News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances, existing methods struggle with three key challenges: (1) incomplete information coverage, (2) weak cross-modal alignment

Cited by 0SourcePDFScholar
2026

Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force Fields

ICLR 2026poster

Predicting physical dynamics from raw visual data remains a major challenge in AI. While recent video generation models have achieved impressive visual quality, they still cannot consistently generate physically plausible videos due to a lack of modeling of physical laws. Recent approaches combining…

Cited by 0SourcecodeScholar
2026

Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation

CVPR 2026

Group Relative Policy Optimization (GRPO) has emerged as an effective and lightweight framework for post-training visual generative models. However, its performance is fundamentally limited by the ambiguity of textual-visual correspondence: a single prompt may validly describe diverse visual outputs

Cited by 0SourceScholar
2026

LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization

ICLR 2026poster

Generating coherent and communicative visual sequences, such as image sequences and videos, remains a significant challenge for current multimodal systems. Despite advances in visual quality and the integration of world knowledge, existing models still struggle to maintain logical flow, often result…

Cited by 0SourceScholar
2026

MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation

ICLR 2026poster

Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits generalization and discards essential 4D information for open-world animation. To address this, we propose MTVCraft (Mot…

Cited by 0SourcecodeScholar
2026

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

ICML 2026poster

Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static heuristics or stochastic search, rendering them brittle against advanced safety alignment. To address this, we introduce…

Cited by 0SourceScholar
2026

Neural Force Field: Few-shot Learning of Generalized Physical Reasoning

ICLR 2026poster

Physical reasoning is a remarkable human ability that enables rapid learning and generalization from limited experience. Current AI models, despite extensive training, still struggle to achieve similar generalization, especially in Out-of-distribution (OOD) settings. This limitation stems from their…

Cited by 0SourcecodeScholar
2026

Not All Documents Are What You Need for Extracting Instruction Tuning Data

ICLR 2026poster

Instruction tuning improves the LLMs performance but depends on high-quality training data. Recently, LLMs have been used to synthesize data, enhancing training with seeds like question-answer (QA) pairs. However, this synthesis often results in instruction examples similar to the seeds, lacking div…

Cited by 0SourceScholar
2026

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

AAAI 2026technical

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, whi

Cited by 0SourcePDFScholar
2026

Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs) and is now being applied to Vision-Language Models (VLMs). However, vanilla RLVR for VLMs verifies only the final textual output, critically neglecting the foun

Cited by 0SourcecodeScholar
2026

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation

CVPR 2026

Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, existing GRPO pipelines rely on a single scalar reward per sample, treating each im

Cited by 0SourceScholar
2026

Signal Structure-Aware Gaussian Splatting for Large-Scale Scene Reconstruction

ICLR 2026poster

3D Gaussian Splatting has demonstrated remarkable potential in novel view synthesis. In contrast to small-scale scenes, large-scale scenes inevitably contain sparsely observed regions with excessively sparse initial points. In this case, supervising Gaussians initialized from low-frequency sparse po…

Cited by 0SourceScholar
2026

Streaming Generated Gaussian Process Experts for Online Learning and Control

AAAI 2026technical

Gaussian Processes (GPs), as a nonparametric learning method, offer flexible modeling capabilities and calibrated uncertainty quantification for function approximations. Additionally, GPs support online learning by efficiently incorporating new data with polynomial-time computation, making them well

Cited by 0SourcePDFScholar
2026

SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls

CVPR 2026

Recent advances in text-to-video diffusion models have enabled high-fidelity and temporally coherent video synthesis. However, current models are predominantly optimized for single-event generation. When handling multi-event prompts, without explicit temporal grounding, such models often produce ble

Cited by 0SourcecodeScholar
2026

Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control

CVPR 2026

Video diffusion models have rich world priors, but their use in spatial tasks is limited by poor control, spatial-temporal inconsistent results, and entangled scene-camera dynamics. Current approaches, such as per-task fine-tuning or post-process warping, often introduce visual artifacts, fail to ge

Cited by 0SourcecodeScholar
2026

Tea-Adapter: Teacher Adapter for Efficient Conditional Generation

CVPR 2026

We propose Tea-Adapter, a plug-and-play adapter designed to efficiently integrate conditional knowledge from a smaller teacher model into a larger student video diffusion model. Existing controllable video DiT methods face critical challenges: full fine-tuning of billion-parameter models is extremel

Cited by 0SourceScholar
2026

TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction

CVPR 2026

We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details

Cited by 0SourceScholar
2026

TsLLM: Augmenting LLMs for General Time Series Understanding and Prediction

ICML 2026poster

Time series data is fundamental to decision-making across many domains including healthcare, finance, power systems, and logistics. However, analyzing this data correctly often requires incorporating unstructured contextual information, answering domain-specific questions, and generating natural lan…

Cited by 0SourceScholar
2026

VQEzy: AN OPEN-SOURCE DATASET FOR PARAMETER INITIALIZATION IN VARIATIONAL QUANTUM EIGENSOLVERS

ICASSP 2026oral

Variational Quantum Eigensolvers (VQEs) are a leading class of noisy intermediate-scale quantum (NISQ) algorithms, whose performance is highly sensitive to parameter initialization. Although recent machine learning-based initialization methods have achieved state-of-the-art performance, their progre…

Cited by 0SourcePDFScholar
2026

ViStoryBench: Comprehensive Benchmark Suite for Story Visualization

CVPR 2026

Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references. Despite progress in generative models, existing benchmarks remain narrow in scope, often limited to short prompts, lacking character references, or single-image c

Cited by 0SourcecodeScholar
2026

When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models

AAAI 2026technical

Multimodal Large Language Models (MLLMs) are susceptible to the implicit reasoning risk, wherein innocuous unimodal inputs synergistically assemble into risky multimodal data that produce harmful outputs. We attribute this vulnerability to the difficulty of MLLMs maintaining safety alignment through

Cited by 0SourcePDFScholar
2025

A Multi-Modal Fusion-Based 3D Multi-Object Tracking Framework With Joint Detection

RA-L 2025

In the classical tracking-by-detection (TBD) paradigm, detection and tracking are separately and sequentially conducted, and data association must be properly performed to achieve satisfactory tracking performance. In this letter, a new multi-object tracking framework is proposed, which integrates o

Cited by 18SourceScholar
2025

Adaptive Stochastic Coefficients for Accelerating Diffusion Sampling

NeurIPS 2025poster

Diffusion-based generative processes, formulated as differential equation solving, frequently balance computational speed with sample quality. Our theoretical investigation of ODE- and SDE-based solvers reveals complementary weaknesses: ODE solvers accumulate irreducible gradient error along de…

Cited by 0SourcecodeScholar
2025

Adversarial Locomotion and Motion Imitation for Humanoid Policy Learning

NeurIPS 2025poster

Humans exhibit diverse and expressive whole-body movements. However, attaining human-like whole-body coordination in humanoid robots remains challenging, as conventional approaches that mimic whole-body motions often neglect the distinct roles of upper and lower body. This oversight leads to computa…

Cited by 0SourcecodeScholar
2025

An Automatic Method to Estimate Correctness of RAG

COLING 2025industry

In sectors in where data quality is critical, like finance and healthcare, it is crucial to have confidence in not only the outputs generated by retrieval-augmented generation (RAG) models but also the process followed by the model while arriving at the output. Existing methods, such as hallucinatio…

Cited by 2SourcePDFScholar
2025

BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals

NeurIPS 2025poster

Electroencephalography (EEG) and magnetoencephalography (MEG) measure neural activity non-invasively by capturing electromagnetic fields generated by dendritic currents. Although rooted in the same biophysics, EEG and MEG exhibit distinct signal patterns, further complicated by variations in sensor…

Cited by 0SourcecodeScholar
2025

CADCrafter: Generating Computer-Aided Design Models from Unconstrained Images

CVPR 2025poster

Creating CAD digital twins from the physical world is crucial for manufacturing, design, and simulation. However, current methods typically rely on costly 3D scanning with labor-intensive post-processing. To provide a user-friendly design process, we explore the problem of reverse engineering from u…

Cited by 3SourcePDFScholar
2025

CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs

NeurIPS 2025poster

Speculative decoding has become a widely adopted as an effective technique for lossless inference acceleration when deploying large language models (LLMs). While on-the-fly self-speculative methods offer seamless integration and broad utility, they often fall short of the speed gains achieved by met…

Cited by 0SourceScholar
2025

Creative Agents: Empowering Agents with Imagination for Creative Tasks

UAI 2025

We study building embodied agents for open-ended creative tasks. While existing methods build instruction-following agents that can perform diverse open-ended tasks, none of them demonstrates creativity – the ability to give novel and diverse solutions implicit in the language instructions. This lim

2025

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

NeurIPS 2025poster

Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the…

Cited by 0SourceScholar
2025

Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models

ICCV 2025poster

Diffusion models (DMs) have achieved state-of-the-art generative performance but suffer from high sampling latency due to their sequential denoising nature. Existing solver-based acceleration methods often face image quality degradation under a low-latency budget. In this paper, we propose the Ensem…

2025

Divide, Optimize, Merge: Scalable Fine-Grained Generative Optimization for LLM Agents

EMNLP 2025

LLM-based optimization has shown remarkable potential in improving agentic systems. However, the conventional approach of prompting LLM-based generative optimizer with the trajectories on the whole training dataset in a single pass becomes untenable as datasets grow, leading to context window overfl

Cited by 0SourcePDFScholar
2025

DualOpt: A Dual Divide-and-Optimize Algorithm for the Large-scale Traveling Salesman Problem

AAAI 2025technical

This paper proposes a dual divide-and-optimize algorithm (DualOpt) for solving the large-scale traveling salesman problem (TSP). DualOpt combines two complementary strategies to improve both solution quality and computational efficiency. The first strategy is a grid-based divide-and-conquer procedur…

2025

Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration

ACL 2025long

Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tack…

Cited by 0SourcePDFScholar
2025

Enhanced Kinematic Calibration of a 4PPa-2PaR Parallel Manipulator with Subchains

IROS 2025

This paper proposes an innovative virtual chain-based kinematic calibration for the 4PPa-2PaR parallel manipulators with subchain architectures. Conventional calibration methods for such architectures suffer from inherent limitations due to coupled parameter constraints and restricted solution space

Cited by 0SourceScholar
2025

From Weight-Based to State-Based Fine-Tuning: Further Memory Reduction on LoRA with Parallel Control

ICML 2025oral

The LoRA method has achieved notable success in reducing GPU memory usage by applying low-rank updates to weight matrices. Yet, one simple question remains: can we push this reduction even further? Furthermore, is it possible to achieve this while improving performance and reducing computation time?…

Cited by 0SourcePDFScholar
2025

Generalized and Invariant Single-Neuron In-Vivo Activity Representation Learning

NeurIPS 2025poster

In computational neuroscience, models representing single-neuron in-vivo activity have become essential for understanding the functional identities of individual neurons. These models, such as implicit representation methods based on Transformer architectures, contrastive learning frameworks, and va…

Cited by 0SourceScholar
2025

Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization

NeurIPS 2025poster

Recent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noisy labels, facing limitations such as high computational costs, heavy hyperparameter tuning process, and coarse-grained…

Cited by 0SourcecodeScholar
2025

Harnessing Diversity for Important Data Selection in Pretraining Large Language Models

ICLR 2025spotlight

Data selection is of great significance in pretraining large language models, given the variation in quality within the large-scale available training corpora. To achieve this, researchers are currently investigating the use of data influence to measure the importance of data instances, $i.e.,$ a…

Cited by 8SourcePDFScholar
2025

Heterogeneous Adversarial Play in Interactive Environments

NeurIPS 2025poster

Self-play constitutes a fundamental paradigm for autonomous skill acquisition, whereby agents iteratively enhance their capabilities through self-directed environmental exploration. Conventional self-play frameworks exploit agent symmetry within zero-sum competitive settings, yet this approach prove…

Cited by 0SourceScholar
2025

Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models

AAAI 2025technical

Diffusion models have revitalized the image generation domain, playing crucial roles in both academic research and artistic expression. With the emergence of new diffusion models, assessing the performance of text-to-image models has become increasingly important. Current metrics focus on directly m…

2025

KINDLE: Knowledge-Guided Distillation for Prior-Free Gene Regulatory Network Inference

NeurIPS 2025poster

Gene regulatory network (GRN) inference serves as a cornerstone for deciphering cellular decision-making processes. Early approaches rely exclusively on gene expression data, thus their predictive power remain fundamentally constrained by the vast combinatorial space of potential gene-gene interacti…

Cited by 0SourceScholar
2025

MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh Tokenization

ICCV 2025poster

Meshes are the de facto 3D representation in the industry but are labor-intensive to produce. Recently, a line of research has focused on autoregressively generating meshes. This approach processes meshes into a sequence composed of vertices and then generates them vertex by vertex, similar to how a…

2025

MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers

ICLR 2025poster

Recently, 3D assets created via reconstruction and generation have matched the quality of manually crafted assets, highlighting their potential for replacement. However, this potential is largely unrealized because these assets always need to be converted to meshes for 3D industry applications, and…

2025

Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models

ACL 2025long

The composition of pre-training datasets for large language models (LLMs) remains largely undisclosed, hindering transparency and efforts to optimize data quality—a critical driver of model performance. Current data selection methods, such as natural language quality assessments, diversity-based fil…

2025

Model-Free Offline Reinforcement Learning with Enhanced Robustness

ICLR 2025poster

Offline reinforcement learning (RL) has gained considerable attention for its ability to learn policies from pre-collected data without real-time interaction, which makes it particularly useful for high-risk applications. However, due to its reliance on offline datasets, existing works inevitably in…

Cited by 0SourcePDFScholar
2025

Monocular Depth Estimation and Segmentation for Transparent Object with Iterative Semantic and Geometric Fusion

ICRA 2025

Transparent object perception is indispensable for numerous robotic tasks. However, accurately segmenting and estimating the depth of transparent objects remain challenging due to complex optical properties. Existing methods primarily delve into only one task using extra inputs or specialized sensor

Cited by 8SourcecodeScholar
2025

MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent

ICCV 2025poster

We propose MotionAgent, enabling fine-grained motion control for text-guided image-to-video generation. The key technique is the motion field agent that converts motion information in text prompts into explicit motion fields, providing flexible and precise motion guidance. Specifically, the agent ex…

2025

NFIG: Multi-Scale Autoregressive Image Generation via Frequency Ordering

NeurIPS 2025poster

Autoregressive models have achieved significant success in image generation. However, unlike the inherent hierarchical structure of image information in the spectral domain, standard autoregressive methods typically generate pixels sequentially in a fixed spatial order. To better leverage this spect…

Cited by 0SourceScholar
2025

Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning

ICML 2025poster

Transfer reinforcement learning aims to derive a near-optimal policy for a target environment with limited data by leveraging abundant data from related source domains. However, it faces two key challenges: the lack of performance guarantees for the transferred policy, which can lead to undesired ac…

Cited by 0SourcePDFScholar
2025

Project-Probe-Aggregate: Efficient Fine-Tuning for Group Robustness

CVPR 2025highlight

While image-text foundation models have succeeded across diverse downstream tasks, they still face challenges in the presence of spurious correlations between the input and label. To address this issue, we propose a simple three-step approach-Project-Probe-Aggregate (PPA)-that enables parameter-effi…

Cited by 0SourcePDFScholar
2025

Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense Game

CVPR 2025poster

Multi-exit neural networks represent a promising approach to enhancing model inference efficiency, yet like common neural networks, they suffer from significantly reduced robustness against adversarial attacks. While some defense methods have been raised to strengthen the adversarial robustness of m…

Cited by 0SourcePDFScholar
2025

RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

NeurIPS 2025poster

Spatial referring is a fundamental capability of embodied robots to interact with the 3D physical world. However, even with the powerful pretrained VLMs, recent approaches are still not qualified to accurately understand the complex 3D scenes and dynamically reason about the instruction-indicated lo…

Cited by 0SourceScholar
2025

Semantic Graph Embedded Energy Minimization Learning for Scene Graph Generation

ICASSP 2025accepted

The performance of current scene graph generation models is affected by training with cross-entropy loss, exacerbating the problem of prediction bias stemming from biased training data. Energy-based model adopts a learning method for joint image and scene graph to alleviate this challenge. However,…

Cited by 0SourceScholar
2025

Sensor-Based Adaptive Robust Torque Control for Flexible Joints

RA-L 2025

Achieving fast transient response and high steady-state tracking accuracy of joint torque for flexible joint systems with dynamic constraints, various parameter uncertainties, and uncertain nonlinearities is always challenging. To this end, a torque control scheme based on multi-torque sensors is pr

Cited by 2SourceScholar
2025

Sim4Rec: Data-Free Model Extraction Attack on Sequential Recommendation

AAAI 2025technical

Model extraction attack shows promising performance in revealing sequential recommendation (SeqRec) robustness, e.g., as an upstream task of transfer-based attack to provide optimization feedback for downstream attacks. However, existing work either heavily relies on impractical prior knowledge or h…

Cited by 0SourcePDFScholar
2025

Style Nursing with Spatial and Semantic Guidance for Zero-Shot Traffic Scene Style Transfer

AAAI 2025technical

Recent advances in text-to-image diffusion models have shown an outstanding ability in zero-shot style transfer. However, existing methods often struggle to balance preserving the semantic content of the input image and faithfully transferring the target style in line with the edit prompt. Especiall…

Cited by 0SourcePDFScholar
2025

StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

CVPR 2025poster

Text-driven style transfer aims to merge the style of a reference image with content described by a text prompt. Recent advancements in text-to-image models have improved the nuance of style transformations, yet significant challenges remain, particularly with overfitting to reference styles, limit…

Cited by 0SourcePDFScholar
2025

SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control

AAAI 2025technical

Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production…

Cited by 11SourcePDFScholar
2025

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond

NeurIPS 2025poster

Recent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for…

Cited by 0SourcecodeScholar
2025

Synthesizing Images on Perceptual Boundaries of ANNs for Uncovering and Manipulating Human Perceptual Variability

ICML 2025poster

Human decision-making in cognitive tasks and daily life exhibits considerable variability, shaped by factors such as task difficulty, individual preferences, and personal experiences. Understanding this variability across individuals is essential for uncovering the perceptual and decision-making me…

Cited by 0SourcePDFScholar
2025

Towards Reliable LLM-based Robots Planning via Combined Uncertainty Estimation

NeurIPS 2025poster

Large language models (LLMs) demonstrate advanced reasoning abilities, enabling robots to understand natural language instructions and generate high-level plans with appropriate grounding. However, LLM hallucinations present a significant challenge, often leading to overconfident yet potentially mis…

Cited by 0SourcecodeScholar
2025

UniScene: Unified Occupancy-centric Driving Scene Generation

CVPR 2025poster

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles…

2025

VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking

NeurIPS 2025poster

Edge deployment of large Vision-Language Models (VLMs) increasingly relies on flash-based weight offloading, where activation sparsification is used to reduce I/O overhead. However, conventional sparsification remains model-centric, selecting neurons solely by activation magnitude and neglecting how…

Cited by 0SourceScholar
2025

Video-Bench: Human-Aligned Video Generation Benchmark

CVPR 2025poster

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories: traditional benchmarks, which use metrics and embeddings to evaluate…

2024

A Novel Miniature Flexible Instrument With Unfolding and Decoupling Design for Endoscopic Surgery

RA-L 2024

Nowadays, gastrointestinal cancer has widely impacted people's health worldwide due to its high mortality rate. Early treatment of gastrointestinal cancer by endoscopic procedure can greatly increase survival rates of patients. Nevertheless, current flexible endoscopic instruments lack of degree of

Cited by 7SourceScholar
2024

A Novel Robot Platform With Decoupled Stiffness Control for Endoscopic Surgery

RA-L 2024

Endoscopic robot has garnered significant attention for its ability to offer auxiliary traction and precise maneuverability. However, the low stiffness of its insertion tube makes it susceptible to deformation, which poses great challenges for precise control. Existing variable stiffness technologie

Cited by 3SourceScholar
2024

Adaptive Hardness Negative Sampling for Collaborative Filtering

AAAI 2024technical

Negative sampling is essential for implicit collaborative filtering to provide proper negative training signals so as to achieve desirable performance. We experimentally unveil a common limitation of all existing negative sampling methods that they can only select negative samples of a fixed hardnes…

2024

Bias-aware Boolean Matrix Factorization Using Disentangled Representation Learning

UAI 2024poster

Boolean matrix factorization (BMF) has been widely utilized in fields such as recommendation systems, graph learning, text mining, and -omics data analysis. Traditional BMF methods decompose a binary matrix into the Boolean product of two lower-rank Boolean matrices plus homoscedastic random errors.…

2024

Do Finetti: On Causal Effects for Exchangeable Data

NeurIPS 2024oral

We study causal effect estimation in a setting where the data are not i.i.d.$\ $(independent and identically distributed). We focus on exchangeable data satisfying an assumption of independent causal mechanisms. Traditional causal effect estimation frameworks, e.g., relying on structural causal mode…

Cited by 1SourcePDFScholar
2024

Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

ACL 2024findings

The remarkable multimodal capabilities demonstrated by OpenAI’s GPT-4 have sparked significant interest in the development of multimodal Large Language Models (LLMs). A primary research objective of such models is to align visual and textual modalities effectively while comprehending human instructi…

2024

FipTR: A Simple yet Effective Transformer Framework for Future Instance Prediction in Autonomous Driving

ECCV 2024poster

"The future instance prediction from a Bird’s Eye View(BEV) perspective is a vital component in autonomous driving, which involves future instance segmentation and instance motion prediction. Existing methods usually rely on a redundant and complex pipeline which requires multiple auxiliary outputs…

2024

GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting

CVPR 2024poster

3D editing plays a crucial role in many areas such as gaming and virtual reality. Traditional 3D editing methods which rely on representations like meshes and point clouds often fall short in realistically depicting complex scenes. On the other hand methods based on implicit 3D representations like…

2024

IT3D: Improved Text-to-3D Generation with Explicit View Synthesis

AAAI 2024technical

Recent strides in Text-to-3D techniques have been propelled by distilling knowledge from powerful large text-to-image diffusion models (LDMs). Nonetheless, existing Text-to-3D approaches often grapple with challenges such as over-saturation, inadequate detailing, and unrealistic outputs. This study…

2024

LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding Reasoning and Planning

CVPR 2024poster

Recent progress in Large Multimodal Models (LMM) has opened up great possibilities for various applications in the field of human-machine interactions. However developing LMMs that can comprehend reason and plan in complex and diverse 3D environments remains a challenging topic especially considerin…

2024

Learn to Optimize Denoising Scores: A Unified and Improved Diffusion Prior for 3D Generation

ECCV 2024poster

"In this paper, we propose a unified framework aimed at enhancing the diffusion priors for 3D generation tasks. Despite the critical importance of these tasks, existing methodologies often struggle to generate high-caliber results. We begin by examining the inherent limitations in previous diffusion…

2024

Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models

NeurIPS 2024poster

As Archimedes famously said, ``Give me a lever long enough and a fulcrum on which to place it, and I shall move the world'', in this study, we propose to use a tiny Language Model (LM), \eg, a Transformer with 67M parameters, to lever much larger Vision-Language Models (LVLMs) with 9B parameters. Sp…

Cited by 8SourcePDFScholar
2024

M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions

ECCV 2024poster

"Recently, the understanding of the 3D world has garnered increased attention, facilitating autonomous agents to perform further decision-making. However, the majority of existing 3D vision-language datasets and methods are often limited to specific tasks, limiting their applicability in diverse sce…

Cited by 0SourcePDFScholar
2024

MotionChain: Conversational Motion Controllers via Multimodal Prompts

ECCV 2024poster

"Recent advancements in language models have demonstrated their adeptness in conducting multi-turn dialogues and retaining conversational context. However, this proficiency remains largely unexplored in other multimodal generative models, particularly in human motion models. By integrating multi-tur…

2024

Panacea: Panoramic and Controllable Video Generation for Autonomous Driving

CVPR 2024poster

The field of autonomous driving increasingly demands high-quality annotated training data. In this paper we propose Panacea an innovative approach to generate panoramic and controllable videos in driving scenarios capable of yielding an unlimited numbers of diverse annotated samples pivotal for auto…

Cited by 50SourcePDFScholar
2024

Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object Detection

ECCV 2024poster

"Multi-view 3D object detection systems often struggle with generating precise predictions due to the challenges in estimating depth from images, increasing redundant and incorrect detections. Our paper presents Ray Denoising, an innovative method that enhances detection accuracy by strategically sa…

2024

Reinforcement Learning for Athletic Intelligence: Lessons from the 1st “AI Olympics with RealAIGym” Competition

IJCAI 2024poster

As artificial intelligence gains new capabilities, it becomes important to evaluate it on real-world tasks. In particular, the fields of robotics and reinforcement learning (RL) are lacking in standardized benchmarking tasks on real hardware. To facilitate reproducibility and stimulate algorithmi…

Cited by 11SourcePDFScholar
2024

Stream Query Denoising for Vectorized HD-Map Construction

ECCV 2024poster

"This paper introduces the Stream Query Denoising (SQD) strategy, a novel and general approach for high-definition map (HD-map) construction. SQD is designed to improve the modeling capability of map elements by learning temporal consistency. Specifically, SQD involves the process of denoising the q…

Cited by 24SourcePDFScholar
2024

Uncertainty-Guided Physics-Driven Deep Learning Reconstruction via Cyclic Measurement Consistency

ICASSP 2024accepted

Physics-driven deep learning (PD-DL) techniques have recently emerged as a powerful means for improved computational imaging, including in MRI applications. These methods use the physics information by incorporating the known forward model for data fidelity, while performing regularization using neu…

Cited by 0SourceScholar
2023

BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset

NeurIPS 2023poster

In this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of helpfulness and harmlessness for question-answering pairs, thus offering distinct perspectives on these crucial attribu…

Cited by 400SourcePDFScholar
2023

Constrained Policy Optimization with Explicit Behavior Density For Offline Reinforcement Learning

NeurIPS 2023poster

Due to the inability to interact with the environment, offline reinforcement learning (RL) methods face the challenge of estimating the Out-of-Distribution (OOD) points. Existing methods for addressing this issue either control policy to exclude the OOD action or make the $Q$ function pessimistic. H…

2023

Cross-Modal Monocular Localization in Prior LiDAR Maps Utilizing Semantic Consistency

ICRA 2023poster

Visual localization for mobile robots and intelligent vehicles in prior LiDAR maps can achieve high accuracy and low cost. However, algorithms for finding the cross-modal correspondences between images and LiDAR map points are not yet stable. In this paper, we propose a monocular visual localization…

Cited by 14SourceScholar
2023

DPF-Net: Combining Explicit Shape Priors in Deformable Primitive Field for Unsupervised Structural Reconstruction of 3D Objects

ICCV 2023poster

Unsupervised methods for reconstructing structures face significant challenges in capturing the geometric details with consistent structures among diverse shapes of the same category. To address this issue, we present a novel unsupervised structural reconstruction method, named DPF-Net, based on a n…

Cited by 9PDFScholar
2023

Discrepant and Multi-Instance Proxies for Unsupervised Person Re-Identification

ICCV 2023poster

Most recent unsupervised person re-identification methods maintain a cluster uni-proxy for contrastive learning. However, due to the intra-class variance and inter-class similarity, the cluster uni-proxy is prone to be biased and confused with similar classes, resulting in the learned features lacki…

Cited by 39PDFScholar
2023

ESSAformer: Efficient Transformer for Hyperspectral Image Super-resolution

ICCV 2023poster

Single hyperspectral image super-resolution (single-HSI-SR) aims to restore a high-resolution hyperspectral image from a low-resolution observation. However, the prevailing CNN-based approaches have shown limitations in building long-range dependencies and capturing interaction information between s…

Cited by 83PDFcodeScholar
2023

End-to-End Vectorized HD-Map Construction With Piecewise Bezier Curve

CVPR 2023poster

Vectorized high-definition map (HD-map) construction, which focuses on the perception of centimeter-level environmental information, has attracted significant research interest in the autonomous driving community. Most existing approaches first obtain rasterized map with the segmentation-based pipel…

2023

Evaluating and Inducing Personality in Pre-trained Language Models

NeurIPS 2023spotlight

Standardized and quantified evaluation of machine behaviors is a crux of understanding LLMs. In this study, we draw inspiration from psychometric studies by leveraging human personality theory as a tool for studying machine behaviors. Originating as a philosophical quest for human behaviors, the stu…

Cited by 143SourcePDFScholar
2023

Generative Gradient Inversion via Over-Parameterized Networks in Federated Learning

ICCV 2023poster

Federated learning has gained recognitions as a secure approach for safeguarding local private data in collaborative learning. But the advent of gradient inversion research has posed significant challenges to this premise by enabling a third-party to recover groundtruth images via gradients. While p…

Cited by 13PDFcodeScholar
2023

Label-Guided Knowledge Distillation for Continual Semantic Segmentation on 2D Images and 3D Point Clouds

ICCV 2023poster

Continual semantic segmentation (CSS) aims to extend an existing model to tackle unseen tasks while retaining its old knowledge. Naively fine-tuning the old model on new data leads to catastrophic forgetting. A common solution is knowledge distillation (KD), where the output distribution of the new…

Cited by 17PDFcodeScholar
2023

MEWL: Few-shot multimodal word learning with referential uncertainty

ICML 2023poster

Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning…

2023

MacFormer: Map-Agent Coupled Transformer for Real-Time and Robust Trajectory Prediction

RA-L 2023

Predicting the future behavior of agents is a fundamental task in autonomous vehicle domains. Accurate prediction relies on comprehending the surrounding map, which significantly regularizes agent behaviors. However, existing methods have limitations in exploiting the map and exhibit a strong depend

Cited by 79SourceScholar
2023

Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image

ICCV 2023poster

Reconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA) monocular metric depth estimation methods can only handle a sin…

Cited by 189PDFcodeScholar
2023

On the Perils of Cascading Robust Classifiers

ICLR 2023poster

Ensembling certifiably robust neural networks is a promising approach for improving the \emph{certified robust accuracy} of neural models. Black-box ensembles that assume only query-access to the constituent models (and their robustness certifiers) during prediction are particularly attractive due…

2023

PivotNet: Vectorized Pivot Learning for End-to-end HD Map Construction

ICCV 2023poster

Vectorized high-definition map online construction has garnered considerable attention in the field of autonomous driving research. Most existing approaches model changeable map elements using a fixed number of points, or predict local maps in a two-stage autoregressive manner, which may miss essent…

Cited by 84PDFcodeScholar
2023

Robust Geometry-Preserving Depth Estimation Using Differentiable Rendering

ICCV 2023poster

In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate for mix-dataset training, enhancing generalization across dive…

Cited by 6PDFScholar
2023

Sora: Scalable Black-Box Reachability Analyser on Neural Networks

ICASSP 2023accepted

The vulnerability of deep neural networks (DNNs) to input perturbations has posed a significant challenge. Recent work on robustness verification of DNNs not only lacks scalability but also requires severe restrictions on the architecture (layers, activation functions, etc.). To address these limita…

Cited by 0SourceScholar
2023

SustainGym: Reinforcement Learning Environments for Sustainable Energy Systems

NeurIPS 2023poster

The lack of standardized benchmarks for reinforcement learning (RL) in sustainability applications has made it difficult to both track progress on specific domains and identify bottlenecks for researchers to focus their efforts. In this paper, we present SustainGym, a suite of five environments desi…

2023

Virtual Passive-Joint Space Based Time-Optimal Trajectory Planning for a 4-DOF Parallel Manipulator

RA-L 2023

The 4-DOF (3T1R) 4 <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">P</u> Pa-2PaR parallel manipulator is developed for high-speed pick-and-place operations. However, conventional trajectory planning methods in either active-joint space or Cartesia

Cited by 3SourceScholar
2023

X-VoE: Measuring eXplanatory Violation of Expectation in Physical Events

ICCV 2023oral

Intuitive physics is pivotal for human understanding of the physical world, enabling prediction and interpretation of events even in infancy. Nonetheless, replicating this level of intuitive physics in artificial intelligence (AI) remains a formidable challenge. This study introduces X-VoE, a compre…

Cited by 4PDFcodeScholar
2022

A "Look-Backward-and-Forward" Adaptation Strategy for Assessing Parameter Estimation Error of Human Motion Prediction Model

RA-L 2022

The prediction of human motion is essential for safe human-robot collaboration (HRC). For existing prediction methods based on adaptive neural network (NN) models, estimation errors (EEs) of model parameters are directly coupled with prior EEs of trajectories. This results in poor assessment of the

Cited by 2SourceScholar
2022

Bias aware probabilistic Boolean matrix factorization

UAI 2022poster

Boolean matrix factorization (BMF) is a combinatorial problem arising from a wide range of applications including recommendation system, collaborative filtering, and dimensionality reduction. Currently, the noise model of existing BMF methods is often assumed to be homoscedastic; however, in real wo…

2022

Causal Inference with Non-IID Data using Linear Graphical Models

NeurIPS 2022accept

Traditional causal inference techniques assume data are independent and identically distributed (IID) and thus ignores interactions among units. However, a unit’s treatment may affect another unit's outcome (interference), a unit’s treatment may be correlated with another unit’s outcome, or a unit’…

Cited by 20SourcePDFScholar
2022

Design and Modeling of a Compound Twisted and Coiled Actuator Based on Spandex Fibers and an SMA Skeleton

RA-L 2022

Twisted and Coiled Actuators (TCAs) are a class of new artificial muscles for flexible actuations. However, conventional TCAs based on nylon fibers commonly require high driving temperature, which limits their applications. Although the TCAs based on spandex fibers can produce high strain under low

Cited by 11SourceScholar
2022

Empathetic and Emotionally Positive Conversation Systems with an Emotion-specific Query-Response Memory

EMNLP 2022finding

Emotional conversation systems generate responses for the input queries considering the speaker’s emotions in a conversation. Existing emotional conversation systems output emotional responses according to either a given emotion or the user’s emotion reflected in the input queries. Following a given…

2022

Grasping State Analysis of Soft Manipulator Based on Flexible Tactile Sensor Array

IROS 2022poster

Although the grasping state analysis is vital in the study of manipulators, the grasping state analysis of soft manipulators as an independent research topic is not much so far. This paper proposes a novel pneumatic soft manipulator with a flexible tactile sensor array (SM-FTSA). The flexible tactil…

Cited by 2SourceScholar
2022

Hierarchical Normalization for Robust Monocular Depth Estimation

NeurIPS 2022accept

In this paper, we address monocular depth estimation with deep neural networks. To enable training of deep monocular estimation models with various sources of datasets, state-of-the-art methods adopt image-level normalization strategies to generate affine-invariant depth representations. However, le…

Cited by 36SourcePDFScholar
2022

KD-MVS: Knowledge Distillation Based Self-Supervised Learning for Multi-View Stereo

ECCV 2022poster

"Supervised multi-view stereo (MVS) methods have achieved remarkable progress in terms of reconstruction quality, but suffer from the challenge of collecting large-scale ground-truth depth. In this paper, we propose a novel self-supervised training pipeline for MVS based on knowledge distillation, t…

2022

Learning Algebraic Representation for Systematic Generalization in Abstract Reasoning

ECCV 2022poster

"Is intelligence realized by connectionist or classicist? While connectionist approaches have achieved superhuman performance, there has been growing evidence that such task-specific superiority is particularly fragile in systematic generalization. This observation lies in the central debate between…

Cited by 37SourcePDFScholar
2022

Multi-Centroid Representation Network for Domain Adaptive Person Re-ID

AAAI 2022technical

Recently, many approaches tackle the Unsupervised Domain Adaptive person re-identification (UDA re-ID) problem through pseudo-label-based contrastive learning. During training, a uni-centroid representation is obtained by simply averaging all the instance features from a cluster with the same pseudo…

Cited by 72SourcePDFScholar
2022

Provable Second-Order Riemannian Gauss-Newton Method for Low-Rank Tensor Estimation ‖

ICASSP 2022accepted

In this paper, we consider the estimation of a low Tucker rank tensor from a number of noisy linear measurements. We propose a Riemannian Gauss-Newton (RGN) method with fast implementations for low Tucker rank tensor estimation. Different from the generic (super)linear convergence guarantee of RGN i…

Cited by 0SourceScholar
2022

RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior

CVPR 2022poster

In this work, we focus on scene flow learning on point clouds in a self-supervised manner. A real-world scene can be well modeled as a collection of rigidly moving parts, therefore its scene flow can be represented as a combination of rigid motion of each part. Inspired by this observation, we propo…

Cited by 65PDFScholar
2022

Sobolev Training for Implicit Neural Representations with Approximated Image Derivatives

ECCV 2022poster

"Recently, Implicit Neural Representations (INRs) parameterized by neural networks have emerged as a powerful and promising tool to represent different kinds of signals due to its continuous, differentiable properties, showing superiorities to classical discretized representations. However, the trai…

2021

AVP-Loc: Surround View Localization and Relocalization Based on HD Vector Map for Automated Valet Parking

IROS 2021poster

Localization is a crucial prerequisite for automated valet parking, in which a vehicle is required to navigate itself in a GPS-denied parking lot. Traditional visual localization methods usually build a feature map and use it for future localizations. However, the feature map is not robust to change…

Cited by 13SourceScholar
2021

Abstract Spatial-Temporal Reasoning via Probabilistic Abduction and Execution

CVPR 2021poster

Spatial-temporal reasoning is a challenging task in Artificial Intelligence (AI) due to its demanding but unique nature: a theoretic requirement on representing and reasoning based on spatial-temporal knowledge in mind, and an applied requirement on a high-level cognitive system capable of navigatin…

Cited by 75PDFScholar
2021

Binocular Mutual Learning for Improving Few-Shot Classification

ICCV 2021poster

Most of the few-shot learning methods learn to transfer knowledge from datasets with abundant labeled data (i.e., the base set). From the perspective of class space on base set, existing methods either focus on utilizing all classes under a global view by normal pretraining, or pay more attention to…

Cited by 113PDFcodeScholar
2021

Congestion-aware Multi-agent Trajectory Prediction for Collision Avoidance

ICRA 2021poster

Predicting agents’ future trajectories plays a crucial role in modern AI systems, yet it is challenging due to intricate interactions exhibited in multi-agent systems, especially when it comes to collision avoidance. To address this challenge, we propose to learn congestion patterns as contextual cu…

Cited by 51SourcecodeScholar
2021

DT-Loc: Monocular Visual Localization on HD Vector Map Using Distance Transforms of 2D Semantic Detections

IROS 2021poster

Localizing a vehicle on a prebuilt HD vector map is a prerequisite for many autonomous driving applications. Existing visual localization approaches usually require a separate local feature layer to function. The separate localization layer suffers from the robustness issue inherited from the local…

Cited by 11SourceScholar
2021

DeFRCN: Decoupled Faster R-CNN for Few-Shot Object Detection

ICCV 2021poster

Few-shot object detection, which aims at detecting novel objects rapidly from extremely few annotated examples of previously unseen classes, has attracted significant research interest in the community. Most existing approaches employ the Faster R-CNN as basic detection framework, yet, due to the la…

Cited by 347PDFcodeScholar
2021

Dynamic Metric Learning: Towards a Scalable Metric Space To Accommodate Multiple Semantic Scales

CVPR 2021poster

This paper introduces a new fundamental characteristics, i.e., the dynamic range, from real-world metric tools to deep visual recognition. In metrology, the dynamic range is a basic quality of a metric tool, indicating its flexibility to accommodate various scales. Larger dynamic range offers higher…

Cited by 20PDFcodeScholar
2021

End-to-End Human Object Interaction Detection With HOI Transformer

CVPR 2021poster

We propose HOI Transformer to tackle human object interaction (HOI) detection in an end-to-end manner. Current approaches either decouple HOI task into separated stages of object detection and interaction classification or introduce surrogate interaction problem. In contrast, our method, named HOI T…

Cited by 266PDFcodeScholar
2021

FSCE: Few-Shot Object Detection via Contrastive Proposal Encoding

CVPR 2021poster

Emerging interests have been brought to recognize previously unseen objects given very few training examples, known as few-shot object detection (FSOD). Recent researches demonstrate that good feature embedding is the key to reach favorable few-shot learning performance. We observe object proposals…

Cited by 531PDFcodeScholar
2021

Few-Shot Incremental Learning With Continually Evolved Classifiers

CVPR 2021poster

Few-shot class-incremental learning (FSCIL) aims to design machine learning algorithms that can continually learn new concepts from a few data points, without forgetting knowledge of old classes. The difficulty lies in that limited data from new classes not only lead to significant overfitting issue…

Cited by 395PDFScholar
2021

IDM: An Intermediate Domain Module for Domain Adaptive Person Re-ID

ICCV 2021poster

Unsupervised domain adaptive person re-identification (UDA re-ID) aims at transferring the labeled source domain's knowledge to improve the model's discriminability on the unlabeled target domain. From a novel perspective, we argue that the bridging between the source and target domains can be utili…

Cited by 164PDFcodeScholar
2021

Meta Navigator: Search for a Good Adaptation Policy for Few-Shot Learning

ICCV 2021poster

Few-shot learning aims to adapt knowledge learned from previous tasks to novel tasks with only a limited amount of labeled data. Research literature on few-shot learning exhibits great diversity, while different algorithms often excel at different few-shot learning scenarios. It is therefore tricky…

Cited by 60PDFScholar
2021

Point Cloud Segmentation via Edge-fused Local Graph Learning

ICRA 2021poster

Traditional convolution for capturing local structures and relationships remains a key technical limit in 3D semantic segmentation, which neglects the certain influence of the adjacent points on the central point in the disordered local point clouds. In this paper, we propose a novel joint-edge grap…

Cited by 4SourceScholar
2021

Robust LiDAR Localization on an HD Vector Map without a Separate Localization Layer

IROS 2021poster

Many autonomous driving applications nowadays come along with a prebuilt vector map for routing and planning purposes. In order to localize on this map, traditional LiDAR localization methods usually require a separate localization layer to function. On one hand, the separate layer occupies large st…

Cited by 11SourceScholar
2021

Spatial Ensemble: a Novel Model Smoothing Mechanism for Student-Teacher Framework

NeurIPS 2021poster

Model smoothing is of central importance for obtaining a reliable teacher model in the student-teacher framework, where the teacher generates surrogate supervision signals to train the student. A popular model smoothing method is the Temporal Moving Average (TMA), which continuously averages the tea…

2021

Temporal Knowledge Consistency for Unsupervised Visual Representation Learning

ICCV 2021poster

The instance discrimination paradigm has become dominant in unsupervised learning. It always adopts a teacher-student framework, in which the teacher provides embedded knowledge as a supervision signal for the student. The student learns meaningful representations by enforcing instance spatial consi…

Cited by 13PDFcodeScholar
2020

Circle Loss: A Unified Perspective of Pair Similarity Optimization

CVPR 2020oral

This paper provides a pair similarity optimization viewpoint on deep feature learning, aiming to maximize the within-class similarity s_p and minimize the between-class similarity s_n. We find a majority of loss functions, including the triplet loss and the softmax cross-entropy loss, embed s_n and…

Cited by 1174PDFScholar
2020

Conditional Gaussian Distribution Learning for Open Set Recognition

CVPR 2020poster

Deep neural networks have achieved state-of-the-art performance in a wide range of recognition/classification tasks. However, when applying deep learning to real-world applications, there are still multiple challenges. A typical challenge is that unknown samples may be fed into the system during the…

Cited by 329PDFScholar
2020

DeepEMD: Few-Shot Image Classification With Differentiable Earth Mover's Distance and Structured Classifiers

CVPR 2020oral

In this paper, we address the few-shot classification task from a new perspective of optimal matching between image regions. We adopt the Earth Mover's Distance (EMD) as a metric to compute a structural distance between dense image representations to determine image relevance. The EMD generates the…

Cited by 1005PDFScholar
2020

Geometric All-way Boolean Tensor Decomposition

NeurIPS 2020poster

Boolean tensor has been broadly utilized in representing high dimensional logical data collected on spatial, temporal and/or other relational domains. Boolean Tensor Decomposition (BTD) factorizes a binary tensor into the Boolean sum of multiple rank-1 tensors, which is an NP-hard problem. Existing…

2020

Iterative Distance-Aware Similarity Matrix Convolution with Mutual-Supervised Point Elimination for Efficient Point Cloud Registration

ECCV 2020poster

In this paper, we propose a novel learning-based pipeline for partially overlapping 3D point cloud registration. The proposed model includes an iterative distance-aware similarity matrix convolution module to incorporate information from both the feature and Euclidean space into the pairwise point m…

2020

Learning Disentangled Representations of Videos with Missing Data

NeurIPS 2020poster

Missing data poses significant challenges while learning representations of video sequences. We present Disentangled Imputed Video autoEncoder (DIVE), a deep generative model that imputes and predicts future video frames in the presence of missing data. Specifically, DIVE introduces a missingness la…

2019

CANet: Class-Agnostic Segmentation Networks With Iterative Refinement and Attentive Few-Shot Learning

CVPR 2019poster

Recent progress in semantic segmentation is driven by deep Convolutional Neural Networks and large-scale labeled image datasets. However, data labeling for pixel-wise segmentation is tedious and costly. Moreover, a trained model can only make predictions within a set of pre-defined classes. In this…

Cited by 747PDFScholar
2019

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

CVPR 2019poster

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from appearance and motion features; 2) a redesigned question memory which helps under…

Cited by 342PDFcodeScholar
2019

Learning Perceptual Inference by Contrasting

NeurIPS 2019spotlight

“Thinking in pictures,” [1] i.e., spatial-temporal reasoning, effortless and instantaneous for humans, is believed to be a significant ability to perform logical induction and a crucial factor in the intellectual history of technology development. Modern Artificial Intelligence (AI), fueled by massi…

2019

Learning Safe Unlabeled Multi-Robot Planning with Motion Constraints

IROS 2019poster

In this paper, we present a learning approach to goal assignment and trajectory planning for unlabeled robots operating in 2D, obstacle-filled workspaces. More specifically, we tackle the unlabeled multi-robot motion planning problem with motion constraints as a multi-agent reinforcement learning pr…

Cited by 41SourceScholar
2019

Learning Virtual Grasp with Failed Demonstrations via Bayesian Inverse Reinforcement Learning

IROS 2019poster

We propose Bayesian Inverse Reinforcement Learning with Failure (BIRLF), which makes use of failed demonstrations that were often ignored or filtered in previous methods due to the difficulties to incorporate them in addition to the successful ones. Specifically, we leverage halfspaces derived from…

Cited by 27SourceScholar
2019

Perceive Where to Focus: Learning Visibility-Aware Part-Level Features for Partial Person Re-Identification

CVPR 2019poster

This paper considers a realistic problem in person re-identification (re-ID) task, i.e., partial re-ID. Under partial re-ID scenario, the images may contain a partial observation of a pedestrian. If we directly compare a partial pedestrian image with a holistic one, the extreme spatial misalignment…

Cited by 460PDFcodeScholar
2019

Pyramid Graph Networks With Connection Attentions for Region-Based One-Shot Semantic Segmentation

ICCV 2019poster

One-shot image segmentation aims to undertake the segmentation task of a novel class with only one training image available. The difficulty lies in that image segmentation has structured data representations, which yields a many-to-many message passing problem. Previous methods often simplify it to…

Cited by 388PDFScholar
2019

RAVEN: A Dataset for Relational and Analogical Visual REasoNing

CVPR 2019poster

Dramatic progress has been witnessed in basic vision tasks involving low-level perception, such as object recognition, detection, and tracking. Unfortunately, there is still enormous performance gap between artificial vision systems and human intelligence in terms of higher-level vision problems, es…

Cited by 357PDFScholar
2019

Re-ID Driven Localization Refinement for Person Search

ICCV 2019poster

Person search aims at localizing and identifying a query person from a gallery of uncropped scene images. Different from person re-identification (re-ID), its performance also depends on the localization accuracy of a pedestrian detector. The state-of-the-art methods train the detector individually,…

Cited by 162PDFcodeScholar
2019

Vehicle Re-Identification With Viewpoint-Aware Metric Learning

ICCV 2019poster

This paper considers vehicle re-identification (re-ID) problem. The extreme viewpoint variation (up to 180 degrees) poses great challenges for existing approaches. Inspired by the behavior in human's recognition process, we propose a novel viewpoint-aware metric learning approach. It learns two metr…

Cited by 253PDFcodeScholar
2016

Joint Multiview Segmentation and Localization of RGB-D Images Using Depth-Induced Silhouette Consistency

CVPR 2016poster

In this paper, we propose an RGB-D camera localization approach which takes an effective geometry constraint, i.e. silhouette consistency, into consideration. Unlike existing approaches which usually assume the silhouettes are provided, we consider more practical scenarios and generate the silhouett…

Cited by 7PDFScholar
2015

Adaptive human-centered representation for activity recognition of multiple individuals from 3D point cloud sequences

ICRA 2015poster

Activity recognition of multi-individuals (ARMI) within a group, which is essential to practical human-centered robotics applications such as childhood education, is a particularly challenging and previously not well studied problem. We present a novel adaptive human-centered (AdHuC) representation…

Cited by 15SourceScholar
2015

MeshStereo: A Global Stereo Model With Mesh Alignment Regularization for View Interpolation

ICCV 2015oral

We present a novel global stereo model designed for view interpolation. Unlike existing stereo models which only output a disparity map, our model is able to output a 3D triangular mesh, which can be directly used for view interpolation. To this aim, we partition the input stereo images into 2D tria…

Cited by 206PDFScholar
2015

Query Adaptive Similarity Measure for RGB-D Object Recognition

ICCV 2015poster

This paper studies the problem of improving the top-1 accuracy of RGB-D object recognition. Despite of the impressive top-5 accuracies achieved by existing methods, their top-1 accuracies are not very satisfactory. The reasons are in two-fold: (1) existing similarity measures are sensitive to object…

Cited by 18PDFScholar