← Search

Chen Zhao

120 accepted papers

2026

APIC: Orthogonalized Neuro-Symbolic Modeling for Nonlinear Dissipative Dynamics

ICML 2026poster

Current data-driven scientific modeling struggles with a functional dichotomy: neural operators exhibit spectral bias in high-frequency regimes, while physics-constrained paradigms suffer from optimization pathologies. To bridge this gap, we propose Adaptive Physics-Informed Computing (APIC), a neur…

Cited by 0SourceScholar
2026

Does FLUX Already Know How to Perform Physically Plausible Image Composition?

ICLR 2026poster

Image composition aims to seamlessly insert a user-specified object into a new scene, but existing models struggle with complex lighting (e.g., accurate shadows, water reflections) and diverse, high-resolution inputs. Modern text-to-image diffusion models (e.g., SD3.5, FLUX) already encode essential…

Cited by 0SourcecodeScholar
2026

GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection

CVPR 2026

Hand-Object Interaction (HOI) remains a core challenge in digital human video synthesis, where models must generate physically plausible contact and preserve object identity across frames. Although recent HOI reenactment approaches have achieved progress, they are typically trained and evaluated in-

Cited by 0SourceScholar
2026

LEMD: Latent Environment Extrapolation and Message Disentanglement for Dynamic Graph Under Distribution Shift

IJCAI 2026

Dynamic graph neural networks (DyGNNs) are widely used to model evolving interactions, but may fail under data distribution shift. Due to limited and unreliable interventions and insufficient disentanglement, the existing dynamic graph domain generalization approaches lead to suboptimal results. We

Cited by 0Scholar
2026

LLM-Enhanced Energy Contrastive Learning for Out-of-Distribution Detection in Text-Attributed Graphs

AAAI 2026technical

Text-attributed graphs, where nodes are enriched with textual attributes, have become a powerful tool for modeling real-world networks such as citation, social, and transaction networks. However, existing methods for learning from these graphs often assume that the distributions of training and test

Cited by 0SourcePDFScholar
2026

LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

ICML 2026poster

Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a formidable challenge due to the compounded difficulties of motion modeling, semantic planning, and detail synthesis. To address these limitations, we propo…

Cited by 0SourceScholar
2026

Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting

ICML 2026poster

Accurate precipitation nowcasting is vital for disaster mitigation, but deep learning methods suffer a key trade-off: regression models produce over-smoothed, spectrally decaying predictions that blur convective details and violate turbulence power laws; diffusion models generate realistic yet unanc…

Cited by 0SourceScholar
2026

MARLIN: Multi-Agent Reinforcement Learning for Incremental DAG Discovery

AAAI 2026technical

Uncovering causal structures from observational data is crucial for understanding complex systems and making informed decisions. While reinforcement learning (RL) has shown promise in identifying these structures in the form of a directed acyclic graph (DAG), existing methods often lack efficiency,

Cited by 0SourcePDFScholar
2026

MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval

ICLR 2026poster

We introduce MRMR, the first expert-level multidisciplinary multimodal retrieval benchmark requiring intensive reasoning. MRMR contains 1,435 queries spanning 23 domains, with positive documents carefully verified by human experts. Compared to prior benchmarks, MRMR introduces three key advancements…

Cited by 0SourceScholar
2026

Out-of-Distribution Detection with Positive and Negative Prompt Supervision Using Large Language Models

AAAI 2026technical

Out-of-distribution (OOD) detection is committed to delineating the classification boundaries between in-distribution (ID) and OOD images. Recent advances in vision-language models (VLMs) have demonstrated remarkable OOD detection performance by integrating both visual and textual modalities. In thi

Cited by 0SourcePDFScholar
2026

PURE: Purging Unrelated Representations for Content-Agnostic Forgery Detection

IJCAI 2026

Existing AI-generated image (AIGI) detectors perform well in-domain but degrade severely under distribution shift. We observe that this failure is mainly caused by content shortcuts, where detectors spuriously couple forgery artifacts with semantic content, such as object categories or demographic a

Cited by 0Scholar
2026

PhysGen: Physically Grounded 3D Shape Generation for Industrial Design

CVPR 2026

Existing generative models for 3D shapes can synthesize high-fidelity and visually plausible shapes. For certain classes of shapes that have undergone an engineering design process, the realism of the shape is tightly coupled with the underlying physical properties, e.g., aerodynamic efficiency for

Cited by 0SourcecodeScholar
2026

QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation

AAAI 2026technical

Developing high-performance GPU kernels is critical for AI and scientific computing, but remains challenging due to its reliance on expert crafting and poor portability. While large language models (LLMs) offer promise for automation, both general-purpose and finetuned LLMs suffer from two fundament

Cited by 0SourcePDFScholar
2026

RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations

CVPR 2026

Humans perceive the 3D world from limited 2D observations. While recent feed-forward generalizable 3D reconstruction models can recover structures from sparse images, they typically represent only observed regions, leaving unseen geometry unmodeled. This raises a fundamental question: Can we infer c

Cited by 0SourceScholar
2026

SkillGen: Learning Domain Skills for In-Context Sequential Decision Making

AAAI 2026technical

Large language models (LLMs) are increasingly applied to sequential decision-making through in-context learning (ICL), yet their effectiveness is highly sensitive to prompt quality. Effective prompts should meet three principles: focus on decision-critical information, provide step-level granularity

Cited by 0SourcePDFScholar
2026

YOLO-IOD: Towards Real Time Incremental Object Detection

AAAI 2026technical

Current methodologies for incremental object detection (IOD) primarily rely on Faster R-CNN or DETR series detectors; however, these approaches do not accommodate the real-time YOLO detection frameworks. In this paper, we first identify three primary types of knowledge conflicts that contribute to c

Cited by 0SourcePDFScholar
2026

b-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment

CVPR 2026

CLIP achieves strong zero-shot image-text retrieval by aligning global vision and text representations, yet it falls behind on fine-grained tasks even when fine-tuned on long, detailed captions. In this work, we propose b-CLIP, a multi-granular text-conditioned contrastive learning framework designe

Cited by 0SourcecodeScholar
2025

A Novel Sparse Active Online Learning Framework for Fast and Accurate Streaming Anomaly Detection Over Data Streams

IJCAI 2025

Online Anomaly Detection (OAD) is critical for identifying rare yet important data points in large, dynamic, and complex data streams. A key challenge lies in achieving accurate and consistent detection of anomalies while maintaining computational and memory efficiency. Conventional OAD approaches,

Cited by 0SourcePDFScholar
2025

Are Multimodal LLMs Robust Against Adversarial Perturbations? RoMMath: A Systematic Evaluation on Multimodal Math Reasoning

NAACL 2025long

We introduce RoMMath, the first benchmark designed to evaluate the capabilities and robustness of multimodal large language models (MLLMs) in handling multimodal math reasoning, particularly when faced with adversarial perturbations. RoMMath consists of 4,800 expert-annotated examples, including an…

2025

Auto-Regressively Generating Multi-View Consistent Images

ICCV 2025poster

Generating multi-view images from human instructions is crucial for 3D content creation. The primary challenges involve maintaining consistency across multiple views and effectively synthesizing shapes and textures under diverse conditions. In this paper, we propose the Multi-View AutoRegressive (MV…

2025

BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding

CVPR 2025poster

Large video-language models (VLMs) have demonstrated promising progress in various video understanding tasks. However, their effectiveness in long-form video analysis is constrained by limited context windows. Traditional approaches, such as uniform frame sampling, often inevitably allocate resource…

2025

BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation

ICCV 2025poster

This paper presents a generalizable RGB-based approach for object pose estimation, specifically designed to address challenges in sparse-view settings. While existing methods can estimate the poses of unseen objects, their generalization ability remains limited in scenarios involving occlusions and…

Cited by 0SourcePDFScholar
2025

DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes

ICRA 2025

Novel-view synthesis approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, making them unsuitable where computational resources are limited. Additionally, few-shot methods often struggle with poor reconst

Cited by 5SourcecodeScholar
2025

Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective

EMNLP 2025

Large language model (LLM)-based embedding models, benefiting from large scale pre-training and post-training, have begun to surpass BERT and T5-based models on general-purpose text embedding tasks such as document retrieval. However, a fundamental limitation of LLM embeddings lies in the unidirecti

2025

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking

AAAI 2025technical

Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal re…

2025

FADE: Towards Fairness-aware Data Generation for Domain Generalization via Classifier-Guided Score-based Diffusion Models

IJCAI 2025

Fairness-aware domain generalization (FairDG) has emerged as a critical challenge for deploying trustworthy AI systems, particularly in scenarios involving distribution shifts. Traditional methods for addressing fairness have failed in domain generalization due to their lack of consideration for dis

Cited by 0SourcePDFScholar
2025

FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering

EMNLP 2025

Large Language Models (LLMs) frequently hallucinate to long-form questions, producing plausible yet factually incorrect answers. A common mitigation strategy is to provide attribution to LLM outputs. However, existing benchmarks primarily focus on simple attribution that retrieves supporting textual

2025

FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain

EMNLP 2025

Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property. This paper introduces FinTrust, a comprehensive benchmark specifically designed for evaluat

2025

From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspective

CVPR 2025poster

Ultra-high-definition (UHD) image restoration faces significant challenges due to its high resolution, complex content, and intricate details. To cope with these challenges, we analyze the restoration process in depth through a progressive spectral perspective, and deconstruct the complex UHD restor…

2025

GDDA: Semantic OOD Detection on Graphs under Covariate Shift via Score-Based Diffusion Models

ICASSP 2025accepted

Out-of-distribution (OOD) detection poses a signifi-cant challenge for Graph Neural Networks (GNNs), particularly in open-world scenarios with varying distribution shifts. Most existing OOD detection methods on graphs primarily focus on identifying instances in test data domains caused by either sem…

Cited by 0SourceScholar
2025

Gaussian-LIC: Real-Time Photo-Realistic SLAM with Gaussian Splatting and LiDAR-Inertial-Camera Fusion

ICRA 2025

In this paper, we present a real-time photo-realistic SLAM method based on marrying Gaussian Splatting with LiDAR-Inertial-Camera SLAM. Most existing radiance-field-based SLAM systems mainly focus on bounded indoor environments, equipped with RGB-D or RGB sensors. However, they are prone to decline

Cited by 28SourcecodeScholar
2025

High temperature sterilization resistant and enclosed three-axial force-sensing surgical instrument integrated with step-reduced FBG

IROS 2025

The Fiber Bragg grating (FBG) three-axial force sensor provides force feedback for an endoscopic surgical robot, reducing operational difficulty and risks. However, the packaging method of the optical fiber sensor demonstrates limited adaptability to both high-temperature sterilization environments

Cited by 0SourceScholar
2025

HyperKAN: Hypergraph Representation Learning with Kolmogorov-Arnold Networks

ICASSP 2025accepted

Hypergraph representation learning has garnered increasing attention across various domains due to its capability to model high-order relationships. Traditional methods often rely on hypergraph neural networks (HNNs) employing messagepassing mechanisms to aggregate vertex and hyperedge features. How…

Cited by 0SourceScholar
2025

HyperSF: A Hypergraph Representation Learning Method Based on Structural Fusion

ICASSP 2025accepted

Hypergraph Neural Networks (HNNs) have recently gained attention as a powerful approach for capturing high-order correlations through hypergraph-structured encoding and learning techniques. However, despite their potential, existing HNN methods often encounter over-smoothing issues, which limit thei…

Cited by 0SourceScholar
2025

Let The Jury Decide: Fair Demonstration Selection for In-Context Learning through Incremental Greedy Evaluation

ACL 2025finding

Large Language Models (LLMs) are powerful in-context learners, achieving strong performance with just a few high-quality demonstrations. However, fairness concerns arise in many in-context classification tasks, especially when predictions involve sensitive attributes. To address this, we propose JUD…

2025

LimRank: Less is More for Reasoning-Intensive Information Reranking

EMNLP 2025

Existing approaches typically rely on large-scale fine-tuning to adapt LLMs for information reranking tasks, which is computationally expensive. In this work, we demonstrate that modern LLMs can be effectively adapted using only minimal, high-quality supervision. To enable this, we design LIMRANK-SY

Cited by 0SourcePDFScholar
2025

MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search

EMNLP 2025

We introduce MCTS-RAG, a novel approach that enhances the reasoning capabilities of small language models on knowledge-intensive tasks by leveraging retrieval-augmented generation (RAG) to provide relevant context and Monte Carlo Tree Search (MCTS) to refine reasoning paths. MCTS-RAG dynamically int

2025

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

CVPR 2025poster

We introduce MMVU, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. MMVU includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. C…

2025

MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering

EMNLP 2025

Understanding temporal concepts and answering time-sensitive questions is crucial yet a challenging task for question-answering systems powered by large language models (LLMs). Existing approaches either update the parametric knowledge of LLMs with new facts, which is resource-intensive and often im

2025

Metric-Agnostic Continual Learning for Sustainable Group Fairness

AAAI 2025technical

Group Fairness-aware Continual Learning (GFCL) aims to eradicate discriminatory predictions against certain demographic groups in a sequence of diverse learning tasks. This paper explores an even more challenging GFCL problem – how to sustain a fair classifier across a sequence of tasks with covaria…

2025

Multi-View Unsupervised Column Subset Selection via Combinatorial Search (Student Abstract)

AAAI 2025technical

Given a data matrix, unsupervised column subset selection refers to the problem of identifying a subset of columns that can be used to linearly approximate the original data matrix. This problem has many applications, such as feature selection and representative selection, but solving it optimally i…

Cited by 0SourcePDFScholar
2025

Physics: Benchmarking Foundation Models on University-Level Physics Problem Solving

ACL 2025finding

We introduce Physics, a comprehensive benchmark for university-level physics problem solving. It contains 1,297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics.Each probl…

2025

QiMeng-TensorOp: One-Line Prompt is Enough for High-Performance Tensor Operator Generation with Hardware Primitives

IJCAI 2025

Computation-intensive tensor operators constitute over 90% of the computations in Large Language Models (LLMs) and Deep Neural Networks. Automatically and efficiently generating high-performance tensor operators with hardware primitives is crucial for diverse and ever-evolving hardware architectures

Cited by 0SourcePDFScholar
2025

RPDR: A Round-trip Prediction-Based Data Augmentation Framework for Long-Tail Question Answering

EMNLP 2025

Long-tail question answering presents significant challenges for large language models (LLMs) due to their limited ability to acquire and accurately recall less common knowledge. Retrieval-augmented generation (RAG) systems have shown great promise in mitigating this limitation by integrating extern

2025

SEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object Detection

CVPR 2025poster

Domain adaptive object detection (DAOD) aims to generalize detectors trained on an annotated source domain to an unlabelled target domain. Traditional works focus on aligning visual features between domains to extract domain-invariant knowledge, and recent VLM-based DAOD methods leverage semantic in…

Cited by 0SourcePDFScholar
2025

SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning

CVPR 2025poster

Masked video modeling, such as VideoMAE, is an effective paradigm for video self-supervised learning (SSL). However, they are primarily based on reconstructing pixel level details on natural videos which have substantial temporal redundancy, limiting their capability for semantic representation and…

2025

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution

ICCV 2025poster

Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively.…

Cited by 0SourcePDFScholar
2025

SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks

NeurIPS 2025spotlight

We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evalua…

Cited by 0SourceScholar
2025

Self-Ensembling Gaussian Splatting for Few-Shot Novel View Synthesis

ICCV 2025poster

3D Gaussian Splatting (3DGS) has demonstrated remarkable effectiveness in novel view synthesis (NVS). However, 3DGS tends to overfit when trained with sparse views, limiting its generalization to novel viewpoints. In this paper, we address this overfitting issue by introducing Self-Ensembling Gaussi…

2025

SolverLLM: Leveraging Test-Time Scaling for Optimization Problem via LLM-Guided Search

NeurIPS 2025poster

Large Language Models (LLMs) offer promising capabilities for tackling complex reasoning tasks, including optimization problems. However, existing methods either rely on prompt engineering, which leads to poor generalization across problem types, or require costly supervised training. We introduce S…

Cited by 0SourceScholar
2025

Splatter-360: Generalizable 360 Gaussian Splatting for Wide-baseline Panoramic Images

CVPR 2025poster

Wide-baseline panoramic images are frequently used in applications like VR and simulations to minimize capturing labor costs and storage needs. However, synthesizing novel views from these panoramic images in real time remains a significant challenge, especially due to panoramic imagery's high resol…

2025

SportReason: Evaluating Retrieval-Augmented Reasoning across Tables and Text for Sports Question Answering

EMNLP 2025

We present SportReason, a benchmark for retrieval-augmented reasoning on numerical sports questions. Unlike existing benchmarks limited to one or two evidence units, SportReason requires combining and reasoning across free-text, structured tables, and semi-structured infoboxes. We provide 3,000 huma

2025

TexGarment: Consistent Garment UV Texture Generation via Efficient 3D Structure-Guided Diffusion Transformer

CVPR 2025poster

This paper introduces TexGarment, an efficient method for synthesizing high-quality, 3D-consistent garment textures in UV space. Traditional approaches based on 2D-to-3D mapping often suffer from 3D inconsistency, while methods learning from limited 3D data lack sufficient texture diversity. These l…

Cited by 0SourcePDFScholar
2025

TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting

CVPR 2025poster

Physically Based Rendering (PBR) materials play a crucial role in modern graphics, enabling photorealistic rendering across diverse environment maps. Developing an effective and efficient algorithm that is capable of automatically generating high-quality PBR materials rather than RGB texture for 3D…

2025

UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset

NeurIPS 2025poster

Ultra-high-resolution (UHR) text-to-image (T2I) generation has seen notable progress. However, two key challenges remain : 1) the absence of a large-scale high-quality UHR T2I dataset, and (2) the neglect of tailored training strategies for fine-grained detail synthesis in UHR scenarios. To tackle t…

Cited by 0SourcecodeScholar
2025

VDG: Vision-Only Dynamic Gaussian for Driving Simulation

RA-L 2025

Recent advances in dynamic Gaussian splatting have significantly improved scene reconstruction and novel-view synthesis. However, existing methods often rely on pre-computed camera poses and Gaussian initialization using Structure from Motion (SfM) or other costly sensors, limiting their scalability

Cited by 23SourceScholar
2024

3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose Estimation

ICLR 2024poster

Prior methods that tackle the problem of generalizable object pose estimation highly rely on having dense views of the unseen object. By contrast, we address the scenario where only a single reference view of the object is available. Our goal then is to estimate the relative object pose between this…

Cited by 10SourcePDFScholar
2024

BPDO: Boundary Points Dynamic Optimization for Arbitrary Shape Scene Text Detection

ICASSP 2024accepted

Arbitrary shape scene text detection is of great importance in scene understanding tasks. Due to the complexity and diversity of text in natural scenes, existing scene text algorithms have limited accuracy for detecting arbitrary shape text. In this paper, we propose a novel arbitrary shape scene te…

Cited by 0SourceScholar
2024

DVMNet: Computing Relative Pose for Unseen Objects Beyond Hypotheses

CVPR 2024poster

Determining the relative pose of an object between two images is pivotal to the success of generalizable object pose estimation. Existing approaches typically approximate the continuous pose representation with a large number of discrete pose hypotheses which incurs a computationally expensive proce…

2024

Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

ICML 2024spotlight

Large language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs process different inputs? To answer these questions, we propose to evaluate $\textbf{counterfactual simulatability}$ of nat…

Cited by 57SourcePDFScholar
2024

Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning

CVPR 2024poster

Large pretrained models are increasingly crucial in modern computer vision tasks. These models are typically used in downstream tasks by end-to-end finetuning which is highly memory-intensive for tasks with high-resolution data e.g. video understanding small object detection and point cloud analysis…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Elliptical torus-based Six-axis FBG Force Sensor with In-situ Calibration for Condition Monitoring of Orthopedic Surgical Robot*

ICRA 2024poster

Six-axis force/moment (6-A F/M) sensors make surgical robots effectively sense intraoperative force feedback and drilling status information, reducing the operating challenges and psychological burden of doctors, which also improves the quality and safety of surgery. However, it is difficult for cur…

Cited by 1SourceScholar
2024

End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames

CVPR 2024poster

Recently temporal action detection (TAD) has seen significant performance improvement with end-to-end training. However due to the memory bottleneck only models with limited scales and limited data volumes can afford end-to-end training which inevitably restricts TAD performance. In this paper we re…

2024

FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents

EMNLP 2024main

We introduce FinDVer, a comprehensive benchmark specifically designed to evaluate the explainable claim verification capabilities of LLMs in the context of understanding and analyzing long, hybrid-content financial documents. FinDVer contains 4,000 expert-annotated examples across four subsets, each…

2024

FinanceMATH: Knowledge-Intensive Math Reasoning in Finance Domains

ACL 2024long

We introduce FinanceMath, a novel benchmark designed to evaluate LLMs' capabilities in solving knowledge-intensive math reasoning problems. Compared to prior works, this study features three core advancements. First, FinanceMath includes 1,200 problems with a hybrid of textual and tabular content. T…

2024

GGRt: Towards Generalizable 3D Gaussians without Pose Priors in Real-Time

ECCV 2024poster

"This paper presents GGRt, a novel approach to generalizable novel view synthesis that alleviates the need for real camera poses, complexity in processing high-resolution images, and lengthy optimization processes, thus facilitating stronger applicability of 3D Gaussian Splatting (3D-GS) in real-wor…

2024

HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields

CVPR 2024poster

Human hands are highly articulated and versatile at handling objects. Jointly estimating the 3D poses of a hand and the object it manipulates from a monocular camera is challenging due to frequent occlusions. Thus existing methods often rely on intermediate 3D shape representations to increase perfo…

2024

Large Language Models Help Humans Verify Truthfulness – Except When They Are Convincingly Wrong

NAACL 2024long

Large Language Models (LLMs) are increasingly used for accessing information on the web. Their truthfulness and factuality are thus of great interest. To help users make the right decisions about the information they get, LLMs should not only provide information but also help users fact-check it. We…

Cited by 36SourcePDFScholar
2024

OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding

NeurIPS 2024poster

This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) that possesses the capability for 3D point-level open vocabulary understanding. Our primary motivation stems from observing that existing 3DGS-based open vocabulary methods mainly focus on 2D pixel-level parsing. Thes…

2024

Parallel Structures in Pre-training Data Yield In-Context Learning

ACL 2024long

Pre-trained language models (LMs) are capable of in-context learning (ICL): they can adapt to a task with only a few examples given in the prompt without any parameter update. However, it is unclear where this capability comes from as there is a stark distribution shift between pre-training text and…

Cited by 13SourcePDFScholar
2024

Proof-of-Concept Development of the Distal Module of a Cystoscope Transurethral Continuum Surgical Robotic System

RA-L 2024

Transurethral resection of bladder tumor (TURBT) is the typical procedure for non-muscle invasive bladder tumors. However, current TURBT using rigid surgical tools can hardly handle en bloc resection of bladder tumor and anterior tumor resection. This letter hence proposes a teleoperation-based cyst

Cited by 5SourceScholar
2024

Supervised Algorithmic Fairness in Distribution Shifts: A Survey

IJCAI 2024poster

Supervised fairness-aware machine learning under distribution shifts is an emerging field that addresses the challenge of maintaining equitable and unbiased predictions when faced with changes in data distributions from source to target domains. In real-world applications, machine learning models a…

Cited by 12SourcePDFScholar
2024

SynTQA: Synergistic Table-based Question Answering via Mixture of Text-to-SQL and E2E TQA

EMNLP 2024finding

Text-to-SQL parsing and end-to-end question answering (E2E TQA) are two main approaches for Table-based Question Answering task. Despite success on multiple benchmarks, they have yet to be compared and their synergy remains unexplored. In this paper, we identify different strengths and weaknesses th…

2024

TaPERA: Enhancing Faithfulness and Interpretability in Long-Form Table QA by Content Planning and Execution-based Reasoning

ACL 2024long

Long-form Table Question Answering (LFTQA) requires systems to generate paragraph long and complex answers to questions over tabular data. While Large language models based systems have made significant progress, it often hallucinates, especially when the task involves complex reasoning over tables.…

2024

TexOct: Generating Textures of 3D Models with Octree-based Diffusion

CVPR 2024poster

This paper focuses on synthesizing high-quality and complete textures directly on the surface of 3D models within 3D space. 2D diffusion-based methods face challenges in generating 2D texture maps due to the infinite possibilities of UV mapping for a given 3D mesh. Utilizing point clouds helps circu…

Cited by 1SourcePDFScholar
2024

Text Region Multiple Information Perception Network for Scene Text Detection

ICASSP 2024accepted

Segmentation-based scene text detection algorithms can handle arbitrary shape scene texts and have strong robustness and adaptability, so it has attracted wide attention. Existing segmentation-based scene text detection algorithms usually only segment the pixels in the center region of the text, whi…

Cited by 0SourceScholar
2024

Toward Sufficient Spatial-Frequency Interaction for Gradient-Aware Underwater Image Enhancement

ICASSP 2024accepted

Underwater images suffer from complex and diverse degradation, which inevitably affects the performance of underwater visual tasks. However, most existing learning-based underwater image enhancement (UIE) methods mainly restore such degradations in the spatial domain, and rarely pay attention to the…

Cited by 0SourceScholar
2024

Towards Automated Movie Trailer Generation

CVPR 2024poster

Movie trailers are an essential tool for promoting films and attracting audiences. However the process of creating trailers can be time-consuming and expensive. To streamline this process we propose an automatic trailer generation framework that generates plausible trailers from a full movie by auto…

Cited by 3SourcePDFScholar
2024

Towards Counterfactual Fairness-aware Domain Generalization in Changing Environments

IJCAI 2024poster

Recognizing domain generalization as a commonplace challenge in machine learning, data distribution might progressively evolve across a continuum of sequential domains in practical scenarios. While current methodologies primarily concentrate on bolstering model effectiveness within these new domains…

Cited by 3SourcePDFScholar
2024

Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration

CVPR 2024poster

Underwater images are subject to intricate and diverse degradation inevitably affecting the effectiveness of underwater visual tasks. However most approaches primarily operate in the raw pixel space of images which limits the exploration of the frequency characteristics of underwater images leading…

2024

Your Co-Workers Matter: Evaluating Collaborative Capabilities of Language Models in Blocks World

ACL 2024findings

Language agents that interact with the world on their own have great potential for automating digital tasks. While large language model (LLM) agents have made progress in understanding and executing tasks such as textual games and webpage control, many real-world tasks also require collaboration wit…

2023

A Unified Continual Learning Framework with General Parameter-Efficient Tuning

ICCV 2023poster

The "pre-training - downstream adaptation" presents both new opportunities and challenges for Continual Learning (CL). Although the recent state-of-the-art in CL is achieved through Parameter-Efficient-Tuning (PET) adaptation paradigm, only prompt has been explored, limiting its application to Trans…

Cited by 118PDFcodeScholar
2023

EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual Queries

ICCV 2023oral

With the recent advances in video and 3D understanding, novel 4D spatio-temporal methods fusing both concepts have emerged. Towards this direction, the Ego4D Episodic Memory Benchmark proposed a task for Visual Queries with 3D Localization (VQ3D). Given an egocentric video clip and an image crop dep…

Cited by 21PDFcodeScholar
2023

Fair Representation Learning for Recommendation: A Mutual Information Perspective

AAAI 2023technical

Recommender systems have been widely used in recent years. By exploiting historical user-item interactions, recommender systems can model personalized potential interests of users and have been widely applied to a wide range of scenarios. Despite their impressive performance, most of them may be sub…

Cited by 20SourcePDFScholar
2023

FreeDoM: Training-Free Energy-Guided Conditional Diffusion Model

ICCV 2023poster

Recently, conditional diffusion models have gained popularity in numerous applications due to their exceptional generation ability. However, many existing methods are training-required. They need to train a time-dependent classifier or a condition-dependent score estimator, which increases the cost…

Cited by 152PDFcodeScholar
2023

Getting MoRE out of Mixture of Language Model Reasoning Experts

EMNLP 2023long findings

While recent large language models (LLMs) improve on various question answering (QA) datasets, it remains difficult for a single model to generalize across question types that require distinct reasoning abilities. We provide empirical evidence that state-of-the-art LLMs suffer from poor generalizabi…

Cited by 0SourceScholar
2023

Large-Capacity and Flexible Video Steganography via Invertible Neural Network

CVPR 2023poster

Video steganography is the art of unobtrusively concealing secret data in a cover video and then recovering the secret data through a decoding protocol at the receiver end. Although several attempts have been made, most of them are limited to low-capacity and fixed steganography. To rectify these we…

2023

Multi-Label Temporal Evidential Neural Networks for Early Event Detection

ICASSP 2023accepted

Early event detection aims to detect events even before the event is complete. However, most of the existing methods focus on an event with a single label but fail to be applied to cases with multiple labels. Another non-negligible issue for early event detection is a prediction with overconfidence…

Cited by 0SourceScholar
2023

Open Set Action Recognition via Multi-Label Evidential Learning

CVPR 2023poster

Existing methods for open set action recognition focus on novelty detection that assumes video clips show a single action, which is unrealistic in the real world. We propose a new method for open set action recognition and novelty detection via MUlti-Label Evidential learning (MULE), that goes beyon…

2023

Re2TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action Localization

CVPR 2023poster

Temporal action localization (TAL) requires long-form reasoning to predict actions of various durations and complex content. Given limited GPU memory, training TAL end to end (i.e., from videos to predictions) on long videos is a significant challenge. Most methods can only train on pre-extracted fe…

2023

RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations

ACL 2023long

Despite significant progress having been made in question answering on tabular data (Table QA), it’s unclear whether, and to what extent existing Table QA models are robust to task-specific perturbations, e.g., replacing key question entities or shuffling table columns. To systematically study the r…

2022

A Nested Bi-level Optimization Framework for Robust Few Shot Learning

AAAI 2022technical

Model-Agnostic Meta-Learning (MAML), a popular gradient-based meta-learning framework, assumes that the contribution of each task or instance to the meta-learner is equal.Hence, it fails to address the domain shift between base and novel classes in few-shot learning. In this work, we propose a novel…

2022

Bridging the Generalization Gap in Text-to-SQL Parsing with Schema Expansion

ACL 2022long

Text-to-SQL parsers map natural language questions to programs that are executable over tables to generate answers, and are typically evaluated on large-scale datasets like Spider (Yu et al., 2018). We argue that existing benchmarks fail to capture a certain out-of-domain generalization problem that…

Cited by 17SourcePDFScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Fusing Local Similarities for Retrieval-Based 3D Orientation Estimation of Unseen Objects

ECCV 2022poster

"In this paper, we tackle the task of estimating the 3D orientation of previously-unseen objects from monocular images. This task contrasts with the one considered by most existing deep learning methods which typically assume that the testing objects have been observed during training. To handle the…

2022

MAD: A Scalable Dataset for Language Grounding in Videos From Movie Audio Descriptions

CVPR 2022poster

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of these datasets for the video-language grounding task. Recen…

Cited by 123PDFcodeScholar
2022

R-DFCIL: Relation-Guided Representation Learning for Data-Free Class Incremental Learning

ECCV 2022poster

"Class-Incremental Learning (CIL) struggles with catastrophic forgetting when learning new knowledge, and Data-Free CIL (DFCIL) is even more challenging without access to the training data of previously learned classes. Though recent DFCIL works introduce techniques such as model inversion to synthe…

2022

Re-Examining Calibration: The Case of Question Answering

EMNLP 2022finding

For users to trust model predictions, they need to understand model outputs, particularly their confidence — calibration aims to adjust (calibrate) models’ confidence to match expected accuracy. We argue that the traditional calibration evaluation does not promote effective calibrations: for example…

2022

Unsupervised Learning of 3D Semantic Keypoints with Mutual Reconstruction

ECCV 2022poster

"Semantic 3D keypoints are category-level semantic consistent points on 3D objects. Detecting 3D semantic keypoints is a foundation for a number of 3D vision tasks but remains challenging, due to the ambiguity of semantic information, especially when the objects are represented by unordered 3D point…

2021

Distantly-Supervised Dense Retrieval Enables Open-Domain Question Answering without Evidence Annotation

EMNLP 2021main

Open-domain question answering answers a question based on evidence retrieved from a large corpus. State-of-the-art neural approaches require intermediate evidence annotations for training. However, such intermediate annotations are expensive, and methods that rely on them cannot transfer to the mor…

2021

Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval

NAACL 2021long

Complex question answering often requires finding a reasoning chain that consists of multiple evidence pieces. Current approaches incorporate the strengths of structured knowledge and unstructured text, assuming text corpora is semi-structured. Building on dense retrieval methods, we propose a new m…

2021

Progressive Correspondence Pruning by Consensus Learning

ICCV 2021poster

Correspondence pruning aims to correctly remove false matches (outliers) from an initial set of putative correspondences. The selection is challenging since putative matches are typically extremely unbalanced, largely dominated by outliers, and the random distribution of such outliers further compli…

Cited by 91PDFScholar
2021

What’s in a Name? Answer Equivalence For Open-Domain Question Answering

EMNLP 2021main

A flaw in QA evaluation is that annotations often only provide one gold answer. Thus, model predictions semantically equivalent to the answer but superficially different are considered incorrect. This work explores mining alias entities from knowledge bases and using them as additional gold answers…

2020

Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization

CVPR 2020poster

Fine-grained visual categorization (FGVC) is an important but challenging task due to high intra-class variances and low inter-class variances caused by deformation, occlusion, illumination, etc. An attention convolutional binary neural tree architecture is presented to address those problems for we…

Cited by 276PDFScholar
2020

G-TAD: Sub-Graph Localization for Temporal Action Detection

CVPR 2020poster

Temporal action detection is a fundamental yet challenging task in video understanding. Video context is a critical cue to effectively detect actions, but current works mainly focus on temporal context, while neglecting semantic context as well as other important context properties. In this work, we…

Cited by 604PDFcodeScholar
2020

Learning Semantic Neural Tree for Human Parsing

ECCV 2020poster

In this paper, we design a novel semantic neural tree for human parsing, which uses a tree architecture to encode physiological structure of human body, and design a coarse to fine process in a cascade manner to generate accurate results. Specifically, the semantic neural tree is designed to segment…

Cited by 71SourcePDFScholar
2020

Sparse-to-Dense Depth Completion Revisited: Sampling Strategy and Graph Construction

ECCV 2020poster

Depth completion is a widely studied problem of predicting a dense depth map from a sparse set of measurements and a single RGB image. In this work, we approach this problem by addressing two issues that have been under-researched in the open literature: sampling strategy (data term) and graph const…

Cited by 46SourcePDFScholar
2020

Transformer-XH: Multi-Evidence Reasoning with eXtra Hop Attention

ICLR 2020poster

Transformers have achieved new heights modeling natural language as a sequence of text tokens. However, in many real world scenarios, textual data inherently exhibits structures beyond a linear sequence such as trees and graphs; many tasks require reasoning with evidence scattered across multiple pi…

Cited by 132SourcecodeScholar
2017

Regulating surface traction of a soft robot through electrostatic adhesion control

IROS 2017poster

This paper reports the electrostatic regulation of surface traction of a quadruped soft robot to improve its locomotion efficiency. The soft robot, containing five pneumatic channel networks (PneuNets) in different parts of its body, is actuated to achieve undulated locomotion. Electrostatic adhesio…

Cited by 17SourceScholar