← Search

Jiaxin GUO

25 accepted papers

2026

BlurPoint: Efficient Motion Blur Aware Student-Teacher Local Feature Learning

ICRA 2026poster

Local feature detection and description serve as the foundation for many 3D vision tasks. However, most existing algorithms rely on sharp images, resulting in degraded performance when motion blur occurs due to long exposure. To tackle this challenge, we propose an effective end-to-end model that jo…

Cited by 0Scholar
2026

DispViT: Direct Stereo Disparity Regression with a Single-Stream Vision Transformer

ICLR 2026poster

Deep stereo disparity estimation has long been dominated by a \textbf{matching-centric paradigm}, built on constructing cost volumes and iteratively refining local correspondences. Despite its success, this paradigm exhibits an intrinsic vulnerability: visual ambiguities from occlusion or non-Lamber…

Cited by 0SourcecodeScholar
2026

EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation

ICML 2026poster

This paper introduces EPS3D, a new end-to-end feed-forward framework for open-vocabulary 3D panoptic segmentation. Unlike existing methods relying on additional preprocessing, we design an end-to-end architecture, with a distillation-based training strategy on diverse 3D scenes to predict 3D-aware s…

Cited by 0SourceScholar
2026

Mem-T: Densifying Rewards for Long-Horizon Memory Agents

ICML 2026poster

Memory agents, which depart from predefined memory-processing pipelines by endogenously managing the processing, storage, and retrieval of memories, have garnered increasing attention for their autonomy and adaptability. However, existing training paradigms remain constrained: agents often traverse …

Cited by 0SourceScholar
2026

Mitigating Error Accumulation in Knowledge Editing for Multi-Hop Question Answering

AAAI 2026technical

Knowledge editing (KE) has emerged as an effective approach for updating factual information in large language models (LLMs) without the need for full retraining. Most of the existing methods for addressing the "ripple effect" in KE adopt a chain-structured reasoning process, making them vulnerable

Cited by 0SourcePDFScholar
2026

SegMem-RAG: Adaptive Memory for Retrieval-Augmented Generation in Open-Ended Knowledge Environments

AAAI 2026technical

Retrieval-Augmented Generation (RAG) improves the factual accuracy of large language models by grounding responses in external content. However, most RAG systems assume access to static and well-organized corpora with fixed retrieval logic. In practice, real-world sources are heterogeneous and unlab

Cited by 0SourcePDFScholar
2025

AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders

NeurIPS 2025spotlight

Speculative Decoding (SD) accelerates large language model inference by employing a small draft model to generate predictions, which are then verified by a larger target model. The effectiveness of SD hinges on the alignment between these models, which is typically enhanced by Knowledge Distillation…

Cited by 0SourcecodeScholar
2025

BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

ICCV 2025poster

Monocular and stereo depth estimation offer complementary strengths: monocular methods capture rich contextual priors but lack geometric precision, while stereo approaches leverage epipolar geometry yet struggle with ambiguities such as reflective or textureless surfaces. Despite post-hoc synergies,…

2025

Combining the Best of Both Worlds: A Method for Hybrid NMT and LLM Translation

ACL 2025finding

Large language model (LLM) shows promising performances in a variety of downstream tasks, such as machine translation (MT). However, using LLMs for translation suffers from high computational costs and significant latency. Based on our evaluation, in most cases, translations using LLMs are comparabl…

2025

Design and Kinematics for the Cystoscope of a Transurethral Continuum Surgical Robotic System

IROS 2025

To achieve en bloc resection of bladder tumor and the anterior tumor resection in transurethral resection of bladder tumor (TURBT), a cystoscope transurethral continuum robotic system has been proposed. A continuum cystoscope in the system needs to bend more than 180° and its base has translation, a

Cited by 0SourceScholar
2025

Enhancing Large Language Models for Document-Level Translation Post-Editing Using Monolingual Data

COLING 2025main

The translation capabilities of neural machine translation (NMT) models based on the encoder-decoder framework are extremely potent. Although Large Language Models (LLMs) have achieved remarkable results in many tasks, they have not reached state-of-the-art performance in NMT. However, traditional N…

2025

Generative Annotation for ASR Named Entity Correction

EMNLP 2025

End-to-end automatic speech recognition systems often fail to transcribe domain-speciffcnamed entities, causing catastrophic failuresin downstream tasks. Numerous fast and lightweight named entity correction (NEC) models have been proposed in recent years. These models, mainly leveraging phonetic-le

2025

M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models

EMNLP 2025

With the widespread application of Large Language Models (LLMs) in the field of Natural Language Processing (NLP), enhancing their performance has become a research hotspot. This paper presents a novel multi-prompt ensemble decoding approach designed to bolster the generation quality of LLMs by leve

2025

M³GQA: A Multi-Entity Multi-Hop Multi-Setting Graph Question Answering Benchmark

ACL 2025long

Recently, GraphRAG systems have achieved remarkable progress in enhancing the performance and reliability of large language models (LLMs). However, most previous benchmarks are template-based and primarily focus on few-entity queries, which are monotypic and simplistic, failing to offer comprehensiv…

2024

A Novel Paradigm Boosting Translation Capabilities of Large Language Models

NAACL 2024findings

This paper presents a study on strategies to enhance the translation capabilities of large language models (LLMs) in the context of machine translation (MT) tasks. The paper proposes a novel paradigm consisting of three stages: Secondary Pre-training using Extensive Monolingual Data, Continual Pre-t…

Cited by 17SourcePDFScholar
2024

Ada-Tracker: Soft Tissue Tracking via Inter-Frame and Adaptive-template Matching

ICRA 2024poster

Soft tissue tracking is crucial for computer-assisted interventions. Existing approaches mainly rely on extracting discriminative features from the template and videos to recover corresponding matches. However, it is difficult to adopt these techniques in surgical scenes, where tissues are changing…

Cited by 2SourcecodeScholar
2023

INarIG: Iterative Non-autoregressive Instruct Generation Model For Word-Level Auto Completion

EMNLP 2023long findings

Computer-aided translation (CAT) aims to enhance human translation efficiency and is still important in scenarios where machine translation cannot meet quality requirements. One fundamental task within this field is Word-Level Auto Completion (WLAC). WLAC predicts a target word given a source senten…

Cited by 0SourceScholar
2023

Text Style Transfer Back-Translation

ACL 2023long

Back Translation (BT) is widely used in the field of machine translation, as it has been proved effective for enhancing translation quality. However, BT mainly improves the translation of inputs that share a similar style (to be more specific, translation-liked inputs), since the source side of BT d…

2023

UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction

ICASSP 2023accepted

Error correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works usually adopt end-to-end models and has strong dependency on Pseudo Paired Data and Original Paired Data. But when only p…

Cited by 0SourceScholar
2022

A Visual Navigation Perspective for Category-Level Object Pose Estimation

ECCV 2022poster

"This paper studies category-level object pose estimation based on a single monocular image. Recent advances in pose-aware generative models have paved the way for addressing this challenging task using analysis-by-synthesis. The idea is to sequentially update a set of latent variables,e.g., pose, s…

2022

Capture Human Disagreement Distributions by Calibrated Networks for Natural Language Inference

ACL 2022findings

Natural Language Inference (NLI) datasets contain examples with highly ambiguous labels due to its subjectivity. Several recent efforts have been made to acknowledge and embrace the existence of ambiguity, and explore how to capture the human disagreement distribution. In contrast with directly lear…

Cited by 10SourcePDFScholar
2021

PREGAN: Pose Randomization and Estimation for Weakly Paired Image Style Translation

RA-L 2021

Utilizing the trained model under different conditions without data annotation is attractive for robot applications. Towards this goal, one class of methods is to translate the image style from another environment to the one on which models are trained. In this letter, we propose a weakly-paired set

Cited by 1SourcecodeScholar