← Search

Zhixuan Liu

13 accepted papers

2026

Native Reasoning Models: Training Language Models to Reason on Unverifiable Data

ICLR 2026poster

The dominant paradigm for training large reasoning models—combining Supervised Fine-Tuning (SFT) with Reinforcement Learning with Verifiable Rewards (RLVR)—is fundamentally constrained by its reliance on high-quality, human-annotated reasoning data and external verifiers. This dependency incurs sign…

Cited by 0SourceScholar
2026

STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation

ICRA 2026poster

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs to navigation presents two key challenges: effectively parsing and structuring complex environment information and determ…

2026

Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory

AAAI 2026technical

Reasoning Video Object Segmentation (ReasonVOS) is a challenging task that requires stable object segmentation across video sequences using implicit and complex textual inputs. Previous methods fine-tune Multimodal Large Language Models (MLLMs) to produce segmentation outputs, which demand substanti

Cited by 0SourcePDFScholar
2025

MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room Environments

ICCV 2025poster

We introduce a diffusion-based approach for generating privacy-preserving digital twins of multi-room indoor environments from depth images only. Central to our approach is a novel Multi-view Overlapped Scene Alignment with Implicit Consistency (MOSAIC) model that explicitly considers cross-view dep…

Cited by 0SourcePDFScholar
2024

Inference-Time Language Model Alignment via Integrated Value Guidance

EMNLP 2024finding

Large language models are typically fine-tuned to align with human preferences, but tuning large models is computationally intensive and complex. In this work, we introduce **Integrated Value Guidance (IVG)**, a method that uses implicit and explicit value functions to guide language model decoding…

Cited by 6SourcePDFScholar
2024

SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation

CVPR 2024poster

Accurate representation in media is known to improve the well-being of the people who consume it. Generative image models trained on large web-crawled datasets such as LAION are known to produce images with harmful stereotypes and misrepresentations of cultures. We improve inclusive representation i…

Cited by 17SourcePDFScholar
2024

Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models

NeurIPS 2024poster

Large language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce $\textit{weak-to-strong search}$, framing the alignment of a large language model as a test-time greedy search to maximize the log-…

2023

Grasp Region Exploration for 7-DoF Robotic Grasping in Cluttered Scenes

IROS 2023poster

Robotic grasping is a fundamental skill for robots, but it is quite challenging in cluttered scenes. In cluttered scenes, the precise prediction of high-quality grasp configurations such as rotation and grasping width while avoiding collisions is essential. To accomplish this, the grasp detection mo…

Cited by 7SourceScholar
2022

A Neural-Symbolic Approach to Natural Language Understanding

EMNLP 2022finding

Deep neural networks, empowered by pre-trained language models, have achieved remarkable results in natural language understanding (NLU) tasks. However, their performances can drastically deteriorate when logical reasoning is needed. This is because NLU in principle depends on not only analogical re…

2022

StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation

IJCAI 2022poster

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however, current methods lack artistic control of the style of image to be generated. We present an approach for generating styl…

2022

TransGrasp: A Multi-Scale Hierarchical Point Transformer for 7-DoF Grasp Detection

ICRA 2022poster

Robotic grasping pose detection that predicts the configuration of the robotic gripper for object grasping is fundamental in robot manipulation. Based on point clouds, most of the existing methods predict grasp pose with the hierarchical PointNet++ backbone, while the non-local geometric information…

Cited by 29SourceScholar