← Search

Weixing Chen

11 accepted papers

2026

AlphaAgentEvo: Evolution-Oriented Alpha Mining via Self-Evolving Agentic Reinforcement Learning

ICLR 2026poster

Alpha mining seeks to identify predictive alpha factors that generate excess returns beyond the market from a vast and noisy search space; however, existing approaches struggle to facilitate the systematic evolution of alphas. Traditional methods, such as genetic programming, are unable to interpret…

Cited by 0SourceScholar
2026

DDP-WM: Disentangled Dynamics Prediction for Efficient World Models

ICML 2026poster

World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformer-based models significantly hinders real-time deployment. To address this efficiency-performance bottleneck, we introduce DDP-WM, a novel world model centered on t…

Cited by 0SourceScholar
2026

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

CVPR 2026

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlusions and during cross-task transfer. To address these challenges, we propose Task-aware Virtual View Exploration (TVVE)

Cited by 0SourcecodeScholar
2026

PhyScene3D: Physically Consistent 3D Interactive Tabletop Scene Generation

ICML 2026poster

Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense object hierarchies and irregular affordances. Existing methods, ranging from decoupled symbolic solvers to end-to-end regress…

Cited by 0SourceScholar
2025

A Multiscale Frequency Domain Causal Framework for Enhanced Pathological Analysis

ICLR 2025poster

Multiple Instance Learning (MIL) in digital pathology Whole Slide Image (WSI) analysis has shown significant progress. However, due to data bias and unobservable confounders, this paradigm still faces challenges in terms of performance and interpretability. Existing MIL methods might identify patche…

2025

Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering

ICCV 2025poster

Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, and perform multi-step reasoning to answer questions. However, current EQA approaches suffer from critical limitations in…

2025

Cross-modal Causal Relation Alignment for Video Question Grounding

CVPR 2025highlight

Video question grounding (VideoQG) requires models to answer the questions and simultaneously infer the relevant video segments to support the answers. However, existing VideoQG methods usually suffer from spurious cross-modal correlations, leading to a failure to identify the dominant visual scenes…

2025

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

CVPR 2025poster

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point c…

2025

TWLHex: A Biologically Inspired Multi-Morphology Transformable Wheel-Legged Hexapod

RA-L 2025

This paper presents a novel biologically inspired transformable wheel-legged hexapod, TWLHex, which is capable of actively switching among three morphologies: leg, wheel, and spoke. The mechanical design and kinematic models for transformable mechanism (TranMech) and wheel-spoke leg mechanism (WSLeg

Cited by 1SourceScholar
2025

Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

CVPR 2025poster

Existing Vision-Language Navigation (VLN) methods primarily focus on single-stage navigation, limiting their effectiveness in multi-stage and long-horizon tasks within complex and dynamic environments. To address these limitations, we propose a novel VLN task, named Long-Horizon Vision-Language Navi…

Cited by 5SourcePDFScholar