Shouwei Ruan

Ph.D. Candidate @ Institute of AI, Beihang University (BUAA)
Visiting Ph.D. Student @ TSAIL, Tsinghua University
01Biography
Hi! I am Shouwei Ruan (阮受炜), a final-year Ph.D. student at the Institute of Artificial Intelligence, Beihang University (BUAA), and a member of the ROSE Vision Lab led by Prof. Xingxing Wei (韦星星). Before that, I received my B.E. in Intelligent Science and Technology from Xidian University (XDU).
I am passionate about bringing AI into the physical world — how machines perceive, understand, and interact with 3D space. My research therefore centers on multimodal spatial intelligence (viewpoint-invariant visual perception and spatial reasoning, e.g., ViewFool, Omniview-Tuning, World2Mind) and extends to biologically-inspired embodied navigation, manipulation, and world modeling (e.g., BSC-Nav, AlloSpatial). My long-term vision is persistent intelligence: agents that operate in the physical world over long horizons and keep refining their cognition through experience. Going forward, I will focus on embodied intelligence and world-model-based policies.
I am fortunate to be jointly supervised by Prof. Hang Su (苏航) at TSAIL, Tsinghua University, and to have been mentored by Prof. Yinpeng Dong (董胤蓬) in the early years of my Ph.D. I have also worked as a research intern at the Foundation Model Department of Huawei 2012 Labs, RealAI (瑞莱智慧), and Dell Technologies.
I enjoy taking on challenging research problems. Lately, though, I have been trying to slow down, think more deliberately, and leave room for a healthier work–life balance — my short-term goals are simply to restore a good sleep schedule and to build muscle (hopefully sticking to the gym three times a week 🏋️).
Feel free to reach out by email if you are interested in discussing or collaborating.
02News
- 2026.07Our T2I safety work UniNDM (extended version of NDM) has been accepted by IEEE TPAMI.
- 2026.06🎉 BSC-Nav — Brain-inspired spatial intelligence for embodied agents — is published in Nature Communications (vol. 17, 8061). Code is available on GitHub.
- 2026.06New preprint AlloSpatial: an agentic harness framework that internalizes allocentric spatial reasoning into foundation models via RL. [arXiv]
- 2026.06Mind over Space (Video2Mental benchmark & NavMind) is presented at the CVPR 2026 Workshop. [paper]
- 2026.03New preprint World2Mind: a training-free cognition toolkit for allocentric spatial reasoning that boosts GPT-5.2-class models by 5–18%. [arXiv]
- 2025.09Manifold Steering (mitigating overthinking in large reasoning models) has been accepted by NeurIPS 2025.
- 2025.08Released BSC-Nav preprint: From reactive to cognitive: brain-inspired spatial intelligence for embodied agents. [arXiv]
- 2025.07NDM has been accepted by ACM MM 2025.
- 2025.06🔥 3 papers accepted by ICCV 2025 — AdvDreamer, ITA, and SI-Attack.
- 2025.05Breaking the Ceiling (jailbreak strategy space) has been accepted by ACL 2025 Findings.
- 2025.01DOPatch — distributionally location-aware adversarial patches for facial images — has been accepted by IEEE TPAMI.
- 2024.10Gave a TechBeat talk: 《探索视觉感知的3D视角鲁棒性》. [link]
- 2024.07Omniview-Tuning has been selected as an Oral at ECCV 2024. [project page]
- 2024.02TT3D has been accepted by CVPR 2024.
- 2023.07VIAT has been accepted by ICCV 2023.
- 2022.09ViewFool has been accepted by NeurIPS 2022.
03Education


04Publications Google Scholar ↗


Sorted by year, then by importance. * equal contribution, † corresponding author. Highlighted cards are first-author / co-first-author works.

🔥 Brain-inspired Spatial Intelligence for Embodied Agents (BSC-Nav)

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models

Mind over Space: Can Multimodal Large Language Models Mentally Navigate?

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation

Improving Safety Alignment via Balanced Direct Preference Optimization

WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts

🔥 AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?

🔥 When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack

Distributionally Location-Aware Transferable Adversarial Patches for Facial Images

The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation

Eva-VLA: Evaluating Vision-Language-Action Models' Robustness Under Real-World Physical Variations

Mitigating Overthinking in Large Reasoning Models via Manifold Steering


NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space

MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models

OODFace: Benchmarking Robustness of Face Recognition under Common Corruptions and Appearance Variations

🔥 Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models

Towards Transferable Targeted 3D Adversarial Attack in the Physical World

DIFFender: Diffusion-based Adversarial Defense Against Patch Attacks

Exploring the Robustness of Decision-Level Through Adversarial Attacks on LLM-Based Embodied Models

🔥 Towards Viewpoint-Invariant Visual Recognition via Adversarial Training

🔥 ViewFool: Evaluating the Robustness of Visual Recognition to Adversarial Viewpoints
05Research & Visiting Experience





06Academic Services
Reviewer
- IEEE TPAMI · IJCV · IEEE TMM
- NeurIPS 2025 · ICCV 2025 · CVPR 2025/2026 · ICLR 2025/2026
Affiliations
- Mobile Researcher, QiYuan Lab (启元实验室), Beijing.
- Contracted Mentor, DeepShare (深度之眼), Shanghai.
Teaching Assistant / Lectures
- Multi-modal Large Language Models: Basic Principles and Applications (多模态大模型:基础原理与前沿应用) — Graduate Course on Pattern Recognition, BUAA, 2023. [Slides]
- Large Language Models (大语言模型技术介绍) — Graduate Course on Pattern Recognition, BUAA, 2024. [Slides]
- Adversarial Attack Methods (对抗攻击方法) — Graduate Course on AI Security and Ethics, BUAA, 2023. [Slides]
07Talks & Reports
- 2024.10 《Talk|北京航空航天大学阮受炜:探索视觉感知的3D视角鲁棒性》 @ TechBeat — sharing our works on viewpoint robustness and invariance (ViewFool, VIAT, Omniview-Tuning).
- 2022.05 《国奖风采展|阮受炜:执着求索,破而后立》 @ XDU
- 2021.11 《润物耕心|以我之肩,载你之梦——竹园3号书院科研竞赛学长经验分享会》 @ XDU
08Awards & Honors selected
2021 · China College Students' "Internet+" Innovation & Entrepreneurship Competition
2021 · Intern Program Final Presentation @ Dell Technologies
2021 · Mathematical Contest in Modeling (MCM)
- National Scholarship (国家奖学金), Xidian University, 2021.
- Outstanding Graduate & Innovation Model (西电创新楷模), Xidian University, 2022.
09Highlight Moments





