Shouwei Ruan

阮受炜

Ph.D. Candidate @ Institute of AI, Beihang University (BUAA)

Visiting Ph.D. Student @ TSAIL, Tsinghua University

Spatial IntelligenceEmbodied AI & World ModelsTrustworthy VLMs
Google Scholar citations GitHub stars GitHub followers
Shouwei Ruan
On my 26th birthday · 📷 by Qiyang Zhang (my girl)

01Biography

Hi! I am Shouwei Ruan (阮受炜), a final-year Ph.D. student at the Institute of Artificial Intelligence, Beihang University (BUAA), and a member of the ROSE Vision Lab led by Prof. Xingxing Wei (韦星星). Before that, I received my B.E. in Intelligent Science and Technology from Xidian University (XDU).

I am passionate about bringing AI into the physical world — how machines perceive, understand, and interact with 3D space. My research therefore centers on multimodal spatial intelligence (viewpoint-invariant visual perception and spatial reasoning, e.g., ViewFool, Omniview-Tuning, World2Mind) and extends to biologically-inspired embodied navigation, manipulation, and world modeling (e.g., BSC-Nav, AlloSpatial). My long-term vision is persistent intelligence: agents that operate in the physical world over long horizons and keep refining their cognition through experience. Going forward, I will focus on embodied intelligence and world-model-based policies.

I am fortunate to be jointly supervised by Prof. Hang Su (苏航) at TSAIL, Tsinghua University, and to have been mentored by Prof. Yinpeng Dong (董胤蓬) in the early years of my Ph.D. I have also worked as a research intern at the Foundation Model Department of Huawei 2012 Labs, RealAI (瑞莱智慧), and Dell Technologies.

I enjoy taking on challenging research problems. Lately, though, I have been trying to slow down, think more deliberately, and leave room for a healthier work–life balance — my short-term goals are simply to restore a good sleep schedule and to build muscle (hopefully sticking to the gym three times a week 🏋️).

Feel free to reach out by email if you are interested in discussing or collaborating.

02News

  • 2026.07Our T2I safety work UniNDM (extended version of NDM) has been accepted by IEEE TPAMI.
  • 2026.06🎉 BSC-NavBrain-inspired spatial intelligence for embodied agents — is published in Nature Communications (vol. 17, 8061). Code is available on GitHub.
  • 2026.06New preprint AlloSpatial: an agentic harness framework that internalizes allocentric spatial reasoning into foundation models via RL. [arXiv]
  • 2026.06Mind over Space (Video2Mental benchmark & NavMind) is presented at the CVPR 2026 Workshop. [paper]
  • 2026.03New preprint World2Mind: a training-free cognition toolkit for allocentric spatial reasoning that boosts GPT-5.2-class models by 5–18%. [arXiv]
  • 2025.09Manifold Steering (mitigating overthinking in large reasoning models) has been accepted by NeurIPS 2025.
  • 2025.08Released BSC-Nav preprint: From reactive to cognitive: brain-inspired spatial intelligence for embodied agents. [arXiv]
  • 2025.07NDM has been accepted by ACM MM 2025.
  • 2025.06🔥 3 papers accepted by ICCV 2025AdvDreamer, ITA, and SI-Attack.
  • 2025.05Breaking the Ceiling (jailbreak strategy space) has been accepted by ACL 2025 Findings.
  • 2025.01DOPatch — distributionally location-aware adversarial patches for facial images — has been accepted by IEEE TPAMI.
  • 2024.10Gave a TechBeat talk: 《探索视觉感知的3D视角鲁棒性》. [link]
  • 2024.07Omniview-Tuning has been selected as an Oral at ECCV 2024. [project page]
  • 2024.02TT3D has been accepted by CVPR 2024.
  • 2023.07VIAT has been accepted by ICCV 2023.
  • 2022.09ViewFool has been accepted by NeurIPS 2022.
⬆ scrollable

03Education

Beihang University (BUAA), Beijing, China
Ph.D. Candidate in Artificial Intelligence (Computer Science and Technology), ROSE Vision Lab, advised by Prof. Xingxing Wei.
Sep. 2022 – Jun. 2027 (expected)
Xidian University (XDU), Xi'an, China
B.E. in Intelligent Science and Technology. GPA 3.8/4.0, overall ranking 2/310.
Sep. 2018 – Jun. 2022

04Publications Google Scholar ↗

Recent Research Highlights
BSC-Nav · Nature Communications 2026Zero-shot object-goal navigation with brain-inspired cognitive maps.
BSC-Nav · Long-horizon instruction followingLandmark / route / survey knowledge retrieved by semantic goals.
BSC-Nav · Real-world mobile manipulationSpatial memory supports versatile embodied behaviors on a real robot.
Omniview-Tuning
Omniview-Tuning · ECCV 2024 Oral 🎖️
AdvDreamer
AdvDreamer · ICCV 2025
Full List

Sorted by year, then by importance. * equal contribution, corresponding author. Highlighted cards are first-author / co-first-author works.

2026

🔥 Brain-inspired Spatial Intelligence for Embodied Agents (BSC-Nav)

Shouwei Ruan, Liyuan Wang, Caixin Kang, Qihui Zhu, Songming Liu, Xingxing Wei, Hang Su
Nature Communications vol. 17, 8061, 2026 · preprint: From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
BSC-Nav is a unified framework for constructing and leveraging structured spatial memory in embodied agents. It builds allocentric cognitive maps from egocentric trajectories, consolidating landmark, route, and survey knowledge, and dynamically retrieves spatial knowledge aligned with semantic goals. Integrated with MLLMs, BSC-Nav lifts SPL from 17.6% to 44.9% on instance-level navigation and from 42.7% to 53.1% on zero-shot long-horizon instruction following, and supports versatile behaviors on real robots.

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu, Yuxiang Zhang, Jingzhi Li, Yubin Wang, Xingxing Wei
arXiv preprint 2026
AlloSpatial converts egocentric observations into allocentric spatial representations through a cognitive-mapping sandbox and a Spatial Reasoning Harness, and shows the harness can be internalized into open models (e.g., Qwen3-VL) via reinforcement learning, yielding 5–18% gains over proprietary models.

World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models

Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu, Yuxiang Zhang, Hang Su, Yubin Wang
arXiv preprint 2026
A training-free toolkit that equips multimodal foundation models with structured cognitive maps. Using 3D reconstruction and instance segmentation to build an Allocentric-Spatial Tree, it improves frontier models by 5–18% and even lets text-only models tackle complex 3D spatial tasks.

Mind over Space: Can Multimodal Large Language Models Mentally Navigate?

Qihui Zhu, Shouwei Ruan, Xiao Yang, Hao Jiang, Yao Huang, Shiji Zhao, Hanwei Fan, Hang Su, Xingxing Wei
CVPR 2026 Workshop (Viscale)
Introduces Video2Mental, a benchmark requiring MLLMs to abstract hierarchical cognitive maps from long egocentric videos and plan landmark-grounded routes, and NavMind, a reasoning model that internalizes mental navigation with explicit cognitive maps as intermediate representations.

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation

Yao Huang, Yitong Sun, Huanran Chen, Ruochen Zhang, Shouwei Ruan, Ranjie Duan, Maoxun Yuan, Yinpeng Dong, Hui Xue, Xiaochun Cao, Xingxing Wei
IEEE TPAMI 2026 (accepted)

Improving Safety Alignment via Balanced Direct Preference Optimization

Shiji Zhao, Mengyang Wang, Shukun Xiong, Fangzhou Chen, Qihui Zhu, Shouwei Ruan, Yisong Xiao, Ranjie Duan, Xun Chen, Xingxing Wei
arXiv preprint 2026

WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts

Yuxin Meng, Yuhan Suo, Junjie Wang, Yuhan Sun, Yiyao Yu, Ruixu Zhang, Ruining Hu, Yubin Wang, Shouwei Ruan, Bin Wang, Yuxiang Zhang, Yujiu Yang
arXiv preprint 2026
2025

🔥 AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?

Shouwei Ruan*, Hanqing Liu*, Yao Huang*, Xiaoqi Wang, Caixin Kang, Hang Su, Yinpeng Dong, Xingxing Wei
ICCV 2025 International Conference on Computer Vision, Honolulu, USA
AdvDreamer is the first framework to generate physically reproducible adversarial 3D transformation (Adv-3DT) samples from single-view images. With it we build MM3DTBench, the first VQA benchmark for VLMs' 3D-variation robustness, and show that real-world 3D variations pose severe threats across tasks.

🔥 When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack

Hanqing Liu*, Shouwei Ruan*, Yao Huang*, Shiji Zhao, Xingxing Wei
ICCV 2025 International Conference on Computer Vision, Honolulu, USA
Illumination Transformation Attack (ITA) is the first framework to systematically assess VLMs' robustness against illumination changes.

Distributionally Location-Aware Transferable Adversarial Patches for Facial Images

Xingxing Wei, Shouwei Ruan, Yinpeng Dong, Hang Su, Xiaochun Cao
IEEE TPAMI IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
By modeling the adversarial location distribution of patches on a surrogate model and transferring this distributional prior to black-box models, DOPatch enables efficient query-based, location-aware patch attacks and stronger defenses.

The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation

Shouwei Ruan, Zhenyu Wu, Yao Huang, Ruochen Zhang, Yitong Sun, Caixin Kang, Shiji Zhao, Xingxing Wei
arXiv preprint 2025
LibraAlign-100K (dual safety/quality annotations), T2I-SPO (synergistic preference optimization with a composite reward), and UAScore (a unified metric) together break the zero-sum trade-off between safety and quality in T2I models.

Eva-VLA: Evaluating Vision-Language-Action Models' Robustness Under Real-World Physical Variations

Hanqing Liu, Shouwei Ruan, Jiahuan Long, Junqi Wu, Jiacheng Hou, Huili Tang, Tingsong Jiang, Weien Zhou, Wen Yao
arXiv preprint 2025
The first unified framework that formulates uncontrollable physical variations (3D object transformations, illumination, adversarial patches) as continuous optimization problems to discover worst-case scenarios for VLA models; OpenVLA-class models fail over 90% of the time under such perturbations.

Mitigating Overthinking in Large Reasoning Models via Manifold Steering

Yao Huang, Huanran Chen, Shouwei Ruan, Yichi Zhang, Xingxing Wei, Yinpeng Dong
NeurIPS 2025 Advances in Neural Information Processing Systems
Shows overthinking lies on a low-dimensional activation manifold and projects steering directions onto it, cutting output tokens by up to 71% on DeepSeek-R1 distilled models while maintaining or improving accuracy.

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

Shiji Zhao, Ranjie Duan, Fengxiang Wang, Chi Chen, Caixin Kang, Shouwei Ruan, Jialing Tao, YueFeng Chen, Hui Xue, Xingxing Wei
ICCV 2025 International Conference on Computer Vision, Honolulu, USA

NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation

Yitong Sun, Yao Huang, Ruochen Zhang, Huanran Chen, Shouwei Ruan, Ranjie Duan, Xingxing Wei
ACM MM 2025 ACM International Conference on Multimedia

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space

Yao Huang, Yitong Sun, Shouwei Ruan, Yichi Zhang, Yinpeng Dong, Xingxing Wei
ACL 2025 Findings

MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models

Shiji Zhao, Qihui Zhu, Shukun Xiong, Shouwei Ruan, Maoxun Yuan, Jialing Tao, Jiexi Liu, Ranjie Duan, Jie Zhang, Jie Zhang, Xingxing Wei
arXiv preprint 2025

OODFace: Benchmarking Robustness of Face Recognition under Common Corruptions and Appearance Variations

Caixin Kang, Yubo Chen, Shouwei Ruan, Shiji Zhao, Ruochen Zhang, Jiayi Wang, Shan Fu, Xingxing Wei
arXiv preprint 2025
OODFace systematically designs 30 OOD scenarios across 9 categories (common corruptions and appearance variations) and benchmarks 19 face recognition models and 3 commercial APIs.
2024

🔥 Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models

Shouwei Ruan, Yinpeng Dong, Hanqing Liu, Yao Huang, Hang Su, Xingxing Wei
ECCV 2024 Oral European Conference on Computer Vision, Milan, Italy
To make VLP models robust to 3D viewpoint changes without sacrificing their original performance, we build MVCap — over four million multi-view image-text pairs across 100K+ objects — and design Omniview-Tuning (OVT), an efficient fine-tuning framework for viewpoint-invariant representations.

Towards Transferable Targeted 3D Adversarial Attack in the Physical World

Yao Huang, Yinpeng Dong, Shouwei Ruan, Xiao Yang, Hang Su, Xingxing Wei
CVPR 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA
TT3D rapidly reconstructs transferable, targeted 3D textured meshes from a few multi-view images for physical-world attacks.

DIFFender: Diffusion-based Adversarial Defense Against Patch Attacks

Caixin Kang, Yinpeng Dong, Zhengyi Wang, Shouwei Ruan, Yubo Chen, Hang Su, Xingxing Wei
ECCV 2024 European Conference on Computer Vision, Milan, Italy · journal extension: Real-world Adversarial Defense against Patch Attacks based on Diffusion Model (arXiv 2024)
DIFFender leverages a text-guided diffusion model and the Adversarial Anomaly Perception (AAP) phenomenon to detect, locate, and restore adversarial patches.

Exploring the Robustness of Decision-Level Through Adversarial Attacks on LLM-Based Embodied Models

Shuyuan Liu*, Jiawei Chen*, Shouwei Ruan, Hang Su, Zhaoxia Yin
ACM MM 2024 ACM International Conference on Multimedia, Melbourne, Australia
Constructs the Embodied Intelligent Robot Attack Dataset (EIRAD) and devises untargeted / targeted attack strategies to evaluate decision-level robustness of LLM-based embodied models.
2023

🔥 Towards Viewpoint-Invariant Visual Recognition via Adversarial Training

Shouwei Ruan, Yinpeng Dong, Hang Su, Jianteng Peng, Ning Chen, Xingxing Wei
ICCV 2023 International Conference on Computer Vision, Paris, France · extended version: Improving Viewpoint Robustness for Visual Recognition via Adversarial Training (arXiv 2023, with ViewRS certification)
Viewpoint-Invariant Adversarial Training (VIAT) formulates viewpoint robustness as a minimax problem whose inner maximization learns a Gaussian-mixture distribution of adversarial viewpoints via GMVFool. We also release ImageNet-V+, an OOD benchmark with nearly 100K adversarial-viewpoint images.
2022

🔥 ViewFool: Evaluating the Robustness of Visual Recognition to Adversarial Viewpoints

Yinpeng Dong, Shouwei Ruan, Hang Su, Caixin Kang, Xingxing Wei, Jun Zhu
NeurIPS 2022 Advances in Neural Information Processing Systems, New Orleans, USA
ViewFool encodes real-world objects as neural radiance fields (NeRF) and characterizes a distribution of adversarial viewpoints under an entropic regularizer, revealing that visual recognition models are highly vulnerable to viewpoint changes.

05Research & Visiting Experience

TSAIL Group, Tsinghua University (THU), Beijing, China
Visiting Ph.D. Student. Supervised by Prof. Hang Su; closely working with Songming Liu and Liyuan Wang.
2023 – Present
Foundation Model Department, Huawei 2012 Laboratories (华为2012实验室 基础大模型部), China
Research Intern — multimodal foundation models & spatial intelligence.
Jan. 2026 – Sep. 2026
Department of AI Security, RealAI (瑞莱智慧), Beijing, China
Research Intern in AI Security. Supervised by Prof. Hang Su and A/Prof. Yinpeng Dong.
Mar. 2022 – Aug. 2022
Digital Team, Dell Technologies (戴尔科技中国), Xiamen, China
Research Intern in Deep Learning. Mentored by SSE. Xuejin Liang (梁学锦).
Jul. 2021 – Sep. 2021
College of Engineering, Hanyang University (汉阳大学), Seoul, South Korea
Visiting Student in Urban Environment and Universal Design. Mentored by Prof. Yoojin Lee.
Jul. 2019 – Sep. 2019

06Academic Services

Reviewer

  • IEEE TPAMI · IJCV · IEEE TMM
  • NeurIPS 2025 · ICCV 2025 · CVPR 2025/2026 · ICLR 2025/2026

Affiliations

  • Mobile Researcher, QiYuan Lab (启元实验室), Beijing.
  • Contracted Mentor, DeepShare (深度之眼), Shanghai.

Teaching Assistant / Lectures

  • Multi-modal Large Language Models: Basic Principles and Applications (多模态大模型:基础原理与前沿应用) — Graduate Course on Pattern Recognition, BUAA, 2023. [Slides]
  • Large Language Models (大语言模型技术介绍) — Graduate Course on Pattern Recognition, BUAA, 2024. [Slides]
  • Adversarial Attack Methods (对抗攻击方法) — Graduate Course on AI Security and Ethics, BUAA, 2023. [Slides]

07Talks & Reports

08Awards & Honors selected

2021 · China College Students' "Internet+" Innovation & Entrepreneurship Competition
中国国际“互联网+”大学生创新创业大赛
🏆 National Gold Award · 6th place, Main Track Grand Final
With Ning Ji, et al.
2021 · Intern Program Final Presentation @ Dell Technologies
戴尔科技实习生项目
🥉 Third Place
With Chengke Fan · Supervised by SSE. Xuejin Liang
2021 · Mathematical Contest in Modeling (MCM)
美国大学生数学建模竞赛
🏅 Meritorious Winner
With Ning Ji, Ziyue Zhang · Supervised by A/Prof. Shengli Zhang
  • National Scholarship (国家奖学金), Xidian University, 2021.
  • Outstanding Graduate & Innovation Model (西电创新楷模), Xidian University, 2022.

09Highlight Moments

⬅ ➡ scrollable

10Visitors since Aug. 2026

🌍 Where visitors come frompowered by ClustrMaps
👤 Unique visitors
👁 Total page views