Guanzhou Ke · 柯冠舟
Active embodied intelligence
under partial observability
Active perception World models Self-improving agents
I study active embodied intelligence under partial observability, with a focus on active perception, world models, and self-improving agents. My current work is grounded in UAV autonomy and inspection.
Bio
I am Guanzhou Ke, an Embodied AI Researcher at Avant Robotics in Shenzhen, working on active embodied intelligence under partial observability. My current work studies how autonomous drones can actively acquire task-relevant evidence, learn from difficult simulated and real-world scenarios, and improve through an evaluation–data–training loop.
I received my Ph.D. from Beijing Jiaotong University and was a CSC visiting Ph.D. researcher at Singapore Management University. My earlier research on multi-view representation learning and missing-modality completion provides the foundation for studying decision-making under partial observability.
Research agenda
A closed loop for reliable autonomy
Active Perception for UAV Inspection
How should a drone coordinate wide-field cameras, a high-resolution gimbal, and body motion to obtain sufficient evidence for an inspection task?
Self-evolving Simulation and Data Engines
Building evaluation-driven loops that identify failure modes, generate targeted interaction data, and improve navigation and action models.
Reliable Multimodal Intelligence
Learning and reasoning when observations are incomplete, missing, uncertain, or viewpoint-dependent.
Featured systems and research
From evaluation to evidence
Simulation-driven embodied learning loop
Problem. UAV policies need repeatable evaluation across navigation, exploration, object search, and obstacle-avoidance tasks—not only more undirected data.
Individual role. As an Embodied AI Researcher, I work on the simulation, evaluation, and data-engine pipeline that connects hard-case discovery, targeted interaction collection, and model improvement across VLN/VLA, object search, exploration, navigation, and obstacle avoidance.
Team evidence. The resulting team system produces 39 million valid simulated interaction steps per month, improves sampling throughput by 3× on a single RTX 5090, and supports a 0.8B world/action model reporting 74% navigation-and-avoidance success.
Real-world-grounded scenario construction
Problem. Useful simulated environments must preserve real-world coordinates, task executability, and repeatability rather than act as generic text-to-3D assets.
Contribution. Current work studies geospatially aligned scenario construction and the return of evaluation failures into the data loop.
Disclosure boundary. Specific scene sources and named locations remain private.
Active evidence acquisition for UAV inspection
Question. When evidence is incomplete, how should a UAV choose what to inspect next, where to move, and when enough visual evidence has been gathered?
Current direction. Jointly reason over viewpoint, sensing resolution, gimbal control, latency, bandwidth, and safety constraints. No completed benchmark or public system is claimed.
Selected publications
Foundations for partial observability




Selected news
Recent milestones
- Paper accepted at IEEE T-PAMI.
- Recognized as an ICML 2026 Gold Reviewer.
- Paper accepted at ICML 2026.
- Paper accepted to the CVPR 2026 Findings track.
- Knowledge Bridger accepted at CVPR 2025.
Experience and service
Research across systems and learning
Experience
- Avant Robotics, Embodied AI Researcher, Shenzhen · Dec. 2025–present
Active embodied intelligence, world/action models, data engines, and UAV autonomy. - Microsoft Research Asia, Research Intern · Feb.–Oct. 2024
- Institute of Automation, CAS, Research Intern · Jun.–Dec. 2023
- Singapore Management University, CSC Visiting Ph.D. Researcher · Oct. 2024–Oct. 2025
Service
Reviewer for journals including IEEE TMM, T-CSVT, and T-NNLS, and conferences including NeurIPS, CVPR, ICML, AAAI, and ACM MM.
Contact
Let’s exchange ideas
I welcome conversations about UAV autonomy, active perception, simulation and data engines, and reliable multimodal learning.
