中文主页

👋 Profile

I am Guanzhou Ke (柯冠舟), an Embodied AI Researcher at Avant Robotics in Shenzhen. I study active embodied intelligence under partial observability, with a focus on active perception, world models, and self-improving agents. At Avant Robotics, I develop outdoor robot simulation infrastructure and cross-embodiment long-range navigation models, spanning UAVs, quadrupeds, wheeled robots, and humanoids.

My research asks how embodied agents can actively acquire task-relevant evidence, learn from difficult simulated and real-world scenarios, and improve through an evaluation–data–training loop. I received my Ph.D. from Beijing Jiaotong University and was a CSC visiting Ph.D. researcher at Singapore Management University. My earlier work on multi-view representation learning and missing-modality completion provides the foundation for reasoning and action under incomplete observations.

Links: Google Scholar · ORCID · GitHub · ResearchGate · Email

🎯 Research Agenda

  1. Active Perception for UAV Inspection — Ongoing research. How should a drone coordinate wide-field cameras, a high-resolution gimbal, and body motion to decide what, where, and when to observe under viewpoint, bandwidth, latency, and safety constraints?
  2. Cross-embodiment Navigation and Simulation Infrastructure. Build shared outdoor simulation environments and study how robots with different bodies can reuse semantic and place memory, adapt routes to their motion capabilities, and recover from navigation failures.
  3. Reliable Multimodal Intelligence under Partial Observability. Learn and reason when observations are incomplete, missing, uncertain, or viewpoint-dependent, connecting prior multimodal work with embodied reasoning and action.

🚁 Current Systems and Research

Avant-AGSim: Outdoor robot simulation infrastructure

I develop Avant-AGSim, an outdoor navigation simulation engine built on UE5, CARLA, and ProjectAirSim. It brings scenes, dynamics, sensors, and control interfaces into a shared environment for aerial and ground robots, with reusable scene and task configurations for cross-embodiment experiments.

My work includes Python tools for automated data collection, synchronized visual observations, trajectories, and scene annotations, as well as target-visibility checks and data-quality filtering. I validated alignment and target annotations on 800 synchronized aerial–ground sample groups and implemented checks that reject seven categories of data anomalies. The engine supports independent robot control, synchronized observations, fault isolation, and recording and replay; validation included 1,000 repeated loading tests and extended stability testing.

Cross-embodiment long-range navigation: A 4B vision-language model

I am developing a 4B vision-language navigation model that combines language instructions, first-person observation histories, and trajectory actions to model spatial goals and execution paths. Using Avant-AGSim, I construct training data across UAVs, quadrupeds, wheeled robots, and humanoids, covering different viewpoints, motion capabilities, and environmental conditions.

This ongoing research studies shared semantic and place memory while modeling robot-specific traversability separately. I design methods for observation-history compression, task-progress tracking, and failure recovery through blockage detection, alternative-route planning, and memory updates. Outdoor tasks range from 350 m to 1.5 km, with evaluation focused on navigation success, path efficiency, and recovery.

📣 News

  • [05/2026] One paper accepted at IEEE T-PAMI.
  • [05/2026] Recognized as an ICML 2026 Gold Reviewer.
  • [05/2026] One paper accepted at ICML 2026.
  • [03/2026] One paper accepted to the CVPR 2026 Findings track.
  • [06/2025] One paper accepted at ICCV 2025.
  • [02/2025] Knowledge Bridger accepted at CVPR 2025.
  • [12/2024] Two papers accepted at AAAI 2025.
  • [10/2024] Started a one-year CSC visiting Ph.D. appointment at Singapore Management University.
  • [02/2024] Joined Microsoft Research Asia as a research intern.
  • [02/2024] MRDD accepted at CVPR 2024.
  • [10/2023] One paper accepted in Information Fusion.
  • [07/2023] DMRIB accepted at ACM MM 2023.
  • [10/2022] One paper accepted at the ICDM 2022 Workshop.
  • [09/2022] Started Ph.D. studies at Beijing Jiaotong University.
  • [12/2021] One paper accepted at IEEE BigData 2021.

💼 Experience

  • 12/2025 – Present: Embodied AI Researcher
    • Avant Robotics, Shenzhen, China.
    • Avant-AGSim: shared outdoor simulation for aerial and ground robots, automated data generation, and multi-robot control and replay.
    • 4B vision-language navigation: cross-embodiment memory transfer, long-range route planning, and failure recovery.
    • Mentor: Zhenguo Li.
  • 02/2024 – 10/2024: Research Intern
    • Microsoft Research Asia, Shanghai AI/ML Group.
    • Multimodal medical report generation and hallucination mitigation.
    • Mentor: Xinyang Jiang.
  • 06/2023 – 12/2023: Research Intern
    • Institute of Automation, Chinese Academy of Sciences.
    • Multimodal deepfake detection across visual, textual, and audio signals.
    • Mentor: Bo Wang.

🎓 Education

  • Ph.D., Management Science and Engineering, Beijing Jiaotong University, 2022–2026.
  • CSC Visiting Ph.D. Researcher, Computer Science, Singapore Management University, 2024–2025. Advisor: Prof. Shengfeng He.
  • M.S., Systems Engineering, Wuyi University, 2019–2022. Outstanding Thesis Award.
  • B.Eng., Communication Engineering, Wuyi University, 2017–2019. Outstanding Graduate.

📄 Selected Publications

Diagram for How Far Are We from Generating Missing Modalities with Foundation Models?
How Far Are We from Generating Missing Modalities with Foundation Models?
Guanzhou Ke, Bo Wang, Guoqing Chao, Weiming Hu, Shengfeng He
IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI)
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Knowledge Bridger: Towards Training-free Missing Multi-modality Completion Venue banner for Knowledge Bridger: Towards Training-free Missing Multi-modality Completion
Knowledge Bridger: Towards Training-free Missing Multi-modality Completion
Guanzhou Ke, Shengfeng He, Xiao-Li Wang, Bo Wang, Guoqing Chao, Yuanyang Zhang, Xie Yi and HeXing Su
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Rethinking Multi-view Representation Learning via Distilled Disentangling
Rethinking Multi-view Representation Learning via Distilled Disentangling
Guanzhou Ke, Bo Wang, Xiaoli Wang, and Shengfeng He
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Disentangling Multi-view Representations Beyond Inductive Bias
Disentangling Multi-view Representations Beyond Inductive Bias
Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, and Shengfeng He
The 31st ACM International Conference on Multimedia (ACM MM 2023)
Rank: CCF A, MISC: [PDF] [CODE]

📚 Full Publications

For the latest citation record, see Google Scholar.

2026

Diagram for How Far Are We from Generating Missing Modalities with Foundation Models?
How Far Are We from Generating Missing Modalities with Foundation Models?
Guanzhou Ke, Bo Wang, Guoqing Chao, Weiming Hu, Shengfeng He
IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI)
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Reliable Neighborhood-Aware Multi-View Outlier Detection Venue banner for Reliable Neighborhood-Aware Multi-View Outlier Detection
Reliable Neighborhood-Aware Multi-View Outlier Detection
Huijie Ma, Haoyuan Xin, Lei Meng, Guanzhou Ke, Yongyong Chen, Guoqing Chao
International Conference on Machine Learning (ICML), 2026
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for OKGraph: Online Knowledge Graph Probing for Open-vocabulary Recognition Venue banner for OKGraph: Online Knowledge Graph Probing for Open-vocabulary Recognition
OKGraph: Online Knowledge Graph Probing for Open-vocabulary Recognition
Junhui Yin, Zhizhen Cai, Puze Wang, Guanzhou Ke, Jianhua Yang, Man Zhang, Qiang Zhang, and Shengfeng He
IEEE/CVF Conference on Computer Vision and Pattern Recognition Findings Track (CVPR Findings), 2026
Rank: CCF A, MISC: [PDF] [CODE]

2025

Diagram for LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation Learning Venue banner for LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation Learning
LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation Learning
Jiang Yuan, JI Ma, Bo Wang, Guanzhou Ke, Weiming Hu
International Conference on Computer Vision (ICCV), 2025
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Knowledge Bridger: Towards Training-free Missing Multi-modality Completion Venue banner for Knowledge Bridger: Towards Training-free Missing Multi-modality Completion
Knowledge Bridger: Towards Training-free Missing Multi-modality Completion
Guanzhou Ke, Shengfeng He, Xiao-Li Wang, Bo Wang, Guoqing Chao, Yuanyang Zhang, Xie Yi and HeXing Su
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Global-Semantic Alignment Distillation for Partial Multi-view Classification Venue banner for Global-Semantic Alignment Distillation for Partial Multi-view Classification
Global-Semantic Alignment Distillation for Partial Multi-view Classification
Xiao-Li Wang, Anqi Huang, Yongli Wang, Guanzhou Ke, Xiaobin Hong, and Jun Liu
The 39th Annual AAAI Conference on Artificial Intelligence (AAAI)
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Incomplete Multi-view Clustering via Diffusion Contrastive Generation Venue banner for Incomplete Multi-view Clustering via Diffusion Contrastive Generation
Incomplete Multi-view Clustering via Diffusion Contrastive Generation
Yuanyang Zhang, Weiqing Yan, Yijie Lin, Li Yao, Xinhang Wan, Guangyuan Li, Chao Zhang, Guanzhou Ke, and Jie Xu
The 39th Annual AAAI Conference on Artificial Intelligence (AAAI)
Rank: CCF A, MISC: [PDF] [CODE]

2024

Diagram for Rethinking Multi-view Representation Learning via Distilled Disentangling
Rethinking Multi-view Representation Learning via Distilled Disentangling
Guanzhou Ke, Bo Wang, Xiaoli Wang, and Shengfeng He
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
Rank: CCF A, MISC: [PDF] [CODE]

2023

Diagram for Knowledge distillation-driven semi-supervised multi-view classification
Knowledge distillation-driven semi-supervised multi-view classification
Xiaoli Wang, Yongli Wang, Guanzhou Ke, Yupeng Wang, and Xiaobin Hong
Information Fusion
Rank: SCI Q1, MISC: [PDF] [CODE]
Diagram for A Clustering-guided Contrastive Fusion for Multi-view Representation Learning
A Clustering-guided Contrastive Fusion for Multi-view Representation Learning
Guanzhou Ke, Guoqing Chao, Xiaoli Wang, Chenyang Xu, Yongqi Zhu, and Yang Yu
IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)
Rank: CCF B, MISC: [PDF] [CODE]
Diagram for Disentangling Multi-view Representations Beyond Inductive Bias
Disentangling Multi-view Representations Beyond Inductive Bias
Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, and Shengfeng He
The 31st ACM International Conference on Multimedia (ACM MM 2023)
Rank: CCF A, MISC: [PDF] [CODE]

2022

Diagram for MORI-RAN: Multi-view Robust Representation Learning via Hybrid Contrastive Fusion
MORI-RAN: Multi-view Robust Representation Learning via Hybrid Contrastive Fusion
Guanzhou Ke, Yongqi Zhu, and Yang Yu
ICDM workshop
Rank: CCF B, MISC: [PDF] [CODE]
Diagram for Efficient Multi-view Clustering Networks
Efficient Multi-view Clustering Networks
Guanzhou Ke, Zhiyong Hong, Wenhua Yu, Xin Zhang, and Zeyi Liu
Applied Intelligence Springer
Rank: CCF C, MISC: [PDF] [CODE]

2021

Diagram for CONAN: Contrastive Fusion Networks for Multi-view Clustering
CONAN: Contrastive Fusion Networks for Multi-view Clustering
Guanzhou Ke, Zhiyong Hong, Zhiqiang Zeng, Zeyi Liu, Yangjie Sun, and Yannan Xie
IEEE International Conference on Big Data (Big Data)
Rank: CCF C, MISC: [PDF] [CODE]

🏆 Awards

  • Second Prize, “Huawei Cup” National Graduate Mathematical Modeling Competition, 2020, 2021, and 2022.
  • Second Prize, National Finals, Blue Bridge Cup Information Competition Group B, 2018.
  • National Scholarship of China, 2015.

🤝 Academic Service

  • Journals: IEEE Transactions on Multimedia, IEEE Transactions on Circuits and Systems for Video Technology, IEEE Transactions on Neural Networks and Learning Systems, Neurocomputing, and others.
  • Conferences: NeurIPS, CVPR, ICML (Gold Reviewer), AAAI, ACM Multimedia, and others.

📬 CV and Contact

The downloadable English and Chinese CV files are being refreshed to synchronize the August 2026 graduation status, current title, and research positioning. Until then, this homepage is the current public profile.

Email: guanzhouk@gmail.com · Google Scholar · ORCID · GitHub