Hello! I am Jingyu Guo (郭靖宇 in Chinese), a Ph.D. candidate at the University of Melbourne. I am advised by Prof. Mingming Gong. My research interests are 3D computer vision and vision-language-model.

I welcome conversations and collaborations around 3D reconstruction/ generation, and vision-language-action/ navigation. Please feel free to reach out by email.

Past Research

I have worked on 3D representation learning, generative modeling, and computer vision methods for understanding sports broadcast videos.

HUGE-Bench high-level UAV trajectory examples

HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks

Jingyu Guo, Ziye Chen, Ziwen Li, Zhengqing Gao, Jiaxin Huang, Hanlue Zhang, Fengming Huang, Yu Yao, Tongliang Liu, and Mingming Gong

European Conference on Computer Vision (ECCV), 2026

TL;DR: HUGE-Bench evaluates whether UAV agents can follow concise high-level commands and execute safe multi-stage trajectories in digital twin scenes, using process-oriented and collision-aware metrics.

Hyper3D qualitative reconstruction comparison

Hyper3D: Efficient 3D Representation via Hybrid Triplane and Octree Feature for Enhanced 3D Shape Variational Auto-Encoders

Jingyu Guo, Sensen Gao, Jia-Wang Bian, Wanhu Sun, Heliang Zheng, Rongfei Jia, and Mingming Gong

arXiv preprint arXiv:2503.10403, 2025

TL;DR: Hyper3D combines octree-based mesh features with a hybrid triplane-grid latent space, improving 3D shape VAE reconstruction fidelity while keeping latent representations efficient.

Tennis analytics demo

Recognizing a Sequence of Events from Tennis Video Clips: Addressing Timestep Identification and Subtle Class Differences

Zhaoyu Liu, Jingyu Guo, Mo Wang, Ruicong Wang, Kan Jiang, and Jin Song Dong

2023 IEEE 28th Pacific Rim International Symposium on Dependable Computing (PRDC)

TL;DR: We propose an end-to-end network for temporally precise tennis event detection, addressing timestep identification and subtle class differences in broadcast clips.

Work Experience

Intern at AGIBOT, Shenzhen

May 2026 - Present

Vision-language navigation for humanoid robots

Intern at Melsy Tech, Hangzhou

Oct 2025 - May 2026

Embodied UAV research

Intern at Math Magic, Beijing

Oct 2024 - May 2025

Image-to-3D research

Intern at LG Electronics in Korea

Jun 2019 - Jul 2019

CFD analysis on air-conditioning design using ANSYS FLUENT

Photography

I like to travel and take pictures. Click for the gallery view.