Yun Hong (洪运)

I am a B.S. candidate in Computer Science at Fudan University and a former exchange student at The University of Texas at Austin.

I conduct research at MagicLab, Fudan University, under the guidance of Wenchao Ding. My current research focuses on execution feedback and subtask handoffs for world-action models in long-horizon robotic manipulation.

From January to June 2026, I was a Test Development Engineer Intern at ByteDance (Douyin R&D, Live Streaming & Media Stream), where I worked on LLM-powered QA efficiency tools.

Portrait of Yun Hong

Research Interests

I am interested in embodied intelligence, robotics, and machine learning — particularly world models for embodied agents and their role in action evaluation, future prediction, and planning under uncertainty. I also work on imitation learning, visuomotor policy learning, and vision-language-action models.

Education

B.S. Candidate in Computer Science, Shanghai, China

Studying execution feedback and subtask handoffs for world-action models at MagicLab under Wenchao Ding. Previously contributed to risk-map research for autonomous driving, published in IEEE Robotics and Automation Letters (RA-L) in 2026.

Exchange Student in Computer Science, Austin, Texas, USA

GPA: 4.0/4.0 (13 credits); earned A grades in machine learning, computer vision, operating systems, and probability/random processes.

Conducted multi-task active perception research at RobIn Lab under Junhong Xu, contributing to benchmark design and baseline evaluation workflows.

Research Experience

World-Action Models for Closed-Loop Robotic Manipulation

Advised by Wenchao Ding

Developing an execution-feedback framework around LingBot-VA on LIBERO-Long to study how task-progress signals and discrepancies between predicted futures and observations can condition the same policy to revise its actions.

Current experiments focus on subtask handoffs: how arm and gripper configurations, reachability, and streaming observation history affect recovery. I use state replay and controlled interventions to distinguish changes in model outputs from task recovery, and investigate how favorable simulated starting states can be reached through safe, executable actions.

Learning a Unified Risk Map for Autonomous Driving in Partially Observable Environments

Jie Jia, Yaofeng Su, Zeyu Bao, Yun Hong, Bingzhao Gao, Zhongxue Gan, Wenchao Ding
IEEE Robotics and Automation Letters (RA-L), 11(4): 4457–4464, 2026
Paper / arXiv

Contributed to the Transformer-based risk-prediction framework, model training, debugging, and result verification at MagicLab under Wenchao Ding.

Multi-Task Active Perception for Robotic Information Gathering

Research Assistant, advised by Junhong Xu

Built multi-task partially observable simulation benchmarks with unknown object properties and spatial locations. Developed baseline training and evaluation workflows for VLA policies to study imitation learning, reinforcement learning, and IL pretraining followed by RL fine-tuning. Reviewed prior work on manipulation exploration and long-horizon POMDP benchmarks.

ChatGPT-o1 Reproduction and Reflection Ability Improvement

Research Assistant, advised by Prof. Xuanjing Huang

Constructed, evaluated, and optimized reflection datasets to improve model reasoning. Synthesized self-critic data based on Llama3.1-8b-instruct and generated diverse error-to-correct reflection trajectories. Applied tree-search and reverse-reasoning strategies; evaluated on GSM8K.

Internship

Test Development Engineer Intern, Douyin R&D, Live Streaming & Media Stream, Shanghai

Developed three AI-assisted QA tools for routine monitoring, test-case risk checking, and alert summarization. Refactored and expanded internal inspection and test cases, improving monitoring accuracy from 89% to 100% and coverage from 95% to 98% within the evaluated scope.

Projects

Language-Conditioned Diffusion Policy for Robotic Control

Built an end-to-end Vision-Language-Action system using diffusion policy. Integrated CLIP encoder and U-Net DDPM/DDIM architecture for high-fidelity action generation. Implemented 16-step action chunking with a full ManiSkill2 data-train-eval pipeline.

Route Planning and Navigation System

Combined OpenStreetMap with Gaode Maps API for route planning and location search. Implemented hybrid point selection, detailed route display, and multi-road-type support with speed constraints.

Modular Neural Style Transfer Toolkit

Developed a modular Python toolkit for neural style transfer with strong maintainability. Built batch processing and a configurable training pipeline for reproducible experiments. Provided an interactive web interface for in-browser stylization.

Skills

Machine Learning
Supervised & unsupervised learning, reinforcement learning, imitation learning, model evaluation and generalization
Deep Learning
CNN/RNN/LSTM, diffusion models, spatiotemporal modeling, PyTorch-based model design, training, and fine-tuning
LLM & Vision-Language
LLaMA series fine-tuning, VLA models, reasoning and reflection mechanism design
Robotics & Perception
Active perception, multi-task POMDP, information gathering, long-horizon manipulation
Autonomous Driving
Motion planning, trajectory prediction, decision-making, risk-field modeling
Programming
C, C++, Python, LaTeX
Languages
Chinese (native), English (fluent)

Honors & Awards

International
2024: Finalist, Water, sanitation, and hygiene for the prevention and care of neglected tropical diseases (Geneva, Switzerland)
Domestic
2026: National Second Prize, 3rd “Jinlingguang Cup” China Internet Innovation Competition, University Student Innovation and Entrepreneurship Commercialization Track (team project: 星火延生 / Xinghuo Yansheng)
2026: Award recipient, Tencent Rhino-Bird Open-Source Talent Program, OpenCloudOS Security Skill issue (one of three awardees)
2025: Outstanding Work Award, The 5th Meituan Business Analysis Elite Competition
2024: Shanghai Third Prize, China Undergraduate Mathematical Contest in Modeling (CUMCM)
2023: Finalist, Full-stack AI development engineer skills training by NVIDIA