Vision-Language-Action Models
Status monitoring, error recovery, dynamic reasoning, and test-time compute scaling for embodied agents.
I am a Ph.D. student in Computer Science at the University of Sydney, advised by Prof. Chang Xu. My work focuses on Vision-Language-Action models for embodied intelligence, multimodal LLM reliability, efficient diffusion models, and autonomous research agents.
My current research connects embodied agents, multimodal reasoning, model self-monitoring, and efficient generative modeling.
Status monitoring, error recovery, dynamic reasoning, and test-time compute scaling for embodied agents.
Motion-aware VLA systems for conveyor-belt and other latency-sensitive manipulation settings.
Identifying, isolating, and purging hallucination components through self-evolving distillation.
Architecture search, progressive scaling, and research-agent systems that explore and accumulate experience.
A snapshot of recent first-author work across embodied AI, multimodal reliability, and efficient generation.
Equips VLA models with active execution-state monitoring and recovery behavior for robust embodied tasks.
Uses an action critic to allocate deliberation at test time when a task requires more careful reasoning.
Analyzes and removes hallucination-related knowledge components to improve multimodal model reliability.
Uses LLM-guided architecture search to discover stronger diffusion backbones under practical budgets.
Research training across academia, robotics, and industrial AI labs.
Advised by Prof. Chang Xu, working on VLA, embodied intelligence, multimodal models, and generative AI.
Research on VLA state monitoring, error handling, and test-time scaling.
Worked on multimodal hallucination mitigation and lightweight diffusion-based image generation.
Graduated with an average score of 89/100 and multiple first-class innovation scholarships.