Long-horizon agents
How do agents stay coherent, recover from failure, and keep making progress across long, open-ended tasks?
↗Student researcher · Builder · Community organizer
I'm a 19-year-old computer science student at UESTC, working on long-horizon agents, agent failure, and embodied intelligence — from evaluation pipelines to real robots.
What keeps me curious
I care about agents that operate beyond a single prompt — systems that remember, adapt, recover, and eventually act in the real world.
How do agents stay coherent, recover from failure, and keep making progress across long, open-ended tasks?
↗Learning VLA, WAM, world models, and robotic manipulation through hands-on deployment on real robots.
↗Building harnesses and benchmarks that reveal what AI agents can actually do — and where they break.
↗Selected work
Agent failure · 2026
A first-author study of how open-ended agents become trapped by their own narratives and continue acting on stale assumptions.
Automated research · EPIC Lab
Designed an automated research workflow that searched, tested, and improved diffusion-cache ideas under a strict compute budget.
Embodied AI · YesAI Lab
Operating real robotic systems, deploying models, and learning the full manipulation stack from perception to action.
Built in public
Runs agents on disposable code and LaTeX workspaces, records their trajectories, and judges whether the work justified its time, scope, and interaction cost.
An all-in-one learning and practice environment for people getting started with AI coding.
A hands-on workflow for exploring agent-driven cache research and experimentation.
An agent harness and LLM evaluation pipeline for repeatable, evidence-driven experiments.
Contributing tasks and evaluation work to a benchmark for long-horizon AI agents.
ESSAY 04投稿越来越多,然后呢?投稿、评审与职业回报怎样形成一个越转越快的循环;当六万篇都认真,研究、判断与生活又该怎么继续。阅读文章 ↗
NOTES 03具身智能随想My Thoughts on Embodied Intelligence一些关于 VLA/WAM、coding agent 构建环境、具身 Harness、物理与语义 grounding 的阶段性想法。阅读随想 ↗
BENCH 02做个专测防御性写作和编程的 BenchmarkGrowing Bench: Worth It正确还不够。这些工作真的值得做吗?一个测试 Agent 分寸感、并随前沿模型和真实用户经验不断生长的 Living Benchmark。阅读文章 ↗
CHECKPOINT 01大一年终总结:顿觉天地宽暑假一直在忙,一直想写这样的一个总结。见到很大的世界,思绪很多,所以写给自己,作为我 world model 的一个 checkpoint。阅读文章 ↗Learning by doing
Research Intern
Embodied intelligence with Shi Ye and Jingya Wang. VLA, WAM, robotic manipulation, and real-robot deployment.
Research Intern
Agent research with Prof. Linfeng Zhang: automated research for diffusion cache and agent failure in open-ended tasks.
B.Sc. Computer Science · Honors College
Undergraduate student and president of the UESTC AI Society, building a community around learning and making with AI.
Beyond the lab
When I'm not training agents or deploying robots, I'm probably watching tokusatsu, following a new anime, listening to Stray Kids, or starting yet another sandbox world.
Learning how intelligence leaves the screen and enters the physical world.
Let's talk
I'm always happy to talk about agents, robots, research, and things worth building.