In this video I cover SKYFLY, a browser flight simulator built by developer vakovalskii. The interesting question is how to evaluate a development process that combines many coding agents, a large token budget and a short elapsed time—not simply whether the demo looks impressive.

What the reported numbers tell us

The video description reports up to 40 agents working simultaneously, 45 million tokens, and around seven hours of development. It also mentions three exhausted $100 subscription limits. These are reported development figures, not measurements independently reproduced for this page.

Each number answers a different question:

  • Agent count describes concurrency. It does not establish how many agents contributed useful changes or how independently they worked.
  • Elapsed time describes time to a reported result. It is different from total agent-hours and human review time.
  • Token count describes aggregate model traffic. It does not establish a cash price without knowing the models, input/output split, cached tokens and subscription terms.
  • Exhausted subscription limits describe capacity constraints. They should not be converted directly into an exact project cost.

Dividing 45 million tokens by seven elapsed hours gives roughly 6.4 million tokens per hour across the workflow. That is a useful scale indicator, not a per-agent rate or evidence of efficiency.

What is actually in the browser game

The project README describes a JavaScript client using Three.js and Vite, with a Node.js and WebSocket server. It combines geographic terrain, OpenStreetMap buildings and multiplayer interaction.

That matters because this is an integration problem as well as a code-generation problem. Rendering, geographic data, player state and networking have to work together. Producing separate pieces quickly is only one part of producing a usable system.

The repository makes its source available for inspection under a restricted noncommercial licence. Publicly visible source is not blanket permission to reuse, modify or deploy it; consult the repository licence before doing so.

What I would check beyond the demo

For this kind of multi-agent build, I would review four things before treating fast delivery as a repeatable method:

  1. Integration: did the agents have clear component boundaries, and who resolved conflicting changes?
  2. Verification: do rendering, input, reconnects and shared player state work outside the original demo sequence?
  3. Browser constraints: does the experience remain usable on slower hardware, mobile screens and variable networks?
  4. Budget: how much useful progress remained after accounting for retries, repeated context and human review?

These are evaluation questions, not claims that SKYFLY passes or fails those checks. They explain why agent count and speed alone are incomplete measures of a development workflow.

How this connects to my own workflow

In How I Work With AI Agents in 2026, I describe an orchestrator as my main interface and human taste gates as checkpoints for direction and evaluation. This case gives a concrete reason to keep those checkpoints: parallel execution increases the amount of work that can happen before an integration or product decision is reviewed.

For more on context and checkpoints, see AI Workflow and Orchestration.

Watch, inspect and explore

Source note: the development figures come from my video description; the architecture and licensing information come from the project README. This page adds workflow analysis and evaluation questions rather than reproducing the description alone.