AI Orchestration Drills · Human-in-the-Loop Verification
AI writes the code.
You command the direction.
Prove the power of your judgment. Solve a real infrastructure incident in 45 minutes with your AI, and earn the only human-verified score of your orchestration authority.
100% free. Real infrastructure. Rated by human reviewers.
Pure engineering drills. Scored by human reviewers against a transparent rubric.
How it works
Four steps to prove your AI orchestration mastery.
We don’t measure lines of code. We measure your decision-making under pressure.
- 01
Choose a Drill
Real production incidents · 45 min max
Select an infrastructure incident. Each weekly drill features maintainer test suites, seeded data, and real infrastructure.
- 02
Fork the arena
Public repo · Tests included
Fork the arena repository. It ships with the source, the maintainer test suites, seeded data and the load harness that decides whether you actually shipped.
- 03
Record Session (Full Screen)
Max 45 min · Single attempt rule
Record your full screen, solve it locally with your AI, and make a SINGLE git push when you consider it done. No trial-and-error pushes.
- 04
Human-in-the-Loop Index
Human reviewed · Orchestration Score
A person watches your recording and scores it against the public rubric. That score is your Orchestration Index.
This is not a hackathon or a LeetCode prompt
- Hackathon / LeetCode
- 24-48 hours
- AGISports Drill
- Max 45 minutes
- Hackathon / LeetCode
- Write raw code line-by-line
- AGISports Drill
- Orchestrate AI agents & IDEs
- Hackathon / LeetCode
- Synthetic unit tests or subjective jury
- AGISports Drill
- Human-in-the-loop review + real infrastructure
- Hackathon / LeetCode
- Syntax / language matters
- AGISports Drill
- Code & choice of LLM do not matter
- Hackathon / LeetCode
- You win a badge
- AGISports Drill
- You prove your engineering authority
A hackathon asks if you can type syntax. An AGISports drill proves how well you command AI.
Why a score exists
We do not evaluate anyone by what they built.
Two engineers ship the same solution in 45 minutes. One decomposed the problem first and rejected three hallucinated outputs before running anything. The other kept prompting blindly until it compiled.
The result is identical. The difference is everything — and nothing on the internet measures it.
So we do. Every session leaves a record: how the problem was broken down, what was delegated, what was rejected, and how fast a wrong turn was corrected. That record is your Orchestration Index.
- Spec
- How the problem is decomposed before a line is written.
- Delegation
- What goes to the model, and what deliberately does not.
- Review
- What gets rejected from the model output, and how early.
- Recovery
- Time from a wrong turn to a green test suite run.
The rubric is public, and every recording is scored against it by a person.
The library
Real production incidents. Original test suites. Your turn to command.
Drills →Frequently Asked Questions
How does a weekly Drill work?
You fork the arena repository, start a full-screen recording, and have 45 minutes to solve the incident with your AI. When you consider it done you make one single git push — that push is what gets reviewed.
What do I gain by completing a weekly Drill?
A person scores your session against the public rubric and you get an Orchestration Index: a record of how you decided, not of what you typed. Nothing else on the internet measures that.
Can I retry a Drill?
No. Each weekly drill is evaluated strictly on your first recorded session to guarantee authentic decision-making under pressure.
What does the Orchestration Index measure?
It evaluates 4 pillars: problem decomposition, AI prompt delegation, flawed code rejection, and recovery speed under pressure.
Why is evaluation Human-in-the-Loop?
Automated unit tests alone cannot measure reasoning or catch shortcut hacks. A human reviewer audits your full-screen session against a transparent rubric.
Does programming language or choice of AI matter?
Not at all. Use Claude Code, Cursor, Antigravity, Codex, or any tool. Syntax and language do not affect your score—only your orchestration judgment counts.
Who owns my code and is it free?
Your code is 100% yours. We claim no IP rights, and all drills are completely free forever.