TankBenchTANKBENCHCODE · BATTLE · REPEATSNAPSHOT / 23 SEPT 2026

THE PROTOCOL

A benchmark with a paper trail.

TankBench measures how frozen, AI-built Robocode bots perform against each other. Here is exactly how.

01 / THE PROCESS

Build.

Each AI model receives the same task, the same starter robot, the same offline reference pack and the same practice opponents. It works in an isolated container with no internet access, within a fixed budget of tokens, tool calls, practice battles and time, and must submit a working Robocode bot. Practice scores help the model improve its bot; they are not the benchmark.

02 / THE PROCESS

Freeze.

A submitted bot is frozen: its code and identity can never change afterwards. Later attempts by the same model become separate bots, and no bot is edited after seeing competition results.

03 / THE PROCESS

Compete.

In the duel league every bot meets every other bot 2 times, with the starting order reversed, over 35 rounds on a 800 × 600 battlefield. In the melee league bots are drawn into balanced heats of up to 10, each appearing 10 times, over 35 rounds on 1000 × 1000. Every battle starts in a fresh arena, so no bot carries learned data between battles.

04 / THE PROCESS

Score.

Robocode awards points for survival, damage dealt and kills. A bot’s share of the total score in each battle is averaged across all its battles: its APS. An equal share is 50% in a duel, or one tenth in a ten-bot melee heat. Duel and melee are ranked separately.

THE LINEAGE

Built on a quarter-century-old arena.

Robocode began in late 2000 as a personal project by Mat Nelson, who brought it to IBM as an alphaWorks download in July 2001. Programmers write a tank in Java, then let it drive, scan and shoot on its own. Nobody touches a controller once the battle starts. It became open source in 2005, and volunteers led by Flemming N. Larsen have maintained it ever since.

The classic Robocode battle window: four small tanks on a black battlefield, each labelled with its name and energy, with shells in flight and a list of competitors down the right-hand side.
A four-bot melee in classic Robocode. Each tank shows its remaining energy.
Robocode’s Robot Editor showing Java source for a robot, with the File, New menu open offering Robot, JuniorRobot and Java File.
Robocode’s built-in Robot Editor, where bots are written in Java.

Its community built the RoboRumble, a ranking that runs continuously on volunteers’ computers. Clients download bots, run battles and upload the results. More than a thousand human-written bots hold places there, and the LiteRumble server still ranks them today in 1v1, melee and team divisions.

TankBench keeps the rumble’s battle rules. A duel is two bots on an 800 × 600 battlefield for 35 rounds. A melee heat is ten bots on 1000 × 1000 for 35 rounds. Each battle is scored with Robocode’s native scoring, and bots are ranked by APS, their average percentage score. Two things differ. The RoboRumble samples pairings continuously and never really finishes. A TankBench league freezes a fixed schedule and completes it. In the duel league every pairing is played exactly twice, so APS is computed just as the rumble computes it. In melee, TankBench scores each bot’s share of the whole heat, so an even share is 10% rather than 50%.

The combined rating.

Rating = 100 × √((duel APS ÷ 50%) × (melee APS ÷ 10.0%)). 100 means exactly equal-share performance in both formats. Only bots entered in both leagues are ranked. The geometric mean rewards bots that are strong in both formats rather than dominant in one.

Build cost is context, not score.

Each contender page lists the builder model, reasoning setting, build cost, tokens and time. Costs are reported by the inference provider where available; where the provider did not report an exact cost, the reserved maximum is shown as an upper bound (≤). Locally run models have no inference cost. None of this affects rank.

Compare like with like.

Scores are noisy, so every table shows how much a bot’s share varied between battles. Most models have one bot so far: a single attempt, not a verdict on the model. Build rules can change between protocols, and each bot page names the protocol it was built under. Battles run on Robocode 1.11.1 with Eclipse Temurin 21. Tank artwork, palettes and arena scenery are cosmetic.

ABOUT

Who runs TankBench.

TankBench is an independent benchmark published by Lantic Media Inc. It is not affiliated with, sponsored by or reviewed by any AI lab, and no lab sees results before they are published. Every model is built and scored under the same published rules.

Found a mistake, want a model tested, or writing about TankBench? Send us a message.

Model and company names are trademarks of their respective owners and are used only to identify the models tested. Tank artwork is illustrative and not an official identity of any AI lab.

This site sets no cookies of its own. It uses Cloudflare Web Analytics, which counts page views, referring sites, countries and device types without cookies and without identifying or following individual visitors. The replay player remembers your camera and sound choices in your own browser; those choices are never sent to us.

The contact page uses Cloudflare Turnstile to block spam. Turnstile may set cookies needed only for that check; Cloudflare uses it to detect bots, not to identify, profile or track people. What you send is emailed to us, used only to reply, and not stored on this site.