Autonomous RE Agent · Post-Run Report
FLARE-On 13, cleared 9/9 by ilio
01 Run
How the run went
Each challenge was solved inside a single Codex turn — no resume/retry loops fired, and the cyber-content filter never blocked a turn. The agent opened every challenge the same way: read the $flareon playbook, triage the handout, then split into parallel static / dynamic / forensic "lanes" and keep flag submission in the root session. That discipline shows up as a tight, uneventful log: 494 commands, 9 turns, 9 accepted flags.
The two hardest challenges — FlareCalc and crux — took ~23–24 min and 127–148 commands each, together accounting for ~59% of all tokens. The fastest, NeonOutRun, fell in 21 seconds and 2 commands.
02 Tools
Tool usage
Count of shell commands that invoked each tool (a single command can chain several). python, sed/awk and rg are the scripting glue; the reverse-engineering workhorses are kuna, objdump, wasm-tools, ilspycmd and tshark.
03 Kuna
Kuna in focus
51 calls, 6 of 9 challenges
Kuna was the agent's default move for anything native: it reached for the decompiler on every challenge that shipped a compiled binary, and leaned on it hardest for the two long native problems.
- FlareCalc native20
- FlareOn13.doc native payload10
- Threat Invaders native aid10
- GhostStream native4
- ToxicMiner native4
- catthief Rust3
Where it sat out — and why
The three challenges with no kuna calls weren't native RE problems:
- FLAreCAPTCHAHTML/JS — browser only
- cruxbrowser extension — no native binary
- NeonOutRunsolved in 21 s / 2 cmds
On the managed-.NET target (Threat Invaders) the agent paired ilspycmd for IL with kuna for the native helper artifacts — kuna was used for guidance, not as the primary decompiler.
No KUNA_NEED.md gap files were written during the run — the agent did not hit a kuna capability it needed and lacked.
04 Cost
Tokens & cost
Token composition
The whole event moved 64.9 M billable tokens (input + output). Prompt caching absorbed almost all of it — the same growing context is re-sent on every internal step, so 98% of input was a cache hit.
- Cached input63.30 M
- Uncached input1.39 M
- Output (incl. 0.10M reasoning)0.20 M
Illustrative cost
gpt-5.6-sol priority-tier pricing isn't in the repo, so this is an estimate at representative frontier rates — swap in the real rate and recompute from the token lines at left.
- Uncached in · 1.39M × $2.50/M$3.48
- Cached in · 63.30M × $0.25/M$15.82
- Output · 0.20M × $10.00/M$2.04
Note: the harness's own results.tsv "tokens" column sums every usage sub-field, so it double-counts cached input and reasoning (128.3 M across the run). The 64.9 M figure here is the real input + output.
Billable tokens per challenge
05 Trace
Per-challenge trace
| # | Challenge | Type | Wall | Cmds | Kuna | Signature tools |
|---|---|---|---|---|---|---|
| 1 | FLAreCAPTCHA | HTML/JS | 42 s | 5 | — | node, rg |
| 2 | GhostStream | disk image → native | 4.9 m | 32 | 4 | objdump, python, kuna |
| 3 | FlareOn13.doc | Office document | 9.7 m | 62 | 10 | xxd, 7z, kuna, python |
| 4 | ToxicMiner | native | 7.9 m | 45 | 4 | python, kuna, objdump |
| 5 | catthief | Rust + PCAP | 4.1 m | 35 | 3 | tshark, python, kuna |
| 6 | Threat Invaders | .NET + PCAP | 6.3 m | 38 | 10 | ilspycmd, tshark, jq, kuna |
| 7 | FlareCalc | native (C++) | 23.8 m | 148 | 20 | python, kuna, gcc, objdump |
| 8 | crux | browser extension | 22.8 m | 127 | — | node, wasm-tools, python |
| 9 | NeonOutRun | Rust native | 21 s | 2 | — | rg, sed |
Challenge 10 "victory" is not a challenge: its CTFd body is a "Congratulations on completing FLARE-On 13!" note with a prize-shipping form (0 solves, no flag). The agent submitted six evidence-based candidates, all rejected, then wrote HELP.md concluding it is an announcement sentinel and should be excluded from ilio auto.