ilio / flare-on 13post-run report

Autonomous RE Agent · Post-Run Report

FLARE-On 13, cleared 9/9 by ilio

ilio auto — run summary
$ model gpt-5.6-sol · reasoning=high · service_tier=priority
$ date 2026-09-26 · target flare-on13.ctfd.io
$ result 9 / 9 solved · 0 resumes · 0 cyber-filter blocks · 0 wrong-flag retries
$ wall ~80.5 min agent time across the 9 solves
9 / 9
challenges solved & accepted
first-attempt each
80.5 min
total agent wall time
42 s → 24 min per chal
494
shell commands run
across 9 codex turns
51
kuna invocations
6 of 9 challenges
64.9 M
billable tokens (in+out)
98% served from cache
~$21 est.
illustrative total cost
~$2.4 / challenge*

01 Run

How the run went

Each challenge was solved inside a single Codex turn — no resume/retry loops fired, and the cyber-content filter never blocked a turn. The agent opened every challenge the same way: read the $flareon playbook, triage the handout, then split into parallel static / dynamic / forensic "lanes" and keep flag submission in the root session. That discipline shows up as a tight, uneventful log: 494 commands, 9 turns, 9 accepted flags.

The two hardest challenges — FlareCalc and crux — took ~23–24 min and 127–148 commands each, together accounting for ~59% of all tokens. The fastest, NeonOutRun, fell in 21 seconds and 2 commands.

02 Tools

Tool usage

Count of shell commands that invoked each tool (a single command can chain several). python, sed/awk and rg are the scripting glue; the reverse-engineering workhorses are kuna, objdump, wasm-tools, ilspycmd and tshark.

03 Kuna

Kuna in focus

51 calls, 6 of 9 challenges

Kuna was the agent's default move for anything native: it reached for the decompiler on every challenge that shipped a compiled binary, and leaned on it hardest for the two long native problems.

  • FlareCalc native20
  • FlareOn13.doc native payload10
  • Threat Invaders native aid10
  • GhostStream native4
  • ToxicMiner native4
  • catthief Rust3

Where it sat out — and why

The three challenges with no kuna calls weren't native RE problems:

  • FLAreCAPTCHAHTML/JS — browser only
  • cruxbrowser extension — no native binary
  • NeonOutRunsolved in 21 s / 2 cmds

On the managed-.NET target (Threat Invaders) the agent paired ilspycmd for IL with kuna for the native helper artifacts — kuna was used for guidance, not as the primary decompiler.

No KUNA_NEED.md gap files were written during the run — the agent did not hit a kuna capability it needed and lacked.

04 Cost

Tokens & cost

Token composition

The whole event moved 64.9 M billable tokens (input + output). Prompt caching absorbed almost all of it — the same growing context is re-sent on every internal step, so 98% of input was a cache hit.

  • Cached input63.30 M
  • Uncached input1.39 M
  • Output (incl. 0.10M reasoning)0.20 M

Illustrative cost

gpt-5.6-sol priority-tier pricing isn't in the repo, so this is an estimate at representative frontier rates — swap in the real rate and recompute from the token lines at left.

  • Uncached in · 1.39M × $2.50/M$3.48
  • Cached in · 63.30M × $0.25/M$15.82
  • Output · 0.20M × $10.00/M$2.04
Estimated total for all 9 solves
≈ $21.3
≈ $2.4 per challenge · cache reads dominate the bill

Note: the harness's own results.tsv "tokens" column sums every usage sub-field, so it double-counts cached input and reasoning (128.3 M across the run). The 64.9 M figure here is the real input + output.

Billable tokens per challenge

05 Trace

Per-challenge trace

#ChallengeTypeWallCmdsKunaSignature tools
1FLAreCAPTCHAHTML/JS42 s5—node, rg
2GhostStreamdisk image → native4.9 m324objdump, python, kuna
3FlareOn13.docOffice document9.7 m6210xxd, 7z, kuna, python
4ToxicMinernative7.9 m454python, kuna, objdump
5catthiefRust + PCAP4.1 m353tshark, python, kuna
6Threat Invaders.NET + PCAP6.3 m3810ilspycmd, tshark, jq, kuna
7FlareCalcnative (C++)23.8 m14820python, kuna, gcc, objdump
8cruxbrowser extension22.8 m127—node, wasm-tools, python
9NeonOutRunRust native21 s2—rg, sed

Challenge 10 "victory" is not a challenge: its CTFd body is a "Congratulations on completing FLARE-On 13!" note with a prize-shipping form (0 solves, no flag). The agent submitted six evidence-based candidates, all rejected, then wrote HELP.md concluding it is an announcement sentinel and should be excluded from ilio auto.