Generating the CLI templates printed a line for every XML file and started Python once per file in series, so most of the stage was interpreter startup. Per-file output is now replaced by one summary line per stage, and XML preprocessing plus template generation run in parallel across all CPUs. On a 4-core machine both stages are about 2.5-3x faster; generated templates and caches are byte-identical to before.