Weighted score out of 100, averaged over five independent runs per model.
Cost of the five generation runs against mean score. Up and left is better.
Judge scores out of 10, averaged over runs. Cell shading follows the score; the best score in each dimension is set in bold.