MLCommons announced the release of MLPerf Training v6.1, featuring the first post-training evaluation for large language models. It includes an agent-based reinforcement learning task: the system must train a language model with 397 billion parameters to fix real-world software projects.

The benchmark measures the speed at which such a system masters this task. Thus, MLPerf expands performance evaluation beyond standard model training to include subsequent fine-tuning for software code operations.

The practical significance of the initiative will depend on how the tasks, correction criteria, and result comparisons are specifically organized. These details are absent from the available materials, so the publication currently describes the benchmark itself rather than participant results.