ScalePRM achieves 67.5 F1 on ProcessBench using synthetic step-level labels
September 1, 2026
ScalePRM trains process reward models by aggregating multiple independent verifications of reasoning steps to generate synthetic labels without ground truth. Using self-consistency and meta-critique scaling, the method achieves 67.5 F1 on ProcessBench, outperforming reference-guided training that utilizes ground-truth answers.
HOW THIS AFFECTS YOU
●
builderThis provides a path to training more reliable reasoning models using self-generated verification compute.
●
researcherYou can scale process-level supervision without the bottleneck of expensive human step-level annotations.