A single number from 0 to 100 telling you how hard a phrase is to rank for is the standard feature in this category, and it is usually not measured against anything.
Ours was. We set a threshold for how well it had to predict real outcomes before it could ship, tested it twice, and it came in below that threshold both times. So it is computed, stored, and rendered nowhere.
What you get instead is the working: how heavy the top ten actually are, how many of them also chart, and where that sits against every other phrase we track. That is a comparison you can check, rather than a score you have to trust.