I was going to bring this up, too—I did a puzzle hunt once where the team running the hunt tried to quantify each puzzle in terms of “lightbulbs” (how much of a leap of logic does it require?) vs. “hammers” (how much work is it?). My puzzle hunt team actually uses this concept a lot now when we design puzzles, but just the “is this puzzle predominantly lightbulb or predominantly hammer?” part. The actual “quantifying lateral thinking and raw effort on a scale of 1–5 each” part we don’t use, and I think that’s at least in part due to how tough it is to quantify this kind of thing in the first place.