There is a strong temptation to evaluate market analysis by whether following it was profitable. It is intuitive, it is easy to compute, and it teaches you almost nothing about whether the analysis was any good.
The four outcomes
A read can be right and profitable, right and unprofitable, wrong and profitable, or wrong and unprofitable. The middle two happen constantly, and any scoring system that cannot distinguish them will reward luck and punish correctness at roughly the rate the market provides each.
A read that said moves would be large, followed by a large move that went the other way, was correct. Scoring it as a failure because a hypothetical position lost money is scoring the position, not the analysis.
What a read actually claims
This is why the claim has to be stated precisely enough to be scored. Volatility will be elevated is scoreable. This level will be tested is scoreable. Things look constructive is not, and unscoreable claims are where most market commentary lives, for obvious reasons.
The discipline of writing claims that can be marked wrong is most of the value, independent of the marking.
The rule comes first
Decide what counts as right before the outcome is known. Tested means touched within this tolerance during this window. Elevated means above this threshold. Written down in advance, applied to every case.
Rules chosen afterwards will fit the outcome. Not through dishonesty, particularly — it is simply very difficult to choose a threshold after seeing the data without the data influencing the choice.
And keep the neutral band
Some outcomes are genuinely ambiguous, and a scoring system without a neutral category will force them into a bucket and add noise. A read that was neither right nor wrong should be recorded as exactly that.