The worst likely day
A go/no-go decision is a bet on the tail, not on the average. Most risk scores model the average.
Most preflight risk tools hand you a number. Add up the factors — weather, currency, terrain, time of day — cross a threshold, and the answer is no. The number is an average of how the flight is expected to go.
But that is not the decision being made. Nobody cancels a flight because the expected outcome is mildly bad. They cancel because of what happens on the bad end of the distribution: the ceiling drops faster than forecast, the alternate goes below minimums, and the margin that looked fine on the ground is gone. A go/no-go call is a bet on the worst likely day, not an average one.
Borrowing from finance
Finance has language for this. Value at risk asks how bad the bad cases get; conditional value at risk asks what the average of those bad cases actually is — the shape of the tail, not just its edge. That question transfers cleanly to a flight. Simulate the day many times over the uncertainty in the inputs, then look only at the worst slice of outcomes and ask whether you would accept them.
Scored that way, two flights with the same headline number stop looking alike. One is uniformly mediocre. The other is usually fine and occasionally unsurvivable. The average hides exactly the difference you care about.
Show the decomposition or don't bother
The other half is refusing to return a bare score. Every number decomposes into per-factor contributions and an uncertainty band, because a pilot should be able to disagree with the model on a specific input rather than accept or reject a verdict whole.
None of this makes the decision for anyone. It is a research prototype on synthetic inputs — not flight advice, and no substitute for official weather, an instructor, or your own judgment. The point is narrower: if software is going to inform a decision somebody has to live with, it should argue its case rather than announce a conclusion.