Open problems
Questions the field has not settled, published openly so that others can work on them alongside us. Correspondence on any of them is welcome.
A list of things that are wrong, or unresolved, or where the honest answer is that nobody appears to know.
1. How do you deflate a score when the trial count is unrecoverable?
The correction for how many candidates were searched requires knowing how many were searched. Results produced before anyone counted cannot be corrected, only discarded. Almost every published finding in applied forecasting is in this position. Is there any principled recovery, or is discarding genuinely the only option?
2. Can a regime be identified from inside it?
Regime-switching models fit history beautifully and identify the switch retrospectively. Every practical use requires knowing the state now. Is the retrospective advantage irreducible, or an artefact of how these models are usually estimated?
3. What is the effective sample size of an overlapping series?
Twenty-five thousand observations with a two-hundred-period window are not twenty-five thousand independent observations. Rules of thumb exist. A defensible general formula does not appear to.
4. How should forecasts from correlated forecasters be combined?
Simple averaging beats most individuals, and stops working when the individuals read each other. Extremising helps empirically without a satisfying theory behind it. What is the right aggregation when the dependence structure is unknown?
5. Where does the physics analogy stop?
Random matrix theory transfers cleanly because the mathematics does not care what the matrix describes. Critical phenomena transfer less cleanly. Agent-as-spin models may not transfer at all. Is there a criterion for which borrowings are legitimate, other than seeing whether they work?
6. What is the right prior for a novel forecasting question?
Maximum entropy answers this when the constraints are known. Most interesting questions arrive with no reference class and no constraints anyone can state. Reference-class forecasting begs the question of which class.
7. Can calibration and sharpness be traded off explicitly?
The stated principle is to maximise sharpness subject to calibration. In practice both are estimated with error on small samples and the constraint is never exactly satisfied. What is the correct decision rule when neither can be measured precisely?
8. Does anything predict its own failure?
Every model has conditions under which it stops working. Almost none carry a signal that those conditions have arrived. Is a general early-warning statistic possible, or is this necessarily domain-specific?
Where the eight land
The open questions are not scattered evenly. Four of the eight touch Validation, and three touch Forecast evaluation, which says something about where the next useful work sits.
A problem that touches two fields is usually a problem about the boundary between them, which is generally where the interesting difficulty lives.
If any of these has a settled answer, the useful correspondence is a reference rather than an argument.