Data
Sources, and what is wrong with each of them. More results are invalidated here than by any failure of statistics, and almost none of it is a statistics problem.
A series that looks clean is usually a series whose faults have not been looked for. Six recur, and all six are silent.
The six silent faults
The break calendar
Any test spanning one of these dates is testing two different series and calling them one. The engine raises a flag rather than allowing it silently.
The same calendar, as a shape
Seventeen breaks in seventy-eight years, and they do not thin out with time. Any test spanning one of these marks is testing two different series and calling them one.
This is the argument for the break calendar in one picture. Not that breaks happen, which everyone concedes, but that they happen constantly, so a long sample is more exposed than a short one rather than less.
Time
Everything is stored in coordinated universal time to microsecond precision and converted only at the point of display. Where two venues disagree on a closing price, one is designated authoritative and the choice is recorded rather than assumed.
Daylight saving is handled by timezone-aware conversion rather than a fixed offset. The transitions in the United States and the United Kingdom do not coincide, which produces two or three weeks in March and about a week in October where a fixed offset is simply wrong.
Under active resolution
One contamination issue in a derived label set is being worked through, and dependent validation is held until it clears rather than proceeding on an unverified base. Holding it is the control working as designed. It is listed under Data integrity in the theories index as open.
Raw deliveries are immutable. Every derived layer can be rebuilt from them. If that is not true, no result is reproducible and the rest of this is decoration.