Double descent

Make the model bigger. Test error falls, spikes when the model can fit every training point, then falls again.

32
32
4
0.80
0
typical test error this sample noise
Drag model size. The spike is at p = n, where training error hits zero.

What’s going on?

The true pattern uses only k inputs. The model sees the first p of them and fits a straight line to noisy labels. Test error is how wrong that line is on a new input. The black curve is the middle of many random samples, so one unlucky sample does not set the height of the spike.

Left of the spike, error falls until the model can see the pattern, then rises as extra inputs fit noise. At p = n the line can pass through every training label, noise included. Past that, many lines can, and the plot keeps the shortest. More parameters let that line stay smaller, so test error falls again. Ridge penalizes large lines and shaves off the spike. Wider neural nets show the same shape.

On the right, each dot is a new input. Across is the true answer, up is the model’s guess. Dots on the dashed line mean the pattern was recovered. Fitting the noisy training labels is a different job, which is why a perfect training fit can still miss that line.