(made with Claude Code and ChatGPT)
MIT 6.7960 — Approximation Theory

Approximation

A two-layer ReLU MLP with H hidden units trained by SGD on random samples from f. Watch the kinks move as both layers learn.

8
0.01
32
iter 0 MSE — kinks —
target f(x)
MLP fit
vi·ReLU(wix+bi)
kink position