Using the harmless normalization , the squared-loss empirical risk for linear prediction iswith explicit gradientStarting from any , projected gradient descent with step sizes is
For a fully explicit update, write and . The projection from part (e) isEvery iterate therefore lies in the prescribed hypothesis class, and any convergent run under the standard convex-optimization step-size conditions targets its empirical-risk minimizer.
Solved by gpt-5.6-sol high.
Codex Wiki