The Sample Complexity Of ERMs In Stochastic Convex Optimization

November 09, 2023 ยท Declared Dead ยท ๐Ÿ› International Conference on Artificial Intelligence and Statistics

๐Ÿ‘ป CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Daniel Carmon, Roi Livni, Amir Yehudayoff arXiv ID 2311.05398 Category cs.LG: Machine Learning Cross-listed stat.ML Citations 6 Venue International Conference on Artificial Intelligence and Statistics Last Checked 6 months ago
Abstract
Stochastic convex optimization is one of the most well-studied models for learning in modern machine learning. Nevertheless, a central fundamental question in this setup remained unresolved: "How many data points must be observed so that any empirical risk minimizer (ERM) shows good performance on the true population?" This question was proposed by Feldman (2016), who proved that $ฮฉ(\frac{d}ฮต+\frac{1}{ฮต^2})$ data points are necessary (where $d$ is the dimension and $ฮต>0$ is the accuracy parameter). Proving an $ฯ‰(\frac{d}ฮต+\frac{1}{ฮต^2})$ lower bound was left as an open problem. In this work we show that in fact $\tilde{O}(\frac{d}ฮต+\frac{1}{ฮต^2})$ data points are also sufficient. This settles the question and yields a new separation between ERMs and uniform convergence. This sample complexity holds for the classical setup of learning bounded convex Lipschitz functions over the Euclidean unit ball. We further generalize the result and show that a similar upper bound holds for all symmetric convex bodies. The general bound is composed of two terms: (i) a term of the form $\tilde{O}(\frac{d}ฮต)$ with an inverse-linear dependence on the accuracy parameter, and (ii) a term that depends on the statistical complexity of the class of $\textit{linear}$ functions (captured by the Rademacher complexity). The proof builds a mechanism for controlling the behavior of stochastic convex optimization problems.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning

Died the same way โ€” ๐Ÿ‘ป Ghosted