Megadose AI progress, ranked and analyzed.

Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks

· ArXiv · AI/CL/LG ·
The paper tightens generalization bounds for one-hidden-layer ReLU networks where each input activates only a small subset of units.

Li and coauthors bound the Rademacher complexity for width `s`, per-input activation cap `k`, sample size `m`, and weight/bias limits `W,B`, removing a previous explicit dimension factor up to logarithms. Their lower bounds nearly match, showing that changing which units are active across inputs still leaves a width-dependent cost. They also separate zero-bias networks sparse over the whole ball, which collapse to far fewer effective units, from comparable-bias networks that recover the worst-case rate. For a bounded normalized loss, the paper gives agnostic minimax excess-risk bounds of order `min{1, sqrt(s/(km))}` up to logarithms. ArXiv · AI/CL/LG's note

score 4

Categories: Research