
Maybe All Neural Networks Learn the Same Thing
A neat theoretical result suggests that different networks, trained on different data, can converge to the same internal representation, up to a rotation. The interesting part is that SGD itself may be responsible. I rea
Read essay ->





