Dark Knowledge and the Enclosure Problem
A compression trick from 2015 became the subject of a joint NSA, CISA and FBI advisory. The path between those two facts is worth walking slowly. A thing or two about distillation! Copyright: Sanjay Basu The mythical digit Start with the experiment that named the thing. In 2015, Hinton, Vinyals and Dean trained a large net with two hidden layers of 1,200 rectified linear units on MNIST, heavily regularized with dropout and weight constraints, and got 67 test errors. A smaller net with 800 units per layer and no regularization got 146 errors. When that same small net was trained to match the soft target distribution of the large net at a temperature of 20, it got 74 errors. Half the error rate, same architecture, same parameter count, different supervision signal. Then they did the part I keep coming back to. They removed every example of the digit 3 from the transfer set, so the student had never seen a 3 in training. It still made only 206 test errors, 133 of them on the 1,010 t...