Interesting, I had been wondering if you could cycle distillation and then splitting weights to turn a saturated model into a non-saturated with the same parameter count for further training.<p>Train until you stop getting sidnificant improvements. distill to a quarter size model. Expand back up, and continue training.<p>Split the weights so W1+W2 = W with W1 = (W + Randoffset)/2, W2 = (W - RandOffset)/2. Double the width of the layers using the split weights, you get the same result from a 4x size network. My hypothesis is that this has way more scope to train that the model you distilled from.