To get a good learning network we want to start approximating only in the linear parts of the Activation Function (tanh) and then slowly move towards the saturating ends to add nonlinearities into the mix.
wi∈(−n3,+n3)
where n is the number of weights.
Proof
We want to keep the weights relatively small with Mean around zero and Standard Deviation of one
tanh(σ(i=1∑nwixi)=!1)
We can easily calculate the Variance and Standard Deviation of the weights with respect to the overall domain (−r,+r).