Recognizing handwritten digits: six neural networks on MNIST, from 92% to 99.4%
In January and February 2017 I trained six neural networks on one task, reading handwritten digits. Each model changed something about the one before it: the number of layers, the optimizer, the learning rate, the activation function and, in the end, the kind of network. The first model got 758 of the 10,000 test images wrong. The last got 62 wrong.
| Model | Test accuracy | Wrong out of 10,000 | Time to train |
|---|---|---|---|
| One layer, softmax, gradient descent | 92.4% | 758 | 38 s |
| Five layers, sigmoid, Adam | 97.4% | 258 | 35 s |
| Five layers, sigmoid, decaying learning rate | 97.8% | 220 | 37 s |
| Five layers, ReLU, decaying learning rate | 98.1% | 193 | 40 s |
| Convolutional network | 98.9% | 110 | 3 min 13 s |
| Wider convolutional network with batch normalization and dropout | 99.4% | 62 | 15 min |
The first two rows show the best of six learning rates, and the time of that one run. Every model is written in TensorFlow as it was in early 2017. The notebooks, with their charts, are on GitHub.
The dataset and a single neuron
MNIST is a dataset of handwritten digits, each labelled with the number it shows. It is a good entry point to machine learning and deep learning. Every image is 28×28 pixels and grayscale, so an image is a 28×28 matrix in which each cell holds a pixel value from 0 to 255.
Four things about a neuron have to be decided before any code is written.
- Inputs. The inputs are the pixel values of one image, so there are 784 of them (28×28), each with a weight.
- Transfer function. The usual choice is the sum of the inputs, each multiplied by its weight.
- Activation function. There are several, and each suits different cases: softmax, ReLU (the rectified linear unit), sigmoid, tanh (the hyperbolic tangent) and softsign.
- Number of neurons. The answer is one of ten digits, 0 to 9, so the output layer has 10 neurons.
One layer of ten neurons
The first model is a single layer of 10 neurons with softmax as the activation function, trained by gradient descent.
X = tf.placeholder(tf.float32, shape=[None, 784])
W = tf.Variable(tf.zeros([784,10]))
b = tf.Variable(tf.zeros([10]))
# model
Y = tf.nn.softmax(tf.matmul(X, W) + b)
# placeholder for correct answers
Y_ = tf.placeholder(tf.float32, [None, 10])
# loss function
cross_entropy = -tf.reduce_sum(Y_ * tf.log(Y))
# % of correct answers found in batch
is_correct = tf.equal(tf.argmax(Y, 1), tf.argmax(Y_, 1))
accuracy = tf.reduce_mean(tf.cast(is_correct, tf.float32))
optimizer = tf.train.GradientDescentOptimizer(learning_rate)
train_step = optimizer.minimize(cross_entropy)
I ran it with six learning rates to see what each did to the accuracy of the predictions: 0.03, 0.01, 0.007, 0.003, 0.001 and 0.0005. Each run trained on batches of 100 images for 10,000 iterations.
At 0.03 the model does not learn. Of the others, 0.003 gives the best result over the last iterations, and a stable test loss: about 92.4% of the test images right, 758 wrong. A smaller learning rate needs more iterations to get as far, and that costs time.
92% sounds acceptable until the digits come in groups. Suppose the model had to read the amount on a scanned cheque. For a six-digit amount, all six digits are right only about 60% of the time.
That is as far as this model goes. A better result needs a more complex one.
Five layers
The second model changes three things:
- Sigmoid as the activation function.
- Adam as the optimization method.
- Five layers: four hidden layers and one output layer.
It is a feedforward network. To get an idea, I gave the hidden layers 200, 100, 60 and 30 neurons. The output layer keeps its 10.
C0 = 1 # input channel count
C1 = 200 # Neuron count for layer 1
C2 = 100 # Neuron count for layer 2
C3 = 60 # Neuron count for layer 3
C4 = 30 # Neuron count for layer 4
C5 = 10 # Neuron count for layer 5 (digit count 0 to 9)
# weights
W1 = tf.Variable(tf.truncated_normal([784, C1], stddev = 0.1))
W2 = tf.Variable(tf.truncated_normal([C1, C2], stddev = 0.1))
W3 = tf.Variable(tf.truncated_normal([C2, C3], stddev = 0.1))
W4 = tf.Variable(tf.truncated_normal([C3, C4], stddev = 0.1))
W5 = tf.Variable(tf.truncated_normal([C4, C5], stddev = 0.1))
# biases
B1 = tf.Variable(tf.zeros([C1]))
B2 = tf.Variable(tf.zeros([C2]))
B3 = tf.Variable(tf.zeros([C3]))
B4 = tf.Variable(tf.zeros([C4]))
B5 = tf.Variable(tf.zeros([C5]))
# model
X = tf.placeholder(tf.float32, shape=[None, 784])
Y1 = tf.nn.sigmoid(tf.matmul(X , W1) + B1)
Y2 = tf.nn.sigmoid(tf.matmul(Y1, W2) + B2)
Y3 = tf.nn.sigmoid(tf.matmul(Y2, W3) + B3)
Y4 = tf.nn.sigmoid(tf.matmul(Y3, W4) + B4)
Ylogits = tf.matmul(Y4, W5) + B5
Y = tf.nn.softmax(Ylogits)
# placeholder for correct answers
Y_ = tf.placeholder(tf.float32, [None, C5])
cross_entropy = tf.nn.softmax_cross_entropy_with_logits(logits=Ylogits, labels=Y_)
cross_entropy = tf.reduce_mean(cross_entropy) * 100
# accuracy of the trained model, between 0 (worst) and 1 (best)
correct_prediction = tf.equal(tf.argmax(Y, 1), tf.argmax(Y_, 1))
accuracy = tf.reduce_mean(tf.cast(correct_prediction, tf.float32))
train_step = tf.train.AdamOptimizer(learning_rate).minimize(cross_entropy)
With the same six learning rates, 0.003 again gives the best accuracy: about 97.4%, with 258 images wrong. Four hidden layers add about five points of accuracy for a small change to the model.
A learning rate that decays
Both models show how much the learning rate matters: with a poor value, the predictions are not acceptable. Finding a good constant value is trial and error, and it takes time, because many values have to be tried.
Exponential decay is one of the best-known ways around that. The rate is no longer a constant: it starts at one value and falls toward a lowest value as the training goes on. Three numbers define it: the starting rate, the lowest rate and the speed of the decay. Setting them is not hard, because a few tests show their effect.
The five-layer model stays as it is. The learning rate becomes a placeholder:
# paceholder for learing_rate variable
lr = tf.placeholder(tf.float32)
train_step = tf.train.AdamOptimizer(lr).minimize(cross_entropy)
Its value is worked out at every iteration and passed to the training step:
# learning rate decay
max_learning_rate = 0.01
min_learning_rate = 0.0001
decay_speed = 2000.0
learning_rate = min_learning_rate + (max_learning_rate - min_learning_rate) * math.exp(-i/decay_speed)
# train
sess.run(train_step, feed_dict = {X: batch_X, Y_: batch_Y, lr: learning_rate})
With these values the rate is down to 0.0023 in under 3,000 iterations. The model reaches 97.8%, with 220 images wrong. That is a small gain in accuracy. The point is that it took one training run, not six.
ReLU in place of sigmoid
The activation function is another hyperparameter to choose with care, and the way to choose is to try several and see which gives the best result. Every hidden layer so far used sigmoid. In early 2017 many reports were finding that ReLU, the rectified linear unit, did better than the other functions. It is defined as f(x) = max(0, x).
I could not find a mathematical proof of that, but in practice it works better in some cases. So I kept the model and changed only the activation function of the hidden layers. The last layer keeps softmax.
Y1 = tf.nn.relu(tf.matmul(X , W1) + B1)
Y2 = tf.nn.relu(tf.matmul(Y1, W2) + B2)
Y3 = tf.nn.relu(tf.matmul(Y2, W3) + B3)
Y4 = tf.nn.relu(tf.matmul(Y3, W4) + B4)
The biases need a decision too. With ReLU it is common to set them all to zero. Another approach is to give them a small value, such as 0.001, so that every ReLU unit fires at the beginning.
With ReLU the model gets 193 images wrong, against 220 with sigmoid: almost 98.1% in place of 97.8%. Again a small gain.
A convolutional network
A convolutional network is known as the better solution for image recognition. For this step I followed a recorded presentation, which I recommend to anyone who is not familiar with convolutional networks, and I used the same model and hyperparameters as its source.
The model has three convolutional layers, then a fully connected layer, then the 10 outputs. ReLU is the common activation function for the convolutional layers, with softmax on the last layer.
C0 = 1 # input channel count
C1 = 4 # convolutional network channel 1 count
C2 = 8 # convolutional network channel 2 count
C3 = 12 # convolutional network channel 3 count
C4 = 200 # fulley connected layer size
C5 = 10 # output count (digit count 0 to 9)
# weights
W1 = tf.Variable(tf.truncated_normal([5, 5, C0, C1], stddev = 0.1))
W2 = tf.Variable(tf.truncated_normal([5, 5, C1, C2], stddev = 0.1))
W3 = tf.Variable(tf.truncated_normal([4, 4, C2, C3], stddev = 0.1))
W4 = tf.Variable(tf.truncated_normal([7 * 7 * C3, C4], stddev = 0.1))
W5 = tf.Variable(tf.truncated_normal([C4, C5], stddev = 0.1))
# biases
B1 = tf.Variable(tf.ones([C1]) / 10)
B2 = tf.Variable(tf.ones([C2]) / 10)
B3 = tf.Variable(tf.ones([C3]) / 10)
B4 = tf.Variable(tf.ones([C4]) / 10)
B5 = tf.Variable(tf.ones([C5]) / 10)
# model
stride = 1 # output is 28x28
X = tf.placeholder(tf.float32, shape=[None, image_width, image_height, C0])
Y1 = tf.nn.relu(tf.nn.conv2d(X, W1, strides=[1, stride, stride, 1], padding='SAME') + B1)
stride = 2 # output is 14x14
Y2 = tf.nn.relu(tf.nn.conv2d(Y1, W2, strides=[1, stride, stride, 1], padding='SAME') + B2)
stride = 2 # output is 7x7
Y3 = tf.nn.relu(tf.nn.conv2d(Y2, W3, strides=[1, stride, stride, 1], padding='SAME') + B3)
# reshape the output from the third convolution for the fully connected layer
YY = tf.reshape(Y3, shape=[-1, 7 * 7 * C3])
Y4 = tf.nn.relu(tf.matmul(YY, W4) + B4)
Ylogits = tf.matmul(Y4, W5) + B5
Y = tf.nn.softmax(Ylogits)
Trained for 15,000 iterations, the model gets 110 images wrong, an accuracy of about 98.9%. That is a considerable improvement, and it costs time: a little over three minutes of training, where each earlier model took less than one.
Is it enough? Take the cheque again. For an amount under a million, the model reads all six digits correctly about 93% of the time.
Batch normalization and dropout
The method recommended for a better result is batch normalization. The last model applies it, and tunes the convolutional layers at the same time: they grow to 24, 48 and 64 channels. Dropout is added to the fully connected layer, which keeps 75% of its nodes during training.
Each convolutional layer now runs in four steps: the convolution, batch normalization, ReLU and dropout. The notebook leaves dropout switched off for these layers.
stride = 1 # output is 28x28
Y1l = tf.nn.conv2d(X, W1, strides = [1, stride, stride, 1], padding = 'SAME')
Y1bn, update_ema1 = batchnorm(Y1l, tst, iter, B1, convolutional = True)
Y1r = tf.nn.relu(Y1bn)
Y1 = tf.nn.dropout(Y1r, pkeep_conv, compatible_convolutional_noise_shape(Y1r))
After 10,000 iterations, which took about 15 minutes, the model gets 62 of the 10,000 test images wrong: almost 99.4%. The best earlier result was 110. In this run the test loss also converges, which had not happened in any of the earlier models.
Adapted in October 2026 from a series of posts first published on my blog and on LinkedIn in January and February 2017. The code is TensorFlow as it stood then.