Files
erotic-generator/poem_generator.ipynb
T
2018-09-10 23:20:25 +08:00

19 KiB

Perth Machine Learning Group Poem Generator

Introduction

The following code uses GRU to generate poems. It reads through a corpus of poems, and learns sequences of characters, including line breaks and titles.

In short, it observes many sequence of characters, and infers the character that should come next. For instance, it guesses that after 'The cat eat' should come the letter 's'.

Further details will be given with the code.

The code

Data exploration

In [2]:
import tensorflow as tf  # version 1.9 or above
tf.enable_eager_execution()  # Execution of code as it runs in the notebook. Normally, TensorFlow looks up the whole code before execution for efficiency.

import numpy as np
import re
import random
import unidecode
import time
In [3]:
path_to_file = 'poem_corpus.txt'
In [5]:
text = unidecode.unidecode(open(path_to_file).read())
print(text[:500])
          CHRISTMAS NIGHT.


    Be peace on earth, good will to men;
      And let this now our carol be:
      If on the land, or on the sea,
    We still will sing the glad refrain;
      And in the closing light of day
      Good words of peace and cheer will say.

    The Babe that in the manger born
      Has risen high above the star,
      To judge in peace, or judge in war,
    To judge at night or judge at morn.
      The star that told us of his birth
      Has given us joy and lastin

Dataset creation

In [6]:
unique = sorted(set(text))  # unique contains all the unique characters in the corpus

char2idx = {u:i for i, u in enumerate(unique)}  # maps characters to indexes
idx2char = {i:u for i, u in enumerate(unique)}  # maps indexes to characters
In [7]:
max_length = 100  # Maximum length sentence we want per input in the network
vocab_size = len(unique)
embedding_dim = 256  # number of 'meaningful' features to learn. Ex: ['queen', 'king', 'man', 'woman'] has a least 2 embedding dimension: royalty and gender.
units = 1024  # In keras: number of output of a sequence. In short it rem
BATCH_SIZE = 64
BUFFER_SIZE = 10000
In [8]:
input_text = []
target_text = []

for f in range(0, len(text) - max_length, max_length):
    inps = text[f : f + max_length]
    targ = text[f + 1 : f + 1 + max_length]
    input_text.append([char2idx[i] for i in inps])
    target_text.append([char2idx[t] for t in targ])
In [13]:
dataset = tf.data.Dataset.from_tensor_slices((input_text, target_text)).shuffle(BUFFER_SIZE)
dataset = dataset.apply(tf.contrib.data.batch_and_drop_remainder(BATCH_SIZE))
WARNING:tensorflow:From <ipython-input-13-3bbc2611a7c8>:2: batch_and_drop_remainder (from tensorflow.contrib.data.python.ops.batching) is deprecated and will be removed in a future version.
Instructions for updating:
Use `tf.data.Dataset.batch(..., drop_remainder=True)`.

Explaination

In fact, the algorithm does not learn which characters comes next. It analyzes sequences of characters as inputs (ex: 'abcd'), and predicts sequences as outputs (ex: 'bcde').

Why?

During the training phase, it learns more that just the next character. It updates weights for each characters from the input sequence to the output sequence.

Consider the sequences 'abcd', 'bcde', 'cdef', 'defg', the letter "d" is given different weights that depend on the previous sequences

The use of these updates helps predicting better the next sequences and so on. So it learns the next character but also all the weights

The next chunk of code is optional.

In [20]:
# example of input:
print('Given the following sequence: \n\n')
print(''.join(idx2char[input_text[14][i]] for i in range(len(target_text[0]))))
print('\n\n')
print('the network has to learn that a correct continuation is: \n')
# example of output the algorithm has to learn
print(''.join(idx2char[target_text[14][i]] for i in range(len(input_text[0]))))
Given the following sequence: 


ew over the land,
      And the country was wild with glee;
    And she stilled the wave in the stor



the network learns that a correct continuation is: 

w over the land,
      And the country was wild with glee;
    And she stilled the wave in the storm

Model

In [14]:
class Model(tf.keras.Model):
  def __init__(self, vocab_size, embedding_dim, units, batch_size):
    super(Model, self).__init__()
    self.units = units
    self.batch_sz = batch_size
    self.embedding = tf.keras.layers.Embedding(vocab_size, embedding_dim)
    if tf.test.is_gpu_available():
      self.gru = tf.keras.layers.CuDNNGRU(self.units, 
                                          return_sequences=True, 
                                          return_state=True, 
                                          recurrent_initializer='glorot_uniform')
    else:
      self.gru = tf.keras.layers.GRU(self.units, 
                                     return_sequences=True, 
                                     return_state=True, 
                                     recurrent_activation='sigmoid', 
                                     recurrent_initializer='glorot_uniform')
    self.fc = tf.keras.layers.Dense(vocab_size)
        
  def call(self, x, hidden):
    x = self.embedding(x)
    output, states = self.gru(x, initial_state=hidden)
    output = tf.reshape(output, (-1, output.shape[2]))
    x = self.fc(output)
    return x, states
In [15]:
model = Model(vocab_size, embedding_dim, units, BATCH_SIZE)
In [16]:
optimizer = tf.train.AdamOptimizer()
In [17]:
def loss_function(real, preds):
    return tf.losses.sparse_softmax_cross_entropy(labels=real, logits=preds)

Training

In [ ]:
n_epochs = 30

for epoch in range(n_epochs):
    start = time.time()
    hidden = model.reset_states()  # initializes the hidden state at the start of every epoch
    
    for (batch, (inp, target)) in enumerate(dataset):
          with tf.GradientTape() as tape:
              predictions, hidden = model(inp, hidden)  # feeds the hidden state back into the model
              target = tf.reshape(target, (-1, ))  # reshapes for the loss function
              loss = loss_function(target, predictions)
              
          grads = tape.gradient(loss, model.variables)
          optimizer.apply_gradients(zip(grads, model.variables), global_step=tf.train.get_or_create_global_step())

          if batch % 100 == 0:
              print ('Epoch {} Batch {} Loss {:.4f}'.format(epoch + 1, batch, loss))
    
    print ('Epoch {} Loss {:.4f}'.format(epoch + 1, loss))
    print('Time taken for 1 epoch {} sec\n'.format(time.time() - start))
Epoch 1 Batch 0 Loss 4.5975
Epoch 1 Batch 100 Loss 2.1609
Epoch 1 Batch 200 Loss 1.9387
Epoch 1 Loss 1.8163
Time taken for 1 epoch 1007.0127582550049 sec

Epoch 2 Batch 0 Loss 1.7427
Epoch 2 Batch 100 Loss 1.7149
Epoch 2 Batch 200 Loss 1.6851
Epoch 2 Loss 1.6786
Time taken for 1 epoch 1009.5401530265808 sec

Epoch 3 Batch 0 Loss 1.5864
Epoch 3 Batch 100 Loss 1.5735
Epoch 3 Batch 200 Loss 1.5432
Epoch 3 Loss 1.5330
Time taken for 1 epoch 1008.4457356929779 sec

Epoch 4 Batch 0 Loss 1.4756
Epoch 4 Batch 100 Loss 1.5105
Epoch 4 Batch 200 Loss 1.4949
Epoch 4 Loss 1.4980
Time taken for 1 epoch 1010.7342929840088 sec

Epoch 5 Batch 0 Loss 1.3859
Epoch 5 Batch 100 Loss 1.4656
Epoch 5 Batch 200 Loss 1.4051
Epoch 5 Loss 1.4314
Time taken for 1 epoch 1012.4246871471405 sec

Epoch 6 Batch 0 Loss 1.3024
Epoch 6 Batch 100 Loss 1.3982
Epoch 6 Batch 200 Loss 1.3920
Epoch 6 Loss 1.4050
Time taken for 1 epoch 1009.992424249649 sec

Epoch 7 Batch 0 Loss 1.2550
Epoch 7 Batch 100 Loss 1.3588
Epoch 7 Batch 200 Loss 1.3480
Epoch 7 Loss 1.3742
Time taken for 1 epoch 1009.2752959728241 sec

Epoch 8 Batch 0 Loss 1.1944
Epoch 8 Batch 100 Loss 1.3088
Epoch 8 Batch 200 Loss 1.3028
Epoch 8 Loss 1.3164
Time taken for 1 epoch 1007.6811842918396 sec

Epoch 9 Batch 0 Loss 1.1755
Epoch 9 Batch 100 Loss 1.2597
Epoch 9 Batch 200 Loss 1.2338
Epoch 9 Loss 1.2662
Time taken for 1 epoch 1010.1400344371796 sec

Epoch 10 Batch 0 Loss 1.1144
Epoch 10 Batch 100 Loss 1.2354
Epoch 10 Batch 200 Loss 1.2604
Epoch 10 Loss 1.1994
Time taken for 1 epoch 1014.5306894779205 sec

Epoch 11 Batch 0 Loss 1.0661
Epoch 11 Batch 100 Loss 1.1259
Epoch 11 Batch 200 Loss 1.2032
Epoch 11 Loss 1.2087
Time taken for 1 epoch 1015.837060213089 sec

Epoch 12 Batch 0 Loss 1.0235
Epoch 12 Batch 100 Loss 1.0992
Epoch 12 Batch 200 Loss 1.1357
Epoch 12 Loss 1.1567
Time taken for 1 epoch 1012.8789830207825 sec

Epoch 13 Batch 0 Loss 0.9910
Epoch 13 Batch 100 Loss 1.1073
Epoch 13 Batch 200 Loss 1.1040
Epoch 13 Loss 1.1534
Time taken for 1 epoch 1013.3924562931061 sec

Epoch 14 Batch 0 Loss 0.9614
Epoch 14 Batch 100 Loss 1.0112
Epoch 14 Batch 200 Loss 1.1047
Epoch 14 Loss 1.0997
Time taken for 1 epoch 1010.0549929141998 sec

Epoch 15 Batch 0 Loss 0.8986
Epoch 15 Batch 100 Loss 1.0099
Epoch 15 Batch 200 Loss 1.0629
Epoch 15 Loss 1.0550
Time taken for 1 epoch 1010.1194486618042 sec

Epoch 16 Batch 0 Loss 0.8742
Epoch 16 Batch 100 Loss 0.9966
Epoch 16 Batch 200 Loss 1.0665
Epoch 16 Loss 1.0293
Time taken for 1 epoch 1010.7748596668243 sec

Epoch 17 Batch 0 Loss 0.8696
Epoch 17 Batch 100 Loss 0.9391
Epoch 17 Batch 200 Loss 1.0458
Epoch 17 Loss 0.9827
Time taken for 1 epoch 1009.4000136852264 sec

Epoch 18 Batch 0 Loss 0.8229
Epoch 18 Batch 100 Loss 0.9418
Epoch 18 Batch 200 Loss 0.9846
Epoch 18 Loss 0.9915
Time taken for 1 epoch 1018.6969776153564 sec

Epoch 19 Batch 0 Loss 0.8420
Epoch 19 Batch 100 Loss 0.9518
Epoch 19 Batch 200 Loss 0.9679
Epoch 19 Loss 0.9826
Time taken for 1 epoch 1015.187970161438 sec

Epoch 20 Batch 0 Loss 0.8031
Epoch 20 Batch 100 Loss 0.9082
Epoch 20 Batch 200 Loss 0.9899
Epoch 20 Loss 0.9728
Time taken for 1 epoch 1015.405241727829 sec

Epoch 21 Batch 0 Loss 0.8155
Epoch 21 Batch 100 Loss 0.8931
Epoch 21 Batch 200 Loss 0.9593
Epoch 21 Loss 0.9730
Time taken for 1 epoch 1014.1239879131317 sec

Epoch 22 Batch 0 Loss 0.7719
Epoch 22 Batch 100 Loss 0.9102
Epoch 22 Batch 200 Loss 0.9354
Epoch 22 Loss 0.9565
Time taken for 1 epoch 1012.7351298332214 sec

Epoch 23 Batch 0 Loss 0.7540
Epoch 23 Batch 100 Loss 0.8891
Epoch 23 Batch 200 Loss 0.9431
Epoch 23 Loss 0.9452
Time taken for 1 epoch 1015.2740514278412 sec

Epoch 24 Batch 0 Loss 0.7557
Epoch 24 Batch 100 Loss 0.8679
Epoch 24 Batch 200 Loss 0.9325
Epoch 24 Loss 0.9061
Time taken for 1 epoch 1013.387583732605 sec

Epoch 25 Batch 0 Loss 0.7418
Epoch 25 Batch 100 Loss 0.8140
Epoch 25 Batch 200 Loss 0.8961
Epoch 25 Loss 0.9026
Time taken for 1 epoch 1012.0506448745728 sec

Epoch 26 Batch 0 Loss 0.7547
Epoch 26 Batch 100 Loss 0.8344

...

The model was trained on Paperspace. 5 epochs are missing due to an average Internet connecton.

Anyway, it is enough to generate some text with the model.

Text generation

In [38]:
num_generate = 1000  # number of characters to generate
start_string = 'The child'  # beginning of the generated text. TODO: try start_string = ' '

input_eval = [char2idx[s] for s in start_string]  # converts start_string to numbers the model understands
input_eval = tf.expand_dims(input_eval, 0)  # 

text_generated = ''

temperature = 0.97  # the greater, the closer to an observation in the corpus

hidden = [tf.zeros((1, units))]
for i in range(num_generate):
    predictions, hidden = model(input_eval, hidden)  # predictions holds the probabily for each character to be most adequate continuation

    predictions = predictions / temperature  # alters characters' probabilities to be picked (but keeps the order)
    predicted_id = tf.multinomial(tf.exp(predictions), num_samples=1)[0][0].numpy()  # picks the next character for the generated text
    
    input_eval = tf.expand_dims([predicted_id], 0)
    text_generated += idx2char[predicted_id]  # appends

print (start_string + text_generated)
The childhood rose that we may do,
And now the country seat her shall he died,
  And her shadow of regret?

If you can dress your head and stone be forget--
  That were coming home from the corner of her eye.

He did one that ever one to the summer of the rain!

She was wanting from the pine,
  And the song of his lower lay,
While I am sad and sent the truest,
  And one clothe lads that waits
  Of the black men doth clouds the stars,
      And the stars have broken the stars
      Stood on the brook, the bridge is passing through the starry skies.

    When the sun was clear, and the thing to be true
    A blaze in the stream of the steed,
      And the stars have broken the stairs,
    And the clearing and the long bright morning stair,
    And one that makes the dark bells they seemed to say:
     "But one song of the crowd.
            The stars come and the straw,
      And cold and still their welcome home.

    All the cold work that has lured the sea,
    And the clock stood calm and sti

Conclusion

That's promising:

  • It spells words correctly
  • There is some structure (line breaks).
  • Found a punctation rule

Easy-to-fix issue: indents. The corpus itself is inconsitent for that regard. The fact that the model mimics the indents is in fact a good news.

Harder-to-fix issue: Sentences make little sense. Maybe further training will be enough. Also, playing with hyperparameters will help.