{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "## 用N-Gram模型在莎士比亚诗中训练word embedding\n", "N-gram 是计算机语言学和概率论范畴内的概念,是指给定的一段文本中N个项目的序列。\n", "N=1 时 N-gram 又称为 unigram,N=2 称为 bigram,N=3 称为 trigram,以此类推。实际应用通常采用 bigram 和 trigram 进行计算。\n", "本示例在莎士比亚十四行诗上实现了trigram。" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "# 环境\n", "本教程基于paddle2.0-alpha编写,如果您的环境不是本版本,请先安装paddle2.0-alpha。" ] }, { "cell_type": "code", "execution_count": 1, "metadata": {}, "outputs": [ { "data": { "text/plain": [ "'2.0.0-alpha0'" ] }, "execution_count": 1, "metadata": {}, "output_type": "execute_result" } ], "source": [ "import paddle\n", "paddle.__version__" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 数据集&&相关参数\n", "训练数据集采用了莎士比亚十四行诗,CONTEXT_SIZE设为2,意味着是trigram。EMBEDDING_DIM设为10。" ] }, { "cell_type": "code", "execution_count": 57, "metadata": {}, "outputs": [], "source": [ "CONTEXT_SIZE = 2\n", "EMBEDDING_DIM = 10\n", "\n", "test_sentence = \"\"\"When forty winters shall besiege thy brow,\n", "And dig deep trenches in thy beauty's field,\n", "Thy youth's proud livery so gazed on now,\n", "Will be a totter'd weed of small worth held:\n", "Then being asked, where all thy beauty lies,\n", "Where all the treasure of thy lusty days;\n", "To say, within thine own deep sunken eyes,\n", "Were an all-eating shame, and thriftless praise.\n", "How much more praise deserv'd thy beauty's use,\n", "If thou couldst answer 'This fair child of mine\n", "Shall sum my count, and make my old excuse,'\n", "Proving his beauty by succession thine!\n", "This were to be new made when thou art old,\n", "And see thy blood warm when thou feel'st it cold.\"\"\".split()" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 数据预处理\n", "将文本被拆成了元组的形式,格式为(('第一个词', '第二个词'), '第三个词');其中,第三个词就是我们的目标。" ] }, { "cell_type": "code", "execution_count": 58, "metadata": {}, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "[(('When', 'forty'), 'winters'), (('forty', 'winters'), 'shall'), (('winters', 'shall'), 'besiege')]\n" ] } ], "source": [ "trigram = [((test_sentence[i], test_sentence[i + 1]), test_sentence[i + 2])\n", " for i in range(len(test_sentence) - 2)]\n", "\n", "vocab = set(test_sentence)\n", "word_to_idx = {word: i for i, word in enumerate(vocab)}\n", "idx_to_word = {word_to_idx[word]: word for word in word_to_idx}\n", "# 看一下数据集\n", "print(trigram[:3])\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 构建`Dataset`类 加载数据" ] }, { "cell_type": "code", "execution_count": 59, "metadata": {}, "outputs": [], "source": [ "import paddle\n", "class TrainDataset(paddle.io.Dataset):\n", " def __init__(self, tuple_data, vocab):\n", " self.tuple_data = tuple_data\n", " self.vocab = vocab\n", "\n", " def __getitem__(self, idx):\n", " data = list(self.tuple_data[idx][0])\n", " label = list(self.tuple_data[idx][1])\n", " return data, label\n", " \n", " def __len__(self):\n", " return len(self.tuple_data)" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 组网&训练\n", "这里用paddle动态图的方式组网,由于是N-Gram模型,只需要一层`Embedding`与两层`Linear`就可以完成网络模型的构建。" ] }, { "cell_type": "code", "execution_count": 79, "metadata": {}, "outputs": [], "source": [ "import paddle\n", "import numpy as np\n", "class NGramModel(paddle.nn.Layer):\n", " def __init__(self, vocab_size, embedding_dim, context_size):\n", " super(NGramModel, self).__init__()\n", " self.embedding = paddle.nn.Embedding(size=[vocab_size, embedding_dim])\n", " self.linear1 = paddle.nn.Linear(context_size * embedding_dim, 128)\n", " self.linear2 = paddle.nn.Linear(128, vocab_size)\n", "\n", " def forward(self, x):\n", " x = self.embedding(x)\n", " x = paddle.reshape(x, [1, -1])\n", " x = self.linear1(x)\n", " x = paddle.nn.functional.relu(x)\n", " x = self.linear2(x)\n", " x = paddle.nn.functional.softmax(x)\n", " return x" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### 初始化Model,并定义相关的参数。" ] }, { "cell_type": "code", "execution_count": 121, "metadata": {}, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "epoch: 0, loss is: [4.631529]\n", "epoch: 50, loss is: [4.6081576]\n", "epoch: 100, loss is: [4.600631]\n", "epoch: 150, loss is: [4.603069]\n", "epoch: 200, loss is: [4.592647]\n", "epoch: 250, loss is: [4.5626693]\n", "epoch: 300, loss is: [4.513106]\n", "epoch: 350, loss is: [4.4345813]\n", "epoch: 400, loss is: [4.3238697]\n", "epoch: 450, loss is: [4.1728854]\n", "epoch: 500, loss is: [3.9622664]\n", "epoch: 550, loss is: [3.67673]\n", "epoch: 600, loss is: [3.2998457]\n", "epoch: 650, loss is: [2.8206367]\n", "epoch: 700, loss is: [2.2514927]\n", "epoch: 750, loss is: [1.6479329]\n", "epoch: 800, loss is: [1.1147357]\n", "epoch: 850, loss is: [0.73231363]\n", "epoch: 900, loss is: [0.49481753]\n", "epoch: 950, loss is: [0.3504072]\n" ] } ], "source": [ "vocab_size = len(vocab)\n", "embedding_dim = 10\n", "context_size = 2\n", "\n", "paddle.enable_imperative()\n", "losses = []\n", "def train(model):\n", " model.train()\n", " optim = paddle.optimizer.SGD(learning_rate=0.001, parameter_list=model.parameters())\n", " for epoch in range(1000):\n", " # 留最后10组作为预测\n", " for context, target in trigram[:-10]:\n", " context_idxs = list(map(lambda w: word_to_idx[w], context))\n", " x_data = paddle.imperative.to_variable(np.array(context_idxs))\n", " y_data = paddle.imperative.to_variable(np.array([word_to_idx[target]]))\n", " predicts = model(x_data)\n", " # print (predicts)\n", " loss = paddle.nn.functional.cross_entropy(predicts, y_data)\n", " loss.backward()\n", " optim.minimize(loss)\n", " model.clear_gradients()\n", " if epoch % 50 == 0:\n", " print(\"epoch: {}, loss is: {}\".format(epoch, loss.numpy()))\n", " losses.append(loss.numpy())\n", "model = NGramModel(vocab_size, embedding_dim, context_size)\n", "train(model)" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 打印loss下降曲线\n", "通过可视化loss的曲线,可以看到模型训练的效果。" ] }, { "cell_type": "code", "execution_count": 123, "metadata": {}, "outputs": [ { "data": { "text/plain": [ "[]" ] }, "execution_count": 123, "metadata": {}, "output_type": "execute_result" }, { "data": { "image/png": "\n", "text/plain": [ "
" ] }, "metadata": { "needs_background": "light" }, "output_type": "display_data" } ], "source": [ "import matplotlib.pyplot as plt\n", "import matplotlib.ticker as ticker\n", "%matplotlib inline\n", "\n", "plt.figure()\n", "plt.plot(losses)" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## 预测\n", "用训练好的模型进行预测。" ] }, { "cell_type": "code", "execution_count": 127, "metadata": {}, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "the input words is: praise., How\n", "the predict words is: much\n", "the true words is: much\n" ] } ], "source": [ "import random\n", "def test(model):\n", " model.eval()\n", " # 从最后10组数据中随机选取1个\n", " idx = random.randint(len(trigram)-10, len(trigram)-1)\n", " print('the input words is: ' + trigram[idx][0][0] + ', ' + trigram[idx][0][1])\n", " x_data = list(map(lambda w: word_to_idx[w], trigram[idx][0]))\n", " x_data = paddle.imperative.to_variable(np.array(x_data))\n", " predicts = model(x_data)\n", " predicts = predicts.numpy().tolist()[0]\n", " predicts = predicts.index(max(predicts))\n", " print('the predict words is: ' + idx_to_word[predicts])\n", " y_data = trigram[idx][1]\n", " print('the true words is: ' + y_data)\n", "test(model)" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3.7.3 64-bit", "language": "python", "name": "python_defaultSpec_1598180286976" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.7.3-final" } }, "nbformat": 4, "nbformat_minor": 4 }