Merge branch 'dygraph' into dyg_db

00e67e0b · wangjingyeye · GitHub · 872d73da · 9ae52fb3 · 00e67e0b
79 changed file
--- a/README_ch.md
+++ b/README_ch.md
@@ -71,6 +71,8 @@ PaddleOCR旨在打造一套丰富、领先、且实用的OCR工具库，助力
 ## 《动手学OCR》电子书
 - [《动手学OCR》电子书📚](./doc/doc_ch/ocr_book.md)
+## 场景应用
+- PaddleOCR场景应用覆盖通用，制造、金融、交通行业的主要OCR垂类应用，在PP-OCR、PP-Structure的通用能力基础之上，以notebook的形式展示利用场景数据微调、模型优化方法、数据增广等内容，为开发者快速落地OCR应用提供示范与启发。详情可查看[README](./applications)。
 <a name="开源社区"></a>
 ## 开源社区

--- a/applications/PCB字符识别/PCB字符识别.md
+++ b/applications/PCB字符识别/PCB字符识别.md
@@ -206,7 +206,11 @@ Eval.dataset.transforms.DetResizeForTest:  尺寸
        limit_type: 'min'
 ```
-然后执行评估代码
+如需获取已训练模型，请扫码填写问卷，加入PaddleOCR官方交流群获取全部OCR垂类模型下载链接、《动手学OCR》电子书等全套OCR学习资料🎁
+<div align="left">
+<img src="https://ai-studio-static-online.cdn.bcebos.com/dd721099bd50478f9d5fb13d8dd00fad69c22d6848244fd3a1d3980d7fefc63e"  width = "150" height = "150" />
+</div>
+将下载或训练完成的模型放置在对应目录下即可完成模型评估。
 ```python

--- a/applications/README.md
+++ b/applications/README.md
+# 场景应用
+PaddleOCR场景应用覆盖通用，制造、金融、交通行业的主要OCR垂类应用，在PP-OCR、PP-Structure的通用能力基础之上，以notebook的形式展示利用场景数据微调、模型优化方法、数据增广等内容，为开发者快速落地OCR应用提供示范与启发。
+> 如需下载全部垂类模型，可以扫描下方二维码，关注公众号填写问卷后，加入PaddleOCR官方交流群获取20G OCR学习大礼包（内含《动手学OCR》电子书、课程回放视频、前沿论文等重磅资料）
+<div align="center">
+<img src="https://ai-studio-static-online.cdn.bcebos.com/dd721099bd50478f9d5fb13d8dd00fad69c22d6848244fd3a1d3980d7fefc63e"  width = "150" height = "150" />
+</div>
+> 如果您是企业开发者且未在下述场景中找到合适的方案，可以填写[OCR应用合作调研问卷](https://paddle.wjx.cn/vj/QwF7GKw.aspx)，免费与官方团队展开不同层次的合作，包括但不限于问题抽象、确定技术方案、项目答疑、共同研发等。如果您已经使用PaddleOCR落地项目，也可以填写此问卷，与飞桨平台共同宣传推广，提升企业技术品宣。期待您的提交！
+## 通用
+| 类别                                              | 亮点     | 类别       | 亮点         |
+| ------------------------------------------------- | -------- | ---------- | ------------ |
+| [高精度中文识别模型SVTR](./高精度中文识别模型.md) | 新增模型 | 手写体识别 | 新增字形支持 |
+## 制造
+| 类别                                                         | 亮点                           | 类别                                        | 亮点                 |
+| ------------------------------------------------------------ | ------------------------------ | ------------------------------------------- | -------------------- |
+| [数码管识别](./光功率计数码管字符识别/光功率计数码管字符识别.md) | 数码管数据合成、漏识别调优     | 电表识别                                    | 大分辨率图像检测调优 |
+| [液晶屏读数识别](./液晶屏读数识别.md)                        | 检测模型蒸馏、Serving部署      | [PCB文字识别](./PCB字符识别/PCB字符识别.md) | 小尺寸文本检测与识别 |
+| [包装生产日期](./包装生产日期识别.md)                        | 点阵字符合成、过曝过暗文字识别 | 液晶屏缺陷检测                              | 非文字字符识别       |
+## 金融
+| 类别                           | 亮点                     | 类别         | 亮点                  |
+| ------------------------------ | ------------------------ | ------------ | --------------------- |
+| [表单VQA](./多模态表单识别.md) | 多模态通用表单结构化提取 | 通用卡证识别 | 通用结构化提取        |
+| 增值税发票                     | 尽请期待                 | 身份证识别   | 结构化提取、图像阴影  |
+| 印章检测与识别                 | 端到端弯曲文本识别       | 合同比对     | 密集文本检测、NLP串联 |
+## 交通
+| 类别                            | 亮点                           | 类别       | 亮点     |
+| ------------------------------- | ------------------------------ | ---------- | -------- |
+| [车牌识别](./轻量级车牌识别.md) | 多角度图像、轻量模型、端侧部署 | 快递单识别 | 尽请期待 |
+| 驾驶证/行驶证识别               | 尽请期待                       |            |          |
\ No newline at end of file
--- a/applications/光功率计数码管字符识别/corpus/digital.txt
+++ b/applications/光功率计数码管字符识别/corpus/digital.txt
+46.39
+40.08
+89.52
+-71.93
+23.19
+-81.02
+-34.09
+05.87
+-67.80
+-51.56
+-34.58
+37.91
+56.98
+29.01
+-90.13
+35.55
+66.07
+-90.35
+-50.93
+42.42
+21.40
+-30.99
+-71.78
+25.60
+-48.69
+-72.28
+-17.55
+-99.93
+-47.35
+-64.89
+-31.28
+-90.01
+05.17
+30.91
+30.56
+-06.90
+79.05
+67.74
+-32.31
+94.22
+28.75
+51.03
+-58.96
--- a/applications/光功率计数码管字符识别/fonts/DS-DIGI.TTF
+++ b/applications/光功率计数码管字符识别/fonts/DS-DIGI.TTF
--- a/applications/光功率计数码管字符识别/fonts/DS-DIGIB.TTF
+++ b/applications/光功率计数码管字符识别/fonts/DS-DIGIB.TTF
--- a/applications/光功率计数码管字符识别/光功率计数码管字符识别.md
+++ b/applications/光功率计数码管字符识别/光功率计数码管字符识别.md
+# 光功率计数码管字符识别
+本案例将使用OCR技术自动识别光功率计显示屏文字，通过本章您可以掌握：
+- PaddleOCR快速使用
+- 数据合成方法
+- 数据挖掘方法
+- 基于现有数据微调
+## 1. 背景介绍
+光功率计（optical power meter ）是指用于测量绝对光功率或通过一段光纤的光功率相对损耗的仪器。在光纤系统中，测量光功率是最基本的，非常像电子学中的万用表；在光纤测量中，光功率计是重负荷常用表。
+<img src="https://bkimg.cdn.bcebos.com/pic/a08b87d6277f9e2f999f5e3e1c30e924b899f35a?x-bce-process=image/watermark,image_d2F0ZXIvYmFpa2U5Mg==,g_7,xp_5,yp_5/format,f_auto" width="400">
+目前光功率计缺少将数据直接输出的功能，需要人工读数。这一项工作单调重复，如果可以使用机器替代人工，将节约大量成本。针对上述问题，希望通过摄像头拍照->智能读数的方式高效地完成此任务。
+为实现智能读数，通常会采取文本检测+文本识别的方案：
+第一步，使用文本检测模型定位出光功率计中的数字部分；
+第二步，使用文本识别模型获得准确的数字和单位信息。
+本项目主要介绍如何完成第二步文本识别部分，包括：真实评估集的建立、训练数据的合成、基于 PP-OCRv3 和 SVTR_Tiny 两个模型进行训练，以及评估和推理。
+本项目难点如下：
+- 光功率计数码管字符数据较少，难以获取。
+- 数码管中小数点占像素较少，容易漏识别。
+针对以上问题， 本例选用 PP-OCRv3 和 SVTR_Tiny 两个高精度模型训练，同时提供了真实数据挖掘案例和数据合成案例。基于 PP-OCRv3 模型，在构建的真实评估集上精度从 52% 提升至 72%，SVTR_Tiny 模型精度可达到 78.9%。
+aistudio项目链接: [光功率计数码管字符识别](https://aistudio.baidu.com/aistudio/projectdetail/4049044?contributionType=1)
+## 2. PaddleOCR 快速使用
+PaddleOCR 旨在打造一套丰富、领先、且实用的OCR工具库，助力开发者训练出更好的模型，并应用落地。
+![](https://github.com/PaddlePaddle/PaddleOCR/raw/release/2.5/doc/imgs_results/ch_ppocr_mobile_v2.0/test_add_91.jpg)
+官方提供了适用于通用场景的高精轻量模型，首先使用官方提供的 PP-OCRv3 模型预测图片，验证下当前模型在光功率计场景上的效果。
+- 准备环境
+```
+python3 -m pip install -U pip
+python3 -m pip install paddleocr
+```
+- 测试效果
+测试图：
+![](https://ai-studio-static-online.cdn.bcebos.com/8dca91f016884e16ad9216d416da72ea08190f97d87b4be883f15079b7ebab9a)
+```
+paddleocr --lang=ch --det=Fase --image_dir=data
+```
+得到如下测试结果：
+```
+('.7000', 0.6885431408882141)
+```
+发现数字识别较准，然而对负号和小数点识别不准确。 由于PP-OCRv3的训练数据大多为通用场景数据，在特定的场景上效果可能不够好。因此需要基于场景数据进行微调。
+下面就主要介绍如何在光功率计（数码管）场景上微调训练。
+## 3. 开始训练
+### 3.1 数据准备
+特定的工业场景往往很难获取开源的真实数据集，光功率计也是如此。在实际工业场景中，可以通过摄像头采集的方法收集大量真实数据，本例中重点介绍数据合成方法和真实数据挖掘方法，如何利用有限的数据优化模型精度。
+数据集分为两个部分：合成数据，真实数据， 其中合成数据由 text_renderer 工具批量生成得到， 真实数据通过爬虫等方式在百度图片中搜索并使用 PPOCRLabel 标注得到。
+- 合成数据
+本例中数据合成工具使用的是 [text_renderer](https://github.com/Sanster/text_renderer)， 该工具可以合成用于文本识别训练的文本行数据:
+![](https://github.com/oh-my-ocr/text_renderer/raw/master/example_data/effect_layout_image/char_spacing_compact.jpg)
+![](https://github.com/oh-my-ocr/text_renderer/raw/master/example_data/effect_layout_image/color_image.jpg)
+```
+export https_proxy=http://172.19.57.45:3128
+git clone https://github.com/oh-my-ocr/text_renderer
+```
+```
+import os
+python3 setup.py develop
+python3 -m pip install -r docker/requirements.txt
+python3 main.py \
+    --config example_data/example.py \
+    --dataset img \
+    --num_processes 2 \
+    --log_period 10
+```
+给定字体和语料，就可以合成较为丰富样式的文本行数据。 光功率计识别场景，目标是正确识别数码管文本，因此需要收集部分数码管字体，训练语料，用于合成文本识别数据。
+将收集好的语料存放在 example_data 路径下：
+```
+ln -s ./fonts/DS* text_renderer/example_data/font/
+ln -s ./corpus/digital.txt text_renderer/example_data/text/
+```
+修改 text_renderer/example_data/font_list/font_list.txt ,选择需要的字体开始合成：
+```
+python3 main.py \
+    --config example_data/digital_example.py \
+    --dataset img \
+    --num_processes 2 \
+    --log_period 10
+```
+合成图片会被存在目录 text_renderer/example_data/digital/chn_data 下
+查看合成的数据样例：
+![](https://ai-studio-static-online.cdn.bcebos.com/7d5774a273f84efba5b9ce7fd3f86e9ef24b6473e046444db69fa3ca20ac0986)
+- 真实数据挖掘
+模型训练需要使用真实数据作为评价指标，否则很容易过拟合到简单的合成数据中。没有开源数据的情况下，可以利用部分无标注数据+标注工具获得真实数据。
+1. 数据搜集
+使用[爬虫工具](https://github.com/Joeclinton1/google-images-download.git)获得无标注数据
+2. [PPOCRLabel](https://github.com/PaddlePaddle/PaddleOCR/tree/release/2.5/PPOCRLabel) 完成半自动标注
+PPOCRLabel是一款适用于OCR领域的半自动化图形标注工具，内置PP-OCR模型对数据自动标注和重新识别。使用Python3和PyQT5编写，支持矩形框标注、表格标注、不规则文本标注、关键信息标注模式，导出格式可直接用于PaddleOCR检测和识别模型的训练。
+![](https://github.com/PaddlePaddle/PaddleOCR/raw/release/2.5/PPOCRLabel/data/gif/steps_en.gif)
+收集完数据后就可以进行分配了，验证集中一般都是真实数据，训练集中包含合成数据+真实数据。本例中标注了155张图片，其中训练集和验证集的数目为100和55。
+最终 `data` 文件夹应包含以下几部分：
+```
+|-data
+  |- synth_train.txt
+  |- real_train.txt
+  |- real_eval.txt
+  |- synthetic_data
+      |- word_001.png
+      |- word_002.jpg
+      |- word_003.jpg
+      | ...
+  |- real_data
+      |- word_001.png
+      |- word_002.jpg
+      |- word_003.jpg
+      | ...
+  ...
+```
+### 3.2 模型选择
+本案例提供了2种文本识别模型：PP-OCRv3 识别模型 和 SVTR_Tiny：
+[PP-OCRv3 识别模型](https://github.com/PaddlePaddle/PaddleOCR/blob/release/2.5/doc/doc_ch/PP-OCRv3_introduction.md)：PP-OCRv3的识别模块是基于文本识别算法SVTR优化。SVTR不再采用RNN结构，通过引入Transformers结构更加有效地挖掘文本行图像的上下文信息，从而提升文本识别能力。并进行了一系列结构改进加速模型预测。
+[SVTR_Tiny](https://arxiv.org/abs/2205.00159):SVTR提出了一种用于场景文本识别的单视觉模型，该模型在patch-wise image tokenization框架内，完全摒弃了序列建模，在精度具有竞争力的前提下，模型参数量更少，速度更快。
+以上两个策略在自建中文数据集上的精度和速度对比如下：
+| ID | 策略 |  模型大小 | 精度 | 预测耗时（CPU + MKLDNN)|
+|-----|-----|--------|----| --- |
+| 01 | PP-OCRv2 | 8M | 74.8% | 8.54ms |
+| 02 | SVTR_Tiny | 21M | 80.1% | 97ms |
+| 03 | SVTR_LCNet(h32) | 12M | 71.9% | 6.6ms |
+| 04 | SVTR_LCNet(h48) | 12M | 73.98% | 7.6ms |
+| 05 | + GTC | 12M | 75.8% | 7.6ms |
+| 06 | + TextConAug | 12M | 76.3% | 7.6ms |
+| 07 | + TextRotNet | 12M | 76.9% | 7.6ms |
+| 08 | + UDML | 12M | 78.4% | 7.6ms |
+| 09 | + UIM | 12M | 79.4% | 7.6ms |
+### 3.3 开始训练
+首先下载 PaddleOCR 代码库
+```
+git clone -b release/2.5 https://github.com/PaddlePaddle/PaddleOCR.git
+```
+PaddleOCR提供了训练脚本、评估脚本和预测脚本，本节将以 PP-OCRv3 中文识别模型为例：
+**Step1：下载预训练模型**
+首先下载 pretrain model，您可以下载训练好的模型在自定义数据上进行finetune
+```
+cd PaddleOCR/
+# 下载PP-OCRv3 中文预训练模型
+wget -P ./pretrain_models/ https://paddleocr.bj.bcebos.com/PP-OCRv3/chinese/ch_PP-OCRv3_rec_train.tar
+# 解压模型参数
+cd pretrain_models
+tar -xf ch_PP-OCRv3_rec_train.tar && rm -rf ch_PP-OCRv3_rec_train.tar
+```
+**Step2：自定义字典文件**
+接下来需要提供一个字典（{word_dict_name}.txt），使模型在训练时，可以将所有出现的字符映射为字典的索引。
+因此字典需要包含所有希望被正确识别的字符，{word_dict_name}.txt需要写成如下格式，并以 `utf-8` 编码格式保存：
+```
+0
+1
+2
+3
+4
+5
+6
+7
+8
+9
+-
+.
+```
+word_dict.txt 每行有一个单字，将字符与数字索引映射在一起，“3.14” 将被映射成 [3, 11, 1, 4]
+* 内置字典
+PaddleOCR内置了一部分字典，可以按需使用。
+`ppocr/utils/ppocr_keys_v1.txt` 是一个包含6623个字符的中文字典
+`ppocr/utils/ic15_dict.txt` 是一个包含36个字符的英文字典
+* 自定义字典
+内置字典面向通用场景，具体的工业场景中，可能需要识别特殊字符，或者只需识别某几个字符，此时自定义字典会更提升模型精度。例如在光功率计场景中，需要识别数字和单位。
+遍历真实数据标签中的字符，制作字典`digital_dict.txt`如下所示：
+```
+-
+.
+0
+1
+2
+3
+4
+5
+6
+7
+8
+9
+B
+E
+F
+H
+L
+N
+T
+W
+d
+k
+m
+n
+o
+z
+```
+**Step3：修改配置文件**
+为了更好的使用预训练模型，训练推荐使用[ch_PP-OCRv3_rec_distillation.yml](../../configs/rec/PP-OCRv3/ch_PP-OCRv3_rec_distillation.yml)配置文件，并参考下列说明修改配置文件：
+以 `ch_PP-OCRv3_rec_distillation.yml` 为例：
+```
+Global:
+  ...
+  # 添加自定义字典，如修改字典请将路径指向新字典
+  character_dict_path: ppocr/utils/dict/digital_dict.txt
+  ...
+  # 识别空格
+  use_space_char: True
+Optimizer:
+  ...
+  # 添加学习率衰减策略
+  lr:
+    name: Cosine
+    learning_rate: 0.001
+  ...
+...
+Train:
+  dataset:
+    # 数据集格式，支持LMDBDataSet以及SimpleDataSet
+    name: SimpleDataSet
+    # 数据集路径
+    data_dir: ./data/
+    # 训练集标签文件
+    label_file_list:
+    - ./train_data/digital_img/digital_train.txt  #11w
+    - ./train_data/digital_img/real_train.txt     #100
+    - ./train_data/digital_img/dbm_img/dbm.txt    #3w
+    ratio_list:
+    - 0.3
+    - 1.0
+    - 1.0
+    transforms:
+      ...
+      - RecResizeImg:
+          # 修改 image_shape 以适应长文本
+          image_shape: [3, 48, 320]
+      ...
+  loader:
+    ...
+    # 单卡训练的batch_size
+    batch_size_per_card: 256
+    ...
+Eval:
+  dataset:
+    # 数据集格式，支持LMDBDataSet以及SimpleDataSet
+    name: SimpleDataSet
+    # 数据集路径
+    data_dir: ./data
+    # 验证集标签文件
+    label_file_list:
+    - ./train_data/digital_img/real_val.txt
+    transforms:
+      ...
+      - RecResizeImg:
+          # 修改 image_shape 以适应长文本
+          image_shape: [3, 48, 320]
+      ...
+  loader:
+    # 单卡验证的batch_size
+    batch_size_per_card: 256
+    ...
+```
+**注意，训练/预测/评估时的配置文件请务必与训练一致。**
+**Step4：启动训练**
+*如果您安装的是cpu版本，请将配置文件中的 `use_gpu` 字段修改为false*
+```
+# GPU训练 支持单卡，多卡训练
+# 训练数码管数据 训练日志会自动保存为 "{save_model_dir}" 下的train.log
+#单卡训练（训练周期长，不建议）
+python3 tools/train.py -c configs/rec/PP-OCRv3/ch_PP-OCRv3_rec_distillation.yml -o Global.pretrained_model=./pretrain_models/ch_PP-OCRv3_rec_train/best_accuracy
+#多卡训练，通过--gpus参数指定卡号
+python3 -m paddle.distributed.launch --gpus '0,1,2,3'  tools/train.py -c configs/rec/PP-OCRv3/ch_PP-OCRv3_rec_distillation.yml -o Global.pretrained_model=./pretrain_models/en_PP-OCRv3_rec_train/best_accuracy
+```
+PaddleOCR支持训练和评估交替进行, 可以在 `configs/rec/PP-OCRv3/ch_PP-OCRv3_rec_distillation.yml` 中修改 `eval_batch_step` 设置评估频率，默认每500个iter评估一次。评估过程中默认将最佳acc模型，保存为 `output/ch_PP-OCRv3_rec_distill/best_accuracy` 。
+如果验证集很大，测试将会比较耗时，建议减少评估次数，或训练完再进行评估。
+### SVTR_Tiny 训练
+SVTR_Tiny 训练步骤与上面一致，SVTR支持的配置和模型训练权重可以参考[算法介绍文档](https://github.com/PaddlePaddle/PaddleOCR/blob/release/2.5/doc/doc_ch/algorithm_rec_svtr.md)
+**Step1：下载预训练模型**
+```
+# 下载 SVTR_Tiny 中文识别预训练模型和配置文件
+wget https://paddleocr.bj.bcebos.com/PP-OCRv3/chinese/rec_svtr_tiny_none_ctc_ch_train.tar
+# 解压模型参数
+tar -xf rec_svtr_tiny_none_ctc_ch_train.tar && rm -rf rec_svtr_tiny_none_ctc_ch_train.tar
+```
+**Step2：自定义字典文件**
+字典依然使用自定义的 digital_dict.txt
+**Step3：修改配置文件**
+配置文件中对应修改字典路径和数据路径
+**Step4：启动训练**
+```
+## 单卡训练
+python tools/train.py -c rec_svtr_tiny_none_ctc_ch_train/rec_svtr_tiny_6local_6global_stn_ch.yml \
+           -o Global.pretrained_model=./rec_svtr_tiny_none_ctc_ch_train/best_accuracy
+```
+### 3.4 验证效果
+如需获取已训练模型，请扫码填写问卷，加入PaddleOCR官方交流群获取全部OCR垂类模型下载链接、《动手学OCR》电子书等全套OCR学习资料🎁
+<div align="left">
+<img src="https://ai-studio-static-online.cdn.bcebos.com/dd721099bd50478f9d5fb13d8dd00fad69c22d6848244fd3a1d3980d7fefc63e"  width = "150" height = "150" />
+</div>
+将下载或训练完成的模型放置在对应目录下即可完成模型推理
+* 指标评估
+训练中模型参数默认保存在`Global.save_model_dir`目录下。在评估指标时，需要设置`Global.checkpoints`指向保存的参数文件。评估数据集可以通过 `configs/rec/PP-OCRv3/ch_PP-OCRv3_rec_distillation.yml`  修改Eval中的 `label_file_path` 设置。
+```
+# GPU 评估， Global.checkpoints 为待测权重
+python3 -m paddle.distributed.launch --gpus '0' tools/eval.py -c configs/rec/PP-OCRv3/ch_PP-OCRv3_rec_distillation.yml -o Global.checkpoints={path/to/weights}/best_accuracy
+```
+* 测试识别效果
+使用 PaddleOCR 训练好的模型，可以通过以下脚本进行快速预测。
+默认预测图片存储在 `infer_img` 里，通过 `-o Global.checkpoints` 加载训练好的参数文件：
+根据配置文件中设置的 `save_model_dir` 和 `save_epoch_step` 字段，会有以下几种参数被保存下来：
+```
+output/rec/
+├── best_accuracy.pdopt  
+├── best_accuracy.pdparams  
+├── best_accuracy.states  
+├── config.yml  
+├── iter_epoch_3.pdopt  
+├── iter_epoch_3.pdparams  
+├── iter_epoch_3.states  
+├── latest.pdopt  
+├── latest.pdparams  
+├── latest.states  
+└── train.log
+```
+其中 best_accuracy.* 是评估集上的最优模型；iter_epoch_x.* 是以 `save_epoch_step` 为间隔保存下来的模型；latest.* 是最后一个epoch的模型。
+```
+# 预测英文结果
+python3 tools/infer_rec.py -c configs/rec/PP-OCRv3/ch_PP-OCRv3_rec_distillation.yml -o Global.pretrained_model={path/to/weights}/best_accuracy  Global.infer_img=test_digital.png
+```
+预测图片：
+![](https://ai-studio-static-online.cdn.bcebos.com/8dca91f016884e16ad9216d416da72ea08190f97d87b4be883f15079b7ebab9a)
+得到输入图像的预测结果：
+```
+infer_img: test_digital.png
+        result: ('-70.00', 0.9998967)
+```
--- a/applications/包装生产日期识别.md
+++ b/applications/包装生产日期识别.md
--- a/applications/多模态表单识别.md
+++ b/applications/多模态表单识别.md
--- a/applications/液晶屏读数识别.md
+++ b/applications/液晶屏读数识别.md
--- a/applications/高精度中文识别模型.md
+++ b/applications/高精度中文识别模型.md
+# 高精度中文场景文本识别模型SVTR
+## 1. 简介
+PP-OCRv3是百度开源的超轻量级场景文本检测识别模型库，其中超轻量的场景中文识别模型SVTR_LCNet使用了SVTR算法结构。为了保证速度，SVTR_LCNet将SVTR模型的Local Blocks替换为LCNet，使用两层Global Blocks。在中文场景中，PP-OCRv3识别主要使用如下优化策略：
+- GTC：Attention指导CTC训练策略；
+- TextConAug：挖掘文字上下文信息的数据增广策略；
+- TextRotNet：自监督的预训练模型；
+- UDML：联合互学习策略；
+- UIM：无标注数据挖掘方案。
+其中 *UIM：无标注数据挖掘方案* 使用了高精度的SVTR中文模型进行无标注文件的刷库，该模型在PP-OCRv3识别的数据集上训练，精度对比如下表。
+|中文识别算法|模型|UIM|精度|
+| --- | --- | --- |--- |
+|PP-OCRv3|SVTR_LCNet| w/o |78.4%|
+|PP-OCRv3|SVTR_LCNet| w |79.4%|
+|SVTR|SVTR-Tiny|-|82.5%|
+aistudio项目链接: [高精度中文场景文本识别模型SVTR](https://aistudio.baidu.com/aistudio/projectdetail/4263032)
+## 2. SVTR中文模型使用
+### 环境准备
+本任务基于Aistudio完成, 具体环境如下：
+- 操作系统: Linux
+- PaddlePaddle: 2.3
+- PaddleOCR: dygraph
+下载 PaddleOCR代码
+```bash
+git clone -b dygraph https://github.com/PaddlePaddle/PaddleOCR
+```
+安装依赖库
+```bash
+pip install -r PaddleOCR/requirements.txt -i https://mirror.baidu.com/pypi/simple
+```
+### 快速使用
+获取SVTR中文模型文件，请扫码填写问卷，加入PaddleOCR官方交流群获取全部OCR垂类模型下载链接、《动手学OCR》电子书等全套OCR学习资料🎁
+<div align="center">
+<img src="https://ai-studio-static-online.cdn.bcebos.com/dd721099bd50478f9d5fb13d8dd00fad69c22d6848244fd3a1d3980d7fefc63e"  width = "150" height = "150" />
+</div>
+```bash
+# 解压模型文件
+tar xf svtr_ch_high_accuracy.tar
+```
+预测中文文本，以下图为例：
+![](../doc/imgs_words/ch/word_1.jpg)
+预测命令：
+```bash
+# CPU预测
+python tools/infer_rec.py -c configs/rec/rec_svtrnet_ch.yml -o Global.pretrained_model=./svtr_ch_high_accuracy/best_accuracy Global.infer_img=./doc/imgs_words/ch/word_1.jpg Global.use_gpu=False
+# GPU预测
+#python tools/infer_rec.py -c configs/rec/rec_svtrnet_ch.yml -o Global.pretrained_model=./svtr_ch_high_accuracy/best_accuracy Global.infer_img=./doc/imgs_words/ch/word_1.jpg Global.use_gpu=True
+```
+可以看到最后打印结果为
+- result: 韩国小馆    0.9853458404541016
+0.9853458404541016为预测置信度。
+### 推理模型导出与预测
+inference 模型（paddle.jit.save保存的模型） 一般是模型训练，把模型结构和模型参数保存在文件中的固化模型，多用于预测部署场景。 训练过程中保存的模型是checkpoints模型，保存的只有模型的参数，多用于恢复训练等。 与checkpoints模型相比，inference 模型会额外保存模型的结构信息，在预测部署、加速推理上性能优越，灵活方便，适合于实际系统集成。
+运行识别模型转inference模型命令，如下：
+```bash
+python tools/export_model.py -c configs/rec/rec_svtrnet_ch.yml -o Global.pretrained_model=./svtr_ch_high_accuracy/best_accuracy Global.save_inference_dir=./inference/svtr_ch
+```
+转换成功后，在目录下有三个文件：
+```shell
+inference/svtr_ch/
+    ├── inference.pdiparams         # 识别inference模型的参数文件
+    ├── inference.pdiparams.info    # 识别inference模型的参数信息，可忽略
+    └── inference.pdmodel           # 识别inference模型的program文件
+```
+inference模型预测，命令如下：
+```bash
+# CPU预测
+python3 tools/infer/predict_rec.py --image_dir="./doc/imgs_words/ch/word_1.jpg" --rec_algorithm='SVTR' --rec_model_dir=./inference/svtr_ch/ --rec_image_shape='3, 32, 320'  --rec_char_dict_path=ppocr/utils/ppocr_keys_v1.txt --use_gpu=False
+# GPU预测
+#python3 tools/infer/predict_rec.py --image_dir="./doc/imgs_words/ch/word_1.jpg" --rec_algorithm='SVTR' --rec_model_dir=./inference/svtr_ch/ --rec_image_shape='3, 32, 320'  --rec_char_dict_path=ppocr/utils/ppocr_keys_v1.txt --use_gpu=True
+```
+**注意**
+- 使用SVTR算法时，需要指定--rec_algorithm='SVTR'
+- 如果使用自定义字典训练的模型，需要将--rec_char_dict_path=ppocr/utils/ppocr_keys_v1.txt修改为自定义的字典
+- --rec_image_shape='3, 32, 320' 该参数不能去掉
--- a/configs/kie/kie_unet_sdmgr.yml
+++ b/configs/kie/kie_unet_sdmgr.yml
@@ -17,7 +17,7 @@ Global:
  checkpoints:
  save_inference_dir:
  use_visualdl: False
-  class_path: ./train_data/wildreceipt/class_list.txt
+  class_path: &class_path ./train_data/wildreceipt/class_list.txt
  infer_img: ./train_data/wildreceipt/1.txt
  save_res_path: ./output/sdmgr_kie/predicts_kie.txt
  img_scale: [ 1024, 512 ]
@@ -72,6 +72,7 @@ Train:
          order: 'hwc'
      - KieLabelEncode: # Class handling label
          character_dict_path: ./train_data/wildreceipt/dict.txt
+          class_path: *class_path
      - KieResize:
      - ToCHWImage:
      - KeepKeys:
@@ -88,7 +89,6 @@ Eval:
    data_dir: ./train_data/wildreceipt
    label_file_list:
      - ./train_data/wildreceipt/wildreceipt_test.txt
-      # - /paddle/data/PaddleOCR/train_data/wildreceipt/1.txt
    transforms:
      - DecodeImage: # load image
          img_mode: RGB

--- a/configs/rec/rec_mtb_nrtr.yml
+++ b/configs/rec/rec_mtb_nrtr.yml
@@ -9,7 +9,7 @@ Global:
  eval_batch_step: [0, 2000]
  cal_metric_during_train: True
  pretrained_model:
-  checkpoints: 
+  checkpoints:
  save_inference_dir:
  use_visualdl: False
  infer_img: doc/imgs_words_en/word_10.png
@@ -49,7 +49,7 @@ Architecture:
 Loss:
-  name: NRTRLoss
+  name: CELoss
  smoothing: True
 PostProcess:
@@ -68,8 +68,8 @@ Train:
          img_mode: BGR
          channel_first: False
      - NRTRLabelEncode: # Class handling label
-      - NRTRRecResizeImg:
+      - GrayRecResizeImg:
-          image_shape: [100, 32]
+          image_shape: [100, 32] # W H
          resize_type: PIL # PIL or OpenCV
      - KeepKeys:
          keep_keys: ['image', 'label', 'length'] # dataloader will return list in this order
@@ -82,14 +82,14 @@ Train:
 Eval:
  dataset:
    name: LMDBDataSet
-    data_dir: ./train_data/data_lmdb_release/evaluation/
+    data_dir: ./train_data/data_lmdb_release/validation/
    transforms:
      - DecodeImage: # load image
          img_mode: BGR
          channel_first: False
      - NRTRLabelEncode: # Class handling label
-      - NRTRRecResizeImg:
+      - GrayRecResizeImg:
-          image_shape: [100, 32]
+          image_shape: [100, 32] # W H
          resize_type: PIL # PIL or OpenCV
      - KeepKeys:
          keep_keys: ['image', 'label', 'length'] # dataloader will return list in this order
@@ -97,5 +97,5 @@ Eval:
    shuffle: False
    drop_last: False
    batch_size_per_card: 256
-    num_workers: 1
+    num_workers: 4
    use_shared_memory: False
--- a/configs/rec/rec_r45_abinet.yml
+++ b/configs/rec/rec_r45_abinet.yml
+Global:
+  use_gpu: True
+  epoch_num: 10
+  log_smooth_window: 20
+  print_batch_step: 10
+  save_model_dir: ./output/rec/r45_abinet/
+  save_epoch_step: 1
+  # evaluation is run every 2000 iterations
+  eval_batch_step: [0, 2000]
+  cal_metric_during_train: True
+  pretrained_model:
+  checkpoints:
+  save_inference_dir:
+  use_visualdl: False
+  infer_img: doc/imgs_words_en/word_10.png
+  # for data or label process
+  character_dict_path:
+  character_type: en
+  max_text_length: 25
+  infer_mode: False
+  use_space_char: False
+  save_res_path: ./output/rec/predicts_abinet.txt
+Optimizer:
+  name: Adam
+  beta1: 0.9
+  beta2: 0.99
+  clip_norm: 20.0
+  lr:
+    name: Piecewise
+    decay_epochs: [6]
+    values: [0.0001, 0.00001] 
+  regularizer:
+    name: 'L2'
+    factor: 0.
+Architecture:
+  model_type: rec
+  algorithm: ABINet
+  in_channels: 3
+  Transform:
+  Backbone:
+    name: ResNet45
+  Head:
+    name: ABINetHead
+    use_lang: True
+    iter_size: 3
+Loss:
+  name: CELoss
+  ignore_index: &ignore_index 100 # Must be greater than the number of character classes
+PostProcess:
+  name: ABINetLabelDecode
+Metric:
+  name: RecMetric
+  main_indicator: acc
+Train:
+  dataset:
+    name: LMDBDataSet
+    data_dir: ./train_data/data_lmdb_release/training/
+    transforms:
+      - DecodeImage: # load image
+          img_mode: RGB
+          channel_first: False
+      - ABINetRecAug:
+      - ABINetLabelEncode: # Class handling label
+          ignore_index: *ignore_index
+      - ABINetRecResizeImg:
+          image_shape: [3, 32, 128]
+      - KeepKeys:
+          keep_keys: ['image', 'label', 'length'] # dataloader will return list in this order
+  loader:
+    shuffle: True
+    batch_size_per_card: 96
+    drop_last: True
+    num_workers: 4
+Eval:
+  dataset:
+    name: LMDBDataSet
+    data_dir: ./train_data/data_lmdb_release/validation/
+    transforms:
+      - DecodeImage: # load image
+          img_mode: RGB
+          channel_first: False
+      - ABINetLabelEncode: # Class handling label
+          ignore_index: *ignore_index
+      - ABINetRecResizeImg:
+          image_shape: [3, 32, 128]
+      - KeepKeys:
+          keep_keys: ['image', 'label', 'length'] # dataloader will return list in this order
+  loader:
+    shuffle: False
+    drop_last: False
+    batch_size_per_card: 256
+    num_workers: 4
+    use_shared_memory: False
--- a/configs/rec/rec_svtrnet.yml
+++ b/configs/rec/rec_svtrnet.yml
@@ -26,7 +26,7 @@ Optimizer:
  name: AdamW
  beta1: 0.9
  beta2: 0.99
-  epsilon: 0.00000008
+  epsilon: 8.e-8
  weight_decay: 0.05
  no_weight_decay_name: norm pos_embed
  one_dim_param_no_weight_decay: true
@@ -77,14 +77,13 @@ Metric:
 Train:
  dataset:
    name: LMDBDataSet
-    data_dir: ./train_data/data_lmdb_release/training/
+    data_dir: ./train_data/data_lmdb_release/training
    transforms:
      - DecodeImage: # load image
          img_mode: BGR
          channel_first: False
      - CTCLabelEncode: # Class handling label
-      - RecResizeImg:
+      - SVTRRecResizeImg:
-          character_dict_path:
          image_shape: [3, 64, 256]
          padding: False
      - KeepKeys:
@@ -98,14 +97,13 @@ Train:
 Eval:
  dataset:
    name: LMDBDataSet
-    data_dir: ./train_data/data_lmdb_release/validation/
+    data_dir: ./train_data/data_lmdb_release/validation
    transforms:
      - DecodeImage: # load image
          img_mode: BGR
          channel_first: False
      - CTCLabelEncode: # Class handling label
-      - RecResizeImg:
+      - SVTRRecResizeImg:
-          character_dict_path:
          image_shape: [3, 64, 256]
          padding: False
      - KeepKeys:

--- a/configs/rec/rec_svtrnet_ch.yml
+++ b/configs/rec/rec_svtrnet_ch.yml
+Global:
+  use_gpu: true
+  epoch_num: 100
+  log_smooth_window: 20
+  print_batch_step: 10
+  save_model_dir: ./output/rec/svtr_ch_all/
+  save_epoch_step: 10
+  eval_batch_step:
+  - 0
+  - 2000
+  cal_metric_during_train: true
+  pretrained_model: null
+  checkpoints: null
+  save_inference_dir: null
+  use_visualdl: false
+  infer_img: doc/imgs_words/ch/word_1.jpg
+  character_dict_path: ppocr/utils/ppocr_keys_v1.txt
+  max_text_length: 25
+  infer_mode: false
+  use_space_char: true
+  save_res_path: ./output/rec/predicts_svtr_tiny_ch_all.txt
+Optimizer:
+  name: AdamW
+  beta1: 0.9
+  beta2: 0.99
+  epsilon: 8.0e-08
+  weight_decay: 0.05
+  no_weight_decay_name: norm pos_embed
+  one_dim_param_no_weight_decay: true
+  lr:
+    name: Cosine
+    learning_rate: 0.0005
+    warmup_epoch: 2
+Architecture:
+  model_type: rec
+  algorithm: SVTR
+  Transform: null
+  Backbone:
+    name: SVTRNet
+    img_size:
+    - 32
+    - 320
+    out_char_num: 40
+    out_channels: 96
+    patch_merging: Conv
+    embed_dim:
+    - 64
+    - 128
+    - 256
+    depth:
+    - 3
+    - 6
+    - 3
+    num_heads:
+    - 2
+    - 4
+    - 8
+    mixer:
+    - Local
+    - Local
+    - Local
+    - Local
+    - Local
+    - Local
+    - Global
+    - Global
+    - Global
+    - Global
+    - Global
+    - Global
+    local_mixer:
+    - - 7
+      - 11
+    - - 7
+      - 11
+    - - 7
+      - 11
+    last_stage: true
+    prenorm: false
+  Neck:
+    name: SequenceEncoder
+    encoder_type: reshape
+  Head:
+    name: CTCHead
+Loss:
+  name: CTCLoss
+PostProcess:
+  name: CTCLabelDecode
+Metric:
+  name: RecMetric
+  main_indicator: acc
+Train:
+  dataset:
+    name: SimpleDataSet
+    data_dir: ./train_data
+    label_file_list:
+    - ./train_data/train_list.txt
+    ext_op_transform_idx: 1
+    transforms:
+    - DecodeImage:
+        img_mode: BGR
+        channel_first: false
+    - RecConAug:
+        prob: 0.5
+        ext_data_num: 2
+        image_shape:
+        - 32
+        - 320
+        - 3
+    - RecAug: null
+    - CTCLabelEncode: null
+    - SVTRRecResizeImg:
+        image_shape:
+        - 3
+        - 32
+        - 320
+        padding: true
+    - KeepKeys:
+        keep_keys:
+        - image
+        - label
+        - length
+  loader:
+    shuffle: true
+    batch_size_per_card: 256
+    drop_last: true
+    num_workers: 8
+Eval:
+  dataset:
+    name: SimpleDataSet
+    data_dir: ./train_data
+    label_file_list:
+    - ./train_data/val_list.txt
+    transforms:
+    - DecodeImage:
+        img_mode: BGR
+        channel_first: false
+    - CTCLabelEncode: null
+    - SVTRRecResizeImg:
+        image_shape:
+        - 3
+        - 32
+        - 320
+        padding: true
+    - KeepKeys:
+        keep_keys:
+        - image
+        - label
+        - length
+  loader:
+    shuffle: false
+    drop_last: false
+    batch_size_per_card: 256
+    num_workers: 2
+profiler_options: null
--- a/configs/rec/rec_vitstr_none_ce.yml
+++ b/configs/rec/rec_vitstr_none_ce.yml
+Global:
+  use_gpu: True
+  epoch_num: 20
+  log_smooth_window: 20
+  print_batch_step: 10
+  save_model_dir: ./output/rec/vitstr_none_ce/
+  save_epoch_step: 1
+  # evaluation is run every 2000 iterations after the 0th iteration#
+  eval_batch_step: [0, 2000]
+  cal_metric_during_train: True
+  pretrained_model:
+  checkpoints:
+  save_inference_dir:
+  use_visualdl: False
+  infer_img: doc/imgs_words_en/word_10.png
+  # for data or label process
+  character_dict_path: ppocr/utils/EN_symbol_dict.txt
+  max_text_length: 25
+  infer_mode: False
+  use_space_char: False
+  save_res_path: ./output/rec/predicts_vitstr.txt
+Optimizer:
+  name: Adadelta
+  epsilon: 1.e-8
+  rho: 0.95
+  clip_norm: 5.0
+  lr:
+    learning_rate: 1.0
+Architecture:
+  model_type: rec
+  algorithm: ViTSTR
+  in_channels: 1
+  Transform:
+  Backbone:
+    name: ViTSTR
+    scale: tiny
+  Neck:
+    name: SequenceEncoder
+    encoder_type: reshape
+  Head:
+    name: CTCHead
+Loss:
+  name: CELoss
+  with_all: True
+  ignore_index: &ignore_index 0 # Must be zero or greater than the number of character classes
+PostProcess:
+  name: ViTSTRLabelDecode
+Metric:
+  name: RecMetric
+  main_indicator: acc
+Train:
+  dataset:
+    name: LMDBDataSet
+    data_dir: ./train_data/data_lmdb_release/training/
+    transforms:
+      - DecodeImage: # load image
+          img_mode: BGR
+          channel_first: False
+      - ViTSTRLabelEncode: # Class handling label
+          ignore_index: *ignore_index
+      - GrayRecResizeImg:
+          image_shape: [224, 224] # W H
+          resize_type: PIL # PIL or OpenCV
+          inter_type: 'Image.BICUBIC'
+          scale: false
+      - KeepKeys:
+          keep_keys: ['image', 'label', 'length'] # dataloader will return list in this order
+  loader:
+    shuffle: True
+    batch_size_per_card: 48
+    drop_last: True
+    num_workers: 8
+Eval:
+  dataset:
+    name: LMDBDataSet
+    data_dir: ./train_data/data_lmdb_release/validation/
+    transforms:
+      - DecodeImage: # load image
+          img_mode: BGR
+          channel_first: False
+      - ViTSTRLabelEncode: # Class handling label
+          ignore_index: *ignore_index
+      - GrayRecResizeImg:
+          image_shape: [224, 224] # W H
+          resize_type: PIL # PIL or OpenCV
+          inter_type: 'Image.BICUBIC'
+          scale: false
+      - KeepKeys:
+          keep_keys: ['image', 'label', 'length'] # dataloader will return list in this order
+  loader:
+    shuffle: False
+    drop_last: False
+    batch_size_per_card: 256
+    num_workers: 2
--- a/configs/vqa/ser/layoutlm.yml
+++ b/configs/vqa/ser/layoutlm.yml
@@ -43,7 +43,7 @@ Optimizer:
 PostProcess:
  name: VQASerTokenLayoutLMPostProcess
-  class_path: &class_path ppstructure/vqa/labels/labels_ser.txt
+  class_path: &class_path train_data/XFUND/class_list_xfun.txt
 Metric:
  name: VQASerTokenMetric
@@ -54,7 +54,7 @@ Train:
    name: SimpleDataSet
    data_dir: train_data/XFUND/zh_train/image
    label_file_list: 
-      - train_data/XFUND/zh_train/xfun_normalize_train.json
+      - train_data/XFUND/zh_train/train.json
    transforms:
      - DecodeImage: # load image
          img_mode: RGB
@@ -89,7 +89,7 @@ Eval:
    name: SimpleDataSet
    data_dir: train_data/XFUND/zh_val/image
    label_file_list:
-      - train_data/XFUND/zh_val/xfun_normalize_val.json
+      - train_data/XFUND/zh_val/val.json
    transforms:
      - DecodeImage: # load image
          img_mode: RGB

--- a/configs/vqa/ser/layoutlmv2.yml
+++ b/configs/vqa/ser/layoutlmv2.yml
@@ -44,7 +44,7 @@ Optimizer:
 PostProcess:
  name: VQASerTokenLayoutLMPostProcess
-  class_path: &class_path ppstructure/vqa/labels/labels_ser.txt
+  class_path: &class_path train_data/XFUND/class_list_xfun.txt
 Metric:
  name: VQASerTokenMetric
@@ -55,7 +55,7 @@ Train:
    name: SimpleDataSet
    data_dir: train_data/XFUND/zh_train/image
    label_file_list: 
-      - train_data/XFUND/zh_train/xfun_normalize_train.json
+      - train_data/XFUND/zh_train/train.json
    transforms:
      - DecodeImage: # load image
          img_mode: RGB
@@ -90,7 +90,7 @@ Eval:
    name: SimpleDataSet
    data_dir: train_data/XFUND/zh_val/image
    label_file_list:
-      - train_data/XFUND/zh_val/xfun_normalize_val.json
+      - train_data/XFUND/zh_val/val.json
    transforms:
      - DecodeImage: # load image
          img_mode: RGB

--- a/configs/vqa/ser/layoutxlm.yml
+++ b/configs/vqa/ser/layoutxlm.yml
@@ -11,7 +11,7 @@ Global:
  save_inference_dir:
  use_visualdl: False
  seed: 2022
-  infer_img: doc/vqa/input/zh_val_42.jpg
+  infer_img: ppstructure/docs/vqa/input/zh_val_42.jpg
  save_res_path: ./output/ser
 Architecture:
@@ -54,7 +54,7 @@ Train:
    name: SimpleDataSet
    data_dir: train_data/XFUND/zh_train/image
    label_file_list: 
-      - train_data/XFUND/zh_train/xfun_normalize_train.json
+      - train_data/XFUND/zh_train/train.json
    ratio_list: [ 1.0 ]
    transforms:
      - DecodeImage: # load image
@@ -90,7 +90,7 @@ Eval:
    name: SimpleDataSet
    data_dir: train_data/XFUND/zh_val/image
    label_file_list:
-      - train_data/XFUND/zh_val/xfun_normalize_val.json
+      - train_data/XFUND/zh_val/val.json
    transforms:
      - DecodeImage: # load image
          img_mode: RGB

--- a/deploy/lite/config.txt
+++ b/deploy/lite/config.txt
@@ -4,4 +4,5 @@ det_db_box_thresh  0.5
 det_db_unclip_ratio  1.6
 det_db_use_dilate 0
 det_use_polygon_score 1
 use_direction_classify  1
\ No newline at end of file
+rec_image_height  32
\ No newline at end of file
--- a/deploy/lite/crnn_process.cc
+++ b/deploy/lite/crnn_process.cc
@@ -19,25 +19,27 @@
 const std::vector<int> rec_image_shape{3, 32, 320};
-cv::Mat CrnnResizeImg(cv::Mat img, float wh_ratio) {
+cv::Mat CrnnResizeImg(cv::Mat img, float wh_ratio, int rec_image_height) {
  int imgC, imgH, imgW;
  imgC = rec_image_shape[0];
+  imgH = rec_image_height;
  imgW = rec_image_shape[2];
-  imgH = rec_image_shape[1];
-  imgW = int(32 * wh_ratio);
+  imgW = int(imgH * wh_ratio);
-  float ratio = static_cast<float>(img.cols) / static_cast<float>(img.rows);
+  float ratio = float(img.cols) / float(img.rows);
  int resize_w, resize_h;
  if (ceilf(imgH * ratio) > imgW)
    resize_w = imgW;
  else
-    resize_w = static_cast<int>(ceilf(imgH * ratio));
+    resize_w = int(ceilf(imgH * ratio));
-  cv::Mat resize_img;
  cv::resize(img, resize_img, cv::Size(resize_w, imgH), 0.f, 0.f,
             cv::INTER_LINEAR);
+  cv::copyMakeBorder(resize_img, resize_img, 0, 0, 0,
-  return resize_img;
+                     int(imgW - resize_img.cols), cv::BORDER_CONSTANT,
+                     {127, 127, 127});
 }
 std::vector<std::string> ReadDict(std::string path) {

--- a/deploy/lite/crnn_process.h
+++ b/deploy/lite/crnn_process.h
@@ -26,7 +26,7 @@
 #include "opencv2/imgcodecs.hpp"
 #include "opencv2/imgproc.hpp"
-cv::Mat CrnnResizeImg(cv::Mat img, float wh_ratio);
+cv::Mat CrnnResizeImg(cv::Mat img, float wh_ratio, int rec_image_height);
 std::vector<std::string> ReadDict(std::string path);

--- a/deploy/lite/ocr_db_crnn.cc
+++ b/deploy/lite/ocr_db_crnn.cc
@@ -162,7 +162,8 @@ void RunRecModel(std::vector<std::vector<std::vector<int>>> boxes, cv::Mat img,
                 std::vector<std::string> charactor_dict,
                 std::shared_ptr<PaddlePredictor> predictor_cls,
                 int use_direction_classify,
-                 std::vector<double> *times) {
+                 std::vector<double> *times,
+                 int rec_image_height) {
  std::vector<float> mean = {0.5f, 0.5f, 0.5f};
  std::vector<float> scale = {1 / 0.5f, 1 / 0.5f, 1 / 0.5f};
@@ -183,7 +184,7 @@ void RunRecModel(std::vector<std::vector<std::vector<int>>> boxes, cv::Mat img,
    float wh_ratio =
        static_cast<float>(crop_img.cols) / static_cast<float>(crop_img.rows);
-    resize_img = CrnnResizeImg(crop_img, wh_ratio);
+    resize_img = CrnnResizeImg(crop_img, wh_ratio, rec_image_height);
    resize_img.convertTo(resize_img, CV_32FC3, 1 / 255.f);
    const float *dimg = reinterpret_cast<const float *>(resize_img.data);
@@ -444,7 +445,7 @@ void system(char **argv){
  //// load config from txt file
  auto Config = LoadConfigTxt(det_config_path);
  int use_direction_classify = int(Config["use_direction_classify"]);
+  int rec_image_height = int(Config["rec_image_height"]);
  auto charactor_dict = ReadDict(dict_path);
  charactor_dict.insert(charactor_dict.begin(), "#"); // blank char for ctc
  charactor_dict.push_back(" ");
@@ -590,12 +591,16 @@ void rec(int argc, char **argv) {
  std::string batchsize = argv[6];
  std::string img_dir = argv[7];
  std::string dict_path = argv[8];
+  std::string config_path = argv[9];
  if (strcmp(argv[4], "FP32") != 0 && strcmp(argv[4], "INT8") != 0) {
      std::cerr << "Only support FP32 or INT8." << std::endl;
      exit(1);
  }
+  auto Config = LoadConfigTxt(config_path);
+  int rec_image_height = int(Config["rec_image_height"]);
  std::vector<cv::String> cv_all_img_names;
  cv::glob(img_dir, cv_all_img_names);
@@ -630,7 +635,7 @@ void rec(int argc, char **argv) {
    std::vector<float> rec_text_score;
    std::vector<double> times;
    RunRecModel(boxes, srcimg, rec_predictor, rec_text, rec_text_score,
-                charactor_dict, cls_predictor, 0, &times);
+                charactor_dict, cls_predictor, 0, &times, rec_image_height);
    //// print recognized text
    for (int i = 0; i < rec_text.size(); i++) {

--- a/deploy/lite/readme.md
+++ b/deploy/lite/readme.md
@@ -34,7 +34,7 @@ For the compilation process of different development environments, please refer
 ### 1.2 Prepare Paddle-Lite library
 There are two ways to obtain the Paddle-Lite library：
- 1. Download directly, the download link of the Paddle-Lite library is as follows：
+- 1. [Recommended] Download directly, the download link of the Paddle-Lite library is as follows：
      | Platform | Paddle-Lite library download link |
      |---|---|
@@ -43,7 +43,9 @@ There are two ways to obtain the Paddle-Lite library：
      Note: 1. The above Paddle-Lite library is compiled from the Paddle-Lite 2.10 branch. For more information about Paddle-Lite 2.10, please refer to [link](https://github.com/PaddlePaddle/Paddle-Lite/releases/tag/v2.10).
- 2. [Recommended] Compile Paddle-Lite to get the prediction library. The compilation method of Paddle-Lite is as follows：
+    **Note: It is recommended to use paddlelite>=2.10 version of the prediction library, other prediction library versions [download link](https://github.com/PaddlePaddle/Paddle-Lite/tags)**
+- 2. Compile Paddle-Lite to get the prediction library. The compilation method of Paddle-Lite is as follows：
 ```
 git clone https://github.com/PaddlePaddle/Paddle-Lite.git
 cd Paddle-Lite
@@ -104,21 +106,17 @@ If you directly use the model in the above table for deployment, you can skip th
 If the model to be deployed is not in the above table, you need to follow the steps below to obtain the optimized model.
-The `opt` tool can be obtained by compiling Paddle Lite.
+- Step 1: Refer to [document](https://www.paddlepaddle.org.cn/lite/v2.10/user_guides/opt/opt_python.html) to install paddlelite, which is used to convert paddle inference model to paddlelite required for running nb model
 ```
-git clone https://github.com/PaddlePaddle/Paddle-Lite.git
+pip install paddlelite==2.10 # The paddlelite version should be the same as the prediction library version
-cd Paddle-Lite
-git checkout release/v2.10
-./lite/tools/build.sh build_optimize_tool
 ```
+After installation, the following commands can view the help information
-After the compilation is complete, the opt file is located under build.opt/lite/api/, You can view the operating options and usage of opt in the following ways:
 ```
-cd build.opt/lite/api/
+paddle_lite_opt
-./opt
 ```
+Introduction to paddle_lite_opt parameters:
 |Options|Description|
 |---|---|
 |--model_dir|The path of the PaddlePaddle model to be optimized (non-combined form)|
@@ -131,6 +129,8 @@ cd build.opt/lite/api/
 `--model_dir` is suitable for the non-combined mode of the model to be optimized, and the inference model of PaddleOCR is the combined mode, that is, the model structure and model parameters are stored in a single file.
+- Step 2: Use paddle_lite_opt to convert the inference model to the mobile model format.
 The following takes the ultra-lightweight Chinese model of PaddleOCR as an example to introduce the use of the compiled opt file to complete the conversion of the inference model to the Paddle-Lite optimized model
 ```
@@ -240,6 +240,7 @@ det_db_thresh  0.3        # Used to filter the binarized image of DB prediction,
 det_db_box_thresh  0.5    # DDB post-processing filter box threshold, if there is a missing box detected, it can be reduced as appropriate
 det_db_unclip_ratio  1.6  # Indicates the compactness of the text box, the smaller the value, the closer the text box to the text
 use_direction_classify  0  # Whether to use the direction classifier, 0 means not to use, 1 means to use
+rec_image_height  32      # The height of the input image of the recognition model, the PP-OCRv3 model needs to be set to 48, and the PP-OCRv2 model needs to be set to 32
 ```
 5. Run Model on phone
@@ -258,8 +259,15 @@ After the above steps are completed, you can use adb to push the file to the pho
 cd /data/local/tmp/debug
 export LD_LIBRARY_PATH=${PWD}:$LD_LIBRARY_PATH
 # The use of ocr_db_crnn is:
- # ./ocr_db_crnn Detection model file Orientation classifier model file Recognition model file Test image path Dictionary file path
+ # ./ocr_db_crnn Mode Detection model file Orientation classifier model file Recognition model file  Hardware  Precision  Threads Batchsize  Test image path Dictionary file path
- ./ocr_db_crnn ch_PP-OCRv2_det_slim_opt.nb  ch_PP-OCRv2_rec_slim_opt.nb  ch_ppocr_mobile_v2.0_cls_opt.nb  ./11.jpg  ppocr_keys_v1.txt
+ ./ocr_db_crnn system ch_PP-OCRv2_det_slim_opt.nb  ch_PP-OCRv2_rec_slim_opt.nb  ch_ppocr_mobile_v2.0_cls_slim_opt.nb  arm8 INT8 10 1  ./11.jpg  config.txt  ppocr_keys_v1.txt  True
+# precision can be INT8 for quantitative model or FP32 for normal model.
+# Only using detection model
+./ocr_db_crnn  det ch_PP-OCRv2_det_slim_opt.nb arm8 INT8 10 1 ./11.jpg  config.txt
+# Only using recognition model
+./ocr_db_crnn  rec ch_PP-OCRv2_rec_slim_opt.nb arm8 INT8 10 1 word_1.jpg ppocr_keys_v1.txt config.txt
 ```
 If you modify the code, you need to recompile and push to the phone.
@@ -283,3 +291,7 @@ A2: Replace the .jpg test image under ./debug with the image you want to test, a
 Q3: How to package it into the mobile APP?
 A3: This demo aims to provide the core algorithm part that can run OCR on mobile phones. Further, PaddleOCR/deploy/android_demo is an example of encapsulating this demo into a mobile app for reference.
+Q4: When running the demo, an error is reported `Error: This model is not supported, because kernel for 'io_copy' is not supported by Paddle-Lite.`
+A4: The problem is that the installed paddlelite version does not match the downloaded prediction library version. Make sure that the paddleliteopt tool matches your prediction library version, and try to switch to the nb model again.
--- a/deploy/lite/readme_ch.md
+++ b/deploy/lite/readme_ch.md
@@ -8,7 +8,7 @@
    - [2.1 模型优化](#21-模型优化)
    - [2.2 与手机联调](#22-与手机联调)
 - [FAQ](#faq)
 本教程将介绍基于[Paddle Lite](https://github.com/PaddlePaddle/Paddle-Lite) 在移动端部署PaddleOCR超轻量中文检测、识别模型的详细步骤。
@@ -32,7 +32,7 @@ Paddle Lite是飞桨轻量化推理引擎，为手机、IOT端提供高效推理
 ### 1.2 准备预测库
 预测库有两种获取方式：
- 1. 直接下载，预测库下载链接如下：
+- 1. [推荐]直接下载，预测库下载链接如下：
      | 平台 | 预测库下载链接 |
      |---|---|
@@ -41,7 +41,9 @@ Paddle Lite是飞桨轻量化推理引擎，为手机、IOT端提供高效推理
      注：1. 上述预测库为PaddleLite 2.10分支编译得到，有关PaddleLite 2.10 详细信息可参考 [链接](https://github.com/PaddlePaddle/Paddle-Lite/releases/tag/v2.10) 。
- 2. [推荐]编译Paddle-Lite得到预测库，Paddle-Lite的编译方式如下：
+**注：建议使用paddlelite>=2.10版本的预测库，其他预测库版本[下载链接](https://github.com/PaddlePaddle/Paddle-Lite/tags)**
+- 2. 编译Paddle-Lite得到预测库，Paddle-Lite的编译方式如下：
 ```
 git clone https://github.com/PaddlePaddle/Paddle-Lite.git
 cd Paddle-Lite
@@ -102,22 +104,16 @@ Paddle-Lite 提供了多种策略来自动优化原始的模型，其中包括
 如果要部署的模型不在上述表格中，则需要按照如下步骤获得优化后的模型。
-模型优化需要Paddle-Lite的opt可执行文件，可以通过编译Paddle-Lite源码获得，编译步骤如下：
+- 步骤1：参考[文档](https://www.paddlepaddle.org.cn/lite/v2.10/user_guides/opt/opt_python.html)安装paddlelite，用于转换paddle inference model为paddlelite运行所需的nb模型
 ```
-# 如果准备环境时已经clone了Paddle-Lite，则不用重新clone Paddle-Lite
+pip install paddlelite==2.10  # paddlelite版本要与预测库版本一致
-git clone https://github.com/PaddlePaddle/Paddle-Lite.git
-cd Paddle-Lite
-git checkout release/v2.10
-# 启动编译
-./lite/tools/build.sh build_optimize_tool
 ```
+安装完后，如下指令可以查看帮助信息
-编译完成后，opt文件位于`build.opt/lite/api/`下，可通过如下方式查看opt的运行选项和使用方式；
 ```
-cd build.opt/lite/api/
+paddle_lite_opt
-./opt
 ```
+paddle_lite_opt 参数介绍：
 |选项|说明|
 |---|---|
 |--model_dir|待优化的PaddlePaddle模型（非combined形式）的路径|
@@ -130,6 +126,8 @@ cd build.opt/lite/api/
 `--model_dir`适用于待优化的模型是非combined方式，PaddleOCR的inference模型是combined方式，即模型结构和模型参数使用单独一个文件存储。
+- 步骤2：使用paddle_lite_opt将inference模型转换成移动端模型格式。
 下面以PaddleOCR的超轻量中文模型为例，介绍使用编译好的opt文件完成inference模型到Paddle-Lite优化模型的转换。
 ```
@@ -148,7 +146,7 @@ wget  https://paddleocr.bj.bcebos.com/dygraph_v2.0/slim/ch_ppocr_mobile_v2.0_cls
 转换成功后，inference模型目录下会多出`.nb`结尾的文件，即是转换成功的模型文件。
-注意：使用paddle-lite部署时，需要使用opt工具优化后的模型。 opt 工具的输入模型是paddle保存的inference模型
+注意：使用paddle-lite部署时，需要使用opt工具优化后的模型。 opt工具的输入模型是paddle保存的inference模型
 <a name="2.2与手机联调"></a>
 ### 2.2 与手机联调
@@ -234,13 +232,14 @@ ppocr_keys_v1.txt   # 中文字典
 ...
 ```
-2.  `config.txt` 包含了检测器、分类器的超参数，如下：
+2.  `config.txt` 包含了检测器、分类器、识别器的超参数，如下：
 ```
 max_side_len  960         # 输入图像长宽大于960时，等比例缩放图像，使得图像最长边为960
 det_db_thresh  0.3        # 用于过滤DB预测的二值化图像，设置为0.-0.3对结果影响不明显
-det_db_box_thresh  0.5    # DB后处理过滤box的阈值，如果检测存在漏框情况，可酌情减小
+det_db_box_thresh  0.5    # 检测器后处理过滤box的阈值，如果检测存在漏框情况，可酌情减小
 det_db_unclip_ratio  1.6  # 表示文本框的紧致程度，越小则文本框更靠近文本
 use_direction_classify  0  # 是否使用方向分类器，0表示不使用，1表示使用
+rec_image_height  32      # 识别模型输入图像的高度，PP-OCRv3模型设置为48，PP-OCRv2模型需要设置为32
 ```
 5. 启动调试
@@ -259,8 +258,14 @@ use_direction_classify  0  # 是否使用方向分类器，0表示不使用，1
 cd /data/local/tmp/debug
 export LD_LIBRARY_PATH=${PWD}:$LD_LIBRARY_PATH
 # 开始使用，ocr_db_crnn可执行文件的使用方式为:
- # ./ocr_db_crnn  检测模型文件 方向分类器模型文件  识别模型文件  测试图像路径  字典文件路径
+ # ./ocr_db_crnn 预测模式  检测模型文件 方向分类器模型文件  识别模型文件 运行硬件 运行精度 线程数  batchsize  测试图像路径  参数配置路径  字典文件路径 是否使用benchmark参数
- ./ocr_db_crnn ch_PP-OCRv2_det_slim_opt.nb  ch_PP-OCRv2_rec_slim_opt.nb  ch_ppocr_mobile_v2.0_cls_slim_opt.nb  ./11.jpg  ppocr_keys_v1.txt
+ ./ocr_db_crnn system  ch_PP-OCRv2_det_slim_opt.nb  ch_PP-OCRv2_rec_slim_opt.nb  ch_ppocr_mobile_v2.0_cls_slim_opt.nb  arm8 INT8 10 1  ./11.jpg  config.txt  ppocr_keys_v1.txt  True
+# 仅使用文本检测模型，使用方式如下：
+./ocr_db_crnn  det ch_PP-OCRv2_det_slim_opt.nb arm8 INT8 10 1 ./11.jpg  config.txt
+# 仅使用文本识别模型，使用方式如下：
+./ocr_db_crnn  rec ch_PP-OCRv2_rec_slim_opt.nb arm8 INT8 10 1 word_1.jpg ppocr_keys_v1.txt config.txt
 ```
 如果对代码做了修改，则需要重新编译并push到手机上。
@@ -284,3 +289,7 @@ A2：替换debug下的.jpg测试图像为你想要测试的图像，adb push 到
 Q3：如何封装到手机APP中？
 A3：此demo旨在提供能在手机上运行OCR的核心算法部分，PaddleOCR/deploy/android_demo是将这个demo封装到手机app的示例，供参考
+Q4：运行demo时遇到报错`Error: This model is not supported, because kernel for 'io_copy' is not supported by Paddle-Lite.`
+A4：问题是安装的paddlelite版本和下载的预测库版本不匹配，确保paddleliteopt工具和你的预测库版本匹配，重新转nb模型试试。
--- a/doc/doc_ch/algorithm_overview.md
+++ b/doc/doc_ch/algorithm_overview.md
@@ -66,6 +66,8 @@
 - [x]  [SAR](./algorithm_rec_sar.md)
 - [x]  [SEED](./algorithm_rec_seed.md)
 - [x]  [SVTR](./algorithm_rec_svtr.md)
+- [x]  [ViTSTR](./algorithm_rec_vitstr.md)
+- [x]  [ABINet](./algorithm_rec_abinet.md)
 参考[DTRB](https://arxiv.org/abs/1904.01906)[3]文字识别训练和评估流程，使用MJSynth和SynthText两个文字识别数据集训练，在IIIT, SVT, IC03, IC13, IC15, SVTP, CUTE数据集上进行评估，算法效果如下：
@@ -84,6 +86,8 @@
 |SAR|Resnet31| 87.20% | rec_r31_sar | [训练模型](https://paddleocr.bj.bcebos.com/dygraph_v2.1/rec/rec_r31_sar_train.tar) |
 |SEED|Aster_Resnet| 85.35% | rec_resnet_stn_bilstm_att | [训练模型](https://paddleocr.bj.bcebos.com/dygraph_v2.1/rec/rec_resnet_stn_bilstm_att.tar) |
 |SVTR|SVTR-Tiny| 89.25% | rec_svtr_tiny_none_ctc_en | [训练模型](https://paddleocr.bj.bcebos.com/PP-OCRv3/chinese/rec_svtr_tiny_none_ctc_en_train.tar) |
+|ViTSTR|ViTSTR| 79.82% | rec_vitstr_none_ce | [训练模型](https://paddleocr.bj.bcebos.com/rec_vitstr_none_ce_train.tar) |
+|ABINet|Resnet45| 90.75% | rec_r45_abinet | [训练模型](https://paddleocr.bj.bcebos.com/rec_r45_abinet_train.tar) |
 <a name="2"></a>

--- a/doc/doc_ch/algorithm_rec_abinet.md
+++ b/doc/doc_ch/algorithm_rec_abinet.md
+# 场景文本识别算法-ABINet
+- [1. 算法简介](#1)
+- [2. 环境配置](#2)
+- [3. 模型训练、评估、预测](#3)
+    - [3.1 训练](#3-1)
+    - [3.2 评估](#3-2)
+    - [3.3 预测](#3-3)
+- [4. 推理部署](#4)
+    - [4.1 Python推理](#4-1)
+    - [4.2 C++推理](#4-2)
+    - [4.3 Serving服务化部署](#4-3)
+    - [4.4 更多推理部署](#4-4)
+- [5. FAQ](#5)
+<a name="1"></a>
+## 1. 算法简介
+论文信息：
+> [ABINet: Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition](https://openaccess.thecvf.com/content/CVPR2021/papers/Fang_Read_Like_Humans_Autonomous_Bidirectional_and_Iterative_Language_Modeling_for_CVPR_2021_paper.pdf)
+> Shancheng Fang and Hongtao Xie and Yuxin Wang and Zhendong Mao and Yongdong Zhang
+> CVPR, 2021
+<a name="model"></a>
+`ABINet`使用MJSynth和SynthText两个文字识别数据集训练，在IIIT, SVT, IC03, IC13, IC15, SVTP, CUTE数据集上进行评估，算法复现效果如下：
+|模型|骨干网络|配置文件|Acc|下载链接|
+| --- | --- | --- | --- | --- |
+|ABINet|ResNet45|[rec_r45_abinet.yml](../../configs/rec/rec_r45_abinet.yml)|90.75%|[预训练、训练模型](https://paddleocr.bj.bcebos.com/rec_r45_abinet_train.tar)|
+<a name="2"></a>
+## 2. 环境配置
+请先参考[《运行环境准备》](./environment.md)配置PaddleOCR运行环境，参考[《项目克隆》](./clone.md)克隆项目代码。
+<a name="3"></a>
+## 3. 模型训练、评估、预测
+<a name="3-1"></a>
+### 3.1 模型训练
+请参考[文本识别训练教程](./recognition.md)。PaddleOCR对代码进行了模块化，训练`ABINet`识别模型时需要**更换配置文件**为`ABINet`的[配置文件](../../configs/rec/rec_r45_abinet.yml)。
+#### 启动训练
+具体地，在完成数据准备后，便可以启动训练，训练命令如下：
+```shell
+#单卡训练（训练周期长，不建议）
+python3 tools/train.py -c configs/rec/rec_r45_abinet.yml
+#多卡训练，通过--gpus参数指定卡号
+python3 -m paddle.distributed.launch --gpus '0,1,2,3'  tools/train.py -c configs/rec/rec_r45_abinet.yml
+```
+<a name="3-2"></a>
+### 3.2 评估
+可下载已训练完成的[模型文件](#model)，使用如下命令进行评估：
+```shell
+# 注意将pretrained_model的路径设置为本地路径。
+python3 -m paddle.distributed.launch --gpus '0' tools/eval.py -c configs/rec/rec_r45_abinet.yml -o Global.pretrained_model=./rec_r45_abinet_train/best_accuracy
+```
+<a name="3-3"></a>
+### 3.3 预测
+使用如下命令进行单张图片预测：
+```shell
+# 注意将pretrained_model的路径设置为本地路径。
+python3 tools/infer_rec.py -c configs/rec/rec_r45_abinet.yml -o Global.infer_img='./doc/imgs_words_en/word_10.png' Global.pretrained_model=./rec_r45_abinet_train/best_accuracy
+# 预测文件夹下所有图像时，可修改infer_img为文件夹，如 Global.infer_img='./doc/imgs_words_en/'。
+```
+<a name="4"></a>
+## 4. 推理部署
+<a name="4-1"></a>
+### 4.1 Python推理
+首先将训练得到best模型，转换成inference model。这里以训练完成的模型为例（[模型下载地址](https://paddleocr.bj.bcebos.com/rec_r45_abinet_train.tar) )，可以使用如下命令进行转换：
+```shell
+# 注意将pretrained_model的路径设置为本地路径。
+python3 tools/export_model.py -c configs/rec/rec_r45_abinet.yml -o Global.pretrained_model=./rec_r45_abinet_train/best_accuracy Global.save_inference_dir=./inference/rec_r45_abinet/
+```
+**注意：**
+- 如果您是在自己的数据集上训练的模型，并且调整了字典文件，请注意修改配置文件中的`character_dict_path`是否是所需要的字典文件。
+- 如果您修改了训练时的输入大小，请修改`tools/export_model.py`文件中的对应ABINet的`infer_shape`。
+转换成功后，在目录下有三个文件：
+```
+/inference/rec_r45_abinet/
+    ├── inference.pdiparams         # 识别inference模型的参数文件
+    ├── inference.pdiparams.info    # 识别inference模型的参数信息，可忽略
+    └── inference.pdmodel           # 识别inference模型的program文件
+```
+执行如下命令进行模型推理：
+```shell
+python3 tools/infer/predict_rec.py --image_dir='./doc/imgs_words_en/word_10.png' --rec_model_dir='./inference/rec_r45_abinet/' --rec_algorithm='ABINet' --rec_image_shape='3,32,128' --rec_char_dict_path='./ppocr/utils/ic15_dict.txt'
+# 预测文件夹下所有图像时，可修改image_dir为文件夹，如 --image_dir='./doc/imgs_words_en/'。
+```
+![](../imgs_words_en/word_10.png)
+执行命令后，上面图像的预测结果（识别的文本和得分）会打印到屏幕上，示例如下：
+结果如下：
+```shell
+Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9999995231628418)
+```
+**注意**：
+- 训练上述模型采用的图像分辨率是[3，32，128]，需要通过参数`rec_image_shape`设置为您训练时的识别图像形状。
+- 在推理时需要设置参数`rec_char_dict_path`指定字典，如果您修改了字典，请修改该参数为您的字典文件。
+- 如果您修改了预处理方法，需修改`tools/infer/predict_rec.py`中ABINet的预处理为您的预处理方法。
+<a name="4-2"></a>
+### 4.2 C++推理部署
+由于C++预处理后处理还未支持ABINet，所以暂未支持
+<a name="4-3"></a>
+### 4.3 Serving服务化部署
+暂不支持
+<a name="4-4"></a>
+### 4.4 更多推理部署
+暂不支持
+<a name="5"></a>
+## 5. FAQ
+1. MJSynth和SynthText两种数据集来自于[ABINet源repo](https://github.com/FangShancheng/ABINet) 。
+2. 我们使用ABINet作者提供的预训练模型进行finetune训练。
+## 引用
+```bibtex
+@article{Fang2021ABINet,
+  title     = {ABINet: Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition},
+  author    = {Shancheng Fang and Hongtao Xie and Yuxin Wang and Zhendong Mao and Yongdong Zhang},
+  booktitle = {CVPR},
+  year      = {2021},
+  url       = {https://arxiv.org/abs/2103.06495},
+  pages     = {7098-7107}
+}
+```
--- a/doc/doc_ch/algorithm_rec_nrtr.md
+++ b/doc/doc_ch/algorithm_rec_nrtr.md
@@ -12,6 +12,7 @@
    - [4.3 Serving服务化部署](#4-3)
    - [4.4 更多推理部署](#4-4)
 - [5. FAQ](#5)
+- [6. 发行公告](#6)
 <a name="1"></a>
 ## 1. 算法简介
@@ -110,7 +111,7 @@ python3 tools/infer/predict_rec.py --image_dir='./doc/imgs_words_en/word_10.png'
 执行命令后，上面图像的预测结果（识别的文本和得分）会打印到屏幕上，示例如下：
 结果如下：
 ```shell
-Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9265879392623901)
+Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9465042352676392)
 ```
 **注意**：
@@ -140,12 +141,147 @@ Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9265879392623901)
 1. `NRTR`论文中使用Beam搜索进行解码字符，但是速度较慢，这里默认未使用Beam搜索，以贪婪搜索进行解码字符。
+<a name="6"></a>
+## 6. 发行公告
+1. release/2.6更新NRTR代码结构，新版NRTR可加载旧版(release/2.5及之前)模型参数，使用下面示例代码将旧版模型参数转换为新版模型参数：
+```python
+    params = paddle.load('path/' + '.pdparams') # 旧版本参数
+    state_dict = model.state_dict() # 新版模型参数
+    new_state_dict = {}
+    for k1, v1 in state_dict.items():
+        k = k1
+        if 'encoder' in k and 'self_attn' in k and 'qkv' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            q = params[k_para.replace('qkv', 'conv1')].transpose((1, 0, 2, 3))
+            k = params[k_para.replace('qkv', 'conv2')].transpose((1, 0, 2, 3))
+            v = params[k_para.replace('qkv', 'conv3')].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = np.concatenate([q[:, :, 0, 0], k[:, :, 0, 0], v[:, :, 0, 0]], -1)
+        elif 'encoder' in k and 'self_attn' in k and 'qkv' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            q = params[k_para.replace('qkv', 'conv1')]
+            k = params[k_para.replace('qkv', 'conv2')]
+            v = params[k_para.replace('qkv', 'conv3')]
+            new_state_dict[k1] = np.concatenate([q, k, v], -1)
+        elif 'encoder' in k and 'self_attn' in k and 'out_proj' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para]
+        elif 'encoder' in k and 'norm3' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para.replace('norm3', 'norm2')]
+        elif 'encoder' in k and 'norm1' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para]
+        elif 'decoder' in k and 'self_attn' in k and 'qkv' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            q = params[k_para.replace('qkv', 'conv1')].transpose((1, 0, 2, 3))
+            k = params[k_para.replace('qkv', 'conv2')].transpose((1, 0, 2, 3))
+            v = params[k_para.replace('qkv', 'conv3')].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = np.concatenate([q[:, :, 0, 0], k[:, :, 0, 0], v[:, :, 0, 0]], -1)
+        elif 'decoder' in k and 'self_attn' in k and 'qkv' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            q = params[k_para.replace('qkv', 'conv1')]
+            k = params[k_para.replace('qkv', 'conv2')]
+            v = params[k_para.replace('qkv', 'conv3')]
+            new_state_dict[k1] = np.concatenate([q, k, v], -1)
+        elif 'decoder' in k and 'self_attn' in k and 'out_proj' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para]
+        elif 'decoder' in k and 'cross_attn' in k and 'q' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            q = params[k_para.replace('q', 'conv1')].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = q[:, :, 0, 0]
+        elif 'decoder' in k and 'cross_attn' in k and 'q' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            q = params[k_para.replace('q', 'conv1')]
+            new_state_dict[k1] = q
+        elif 'decoder' in k and 'cross_attn' in k and 'kv' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            k = params[k_para.replace('kv', 'conv2')].transpose((1, 0, 2, 3))
+            v = params[k_para.replace('kv', 'conv3')].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = np.concatenate([k[:, :, 0, 0], v[:, :, 0, 0]], -1)
+        elif 'decoder' in k and 'cross_attn' in k and 'kv' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            k = params[k_para.replace('kv', 'conv2')]
+            v = params[k_para.replace('kv', 'conv3')]
+            new_state_dict[k1] = np.concatenate([k, v], -1)
+        elif 'decoder' in k and 'cross_attn' in k and 'out_proj' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            new_state_dict[k1] = params[k_para]
+        elif 'decoder' in k and 'norm' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para]
+        elif 'mlp' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('fc', 'conv')
+            k_para = k_para.replace('mlp.', '')
+            w = params[k_para].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = w[:, :, 0, 0]
+        elif 'mlp' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('fc', 'conv')
+            k_para = k_para.replace('mlp.', '')
+            w = params[k_para]
+            new_state_dict[k1] = w
+        else:
+            new_state_dict[k1] = params[k1]
+        if list(new_state_dict[k1].shape) != list(v1.shape):
+            print(k1)
+    for k, v1 in state_dict.items():
+        if k not in new_state_dict.keys():
+            print(1, k)
+        elif list(new_state_dict[k].shape) != list(v1.shape):
+            print(2, k)
+    model.set_state_dict(new_state_dict)
+    paddle.save(model.state_dict(), 'nrtrnew_from_old_params.pdparams')
+```
+2. 新版相比与旧版，代码结构简洁，推理速度有所提高。
 ## 引用
 ```bibtex
 @article{Sheng2019NRTR,
  title     = {NRTR: A No-Recurrence Sequence-to-Sequence Model For Scene Text Recognition},
-  author    = {Fenfen Sheng and Zhineng Chen andBo Xu},
+  author    = {Fenfen Sheng and Zhineng Chen and Bo Xu},
  booktitle = {ICDAR},
  year      = {2019},
  url       = {http://arxiv.org/abs/1806.00926},

--- a/doc/doc_ch/algorithm_rec_svtr.md
+++ b/doc/doc_ch/algorithm_rec_svtr.md
@@ -111,7 +111,6 @@ python3 tools/export_model.py -c ./rec_svtr_tiny_none_ctc_en_train/rec_svtr_tiny
 **注意：**
 - 如果您是在自己的数据集上训练的模型，并且调整了字典文件，请注意修改配置文件中的`character_dict_path`是否为所正确的字典文件。
- 如果您修改了训练时的输入大小，请修改`tools/export_model.py`文件中的对应SVTR的`infer_shape`。
 转换成功后，在目录下有三个文件：
 ```

--- a/doc/doc_ch/algorithm_rec_vitstr.md
+++ b/doc/doc_ch/algorithm_rec_vitstr.md
+# 场景文本识别算法-ViTSTR
+- [1. 算法简介](#1)
+- [2. 环境配置](#2)
+- [3. 模型训练、评估、预测](#3)
+    - [3.1 训练](#3-1)
+    - [3.2 评估](#3-2)
+    - [3.3 预测](#3-3)
+- [4. 推理部署](#4)
+    - [4.1 Python推理](#4-1)
+    - [4.2 C++推理](#4-2)
+    - [4.3 Serving服务化部署](#4-3)
+    - [4.4 更多推理部署](#4-4)
+- [5. FAQ](#5)
+<a name="1"></a>
+## 1. 算法简介
+论文信息：
+> [Vision Transformer for Fast and Efficient Scene Text Recognition](https://arxiv.org/abs/2105.08582)
+> Rowel Atienza
+> ICDAR, 2021
+<a name="model"></a>
+`ViTSTR`使用MJSynth和SynthText两个文字识别数据集训练，在IIIT, SVT, IC03, IC13, IC15, SVTP, CUTE数据集上进行评估，算法复现效果如下：
+|模型|骨干网络|配置文件|Acc|下载链接|
+| --- | --- | --- | --- | --- |
+|ViTSTR|ViTSTR|[rec_vitstr_none_ce.yml](../../configs/rec/rec_vitstr_none_ce.yml)|79.82%|[训练模型](https://paddleocr.bj.bcebos.com/rec_vitstr_none_ce_train.tar)|
+<a name="2"></a>
+## 2. 环境配置
+请先参考[《运行环境准备》](./environment.md)配置PaddleOCR运行环境，参考[《项目克隆》](./clone.md)克隆项目代码。
+<a name="3"></a>
+## 3. 模型训练、评估、预测
+<a name="3-1"></a>
+### 3.1 模型训练
+请参考[文本识别训练教程](./recognition.md)。PaddleOCR对代码进行了模块化，训练`ViTSTR`识别模型时需要**更换配置文件**为`ViTSTR`的[配置文件](../../configs/rec/rec_vitstr_none_ce.yml)。
+#### 启动训练
+具体地，在完成数据准备后，便可以启动训练，训练命令如下：
+```shell
+#单卡训练（训练周期长，不建议）
+python3 tools/train.py -c configs/rec/rec_vitstr_none_ce.yml
+#多卡训练，通过--gpus参数指定卡号
+python3 -m paddle.distributed.launch --gpus '0,1,2,3'  tools/train.py -c configs/rec/rec_vitstr_none_ce.yml
+```
+<a name="3-2"></a>
+### 3.2 评估
+可下载已训练完成的[模型文件](#model)，使用如下命令进行评估：
+```shell
+# 注意将pretrained_model的路径设置为本地路径。
+python3 -m paddle.distributed.launch --gpus '0' tools/eval.py -c configs/rec/rec_vitstr_none_ce.yml -o Global.pretrained_model=./rec_vitstr_none_ce_train/best_accuracy
+```
+<a name="3-3"></a>
+### 3.3 预测
+使用如下命令进行单张图片预测：
+```shell
+# 注意将pretrained_model的路径设置为本地路径。
+python3 tools/infer_rec.py -c configs/rec/rec_vitstr_none_ce.yml -o Global.infer_img='./doc/imgs_words_en/word_10.png' Global.pretrained_model=./rec_vitstr_none_ce_train/best_accuracy
+# 预测文件夹下所有图像时，可修改infer_img为文件夹，如 Global.infer_img='./doc/imgs_words_en/'。
+```
+<a name="4"></a>
+## 4. 推理部署
+<a name="4-1"></a>
+### 4.1 Python推理
+首先将训练得到best模型，转换成inference model。这里以训练完成的模型为例（[模型下载地址](https://paddleocr.bj.bcebos.com/rec_vitstr_none_ce_train.tar) )，可以使用如下命令进行转换：
+```shell
+# 注意将pretrained_model的路径设置为本地路径。
+python3 tools/export_model.py -c configs/rec/rec_vitstr_none_ce.yml -o Global.pretrained_model=./rec_vitstr_none_ce_train/best_accuracy Global.save_inference_dir=./inference/rec_vitstr/
+```
+**注意：**
+- 如果您是在自己的数据集上训练的模型，并且调整了字典文件，请注意修改配置文件中的`character_dict_path`是否是所需要的字典文件。
+- 如果您修改了训练时的输入大小，请修改`tools/export_model.py`文件中的对应ViTSTR的`infer_shape`。
+转换成功后，在目录下有三个文件：
+```
+/inference/rec_vitstr/
+    ├── inference.pdiparams         # 识别inference模型的参数文件
+    ├── inference.pdiparams.info    # 识别inference模型的参数信息，可忽略
+    └── inference.pdmodel           # 识别inference模型的program文件
+```
+执行如下命令进行模型推理：
+```shell
+python3 tools/infer/predict_rec.py --image_dir='./doc/imgs_words_en/word_10.png' --rec_model_dir='./inference/rec_vitstr/' --rec_algorithm='ViTSTR' --rec_image_shape='1,224,224' --rec_char_dict_path='./ppocr/utils/EN_symbol_dict.txt'
+# 预测文件夹下所有图像时，可修改image_dir为文件夹，如 --image_dir='./doc/imgs_words_en/'。
+```
+![](../imgs_words_en/word_10.png)
+执行命令后，上面图像的预测结果（识别的文本和得分）会打印到屏幕上，示例如下：
+结果如下：
+```shell
+Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9998350143432617)
+```
+**注意**：
+- 训练上述模型采用的图像分辨率是[1，224，224]，需要通过参数`rec_image_shape`设置为您训练时的识别图像形状。
+- 在推理时需要设置参数`rec_char_dict_path`指定字典，如果您修改了字典，请修改该参数为您的字典文件。
+- 如果您修改了预处理方法，需修改`tools/infer/predict_rec.py`中ViTSTR的预处理为您的预处理方法。
+<a name="4-2"></a>
+### 4.2 C++推理部署
+由于C++预处理后处理还未支持ViTSTR，所以暂未支持
+<a name="4-3"></a>
+### 4.3 Serving服务化部署
+暂不支持
+<a name="4-4"></a>
+### 4.4 更多推理部署
+暂不支持
+<a name="5"></a>
+## 5. FAQ
+1. 在`ViTSTR`论文中，使用在ImageNet1k上的预训练权重进行初始化训练，我们在训练未采用预训练权重，最终精度没有变化甚至有所提高。
+2. 我们仅仅复现了`ViTSTR`中的tiny版本，如果需要使用small、base版本，可将[ViTSTR源repo](https://github.com/roatienza/deep-text-recognition-benchmark) 中的预训练权重转为Paddle权重使用。
+## 引用
+```bibtex
+@article{Atienza2021ViTSTR,
+  title     = {Vision Transformer for Fast and Efficient Scene Text Recognition},
+  author    = {Rowel Atienza},
+  booktitle = {ICDAR},
+  year      = {2021},
+  url       = {https://arxiv.org/abs/2105.08582}
+}
+```
--- a/doc/doc_ch/application.md
+++ b/doc/doc_ch/application.md
+# 场景应用
+PaddleOCR场景应用覆盖通用，制造、金融、交通行业的主要OCR垂类应用，在PP-OCR、PP-Structure的通用能力基础之上，以notebook的形式展示利用场景数据微调、模型优化方法、数据增广等内容，为开发者快速落地OCR应用提供示范与启发。
+> 如需下载全部垂类模型，可以扫描下方二维码，关注公众号填写问卷后，加入PaddleOCR官方交流群获取20G OCR学习大礼包（内含《动手学OCR》电子书、课程回放视频、前沿论文等重磅资料）
+<div align="center">
+<img src="https://ai-studio-static-online.cdn.bcebos.com/dd721099bd50478f9d5fb13d8dd00fad69c22d6848244fd3a1d3980d7fefc63e"  width = "150" height = "150" />
+</div>
+> 如果您是企业开发者且未在下述场景中找到合适的方案，可以填写[OCR应用合作调研问卷](https://paddle.wjx.cn/vj/QwF7GKw.aspx)，免费与官方团队展开不同层次的合作，包括但不限于问题抽象、确定技术方案、项目答疑、共同研发等。如果您已经使用PaddleOCR落地项目，也可以填写此问卷，与飞桨平台共同宣传推广，提升企业技术品宣。期待您的提交！
+## 通用
+| 类别                   | 亮点     | 类别       | 亮点         |
+| ---------------------- | -------- | ---------- | ------------ |
+| 高精度中文识别模型SVTR | 新增模型 | 手写体识别 | 新增字形支持 |
+## 制造
+| 类别           | 亮点                           | 类别           | 亮点                 |
+| -------------- | ------------------------------ | -------------- | -------------------- |
+| 数码管识别     | 数码管数据合成、漏识别调优     | 电表识别       | 大分辨率图像检测调优 |
+| 液晶屏读数识别 | 检测模型蒸馏、Serving部署      | PCB文字识别    | 小尺寸文本检测与识别 |
+| 包装生产日期   | 点阵字符合成、过曝过暗文字识别 | 液晶屏缺陷检测 | 非文字形态识别       |
+## 金融
+| 类别           | 亮点                     | 类别         | 亮点                  |
+| -------------- | ------------------------ | ------------ | --------------------- |
+| 表单VQA        | 多模态通用表单结构化提取 | 通用卡证识别 | 通用结构化提取        |
+| 增值税发票     | 尽请期待                 | 身份证识别   | 结构化提取、图像阴影  |
+| 印章检测与识别 | 端到端弯曲文本识别       | 合同比对     | 密集文本检测、NLP串联 |
+## 交通
+| 类别              | 亮点                           | 类别       | 亮点     |
+| ----------------- | ------------------------------ | ---------- | -------- |
+| 车牌识别          | 多角度图像、轻量模型、端侧部署 | 快递单识别 | 尽请期待 |
+| 驾驶证/行驶证识别 | 尽请期待                       |            |          |
\ No newline at end of file
--- a/doc/doc_en/algorithm_overview_en.md
+++ b/doc/doc_en/algorithm_overview_en.md
@@ -65,6 +65,8 @@ Supported text recognition algorithms (Click the link to get the tutorial):
 - [x]  [SAR](./algorithm_rec_sar_en.md)
 - [x]  [SEED](./algorithm_rec_seed_en.md)
 - [x]  [SVTR](./algorithm_rec_svtr_en.md)
+- [x]  [ViTSTR](./algorithm_rec_vitstr_en.md)
+- [x]  [ABINet](./algorithm_rec_abinet_en.md)
 Refer to [DTRB](https://arxiv.org/abs/1904.01906), the training and evaluation result of these above text recognition (using MJSynth and SynthText for training, evaluate on IIIT, SVT, IC03, IC13, IC15, SVTP, CUTE) is as follow:
@@ -83,6 +85,8 @@ Refer to [DTRB](https://arxiv.org/abs/1904.01906), the training and evaluation r
 |SAR|Resnet31| 87.20% | rec_r31_sar | [trained model](https://paddleocr.bj.bcebos.com/dygraph_v2.1/rec/rec_r31_sar_train.tar) |
 |SEED|Aster_Resnet| 85.35% | rec_resnet_stn_bilstm_att | [trained model](https://paddleocr.bj.bcebos.com/dygraph_v2.1/rec/rec_resnet_stn_bilstm_att.tar) |
 |SVTR|SVTR-Tiny| 89.25% | rec_svtr_tiny_none_ctc_en | [trained model](https://paddleocr.bj.bcebos.com/PP-OCRv3/chinese/rec_svtr_tiny_none_ctc_en_train.tar) |
+|ViTSTR|ViTSTR| 79.82% | rec_vitstr_none_ce | [trained model](https://paddleocr.bj.bcebos.com/rec_vitstr_none_none_train.tar) |
+|ABINet|Resnet45| 90.75% | rec_r45_abinet | [trained model](https://paddleocr.bj.bcebos.com/rec_r45_abinet_train.tar) |
 <a name="2"></a>

--- a/doc/doc_en/algorithm_rec_abinet_en.md
+++ b/doc/doc_en/algorithm_rec_abinet_en.md
+# ABINet
+- [1. Introduction](#1)
+- [2. Environment](#2)
+- [3. Model Training / Evaluation / Prediction](#3)
+    - [3.1 Training](#3-1)
+    - [3.2 Evaluation](#3-2)
+    - [3.3 Prediction](#3-3)
+- [4. Inference and Deployment](#4)
+    - [4.1 Python Inference](#4-1)
+    - [4.2 C++ Inference](#4-2)
+    - [4.3 Serving](#4-3)
+    - [4.4 More](#4-4)
+- [5. FAQ](#5)
+<a name="1"></a>
+## 1. Introduction
+Paper:
+> [ABINet: Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition](https://openaccess.thecvf.com/content/CVPR2021/papers/Fang_Read_Like_Humans_Autonomous_Bidirectional_and_Iterative_Language_Modeling_for_CVPR_2021_paper.pdf)
+> Shancheng Fang and Hongtao Xie and Yuxin Wang and Zhendong Mao and Yongdong Zhang
+> CVPR, 2021
+Using MJSynth and SynthText two text recognition datasets for training, and evaluating on IIIT, SVT, IC03, IC13, IC15, SVTP, CUTE datasets, the algorithm reproduction effect is as follows:
+|Model|Backbone|config|Acc|Download link|
+| --- | --- | --- | --- | --- |
+|ABINet|ResNet45|[rec_r45_abinet.yml](../../configs/rec/rec_r45_abinet.yml)|90.75%|[pretrained & trained model](https://paddleocr.bj.bcebos.com/rec_r45_abinet_train.tar)|
+<a name="2"></a>
+## 2. Environment
+Please refer to ["Environment Preparation"](./environment_en.md) to configure the PaddleOCR environment, and refer to ["Project Clone"](./clone_en.md) to clone the project code.
+<a name="3"></a>
+## 3. Model Training / Evaluation / Prediction
+Please refer to [Text Recognition Tutorial](./recognition_en.md). PaddleOCR modularizes the code, and training different recognition models only requires **changing the configuration file**.
+Training:
+Specifically, after the data preparation is completed, the training can be started. The training command is as follows:
+```
+#Single GPU training (long training period, not recommended)
+python3 tools/train.py -c configs/rec/rec_r45_abinet.yml
+#Multi GPU training, specify the gpu number through the --gpus parameter
+python3 -m paddle.distributed.launch --gpus '0,1,2,3'  tools/train.py -c configs/rec/rec_r45_abinet.yml
+```
+Evaluation:
+```
+# GPU evaluation
+python3 -m paddle.distributed.launch --gpus '0' tools/eval.py -c configs/rec/rec_r45_abinet.yml -o Global.pretrained_model={path/to/weights}/best_accuracy
+```
+Prediction:
+```
+# The configuration file used for prediction must match the training
+python3 tools/infer_rec.py -c configs/rec/rec_r45_abinet.yml -o Global.infer_img='./doc/imgs_words_en/word_10.png' Global.pretrained_model=./rec_r45_abinet_train/best_accuracy
+```
+<a name="4"></a>
+## 4. Inference and Deployment
+<a name="4-1"></a>
+### 4.1 Python Inference
+First, the model saved during the ABINet text recognition training process is converted into an inference model. ( [Model download link](https://paddleocr.bj.bcebos.com/rec_r45_abinet_train.tar)) ), you can use the following command to convert:
+```
+python3 tools/export_model.py -c configs/rec/rec_r45_abinet.yml -o Global.pretrained_model=./rec_r45_abinet_train/best_accuracy  Global.save_inference_dir=./inference/rec_r45_abinet
+```
+**Note:**
+- If you are training the model on your own dataset and have modified the dictionary file, please pay attention to modify the `character_dict_path` in the configuration file to the modified dictionary file.
+- If you modified the input size during training, please modify the `infer_shape` corresponding to ABINet in the `tools/export_model.py` file.
+After the conversion is successful, there are three files in the directory:
+```
+/inference/rec_r45_abinet/
+    ├── inference.pdiparams
+    ├── inference.pdiparams.info
+    └── inference.pdmodel
+```
+For ABINet text recognition model inference, the following commands can be executed:
+```
+python3 tools/infer/predict_rec.py --image_dir='./doc/imgs_words_en/word_10.png' --rec_model_dir='./inference/rec_r45_abinet/' --rec_algorithm='ABINet' --rec_image_shape='3,32,128' --rec_char_dict_path='./ppocr/utils/ic15_dict.txt'
+```
+![](../imgs_words_en/word_10.png)
+After executing the command, the prediction result (recognized text and score) of the image above is printed to the screen, an example is as follows:
+The result is as follows:
+```shell
+Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9999995231628418)
+```
+<a name="4-2"></a>
+### 4.2 C++ Inference
+Not supported
+<a name="4-3"></a>
+### 4.3 Serving
+Not supported
+<a name="4-4"></a>
+### 4.4 More
+Not supported
+<a name="5"></a>
+## 5. FAQ
+1. Note that the MJSynth and SynthText datasets come from [ABINet repo](https://github.com/FangShancheng/ABINet).
+2. We use the pre-trained model provided by the ABINet authors for finetune training.
+## Citation
+```bibtex
+@article{Fang2021ABINet,
+  title     = {ABINet: Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition},
+  author    = {Shancheng Fang and Hongtao Xie and Yuxin Wang and Zhendong Mao and Yongdong Zhang},
+  booktitle = {CVPR},
+  year      = {2021},
+  url       = {https://arxiv.org/abs/2103.06495},
+  pages     = {7098-7107}
+}
+```
--- a/doc/doc_en/algorithm_rec_nrtr_en.md
+++ b/doc/doc_en/algorithm_rec_nrtr_en.md
@@ -12,6 +12,7 @@
    - [4.3 Serving](#4-3)
    - [4.4 More](#4-4)
 - [5. FAQ](#5)
+- [6. Release Note](#6)
 <a name="1"></a>
 ## 1. Introduction
@@ -25,7 +26,7 @@ Using MJSynth and SynthText two text recognition datasets for training, and eval
 |Model|Backbone|config|Acc|Download link|
 | --- | --- | --- | --- | --- |
-|NRTR|MTB|[rec_mtb_nrtr.yml](../../configs/rec/rec_mtb_nrtr.yml)|84.21%|[train model](https://paddleocr.bj.bcebos.com/dygraph_v2.0/en/rec_mtb_nrtr_train.tar)|
+|NRTR|MTB|[rec_mtb_nrtr.yml](../../configs/rec/rec_mtb_nrtr.yml)|84.21%|[trained model](https://paddleocr.bj.bcebos.com/dygraph_v2.0/en/rec_mtb_nrtr_train.tar)|
 <a name="2"></a>
 ## 2. Environment
@@ -98,7 +99,7 @@ python3 tools/infer/predict_rec.py --image_dir='./doc/imgs_words_en/word_10.png'
 After executing the command, the prediction result (recognized text and score) of the image above is printed to the screen, an example is as follows:
 The result is as follows:
 ```shell
-Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9265879392623901)
+Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9465042352676392)
 ```
 <a name="4-2"></a>
@@ -121,12 +122,146 @@ Not supported
 1. In the `NRTR` paper, Beam search is used to decode characters, but the speed is slow. Beam search is not used by default here, and greedy search is used to decode characters.
+<a name="6"></a>
+## 6. Release Note
+1. The release/2.6 version updates the NRTR code structure. The new version of NRTR can load the model parameters of the old version (release/2.5 and before), and you may use the following code to convert the old version model parameters to the new version model parameters:
+```python
+    params = paddle.load('path/' + '.pdparams') # the old version parameters
+    state_dict = model.state_dict() # the new version model parameters
+    new_state_dict = {}
+    for k1, v1 in state_dict.items():
+        k = k1
+        if 'encoder' in k and 'self_attn' in k and 'qkv' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            q = params[k_para.replace('qkv', 'conv1')].transpose((1, 0, 2, 3))
+            k = params[k_para.replace('qkv', 'conv2')].transpose((1, 0, 2, 3))
+            v = params[k_para.replace('qkv', 'conv3')].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = np.concatenate([q[:, :, 0, 0], k[:, :, 0, 0], v[:, :, 0, 0]], -1)
+        elif 'encoder' in k and 'self_attn' in k and 'qkv' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            q = params[k_para.replace('qkv', 'conv1')]
+            k = params[k_para.replace('qkv', 'conv2')]
+            v = params[k_para.replace('qkv', 'conv3')]
+            new_state_dict[k1] = np.concatenate([q, k, v], -1)
+        elif 'encoder' in k and 'self_attn' in k and 'out_proj' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para]
+        elif 'encoder' in k and 'norm3' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para.replace('norm3', 'norm2')]
+        elif 'encoder' in k and 'norm1' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para]
+        elif 'decoder' in k and 'self_attn' in k and 'qkv' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            q = params[k_para.replace('qkv', 'conv1')].transpose((1, 0, 2, 3))
+            k = params[k_para.replace('qkv', 'conv2')].transpose((1, 0, 2, 3))
+            v = params[k_para.replace('qkv', 'conv3')].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = np.concatenate([q[:, :, 0, 0], k[:, :, 0, 0], v[:, :, 0, 0]], -1)
+        elif 'decoder' in k and 'self_attn' in k and 'qkv' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            q = params[k_para.replace('qkv', 'conv1')]
+            k = params[k_para.replace('qkv', 'conv2')]
+            v = params[k_para.replace('qkv', 'conv3')]
+            new_state_dict[k1] = np.concatenate([q, k, v], -1)
+        elif 'decoder' in k and 'self_attn' in k and 'out_proj' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para]
+        elif 'decoder' in k and 'cross_attn' in k and 'q' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            q = params[k_para.replace('q', 'conv1')].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = q[:, :, 0, 0]
+        elif 'decoder' in k and 'cross_attn' in k and 'q' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            q = params[k_para.replace('q', 'conv1')]
+            new_state_dict[k1] = q
+        elif 'decoder' in k and 'cross_attn' in k and 'kv' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            k = params[k_para.replace('kv', 'conv2')].transpose((1, 0, 2, 3))
+            v = params[k_para.replace('kv', 'conv3')].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = np.concatenate([k[:, :, 0, 0], v[:, :, 0, 0]], -1)
+        elif 'decoder' in k and 'cross_attn' in k and 'kv' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            k = params[k_para.replace('kv', 'conv2')]
+            v = params[k_para.replace('kv', 'conv3')]
+            new_state_dict[k1] = np.concatenate([k, v], -1)
+        elif 'decoder' in k and 'cross_attn' in k and 'out_proj' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('cross_attn', 'multihead_attn')
+            new_state_dict[k1] = params[k_para]
+        elif 'decoder' in k and 'norm' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            new_state_dict[k1] = params[k_para]
+        elif 'mlp' in k and 'weight' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('fc', 'conv')
+            k_para = k_para.replace('mlp.', '')
+            w = params[k_para].transpose((1, 0, 2, 3))
+            new_state_dict[k1] = w[:, :, 0, 0]
+        elif 'mlp' in k and 'bias' in k:
+            k_para = k[:13] + 'layers.' + k[13:]
+            k_para = k_para.replace('fc', 'conv')
+            k_para = k_para.replace('mlp.', '')
+            w = params[k_para]
+            new_state_dict[k1] = w
+        else:
+            new_state_dict[k1] = params[k1]
+        if list(new_state_dict[k1].shape) != list(v1.shape):
+            print(k1)
+    for k, v1 in state_dict.items():
+        if k not in new_state_dict.keys():
+            print(1, k)
+        elif list(new_state_dict[k].shape) != list(v1.shape):
+            print(2, k)
+    model.set_state_dict(new_state_dict)
+    paddle.save(model.state_dict(), 'nrtrnew_from_old_params.pdparams')
+```
+2. The new version has a clean code structure and improved inference speed compared with the old version.
 ## Citation
 ```bibtex
 @article{Sheng2019NRTR,
  title     = {NRTR: A No-Recurrence Sequence-to-Sequence Model For Scene Text Recognition},
-  author    = {Fenfen Sheng and Zhineng Chen andBo Xu},
+  author    = {Fenfen Sheng and Zhineng Chen and Bo Xu},
  booktitle = {ICDAR},
  year      = {2019},
  url       = {http://arxiv.org/abs/1806.00926},

--- a/doc/doc_en/algorithm_rec_svtr_en.md
+++ b/doc/doc_en/algorithm_rec_svtr_en.md
@@ -88,7 +88,6 @@ python3 tools/export_model.py -c configs/rec/rec_svtrnet.yml -o Global.pretraine
 **Note:**
 - If you are training the model on your own dataset and have modified the dictionary file, please pay attention to modify the `character_dict_path` in the configuration file to the modified dictionary file.
- If you modified the input size during training, please modify the `infer_shape` corresponding to SVTR in the `tools/export_model.py` file.
 After the conversion is successful, there are three files in the directory:
 ```

--- a/doc/doc_en/algorithm_rec_vitstr_en.md
+++ b/doc/doc_en/algorithm_rec_vitstr_en.md
+# ViTSTR
+- [1. Introduction](#1)
+- [2. Environment](#2)
+- [3. Model Training / Evaluation / Prediction](#3)
+    - [3.1 Training](#3-1)
+    - [3.2 Evaluation](#3-2)
+    - [3.3 Prediction](#3-3)
+- [4. Inference and Deployment](#4)
+    - [4.1 Python Inference](#4-1)
+    - [4.2 C++ Inference](#4-2)
+    - [4.3 Serving](#4-3)
+    - [4.4 More](#4-4)
+- [5. FAQ](#5)
+<a name="1"></a>
+## 1. Introduction
+Paper:
+> [Vision Transformer for Fast and Efficient Scene Text Recognition](https://arxiv.org/abs/2105.08582)
+> Rowel Atienza
+> ICDAR, 2021
+Using MJSynth and SynthText two text recognition datasets for training, and evaluating on IIIT, SVT, IC03, IC13, IC15, SVTP, CUTE datasets, the algorithm reproduction effect is as follows:
+|Model|Backbone|config|Acc|Download link|
+| --- | --- | --- | --- | --- |
+|ViTSTR|ViTSTR|[rec_vitstr_none_ce.yml](../../configs/rec/rec_vitstr_none_ce.yml)|79.82%|[trained model](https://paddleocr.bj.bcebos.com/rec_vitstr_none_none_train.tar)|
+<a name="2"></a>
+## 2. Environment
+Please refer to ["Environment Preparation"](./environment_en.md) to configure the PaddleOCR environment, and refer to ["Project Clone"](./clone_en.md) to clone the project code.
+<a name="3"></a>
+## 3. Model Training / Evaluation / Prediction
+Please refer to [Text Recognition Tutorial](./recognition_en.md). PaddleOCR modularizes the code, and training different recognition models only requires **changing the configuration file**.
+Training:
+Specifically, after the data preparation is completed, the training can be started. The training command is as follows:
+```
+#Single GPU training (long training period, not recommended)
+python3 tools/train.py -c configs/rec/rec_vitstr_none_ce.yml
+#Multi GPU training, specify the gpu number through the --gpus parameter
+python3 -m paddle.distributed.launch --gpus '0,1,2,3'  tools/train.py -c configs/rec/rec_vitstr_none_ce.yml
+```
+Evaluation:
+```
+# GPU evaluation
+python3 -m paddle.distributed.launch --gpus '0' tools/eval.py -c configs/rec/rec_vitstr_none_ce.yml -o Global.pretrained_model={path/to/weights}/best_accuracy
+```
+Prediction:
+```
+# The configuration file used for prediction must match the training
+python3 tools/infer_rec.py -c configs/rec/rec_vitstr_none_ce.yml -o Global.infer_img='./doc/imgs_words_en/word_10.png' Global.pretrained_model=./rec_vitstr_none_ce_train/best_accuracy
+```
+<a name="4"></a>
+## 4. Inference and Deployment
+<a name="4-1"></a>
+### 4.1 Python Inference
+First, the model saved during the ViTSTR text recognition training process is converted into an inference model. ( [Model download link](https://paddleocr.bj.bcebos.com/rec_vitstr_none_none_train.tar)) ), you can use the following command to convert:
+```
+python3 tools/export_model.py -c configs/rec/rec_vitstr_none_ce.yml -o Global.pretrained_model=./rec_vitstr_none_ce_train/best_accuracy  Global.save_inference_dir=./inference/rec_vitstr
+```
+**Note:**
+- If you are training the model on your own dataset and have modified the dictionary file, please pay attention to modify the `character_dict_path` in the configuration file to the modified dictionary file.
+- If you modified the input size during training, please modify the `infer_shape` corresponding to ViTSTR in the `tools/export_model.py` file.
+After the conversion is successful, there are three files in the directory:
+```
+/inference/rec_vitstr/
+    ├── inference.pdiparams
+    ├── inference.pdiparams.info
+    └── inference.pdmodel
+```
+For ViTSTR text recognition model inference, the following commands can be executed:
+```
+python3 tools/infer/predict_rec.py --image_dir='./doc/imgs_words_en/word_10.png' --rec_model_dir='./inference/rec_vitstr/' --rec_algorithm='ViTSTR' --rec_image_shape='1,224,224' --rec_char_dict_path='./ppocr/utils/EN_symbol_dict.txt'
+```
+![](../imgs_words_en/word_10.png)
+After executing the command, the prediction result (recognized text and score) of the image above is printed to the screen, an example is as follows:
+The result is as follows:
+```shell
+Predicts of ./doc/imgs_words_en/word_10.png:('pain', 0.9998350143432617)
+```
+<a name="4-2"></a>
+### 4.2 C++ Inference
+Not supported
+<a name="4-3"></a>
+### 4.3 Serving
+Not supported
+<a name="4-4"></a>
+### 4.4 More
+Not supported
+<a name="5"></a>
+## 5. FAQ
+1. In the `ViTSTR` paper, using pre-trained weights on ImageNet1k for initial training, we did not use pre-trained weights in training, and the final accuracy did not change or even improved.
+## Citation
+```bibtex
+@article{Atienza2021ViTSTR,
+  title     = {Vision Transformer for Fast and Efficient Scene Text Recognition},
+  author    = {Rowel Atienza},
+  booktitle = {ICDAR},
+  year      = {2021},
+  url       = {https://arxiv.org/abs/2105.08582}
+}
+```
--- a/ppocr/data/imaug/__init__.py
+++ b/ppocr/data/imaug/__init__.py
@@ -22,8 +22,11 @@ from .make_shrink_map import MakeShrinkMap
 from .random_crop_data import EastRandomCropData, RandomCropImgMask
 from .make_pse_gt import MakePseGt
 from .rec_img_aug import BaseDataAugmentation, RecAug, RecConAug, RecResizeImg, ClsResizeImg, \
-    SRNRecResizeImg, NRTRRecResizeImg, SARRecResizeImg, PRENResizeImg
+                         SRNRecResizeImg, GrayRecResizeImg, SARRecResizeImg, PRENResizeImg, \
+                         ABINetRecResizeImg, SVTRRecResizeImg, ABINetRecAug
 from .ssl_img_aug import SSLRotateResize
 from .randaugment import RandAugment
 from .copy_paste import CopyPaste

--- a/ppocr/data/imaug/abinet_aug.py
+++ b/ppocr/data/imaug/abinet_aug.py
+# copyright (c) 2020 PaddlePaddle Authors. All Rights Reserve.
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+#    http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+"""
+This code is refer from:
+https://github.com/FangShancheng/ABINet/blob/main/transforms.py
+"""
+import math
+import numbers
+import random
+import cv2
+import numpy as np
+from paddle.vision.transforms import Compose, ColorJitter
+def sample_asym(magnitude, size=None):
+    return np.random.beta(1, 4, size) * magnitude
+def sample_sym(magnitude, size=None):
+    return (np.random.beta(4, 4, size=size) - 0.5) * 2 * magnitude
+def sample_uniform(low, high, size=None):
+    return np.random.uniform(low, high, size=size)
+def get_interpolation(type='random'):
+    if type == 'random':
+        choice = [
+            cv2.INTER_NEAREST, cv2.INTER_LINEAR, cv2.INTER_CUBIC, cv2.INTER_AREA
+        ]
+        interpolation = choice[random.randint(0, len(choice) - 1)]
+    elif type == 'nearest':
+        interpolation = cv2.INTER_NEAREST
+    elif type == 'linear':
+        interpolation = cv2.INTER_LINEAR
+    elif type == 'cubic':
+        interpolation = cv2.INTER_CUBIC
+    elif type == 'area':
+        interpolation = cv2.INTER_AREA
+    else:
+        raise TypeError(
+            'Interpolation types only nearest, linear, cubic, area are supported!'
+        )
+    return interpolation
+class CVRandomRotation(object):
+    def __init__(self, degrees=15):
+        assert isinstance(degrees,
+                          numbers.Number), "degree should be a single number."
+        assert degrees >= 0, "degree must be positive."
+        self.degrees = degrees
+    @staticmethod
+    def get_params(degrees):
+        return sample_sym(degrees)
+    def __call__(self, img):
+        angle = self.get_params(self.degrees)
+        src_h, src_w = img.shape[:2]
+        M = cv2.getRotationMatrix2D(
+            center=(src_w / 2, src_h / 2), angle=angle, scale=1.0)
+        abs_cos, abs_sin = abs(M[0, 0]), abs(M[0, 1])
+        dst_w = int(src_h * abs_sin + src_w * abs_cos)
+        dst_h = int(src_h * abs_cos + src_w * abs_sin)
+        M[0, 2] += (dst_w - src_w) / 2
+        M[1, 2] += (dst_h - src_h) / 2
+        flags = get_interpolation()
+        return cv2.warpAffine(
+            img,
+            M, (dst_w, dst_h),
+            flags=flags,
+            borderMode=cv2.BORDER_REPLICATE)
+class CVRandomAffine(object):
+    def __init__(self, degrees, translate=None, scale=None, shear=None):
+        assert isinstance(degrees,
+                          numbers.Number), "degree should be a single number."
+        assert degrees >= 0, "degree must be positive."
+        self.degrees = degrees
+        if translate is not None:
+            assert isinstance(translate, (tuple, list)) and len(translate) == 2, \
+                "translate should be a list or tuple and it must be of length 2."
+            for t in translate:
+                if not (0.0 <= t <= 1.0):
+                    raise ValueError(
+                        "translation values should be between 0 and 1")
+        self.translate = translate
+        if scale is not None:
+            assert isinstance(scale, (tuple, list)) and len(scale) == 2, \
+                "scale should be a list or tuple and it must be of length 2."
+            for s in scale:
+                if s <= 0:
+                    raise ValueError("scale values should be positive")
+        self.scale = scale
+        if shear is not None:
+            if isinstance(shear, numbers.Number):
+                if shear < 0:
+                    raise ValueError(
+                        "If shear is a single number, it must be positive.")
+                self.shear = [shear]
+            else:
+                assert isinstance(shear, (tuple, list)) and (len(shear) == 2), \
+                    "shear should be a list or tuple and it must be of length 2."
+                self.shear = shear
+        else:
+            self.shear = shear
+    def _get_inverse_affine_matrix(self, center, angle, translate, scale,
+                                   shear):
+        # https://github.com/pytorch/vision/blob/v0.4.0/torchvision/transforms/functional.py#L717
+        from numpy import sin, cos, tan
+        if isinstance(shear, numbers.Number):
+            shear = [shear, 0]
+        if not isinstance(shear, (tuple, list)) and len(shear) == 2:
+            raise ValueError(
+                "Shear should be a single value or a tuple/list containing " +
+                "two values. Got {}".format(shear))
+        rot = math.radians(angle)
+        sx, sy = [math.radians(s) for s in shear]
+        cx, cy = center
+        tx, ty = translate
+        # RSS without scaling
+        a = cos(rot - sy) / cos(sy)
+        b = -cos(rot - sy) * tan(sx) / cos(sy) - sin(rot)
+        c = sin(rot - sy) / cos(sy)
+        d = -sin(rot - sy) * tan(sx) / cos(sy) + cos(rot)
+        # Inverted rotation matrix with scale and shear
+        # det([[a, b], [c, d]]) == 1, since det(rotation) = 1 and det(shear) = 1
+        M = [d, -b, 0, -c, a, 0]
+        M = [x / scale for x in M]
+        # Apply inverse of translation and of center translation: RSS^-1 * C^-1 * T^-1
+        M[2] += M[0] * (-cx - tx) + M[1] * (-cy - ty)
+        M[5] += M[3] * (-cx - tx) + M[4] * (-cy - ty)
+        # Apply center translation: C * RSS^-1 * C^-1 * T^-1
+        M[2] += cx
+        M[5] += cy
+        return M
+    @staticmethod
+    def get_params(degrees, translate, scale_ranges, shears, height):
+        angle = sample_sym(degrees)
+        if translate is not None:
+            max_dx = translate[0] * height
+            max_dy = translate[1] * height
+            translations = (np.round(sample_sym(max_dx)),
+                            np.round(sample_sym(max_dy)))
+        else:
+            translations = (0, 0)
+        if scale_ranges is not None:
+            scale = sample_uniform(scale_ranges[0], scale_ranges[1])
+        else:
+            scale = 1.0
+        if shears is not None:
+            if len(shears) == 1:
+                shear = [sample_sym(shears[0]), 0.]
+            elif len(shears) == 2:
+                shear = [sample_sym(shears[0]), sample_sym(shears[1])]
+        else:
+            shear = 0.0
+        return angle, translations, scale, shear
+    def __call__(self, img):
+        src_h, src_w = img.shape[:2]
+        angle, translate, scale, shear = self.get_params(
+            self.degrees, self.translate, self.scale, self.shear, src_h)
+        M = self._get_inverse_affine_matrix((src_w / 2, src_h / 2), angle,
+                                            (0, 0), scale, shear)
+        M = np.array(M).reshape(2, 3)
+        startpoints = [(0, 0), (src_w - 1, 0), (src_w - 1, src_h - 1),
+                       (0, src_h - 1)]
+        project = lambda x, y, a, b, c: int(a * x + b * y + c)
+        endpoints = [(project(x, y, *M[0]), project(x, y, *M[1]))
+                     for x, y in startpoints]
+        rect = cv2.minAreaRect(np.array(endpoints))
+        bbox = cv2.boxPoints(rect).astype(dtype=np.int)
+        max_x, max_y = bbox[:, 0].max(), bbox[:, 1].max()
+        min_x, min_y = bbox[:, 0].min(), bbox[:, 1].min()
+        dst_w = int(max_x - min_x)
+        dst_h = int(max_y - min_y)
+        M[0, 2] += (dst_w - src_w) / 2
+        M[1, 2] += (dst_h - src_h) / 2
+        # add translate
+        dst_w += int(abs(translate[0]))
+        dst_h += int(abs(translate[1]))
+        if translate[0] < 0: M[0, 2] += abs(translate[0])
+        if translate[1] < 0: M[1, 2] += abs(translate[1])
+        flags = get_interpolation()
+        return cv2.warpAffine(
+            img,
+            M, (dst_w, dst_h),
+            flags=flags,
+            borderMode=cv2.BORDER_REPLICATE)
+class CVRandomPerspective(object):
+    def __init__(self, distortion=0.5):
+        self.distortion = distortion
+    def get_params(self, width, height, distortion):
+        offset_h = sample_asym(
+            distortion * height / 2, size=4).astype(dtype=np.int)
+        offset_w = sample_asym(
+            distortion * width / 2, size=4).astype(dtype=np.int)
+        topleft = (offset_w[0], offset_h[0])
+        topright = (width - 1 - offset_w[1], offset_h[1])
+        botright = (width - 1 - offset_w[2], height - 1 - offset_h[2])
+        botleft = (offset_w[3], height - 1 - offset_h[3])
+        startpoints = [(0, 0), (width - 1, 0), (width - 1, height - 1),
+                       (0, height - 1)]
+        endpoints = [topleft, topright, botright, botleft]
+        return np.array(
+            startpoints, dtype=np.float32), np.array(
+                endpoints, dtype=np.float32)
+    def __call__(self, img):
+        height, width = img.shape[:2]
+        startpoints, endpoints = self.get_params(width, height, self.distortion)
+        M = cv2.getPerspectiveTransform(startpoints, endpoints)
+        # TODO: more robust way to crop image
+        rect = cv2.minAreaRect(endpoints)
+        bbox = cv2.boxPoints(rect).astype(dtype=np.int)
+        max_x, max_y = bbox[:, 0].max(), bbox[:, 1].max()
+        min_x, min_y = bbox[:, 0].min(), bbox[:, 1].min()
+        min_x, min_y = max(min_x, 0), max(min_y, 0)
+        flags = get_interpolation()
+        img = cv2.warpPerspective(
+            img,
+            M, (max_x, max_y),
+            flags=flags,
+            borderMode=cv2.BORDER_REPLICATE)
+        img = img[min_y:, min_x:]
+        return img
+class CVRescale(object):
+    def __init__(self, factor=4, base_size=(128, 512)):
+        """ Define image scales using gaussian pyramid and rescale image to target scale.
+        Args:
+            factor: the decayed factor from base size, factor=4 keeps target scale by default.
+            base_size: base size the build the bottom layer of pyramid
+        """
+        if isinstance(factor, numbers.Number):
+            self.factor = round(sample_uniform(0, factor))
+        elif isinstance(factor, (tuple, list)) and len(factor) == 2:
+            self.factor = round(sample_uniform(factor[0], factor[1]))
+        else:
+            raise Exception('factor must be number or list with length 2')
+        # assert factor is valid
+        self.base_h, self.base_w = base_size[:2]
+    def __call__(self, img):
+        if self.factor == 0: return img
+        src_h, src_w = img.shape[:2]
+        cur_w, cur_h = self.base_w, self.base_h
+        scale_img = cv2.resize(
+            img, (cur_w, cur_h), interpolation=get_interpolation())
+        for _ in range(self.factor):
+            scale_img = cv2.pyrDown(scale_img)
+        scale_img = cv2.resize(
+            scale_img, (src_w, src_h), interpolation=get_interpolation())
+        return scale_img
+class CVGaussianNoise(object):
+    def __init__(self, mean=0, var=20):
+        self.mean = mean
+        if isinstance(var, numbers.Number):
+            self.var = max(int(sample_asym(var)), 1)
+        elif isinstance(var, (tuple, list)) and len(var) == 2:
+            self.var = int(sample_uniform(var[0], var[1]))
+        else:
+            raise Exception('degree must be number or list with length 2')
+    def __call__(self, img):
+        noise = np.random.normal(self.mean, self.var**0.5, img.shape)
+        img = np.clip(img + noise, 0, 255).astype(np.uint8)
+        return img
+class CVMotionBlur(object):
+    def __init__(self, degrees=12, angle=90):
+        if isinstance(degrees, numbers.Number):
+            self.degree = max(int(sample_asym(degrees)), 1)
+        elif isinstance(degrees, (tuple, list)) and len(degrees) == 2:
+            self.degree = int(sample_uniform(degrees[0], degrees[1]))
+        else:
+            raise Exception('degree must be number or list with length 2')
+        self.angle = sample_uniform(-angle, angle)
+    def __call__(self, img):
+        M = cv2.getRotationMatrix2D((self.degree // 2, self.degree // 2),
+                                    self.angle, 1)
+        motion_blur_kernel = np.zeros((self.degree, self.degree))
+        motion_blur_kernel[self.degree // 2, :] = 1
+        motion_blur_kernel = cv2.warpAffine(motion_blur_kernel, M,
+                                            (self.degree, self.degree))
+        motion_blur_kernel = motion_blur_kernel / self.degree
+        img = cv2.filter2D(img, -1, motion_blur_kernel)
+        img = np.clip(img, 0, 255).astype(np.uint8)
+        return img
+class CVGeometry(object):
+    def __init__(self,
+                 degrees=15,
+                 translate=(0.3, 0.3),
+                 scale=(0.5, 2.),
+                 shear=(45, 15),
+                 distortion=0.5,
+                 p=0.5):
+        self.p = p
+        type_p = random.random()
+        if type_p < 0.33:
+            self.transforms = CVRandomRotation(degrees=degrees)
+        elif type_p < 0.66:
+            self.transforms = CVRandomAffine(
+                degrees=degrees, translate=translate, scale=scale, shear=shear)
+        else:
+            self.transforms = CVRandomPerspective(distortion=distortion)
+    def __call__(self, img):
+        if random.random() < self.p:
+            return self.transforms(img)
+        else:
+            return img
+class CVDeterioration(object):
+    def __init__(self, var, degrees, factor, p=0.5):
+        self.p = p
+        transforms = []
+        if var is not None:
+            transforms.append(CVGaussianNoise(var=var))
+        if degrees is not None:
+            transforms.append(CVMotionBlur(degrees=degrees))
+        if factor is not None:
+            transforms.append(CVRescale(factor=factor))
+        random.shuffle(transforms)
+        transforms = Compose(transforms)
+        self.transforms = transforms
+    def __call__(self, img):
+        if random.random() < self.p:
+            return self.transforms(img)
+        else:
+            return img
+class CVColorJitter(object):
+    def __init__(self,
+                 brightness=0.5,
+                 contrast=0.5,
+                 saturation=0.5,
+                 hue=0.1,
+                 p=0.5):
+        self.p = p
+        self.transforms = ColorJitter(
+            brightness=brightness,
+            contrast=contrast,
+            saturation=saturation,
+            hue=hue)
+    def __call__(self, img):
+        if random.random() < self.p: return self.transforms(img)
+        else: return img
--- a/ppocr/data/imaug/label_ops.py
+++ b/ppocr/data/imaug/label_ops.py
@@ -157,37 +157,6 @@ class BaseRecLabelEncode(object):
        return text_list
-class NRTRLabelEncode(BaseRecLabelEncode):
-    """ Convert between text-label and text-index """
-    def __init__(self,
-                 max_text_length,
-                 character_dict_path=None,
-                 use_space_char=False,
-                 **kwargs):
-        super(NRTRLabelEncode, self).__init__(
-            max_text_length, character_dict_path, use_space_char)
-    def __call__(self, data):
-        text = data['label']
-        text = self.encode(text)
-        if text is None:
-            return None
-        if len(text) >= self.max_text_len - 1:
-            return None
-        data['length'] = np.array(len(text))
-        text.insert(0, 2)
-        text.append(3)
-        text = text + [0] * (self.max_text_len - len(text))
-        data['label'] = np.array(text)
-        return data
-    def add_special_char(self, dict_character):
-        dict_character = ['blank', '<unk>', '<s>', '</s>'] + dict_character
-        return dict_character
 class CTCLabelEncode(BaseRecLabelEncode):
    """ Convert between text-label and text-index """
@@ -290,15 +259,26 @@ class E2ELabelEncodeTrain(object):
 class KieLabelEncode(object):
-    def __init__(self, character_dict_path, norm=10, directed=False, **kwargs):
+    def __init__(self,
+                 character_dict_path,
+                 class_path,
+                 norm=10,
+                 directed=False,
+                 **kwargs):
        super(KieLabelEncode, self).__init__()
        self.dict = dict({'': 0})
+        self.label2classid_map = dict()
        with open(character_dict_path, 'r', encoding='utf-8') as fr:
            idx = 1
            for line in fr:
                char = line.strip()
                self.dict[char] = idx
                idx += 1
+        with open(class_path, "r") as fin:
+            lines = fin.readlines()
+            for idx, line in enumerate(lines):
+                line = line.strip("\n")
+                self.label2classid_map[line] = idx
        self.norm = norm
        self.directed = directed
@@ -439,7 +419,7 @@ class KieLabelEncode(object):
            text_ind = [self.dict[c] for c in text if c in self.dict]
            text_inds.append(text_ind)
            if 'label' in ann.keys():
-                labels.append(ann['label'])
+                labels.append(self.label2classid_map[ann['label']])
            elif 'key_cls' in ann.keys():
                labels.append(ann['key_cls'])
            else:
@@ -907,15 +887,16 @@ class VQATokenLabelEncode(object):
        for info in ocr_info:
            if train_re:
                # for re
-                if len(info["text"]) == 0:
+                if len(info["transcription"]) == 0:
                    empty_entity.add(info["id"])
                    continue
                id2label[info["id"]] = info["label"]
                relations.extend([tuple(sorted(l)) for l in info["linking"]])
            # smooth_box
+            info["bbox"] = self.trans_poly_to_bbox(info["points"])
            bbox = self._smooth_box(info["bbox"], height, width)
-            text = info["text"]
+            text = info["transcription"]
            encode_res = self.tokenizer.encode(
                text, pad_to_max_seq_len=False, return_attention_mask=True)
@@ -931,7 +912,7 @@ class VQATokenLabelEncode(object):
                label = info['label']
                gt_label = self._parse_label(label, encode_res)
-            # construct entities for re
+# construct entities for re
            if train_re:
                if gt_label[0] != self.label2id_map["O"]:
                    entity_id_to_index_map[info["id"]] = len(entities)
@@ -975,29 +956,29 @@ class VQATokenLabelEncode(object):
            data['entity_id_to_index_map'] = entity_id_to_index_map
        return data
-    def _load_ocr_info(self, data):
+    def trans_poly_to_bbox(self, poly):
-        def trans_poly_to_bbox(poly):
+        x1 = np.min([p[0] for p in poly])
-            x1 = np.min([p[0] for p in poly])
+        x2 = np.max([p[0] for p in poly])
-            x2 = np.max([p[0] for p in poly])
+        y1 = np.min([p[1] for p in poly])
-            y1 = np.min([p[1] for p in poly])
+        y2 = np.max([p[1] for p in poly])
-            y2 = np.max([p[1] for p in poly])
+        return [x1, y1, x2, y2]
-            return [x1, y1, x2, y2]
+    def _load_ocr_info(self, data):
        if self.infer_mode:
            ocr_result = self.ocr_engine.ocr(data['image'], cls=False)
            ocr_info = []
            for res in ocr_result:
                ocr_info.append({
-                    "text": res[1][0],
+                    "transcription": res[1][0],
-                    "bbox": trans_poly_to_bbox(res[0]),
+                    "bbox": self.trans_poly_to_bbox(res[0]),
-                    "poly": res[0],
+                    "points": res[0],
                })
            return ocr_info
        else:
            info = data['label']
            # read text info
            info_dict = json.loads(info)
-            return info_dict["ocr_info"]
+            return info_dict
    def _smooth_box(self, bbox, height, width):
        bbox[0] = int(bbox[0] * 1000.0 / width)
@@ -1008,7 +989,7 @@ class VQATokenLabelEncode(object):
    def _parse_label(self, label, encode_res):
        gt_label = []
-        if label.lower() == "other":
+        if label.lower() in ["other", "others", "ignore"]:
            gt_label.extend([0] * len(encode_res["input_ids"]))
        else:
            gt_label.append(self.label2id_map[("b-" + label).upper()])
@@ -1046,3 +1027,99 @@ class MultiLabelEncode(BaseRecLabelEncode):
        data_out['label_sar'] = sar['label']
        data_out['length'] = ctc['length']
        return data_out
+class NRTRLabelEncode(BaseRecLabelEncode):
+    """ Convert between text-label and text-index """
+    def __init__(self,
+                 max_text_length,
+                 character_dict_path=None,
+                 use_space_char=False,
+                 **kwargs):
+        super(NRTRLabelEncode, self).__init__(
+            max_text_length, character_dict_path, use_space_char)
+    def __call__(self, data):
+        text = data['label']
+        text = self.encode(text)
+        if text is None:
+            return None
+        if len(text) >= self.max_text_len - 1:
+            return None
+        data['length'] = np.array(len(text))
+        text.insert(0, 2)
+        text.append(3)
+        text = text + [0] * (self.max_text_len - len(text))
+        data['label'] = np.array(text)
+        return data
+    def add_special_char(self, dict_character):
+        dict_character = ['blank', '<unk>', '<s>', '</s>'] + dict_character
+        return dict_character
+class ViTSTRLabelEncode(BaseRecLabelEncode):
+    """ Convert between text-label and text-index """
+    def __init__(self,
+                 max_text_length,
+                 character_dict_path=None,
+                 use_space_char=False,
+                 ignore_index=0,
+                 **kwargs):
+        super(ViTSTRLabelEncode, self).__init__(
+            max_text_length, character_dict_path, use_space_char)
+        self.ignore_index = ignore_index
+    def __call__(self, data):
+        text = data['label']
+        text = self.encode(text)
+        if text is None:
+            return None
+        if len(text) >= self.max_text_len:
+            return None
+        data['length'] = np.array(len(text))
+        text.insert(0, self.ignore_index)
+        text.append(1)
+        text = text + [self.ignore_index] * (self.max_text_len + 2 - len(text))
+        data['label'] = np.array(text)
+        return data
+    def add_special_char(self, dict_character):
+        dict_character = ['<s>', '</s>'] + dict_character
+        return dict_character
+class ABINetLabelEncode(BaseRecLabelEncode):
+    """ Convert between text-label and text-index """
+    def __init__(self,
+                 max_text_length,
+                 character_dict_path=None,
+                 use_space_char=False,
+                 ignore_index=100,
+                 **kwargs):
+        super(ABINetLabelEncode, self).__init__(
+            max_text_length, character_dict_path, use_space_char)
+        self.ignore_index = ignore_index
+    def __call__(self, data):
+        text = data['label']
+        text = self.encode(text)
+        if text is None:
+            return None
+        if len(text) >= self.max_text_len:
+            return None
+        data['length'] = np.array(len(text))
+        text.append(0)
+        text = text + [self.ignore_index] * (self.max_text_len + 1 - len(text))
+        data['label'] = np.array(text)
+        return data
+    def add_special_char(self, dict_character):
+        dict_character = ['</s>'] + dict_character
+        return dict_character
--- a/ppocr/data/imaug/operators.py
+++ b/ppocr/data/imaug/operators.py
@@ -67,39 +67,6 @@ class DecodeImage(object):
        return data
-class NRTRDecodeImage(object):
-    """ decode image """
-    def __init__(self, img_mode='RGB', channel_first=False, **kwargs):
-        self.img_mode = img_mode
-        self.channel_first = channel_first
-    def __call__(self, data):
-        img = data['image']
-        if six.PY2:
-            assert type(img) is str and len(
-                img) > 0, "invalid input 'img' in DecodeImage"
-        else:
-            assert type(img) is bytes and len(
-                img) > 0, "invalid input 'img' in DecodeImage"
-        img = np.frombuffer(img, dtype='uint8')
-        img = cv2.imdecode(img, 1)
-        if img is None:
-            return None
-        if self.img_mode == 'GRAY':
-            img = cv2.cvtColor(img, cv2.COLOR_GRAY2BGR)
-        elif self.img_mode == 'RGB':
-            assert img.shape[2] == 3, 'invalid shape of image[%s]' % (img.shape)
-            img = img[:, :, ::-1]
-        img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
-        if self.channel_first:
-            img = img.transpose((2, 0, 1))
-        data['image'] = img
-        return data
 class NormalizeImage(object):
    """ normalize image such as substract mean, divide std
    """

--- a/ppocr/data/imaug/rec_img_aug.py
+++ b/ppocr/data/imaug/rec_img_aug.py
--- a/ppocr/data/imaug/vqa/__init__.py
+++ b/ppocr/data/imaug/vqa/__init__.py
@@ -13,7 +13,12 @@
 # limitations under the License.
 from .token import VQATokenPad, VQASerTokenChunk, VQAReTokenChunk, VQAReTokenRelation
+from .augment import DistortBBox
 __all__ = [
-    'VQATokenPad', 'VQASerTokenChunk', 'VQAReTokenChunk', 'VQAReTokenRelation'
+    'VQATokenPad',
+    'VQASerTokenChunk',
+    'VQAReTokenChunk',
+    'VQAReTokenRelation',
+    'DistortBBox',
 ]
--- a/ppocr/data/imaug/vqa/augment.py
+++ b/ppocr/data/imaug/vqa/augment.py
--- a/ppocr/losses/__init__.py
+++ b/ppocr/losses/__init__.py
--- a/ppocr/losses/rec_ce_loss.py
+++ b/ppocr/losses/rec_ce_loss.py
--- a/ppocr/losses/rec_nrtr_loss.py
+++ b/ppocr/losses/rec_nrtr_loss.py
--- a/ppocr/modeling/backbones/__init__.py
+++ b/ppocr/modeling/backbones/__init__.py
--- a/ppocr/modeling/backbones/rec_resnet_45.py
+++ b/ppocr/modeling/backbones/rec_resnet_45.py
--- a/ppocr/modeling/backbones/rec_svtrnet.py
+++ b/ppocr/modeling/backbones/rec_svtrnet.py
--- a/ppocr/modeling/backbones/rec_vitstr.py
+++ b/ppocr/modeling/backbones/rec_vitstr.py
--- a/ppocr/modeling/heads/__init__.py
+++ b/ppocr/modeling/heads/__init__.py
--- a/ppocr/modeling/heads/multiheadAttention.py
+++ b/ppocr/modeling/heads/multiheadAttention.py
--- a/ppocr/modeling/heads/rec_abinet_head.py
+++ b/ppocr/modeling/heads/rec_abinet_head.py
--- a/ppocr/modeling/heads/rec_nrtr_head.py
+++ b/ppocr/modeling/heads/rec_nrtr_head.py
--- a/ppocr/postprocess/__init__.py
+++ b/ppocr/postprocess/__init__.py
--- a/ppocr/postprocess/rec_postprocess.py
+++ b/ppocr/postprocess/rec_postprocess.py
--- a/ppocr/utils/utility.py
+++ b/ppocr/utils/utility.py
--- a/ppocr/utils/visual.py
+++ b/ppocr/utils/visual.py
--- a/ppstructure/docs/kie.md
+++ b/ppstructure/docs/kie.md
--- a/ppstructure/docs/kie_en.md
+++ b/ppstructure/docs/kie_en.md
--- a/ppstructure/vqa/README.md
+++ b/ppstructure/vqa/README.md
--- a/ppstructure/vqa/README_ch.md
+++ b/ppstructure/vqa/README_ch.md
--- a/ppstructure/vqa/labels/labels_ser.txt
+++ b/ppstructure/vqa/labels/labels_ser.txt
--- a/ppstructure/vqa/tools/trans_xfun_data.py
+++ b/ppstructure/vqa/tools/trans_xfun_data.py
--- a/test_tipc/configs/rec_mtb_nrtr/rec_mtb_nrtr.yml
+++ b/test_tipc/configs/rec_mtb_nrtr/rec_mtb_nrtr.yml
--- a/test_tipc/configs/rec_r45_abinet/rec_r45_abinet.yml
+++ b/test_tipc/configs/rec_r45_abinet/rec_r45_abinet.yml
--- a/test_tipc/configs/rec_r45_abinet/train_infer_python.txt
+++ b/test_tipc/configs/rec_r45_abinet/train_infer_python.txt
--- a/test_tipc/configs/rec_svtrnet/rec_svtrnet.yml
+++ b/test_tipc/configs/rec_svtrnet/rec_svtrnet.yml
--- a/test_tipc/configs/rec_svtrnet/train_infer_python.txt
+++ b/test_tipc/configs/rec_svtrnet/train_infer_python.txt
--- a/test_tipc/configs/rec_vitstr_none_ce/rec_vitstr_none_ce.yml
+++ b/test_tipc/configs/rec_vitstr_none_ce/rec_vitstr_none_ce.yml
--- a/test_tipc/configs/rec_vitstr_none_ce/train_infer_python.txt
+++ b/test_tipc/configs/rec_vitstr_none_ce/train_infer_python.txt
--- a/test_tipc/test_serving_infer_cpp.sh
+++ b/test_tipc/test_serving_infer_cpp.sh
--- a/test_tipc/test_serving_infer_python.sh
+++ b/test_tipc/test_serving_infer_python.sh
--- a/tools/export_model.py
+++ b/tools/export_model.py
--- a/tools/infer/predict_rec.py
+++ b/tools/infer/predict_rec.py
--- a/tools/infer_kie.py
+++ b/tools/infer_kie.py
--- a/tools/infer_vqa_token_ser.py
+++ b/tools/infer_vqa_token_ser.py
--- a/tools/program.py
+++ b/tools/program.py