Monday, October 07, 2019

NPL(一一):NTM

NPL(一一):NTM NTM

2019/10/07

Neural Turing Machine

-----



// The mostly complete chart of Neural Networks, explained

-----



// The mostly complete chart of Neural Networks, explained

-----

-----



// Attention  Attention!

-----


// Attention  Attention!

-----


// The Neural Turing Machine - Aidan Gomez - Medium

-----

-----

References

◎ Paper

# NTM
Graves, Alex, Greg Wayne, and Ivo Danihelka. "Neural turing machines." arXiv preprint arXiv:1410.5401 (2014).
https://arxiv.org/pdf/1410.5401.pdf

-----

◎ 英文參考資料

A Tour of Recurrent Neural Network Algorithms for Deep Learning
https://machinelearningmastery.com/recurrent-neural-network-algorithms-for-deep-learning/

# NTM
Attention  Attention!
https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html 

# NTM
# 204 claps
The Neural Turing Machine - Aidan Gomez - Medium
https://medium.com/@aidangomez/the-neural-turing-machine-79f6e806c0a1

The mostly complete chart of Neural Networks, explained
https://towardsdatascience.com/the-mostly-complete-chart-of-neural-networks-explained-3fb6f2367464

Neural Turing Machines  Perils and Promise
https://blog.talla.com/neural-turing-machines-perils-and-promise

NTM-Lasagne  A Library for Neural Turing Machines in Lasagne
https://medium.com/snips-ai/ntm-lasagne-a-library-for-neural-turing-machines-in-lasagne-2cdce6837315


◎ 簡體中文參考資料

# 概述
一文看尽深度学习RNN:为啥就它适合语音识别、NLP与机器翻译? - Python开发社区 _ CTOLib码库
https://www.ctolib.com/topics-121320.html 

模型那么多,该怎么选择呢?这里有27种神经网络的图解_36氪
https://36kr.com/p/5115489

神经图灵机深度讲解:从图灵机基本概念到可微分神经计算机 _ 机器之心
https://www.jiqizhixin.com/articles/2017-04-11-7
 
# NTM 本質是一個使用外部存儲矩陣進行 attentive interaction 機制的 RNN
神经网络图灵机的通俗解释和详细过程及应用? - 知乎
https://www.zhihu.com/question/42029751

记忆网络之Neural Turing Machines - 知乎
https://zhuanlan.zhihu.com/p/30383994 

Thursday, October 03, 2019

AI 從頭學(三二):VGGNet

AI 從頭學(三二):VGGNet

2019/09/03

前言:

VGGNet 是一個經典的 CNN 模型。所謂經典,就是舊(不一定),而且很好。當然,即使是新的,也會變舊。VGG 的架構很簡潔,所以之前讀的時候,把重點擺在 Weight 與 Momentum。這次為了 ResNet,仔細地讀了一下 VGG,發現有許多細節還蠻重要的。

-----

Summary:

參考文獻的安排依論文 [1], [2]、英文 [3]-[8]、簡體中文 [9]-[16]、繁體中文 [17], [18] 與代碼 [19]-[21] 為順序。

VGGNet [1], [5], [6], [11]-[13] 與 AlexNet 都引用了 PreVGGNet [2]。PreVGGNet 同時加深與加寬網路,但發現加深有效,加寬效果不大。這給了 VGGNet 專注於加深的靈感 [2], [15], [18]。

文章中我們首先回顧 CNN 發展的歷史 [3], [4], [9], [10]。然後依序討論 VGGNet 的設計、訓練、與測試。

最後分析了 Conv3 [3], [12]、Conv1 [16]、與 LRN [7], [8] 對於 VGGNet 的影響。

-----

Outline

1.1. Top 5 Accuracy
1.2. Evolution of CNN

2.1. Structure of VGG
2.2. Conv3
2.3. Pooling

3.1. Training
3.2. Testing

4.1. Single Scale Evaluation
4.2. Multi-Scale Evaluation
4.3. Multi-crop Evalution
4.4. ConvNet Fusion

5.1. Deep
5.2. Conv1
5.3. LRN

-----

1.1. Top 5 Accuracy

-----


Fig. 1.1. Top 5 Accuracy [6]。

-----

卷積神經網路(Convolutional Neural Network,CNN)主要的功用,是判斷圖片的類別,通常以 Top 5 或 Top 1 為模型準確率的判斷標準。Top 5 是前五名其中之一正確即可列為模型判斷成功,Top 1 則只取第一名。

-----

1.2. Evolution of CNN

-----


Fig. 1.2a. Evolution of CNN [3]。

-----

圖1.2a 可以看到,CNN 的發展路線是,網路越來越深,而錯誤率越來越小。2012 年的 AlexNet 是 8 層,錯誤率是 16.4。2013 AlexNet 微調版的 ZFNet,錯誤率是 11.7。2014 年的 VGGNet 是 19 層,錯誤率是 7.3。同年的 GoogLeNet 是 22 層,錯誤率是 6.7。到了 2015 年的 ResNet 是 152 層,錯誤率是 3.57,一舉超過人類專家。

-----


Fig. 1.2b. Evolution of CNN [9]。

-----


Fig. 1.2c. Evolution of CNN [10]。

-----

圖1.2b 是 CNN 的發展歷史。

1998,LeNet 奠定了成熟的 CNN 架構。
2012,AlexNet 擴大 LeNet 的架構,並使用 GPU,成為第一個在大型圖片資料集表現優異的 CNN。
2013,NIN 發展了 1x1 convolution(conv1)。
2014,GoogLeNet(Inception V1)基於 conv1 成功地加深網路。
2014,VGGNet 使用兩個 conv3 組成 conv5,也成功地加深網路。
2015,ResNet 運用 LSTM 的直通架構,一舉將網路加到極深。

CNN 的成功,也連帶引起 Object Detection 與 Semantic Segmentation 種種應用的風行,參考圖 1.2c。

-----

2.1. Structure of VGG

-----


Fig. 2.1a. LeNet [4]。

-----


Fig. 2.1b. AlexNet [4]。

-----


Fig. 2.1c. VGGNet [4]。

-----


Fig. 2.1d. VGGNet Architecture [1]。

-----

VGGNet 的架構參考之前的 LeNet、PreVGGNet、AlexNet、ZFNet、NIN、OverFeat。

AlexNet 是 LeNet 的大型版。由於 ZFNet 的第一層採取較小的卷積核與 Stride,得到較高的辨識率,所以 VGGNet 也朝這個方向繼續前進。另外,對比於 AlexNet,VGGNet 也採用較小的 Pooling。

另外,PreVGGNet 發現加寬網路用處不大,所以 VGGNet 專注於加深網路 [2], [15], [18]。

-----

架構上首先訓練好 A,然後加上 LRN,稱為 A-LRN。

然後把 A 前兩層的一層 Conv3 換成兩層 Conv3,繼續訓練,稱為 B。

然後把 B 後三層的兩層 Conv3 加一層 Conv1,繼續訓練,稱為 C。

然後把 C 的 Conv1 換成 Conv3,繼續訓練,稱為 D。也就是俗稱的 VGG-16(層)。

最後 VGG-19 只有好一點點,但消耗資源較多,所以 VGG-16 比較常用。繼續加深就變差了。

-----

2.2. Conv3

-----


Fig. 2.2. Conv3 [5]。

-----

VGGNet 利用兩層 Conv3 組成 Conv5,有表現力強跟計算較少,這兩個優點。同理,三層 Conv3 可以組成 Conv7。

-----

2.3. Pooling

-----


Fig. 2.3. Pooling [17]。

-----

LeNet 採用 average pooling。

相對於 AlexNet 採用 3x3 且 stride = 2 的 max pooling,有重疊。VGGNet 採用 2x2 且 stride = 2 的 max pooling,無重疊。

-----

3.1. Training

-----


Fig. 3.1a. Single Scale [5]。

-----


Fig. 3.1b. Multi-Scale [17]。

-----

VGGNet 的 training 有 single 跟 multiple 兩種。

所謂 single,有 256 跟 384 兩種。訓練圖片的短邊縮放到 256,然後取中間 224x224。另外一種是把 256 改成 384。

所謂 multiple,則是將短邊隨機縮放為 256 到 512 其中一個大小,然後取中間 224x224。

-----

3.2. Testing

-----

Testing 分為 dense 跟 crop 兩種。

-----


Fig. 3.2a. Dense [5]。

-----


Fig. 3.2b. Dense [11]。

-----


Fig. 3.2c. OverFeat [15]。

-----

Dense 的方式是將最後三層的全連接層改為卷積層,參考圖3.2a 跟 3.2b。

靈感上是來自 OverFeat [15]。全連接層改為卷積層之後,輸入格式大小就不受限定,比 224x224 大,也沒關係。

-----


Fig. 3.2d. Multi-crop [17]。

-----

Multi-crop 是將 250x250 先放大到 280x280,然後取四角跟中間 224x224 下去測試,並取 softmax 後的輸出值平均。

-----

4.1. Single Scale Evaluation

-----


Fig. 4.1a. ConvNet performance at a single test scale [1]。

-----


Fig. 4.1b. ConvNet performance at a single test scale [5]。

-----

參考圖4.1b,通過 Multi-scale Training,我們可以猜測它對於具有不同尺寸的測試圖像而言更準確。

VGG-13 將錯誤率從 9.4% / 9.3%降低到8.8%。
VGG-16 將錯誤率從 8.8% / 8.7%降低到8.1%。
VGG-19 將錯誤率從 9.0% / 8.7%降低到8.0%。

-----

4.2. Multi-Scale Evaluation

-----


Fig. 4.2a. ConvNet performance at multiple test scales [1]。

-----


Fig. 4.2b. ConvNet performance at multiple test scales [5]。

-----

通過 Multi-scale training 而不是 Single-scale training,可以降低錯誤率。

與 Single-scale Training  Single-scale Testing 相比,

VGG-13將錯誤率從 9.4% / 9.3% 降低到 9.2%,
VGG-16將錯誤率從 8.8% / 8.7% 降低到 8.6%,
VGG-19將錯誤率從 9.0% / 8.7% 降低到 8.7 / 8.6%。

-----

通過同時使用多尺度培訓和測試,可以減少錯誤率。

與僅多尺度測試相比,

VGG-13 將錯誤率從 9.2% / 9.2% 降低到 8.2%。
VGG-16 將錯誤率從 8.6% / 8.6% 降低到 7.5%。
VGG-19 將錯誤率從 8.7% / 8.6% 降低到 7.5%。

-----

4.3. Multi-crop Evaluation

-----


Fig. 4.3a. ConvNet evaluation techniques comparison [1]。

-----


Fig. 4.3b. ConvNet evaluation techniques comparison [5]。

-----

藉由平均 dense 與 multi-crop 的結果,VGG-16 與 VGG-19 的錯誤率降到 7.2% 與 7.1%。

-----

4.4. ConvNet Fusion

-----


Fig. 4.4a. Multiple ConvNet fusion results [1]。

-----


Fig. 4.4b. Multiple ConvNet fusion results [5]。

-----

最好的結果達到 6.8%。

-----

5.1. Deep

-----


Fig. 5.1a. VGGNet [3]。

-----



Fig. 5.1b. Conv3 [12]。

-----

總之,VGGNet 採用 Conv3 反覆加深,得到很好的效果,是個簡潔優美而強大的設計。

-----

5.2. Conv1

-----


Fig. 5.2. Conv1 [16]。

-----

Conv1 [16] 的效果沒有 Conv3 好。但 Conv1 其實最主要是降低模型的資料量,這在 GoogLeNet 中被發揮到淋漓盡致 [9]。

-----

5.3. LRN 

-----


Fig. 5.3. LRN [7]。

-----

LRN [7], [8] 在深度學習,最早被使用在 AlexNet 上。很多文章說,VGGNet 證明了 LRN 無效。但舉一個無效的例子是無法「證明」什麼的,只能「說明」LRN 用在 VGGNet 無效。事實上,與 VGGNet 同時的 GoogLeNet 就有使用。

個人認為 LRN 後來沒有被繼續使用的主要原因是 Conv1 的功能與 LRN 類似,但更一般化與效果更好。由於 Conv1 幾乎已經是標準配備,自然不需要 LRN 多此一舉。

-----

結論:

閱讀論文的要點,除了理解其架構,另外要注意的是論文中提到,訓練模型的方法,即 weight decay 與 momentum。

VGGNet 的最大貢獻是「證明」持續加深網路有明顯效果,但到 16、19 已是極限。隔一年發表的 ResNet 加上一個 shortcut,採取殘差的設計,達到 CNN 的 SOTA,State of the Art [3], [4], [9], [10]!

-----

References

◎ 論文

[1] VGGNet
Simonyan, Karen, and Andrew Zisserman. "Very deep convolutional networks for large-scale image recognition." arXiv preprint arXiv:1409.1556 (2014).
https://arxiv.org/pdf/1409.1556.pdf

[2] PreVGGNet
Ciresan, Dan C., et al. "Flexible, high performance convolutional neural networks for image classification." IJCAI Proceedings-International Joint Conference on Artificial Intelligence. Vol. 22. No. 1. 2011.
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.481.4406&rep=rep1&type=pdf

-----

◎ 英文參考資料

# 綜述
# 3.9K claps
[3] CNN Architectures  LeNet, AlexNet, VGG, GoogLeNet, ResNet and more …
https://medium.com/@sidereal/cnns-architectures-lenet-alexnet-vgg-googlenet-resnet-and-more-666091488df5

# 綜述
# 1.1K claps
[4] Illustrated  10 CNN Architectures - Towards Data Science
https://towardsdatascience.com/illustrated-10-cnn-architectures-95d78ace614d

# VGGNet
# 246 claps
[5] Review  VGGNet — 1st Runner-Up (Image Classification), Winner (Localization) in ILSVRC 2014
https://medium.com/coinmonks/paper-review-of-vggnet-1st-runner-up-of-ilsvlc-2014-image-classification-d02355543a11

# VGGNet
[6] Convolutional neural networks on the iPhone with VGGNet
http://machinethink.net/blog/convolutional-neural-networks-on-the-iphone-with-vggnet/

# LRN
[7] What is local response normalization  - Quora
https://www.quora.com/What-is-local-response-normalization

# LRN
# 135 claps
[8] Difference between Local Response Normalization and Batch Normalization
https://towardsdatascience.com/difference-between-local-response-normalization-and-batch-normalization-272308c034ac

-----

◎ 簡體中文參考資料

# 綜述
[9] 深度学习之四大经典CNN技术浅析 _ 硬创公开课 _ 雷锋网
https://www.leiphone.com/news/201702/dgpHuriVJHTPqqtT.html

# 綜述
[10] GitHub - weslynn_AlphaTree-graphic-deep-neural-network  将深度神经网络中的一些模型 进行统一的图示,便于大家对模型的理解
https://github.com/weslynn/AlphaTree-graphic-deep-neural-network

# VGGNet 
[11] 大话CNN经典模型:VGGNet - 雪饼的个人空间 - OSCHINA
https://my.oschina.net/u/876354/blog/1634322

# VGGNet
[12] 【论文阅读】—— VERY DEEP CONVOLUTIONAL NETWORKS FOR LARGE-SCALE IMAGE RECOGNITION _ Nameless rookie
http://vincentho.name/2018/11/29/%E3%80%90%E8%AE%BA%E6%96%87%E9%98%85%E8%AF%BB%E3%80%91%E2%80%94%E2%80%94-VERY-DEEP-CONVOLUTIONAL-NETWORKS-FOR-LARGE-SCALE-IMAGE-RECOGNITION/

[13] VGG论文翻译——中英文对照 _ SnailTyan
http://noahsnail.com/2017/08/17/2017-08-17-VGG%E8%AE%BA%E6%96%87%E7%BF%BB%E8%AF%91%E2%80%94%E2%80%94%E4%B8%AD%E8%8B%B1%E6%96%87%E5%AF%B9%E7%85%A7/

# PreVGGNet
[14] 深度学习论文理解3:Flexible, high performance convolutional neural networks for image classification - whiteinblue的专栏 - CSDN博客
https://blog.csdn.net/whiteinblue/article/details/43149363

# OverFeat
[15] OverFeat Integrated Recognition, Localization and Detection using Convolutional Networks - baobei0112的专栏 - CSDN博客
https://blog.csdn.net/baobei0112/article/details/47775647

# Conv1
[16] CNN网络中的 1 x 1 卷积是什么? - AI小作坊 的博客 - CSDN博客
https://blog.csdn.net/zhangjunhit/article/details/55101559

-----
 
◎ 繁體中文參考資料

# VGGNet
[17] VGG_深度學習_原理 – JT – Medium
https://medium.com/@danjtchen/vgg-%E6%B7%B1%E5%BA%A6%E5%AD%B8%E7%BF%92-%E5%8E%9F%E7%90%86-d31d0aa13d88

# PreVGGNet
[18] [Pytorch Taipei] Paper  Flexible, high performance convolutional neural networks for image classification
https://medium.com/@ChrisChou0426/pytorch-taipei-paper-flexible-high-performance-convolutional-neural-networks-for-image-4153f9495113 

-----

◎ 代碼實作

# PyTorch
[19] torchvision.models.vgg — PyTorch master documentation
https://pytorch.org/docs/stable/_modules/torchvision/models/vgg.html

# PyTorch
[20] vision_vgg.py at master · pytorch_vision · GitHub
https://github.com/pytorch/vision/blob/master/torchvision/models/vgg.py  

# PyTorch
[21] 简单易懂Pytorch实战实例VGG深度网络 - 心之所向 - CSDN博客
https://blog.csdn.net/qq_16234613/article/details/79818370

Wednesday, October 02, 2019

AI 從頭學(一四):Recommender

AI 從頭學(一四):Recommender

2017/03/16

前言:

如何在 Azure 上建立 AI 的個人化新聞推薦系統?本文列出部分參考資料。

底下純屬紙上談兵,若有不足之處,還請不吝指教!

-----


Fig. Recommender(圖片來源:Pixabay)。

-----

Summary:

推薦系統 [1] 可用於個人化新聞推薦 [2]-[6],或是個人化財經新聞推薦 [7], [8]。概念上是在 Hadoop [9]-[30] 上跑 Mahout [31]-[34]。也可以使用記憶體作為暫存的 Spark [35]-[45],在速度上有驚人的提升。實作上則是先用 Sqoop [46] 把 SQL [47] 上的資料搬到 HDFS [48] 上。

如果要在 Azure [49]-[62] 上做,一樣要先把 Azure SQL [63], [64] 的資料先移到 Azure 的 Hadoop,也就是 HDInsight [65]-[67] 上,再跑 Azure 的 Machine Learning Service [68]-[70]。

如要自行將演算法 [2]-[8] 開發成系統,微軟推薦 F# [71]。F# [72]-[84] 屬於 Functional Programming 的程式語言,有簡潔的程式碼、較高的產能等種種好處。

Deep Learning [85]-[91] 也可用來開發推薦系統 [92]-[97],今日頭條是此中翹楚;目前總員工 2,500 名,工程師有 1,500 個,其中 800 名工程師聚焦在演算法、資料分析 [98]。

-----

出版說明:

Collaborative Filtering - RDD-based API - Spark 2.4.4 Documentation
https://spark.apache.org/docs/latest/mllib-collaborative-filtering.html

之前不熟 Big Data 時,幫公司規劃推薦系統,洋洋灑灑找了近百篇資料。其實現在熟了,只要裝好 Spark 再跑一下預設的演算法就可以了:「Collaborative filtering is commonly used for recommender systems.」。較新的作法是使用深度學習,資料也不難找。

-----

References

◎ 01_Recommender

[1] 2015_Recommender Systems Handbook

-----

◎ 02_News

[2] 2013_Personalized News Recommendation Using Ontologies Harvested from the Web

[3] 2013_Mobile Recommender Systems and Their Applications

[4] 2011_News personalization using the CF-IDF semantic recommender

[5] 2010_Personalized news recommendation based on click behavior

[6] 2010_A contextual-bandit approach to personalized news article recommendation

-----

◎ 03_Financial

[7] 2015_Personalized Financial News Recommendation Algorithm Based on Ontology

[8] 2000_Language models for financial news recommendation

-----

◎ 04_Hadoop

[9] 2015_Learning Hadoop 2

[10] 2015_Hadoop,The definitive guide

[11] 2015_Hadoop Essentials

[12] 2015_Hadoop Backup and Recovery Solutions

[13] 2015_Guide to high performance distributed computing, case studies with Hadoop, Scalding and Spark

[14] 2015_Field Guide to Hadoop

[15] 2015_Big data made easy, a working guide to the complete Hadoop toolset

[16] 2015_Big Data Governance, Modern Data Management Principles for Hadoop, NoSQL & Big Data Analytics

[17] 2015_Big Data Forensics, Learning Hadoop Investigations

[18] 2014_Pro Apache Hadoop

[19] 2014_Practical Hadoop security

[20] 2014_Hadoop For Dummies

[21] 2013_Securing Hadoop

[22] 2013_Professional Hadoop Solutions

[23] 2013_Hadoop Real-World Solutions Cookbook

[24] 2013_Hadoop Operations and Cluster Management Cookbook

[25] 2013_Hadoop Cluster Deployment

[26] 2013_Hadoop Beginner's Guide

[27] 2012_Hadoop Operations

[28] 2012_Hadoop in Practice

[29] 2010_Hadoop in Action

[30] 2009_Pro Hadoop

-----

◎ 05_Mahout

[31] 2015_Learning Apache Mahout

[32] 2015_Learning Apache Mahout Classification

[33] 2013_Apache Mahout Cookbook

[34] 2011_Mahout in Action

-----

◎ 06_Spark

[35] 2016_Spark Guide

[36] 2016_Pro Spark Streaming

[37] 2016_High Performance Spark

[38] 2015_Spark for Python Developers

[39] 2015_Spark Core Programming

[40] 2015_Mastering Apache Spark

[41] 2015_Learning Spark

[42] 2015_Getting Started with Apache Spark

[43] 2015_Fast Data Processing with Spark

[44] 2015_Big Data Analytics with Spark

[45] 2015_Advanced Analytics with Spark

-----

◎ 07_Sqoop

[46] 2013_Apache Sqoop Cookbook

-----

◎ 08_SQL

[47] 2013_Microsoft SQL Server 2012 with Hadoop

-----

◎ 09_HDFS

[48] 2010_The Hadoop Distributed File System

-----

◎ 10_Azure

[49] 2015_Microsoft Azure, planning, deploying, and managing your data center in the cloud

[50] 2015_Microsoft Azure Essentials, Fundamentals of Azure

[51] 2015_Microservices, IoT, and Azure, leveraging DevOps and microservice architecture to deliver SaaS solutions

[52] 2015_Hardening Azure Applications

[53] 2015_15 minute Azure Installation, Set up the Microsoft Cloud Server by the Numbers

[54] 2014_Zen of Cloud, Learning Cloud Computing by Examples on Microsoft Azure

[55] 2014_Mastering Hyper-V 2012 R2 With System Center and Windows Azure

[56] 2014_Learning Windows Azure Mobile Services for Windows 8 and Windows Phone 8

[57] 2012_Programming Microsoft's Clouds, Windows Azure and Office 365

[58] 2012_Cloud Architecture Patterns Using Microsoft Azure

[59] 2011_Windows Azure platform

[60] 2011_Azure in Action

[61] 2009_Windows Azure platform

[62] 2009_Introduction to Windows Azure, an introduction to cloud computing using Microsoft Windows Azure

-----

◎ 11_SQL

[63] 2012_Pro SQL Database for Windows Azure, SQL Server in the Cloud

[64] 2010_Pro SQL Azure

-----

◎ 12_HDInsight

[65] 2014_Pro Microsoft HDInsight, Hadoop on Windows

[66] 2014_Microsoft Big Data Solutions

[67] 2013_HDInsight Essentials

-----

◎ 13_ML

[68] 2015_Predictive analytics with Microsoft azure machine learning, build and deploy actionable solutions in minutes

[69] 2015_Data Science in the Cloud with Microsoft Azure Machine Learning and R

[70] 2014_Predictive analytics with Microsoft azure machine learning, build and deploy actionable solutions in minutes

-----

◎ 14_ML_F#

[71] 2015_Machine Learning Projects for _NET Developers

-----

◎ 15_F#

[72] 2016_Beginning F# 4_0

[73] 2015_Learning F# Functional Data Structures and Algorithms

[74] 2015_Expert F# 4_0

[75] 2014_The Book of F#, Breaking Free with Managed Functional Programming

[76] 2013_Windows Phone 7_5 Application Development with F#

[77] 2013_Functional Programming Using F#

[78] 2013_F# for Quantitative Finance

[79] 2013_F# for C# Developers

[80] 2012_Programming F# 3_0

[81] 2012_Expert F# 3_0

[82] 2010_Visual Studio 2010 and _NET 4 Six-in-One

[83] 2010_Expert F# 2_0

[84] 2007_Expert F#

-----

◎ 16_DL_Overview

[85] 2016_Towards Bayesian Deep Learning, A Survey

[86] 2016_Deep Learning on FPGAs, Past, Present, and Future

[87] 2015_Deep learning

[88] 2015_Deep learning in neural networks, An overview

[89] 2014_Deep Learning, Methods and Applications

[90] 2012_Unsupervised feature learning and deep learning, A review and new perspectives

[91] 2009_Learning deep architectures for AI

-----

◎ 17_DL_RS

[92] 2016_Deep neural networks for youtube recommendations

[93] 2016_Collaborative denoising auto-encoders for top-n recommender systems

[94] 2015_Deep collaborative filtering via marginalized denoising auto-encoder

[95] 2015_Collaborative deep learning for recommender systems

[96] 2014_Improving content-based and hybrid music recommendation using deep learning

[97] 2013_Deep content-based music recommendation

[98] 4 年拿下 7,400 萬日活躍用戶,今日頭條已經準備跨入全球市場
http://technews.tw/2016/12/08/toutiao-china-global-market/ 

NLP(一):LSTM

NLP(一):LSTM

2019/09/06

說明:

Recurrent Neural Network (RNN) 跟 Long Short-Term Memory (LSTM) [1]-[3] 都是用來處理時間序列的訊號,譬如 Audio、Speech、Language [4], [5]。由於 RNN 有梯度消失與梯度爆炸的問題,所以 LSTM 被開發出來取代 RNN。由於本質上的缺陷(不能使用 GPU 平行加速),所以雖然 NLP 原本使用 LSTM、GRU 等開發出來的語言模型如 Seq2seq、Attention 等,最後也捨棄了 RNN 系列,而改用全連接層為主的 Transformer,並且取得很好的成果 [8]-[10]。即便如此,還是有更新的 RNN 模型譬如 MGU、SRU 被提出 [11]。

-----


Fig. 1. RNN, [1].

-----

-----


Recurrent Neural Networks and LSTM explained - purnasai gudikandula - Medium

-----




Fig. 3.1b. BPTT algorithm, p. 243, [14].

-----

-----


Recurrent Neural Networks and LSTM explained - purnasai gudikandula - Medium

-----


Recurrent Neural Networks and LSTM explained - purnasai gudikandula - Medium

-----



-----


Fig. 2. LSTM, [1].

-----



// Recurrent Neural Networks and LSTM explained - purnasai gudikandula - Medium

-----


Understanding LSTM and its diagrams - ML Review - Medium

-----

重點在於三個 sigmoid 產生控制訊號。以及兩個 tanh 用來壓縮資料。



Optimizing Recurrent Neural Networks in cuDNN 5


Optimizing Recurrent Neural Networks in cuDNN 5

-----



Fig. 14. Peephole connections [1].

-----


-----



Fig. 15. Coupled forget and input gates [1].

-----
-----



Fig. 16. GRU [1].

-----










-----



-----


// Why LSTM cannot prevent gradient exploding  - Cecile Liu - Medium

-----

References

◎ 論文

# LSTM(Long Short-Term Memory)

Hochreiter, Sepp, and Jürgen Schmidhuber. "Long short-term memory." Neural computation 9.8 (1997): 1735-1780.
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.676.4320&rep=rep1&type=pdf
-----

◎ 英文參考資料

Understanding LSTM Networks -- colah's blog
http://colah.github.io/posts/2015-08-Understanding-LSTMs/ 

The Unreasonable Effectiveness of Recurrent Neural Networks
http://karpathy.github.io/2015/05/21/rnn-effectiveness/
 
# 10.8K claps
The fall of RNN _ LSTM – Towards Data Science
https://towardsdatascience.com/the-fall-of-rnn-lstm-2d1594c74ce0

-----

Written Memories  Understanding, Deriving and Extending the LSTM - R2RT
https://r2rt.com/written-memories-understanding-deriving-and-extending-the-lstm.html

Neural Network Zoo Prequel  Cells and Layers - The Asimov Institute
https://www.asimovinstitute.org/neural-network-zoo-prequel-cells-layers/
 
-----

# 10.1K claps
[] Illustrated Guide to LSTM’s and GRU’s  A step by step explanation
https://towardsdatascience.com/illustrated-guide-to-lstms-and-gru-s-a-step-by-step-explanation-44e9eb85bf21
 
# 7.8K claps
[] Understanding LSTM and its diagrams - ML Review - Medium
https://medium.com/mlreview/understanding-lstm-and-its-diagrams-37e2f46f1714

# 386 claps
[] The magic of LSTM neural networks - DataThings - Medium
https://medium.com/datathings/the-magic-of-lstm-neural-networks-6775e8b540cd
 
# 247 claps
[] Recurrent Neural Networks and LSTM explained - purnasai gudikandula - Medium
https://medium.com/@purnasaigudikandula/recurrent-neural-networks-and-lstm-explained-7f51c7f6bbb9

# 113 claps
[] A deeper understanding of NNets (Part 3) — LSTM and GRU
https://medium.com/@godricglow/a-deeper-understanding-of-nnets-part-3-lstm-and-gru-e557468acb04

# 58 claps
[] Basic understanding of LSTM - Good Audience
https://blog.goodaudience.com/basic-understanding-of-lstm-539f3b013f1e

Optimizing Recurrent Neural Networks in cuDNN 5
https://devblogs.nvidia.com/optimizing-recurrent-neural-networks-cudnn-5/

-----

◎ 簡體中文參考資料

# 梯度消失 梯度爆炸
[] 三次简化一张图:一招理解LSTM_GRU门控机制 _ 机器之心
https://www.jiqizhixin.com/articles/2018-12-18-12

# 梯度消失 梯度爆炸
[] 长短期记忆(LSTM)-tensorflow代码实现 - Jason160918的博客 - CSDN博客
https://blog.csdn.net/Jason160918/article/details/78295423

[] 周志华等提出 RNN 可解释性方法,看看 RNN 内部都干了些什么 _ 机器之心
https://www.jiqizhixin.com/articles/110404 

-----

◎ 繁體中文參考資料

[] 遞歸神經網路和長短期記憶模型 RNN & LSTM · 資料科學・機器・人
https://brohrer.mcknote.com/zh-Hant/how_machine_learning_works/how_rnns_lstm_work.html

# 593 claps
[] 淺談遞歸神經網路 (RNN) 與長短期記憶模型 (LSTM) - TengYuan Chang - Medium
https://medium.com/@tengyuanchang/%E6%B7%BA%E8%AB%87%E9%81%9E%E6%AD%B8%E7%A5%9E%E7%B6%93%E7%B6%B2%E8%B7%AF-rnn-%E8%88%87%E9%95%B7%E7%9F%AD%E6%9C%9F%E8%A8%98%E6%86%B6%E6%A8%A1%E5%9E%8B-lstm-300cbe5efcc3

# 405 claps
[] 速記AI課程-深度學習入門(二) - Gimi Kao - Medium
https://medium.com/@baubibi/%E9%80%9F%E8%A8%98ai%E8%AA%B2%E7%A8%8B-%E6%B7%B1%E5%BA%A6%E5%AD%B8%E7%BF%92%E5%85%A5%E9%96%80-%E4%BA%8C-954b0e473d7f

Why LSTM cannot prevent gradient exploding  - Cecile Liu - Medium
https://medium.com/@CecileLiu/why-lstm-cannot-prevent-gradient-exploding-17fd52c4d772

[翻譯] Understanding LSTM Networks
https://hemingwang.blogspot.com/2019/09/understanding-lstm-networks.html

-----

[] 深入淺出 Deep Learning(三):RNN (LSTM)
http://hemingwang.blogspot.com/2018/02/airnnlstmin-120-mins.html

[] AI從頭學(一九):Recurrent Neural Network
http://hemingwang.blogspot.com/2017/03/airecurrent-neural-network.html 

-----

◎ 代碼實作

[] Sequence Models and Long-Short Term Memory Networks — PyTorch Tutorials 1.2.0 documentation
https://pytorch.org/tutorials/beginner/nlp/sequence_models_tutorial.html

Predict Stock Prices Using RNN  Part 1
https://lilianweng.github.io/lil-log/2017/07/08/predict-stock-prices-using-RNN-part-1.html

Predict Stock Prices Using RNN  Part 2
https://lilianweng.github.io/lil-log/2017/07/22/predict-stock-prices-using-RNN-part-2.html

# 230 claps
[] [Keras] 利用Keras建構LSTM模型,以Stock Prediction 為例 1 - PJ Wang - Medium
https://medium.com/@daniel820710/%E5%88%A9%E7%94%A8keras%E5%BB%BA%E6%A7%8Blstm%E6%A8%A1%E5%9E%8B-%E4%BB%A5stock-prediction-%E7%82%BA%E4%BE%8B-1-67456e0a0b

# 38 claps
[] LSTM_深度學習_股價預測 - Data Scientists Playground - Medium
https://medium.com/data-scientists-playground/lstm-%E6%B7%B1%E5%BA%A6%E5%AD%B8%E7%BF%92-%E8%82%A1%E5%83%B9%E9%A0%90%E6%B8%AC-cd72af64413a