Friday, June 16, 2017

AI 從頭學(三0):Conv1

AI 從頭學(三0):Conv1

2017/06/16

網路上查了一些資料,主要是 inception 的輸出如何接起來 [1]-[4],接起來後,緊接著就是 1 x 1 convolution (conv1) [5], [6]。

Conv1 簡單說就是把圖1右邊壓成一張,然後把它第一張丟掉,補上新的在最後面,這樣可以壓成第二張新的。

這樣會引出一個問題,不重要的要擺前面,因為一張一張一直丟。

果真如此嗎?

其實第一張的資訊在第一次就有一個 w 留下了,在最後訓練完,w 也會有它應得的比重,第一張,其實從來沒被丟掉過...

----- 



Fig. 2.531m [2].



Fig. 3. 53m1 [3].



Fig. 4. 531m [4].

-----

References

[2] 531m
https://image.slidesharecdn.com/googlenet-insights-160920190051/95/googlenet-insights-10-638.jpg?cb=1484768593

[3] 53m1
https://image.slidesharecdn.com/googlenet-insights-160920190051/95/googlenet-insights-11-638.jpg?cb=1484768593

[4] 531m
https://www.e-sciencecentral.org/upload/ijfis/thumb/ijfis-17-026f5.gif

[5] AI從頭學(二八):Network in Network
http://hemingwang.blogspot.tw/2017/06/ainetwork-in-network.html

[6] AI從頭學(二九):GoogLeNet
http://hemingwang.blogspot.tw/2017/06/aigooglenet.html 

Monday, June 12, 2017

AI 從頭學(二九):GoogLeNet

AI 從頭學(二九):GoogLeNet

2017/06/12

前言:

GoogLeNet 裡面,1 x 1 convolution 很重要,也不難瞭解。至於它的理論基礎:稀疏網路的最佳化,這個才是核心。讀到現在,發現數學越來越多了!

Summary:

GoogLeNet 是 Google 向 LeNet 致敬的論文,本文討論 GoogLeNet 的架構 [1]。

構想來自大腦皮層的概念 [2],作法上利用大量的 1 x 1 convolution 降維以使超深的架構運算量降低到可以實現的水平 [3], [4],數學上的理論基礎則來自 [5]。

-----


http://celebnetworth.wiki/wp-content/uploads/2014/04/Google-Net-Worth.jpg
 
-----

Question

Q1: Multiple scales
Q2: Dimension reduction
Q3: Sparsity
Q4: Structure
Q5: Concatenation
Q6: Convolution
Q7: LRN, dropout and softmax

本文以回答七個問題的方式完成 GoogLeNet 的解說。

-----

Q1: Multiple scales

GoogLeNet 提出了 Inception 的概念作法是將 1 x 1, 3 x 3, 5 x 5 三種 convolutions 與 3 x 3 maxpooling 封裝在一個模組中,參考圖1.1a。然後重複使用這種模組進行特徵抽取。

這個概念是來自之前的視覺研究採用了類似大腦皮層的架構 [2],即皮層中對不同 scales 有不同的神經細胞對應,參考圖1.2。採用小的 filters 其中一個考量是運算量較小,另一個考量是由小的 filters 組成大型 filter 有較高的非線性。多種表現良好的大型網路都使用較小的 filters 組成。



Fig. 1.1. Inception module, p. 4 [1].



Fig. 1.2a. Multiple scales, p. 2 [1].



Fig. 1.2b. Multiple scales, p. 414 [2].

-----

Q2: Dimension reduction

由於擴展越到深層,運算量也大幅提升,因此利用 1 x 1 convolution 來降維 [3], [4],參考圖1.3。



Fig. 1.3. Dimension reduction, p. 2 [1].

-----

Q3: Sparsity

其理論基礎來自一個嚴格的數學證明:如果一個資料集的機率分布函數可以表現為一個大型的深層稀疏網路,則藉由分析前一層的激活,把有高度相關輸出的神經元群集在一起,則可以產生一個最佳化的網路拓樸。

這段話算是圖1.4的簡單翻譯,但簡單來說,就是不斷提煉,傳統上 CNN 即是如此 [6], [7],經由反覆的 convolution 與 pooling 萃取特徵。只是 Inception 的架構藉由 Hebbian principle 把不同 scale 的 filters 與 pooling 封裝在一起,然後重複使用。



Fig. 1.4. Sparsity, p. 2 [1].



Fig. 1.5. Hebbian principle. p. 3 [1].



Fig. 1.6a. The Inception architecture started out as a case study for assessing the hypothetical output and covering the hypothesized outcome, p. 3 [1].



Fig. 1.6b. The Inception architecture started out as a case study for assessing the hypothetical output and covering the hypothesized outcome, p. 3 [1].

-----



Fig. 2.1. Module 1, p. 6 [1].



Fig. 2.2. Module 2, p. 6 [1].



Fig. 2.3. Module 3, p. 6 [1].



Fig. 2.4. Module 4, p. 6 [1].



Fig. 2.5. Module 5, p. 6 [1].



Fig. 2.6. Module 6, p. 6 [1].

-----

Q4: Structure

圖2.1到2.6為 GoogLeNet 的細部分解。接下來透過圖3.5說明 GoogLeNet 的架構。

首先,#3x3 reduce 與 #5x5 reduce 分別代表 1 x 1 conv 的數目,降維用,參考圖3.1a。另外,pool proj 也是代表 1 x 1 conv 的數目,也是用來降維(接在 max pooling 之後)。



Fig. 3.1a. Number of 1 x 1 filters, p. 4 [1].



Fig. 3.1b. Pool proj, p. 4 [1].

-----

輸入的圖片大小是224x224,有 RGB 三個 channels,參考圖3.2。論文中並未特別說明處理方式。

經過第一層的 convolution, stride = 2 處理後,大小變為112x112,深度為64(有64個7x7 convolutions)。再經過 max pooling, stride = 2,大小變為 56x56x64。然後經過第二層的處理,大小變為28x28x193,這個會成為 Inception 3(a) 的輸入,參考圖3.3a。

-----

Q5: Concatenation

接下來是很重要的一部份,輸出要串在一起,參考圖3.3b。在 note * - **** 的接法分別是m531、531m、53m1、531m,並不一致。圖3.3c與3.3d則以論文圖表列出的順序進行一些討論。

首先,相關性高的應該在一起,所以 1 x 1 跟 3 x 3 接在一起,3 x 3 跟 5 x 5 接在一起。另外,不重要的應該要擺前面,因為 1 x 1 降維時會沿著前面一路丟掉 feature map。max pooling 的重要性高或低?以它數量較少,擺在後面比較適合,因為物以稀為貴。不過這些都只是很浮掠的猜測。順序到底會不會真的影響結果,也值得探討!




Fig. 3.2. Input image size, p. 4 [1].



Fig. 3.3a. Output size, p. 5 [1].



Fig. 3.3b. Single output vector, p. 5 [1].



Fig. 3.3c. Inception 3.



Fig. 3.3d. Inception modules.

-----

論文中提到:隨著層的提高,3 x 3 跟 5 x 5 convolution 的比例要提高,參考圖3.3e。我們可以看到,在 Inception 3、4、5 內往 top 移時,3 x 3 跟 5 x 5 個數變多,Ratio則未必,參考圖3.3d與3.3f。

在 Inception 3、4、5 的最高層,3b、4e、5b,1 x 1 的 convolutions 用來升維,作用跟 ResNet 中部分的 1 x 1 convolution 是一樣的 [4]。Inception 3、4、5 的最高層,其深度跟下一層接近,參考圖3.3d。




Fig. 3.3e. Ratio of convolutions, p. 5 [1].



Fig. 3.3f. Ratio of convolutions.

-----

Q6: Convolution

從 1 x 1 的降維,到 3 x 3 或 5 x 5 convolution,數目不一定成倍數增加,細節在論文中並沒提到,參考圖 3.4a。這有 LeNet 可以參考,如何從6張圖增加到16張圖,參考圖3.4b與3.4c。

至於 convolution 之後圖會變小,可以用 padding 的方式讓圖維持固定大小以便接起來傳到下一層,可以參考圖3.4d。



Fig. 3.4a. Convolutions, p. 5 [1].



Fig. 3.4b. Convolutions from 6 to 16, p. 7 [6].



Fig. 3.4c. Convolutions from 6 to 16, p. 8 [6].



Fig. 3.4d. Padding, p. 13 [8].

-----

Q7: LRN, dropout and softmax

圖3.5是較完整參數說明。LRN 與 dropout 可參閱之前寫的 AlexNet 介紹 [9], [10]。Softmax 則可參考 [11]。



Fig. 3.5. GoogLeNet incarnation of the Inception architecture, p. 5 [1].

-----

結論:

GoogLeNet 的設計很巧妙,值得細細品味!

-----

References

[1] 2015_Going deeper with convolutions

[2] 2007_Robust object recognition with cortex-like mechanisms

[3] 2014_Network in network

[4] AI從頭學(二八):Network in Network

[5] 2014_Provable bounds for learning some deep representations

[6] 1998_Gradient-Based Learning Applied to Document Recognition

[7] AI從頭學(一二):LeNet

[8] 2016_A guide to convolution arithmetic for deep learning

[9] 2012_Imagenet classification with deep convolutional neural networks

[10] AI從頭學(二七):AlexNet

[11] Lab DRL_04:Caffe網絡定義

Wednesday, May 31, 2017

AI 從頭學(二六):AlexNet

AI 從頭學(二六):AlexNet

2017/05/31

前言:
  
講解 AlexNet,主要是為 Caffe 的實作課程鋪路。

-----

Summary:

AlexNet [1] 扮演承先啟後的角色 [2]-[4],是現代深度學習網路的基礎。相關卷積神經網路的簡介,可以參考之前的文章 [6]-[10]。

本文介紹 AlexNet 使用的一些技巧,包含 LRN [11], [12]、ReLU [13]、Pooling [14]、Dropout [15]-[18]等。

-----

AlexNet 是現代大型卷積神經網路的濫觴 [1],自2012年發表以來,到筆者寫稿時間已經有超過12000的引用,粗估每個月有200次引用。我們大約可以說,LeNet [3] 引起機器學習界對於神經網路的關注,AlexNet 引起學術界對深度學習的關注,而 AlphaGo [4] 則引起全世界的人對人工智慧的好奇。

在以問題進入 AlexNet 之前,先說明何謂 Top-1 與 Top-5 [1]。Top-1 的錯誤率指的是圖片正確的標籤是否被準確預測。Top-5 的錯誤率則是,如果你預測的前五個選項包含正解,則此次的預測就算正確。

下面以七則問題來解釋 AlexNet。

-----

Question

Q1: Structure
Q2: GPU
Q3: Augmentation
Q4: LRN
Q5: ReLU
Q6: Pooling
Q7: Dropout

-----

Q1: Structure

AlexNet 的架構,大致可以想像成比較「巨大」的 LeNet,參考圖1.1 與 圖1.3。包含輸入層、卷積層、池化層、全連接層、輸出層等。另外有些層後面也會使用激活函數。2013年提出的 ZFNet則修改了一些參數,提升不少效能,參考圖1.2 與 圖1.4。

-----


Fig. 1.1. AlexNet [1].



Fig. 1.2. Modification of AlexNet [2].



Fig. 1.3. LeNet [3].



Fig. 1.4. Architectural changes of AlexNet [2].

-----

Q2: GPU

當時 AlexNet 率先使用 GPU 並獲得極大的成功。GPU 使用兩塊,而線條與顏色的 C1 filters 每次都會落在不同的 GPU 上 [7],是值得思考的問題。

-----

Q3: Augmentation

資料擴增有助於降低 overfitting,也就是提升測試時的準確性。由於在 CPU 而不是 GPU 上做,所以對於整體計算時間,可說沒有影響。

第一個方法是 horizontal reflections,也就是左右翻轉。
第二個方法是 altering the intensities of the RGB channels,使用 PCA 改變一點顏色。此處原文並未引用論文,雖然影響不小,對 Top-1 有1%,這邊我們就帶過。

-----

Q4: LRN

Local response normalization (LRN),局部反應標準化。

這點值得一提。

圖2.1的公式有點複雜,所以我們可以先看看它的靈感來源,從最早的圖2.3先看好了。這是一個 V1-like 的模型,也就是它模仿視覺皮層的第一層的反應。主要概念是減去平均值然後再標準化。圖2.3跟圖2.2都是處理對比,圖2.1由於沒有減去平均值,所以他說他是在處理亮度。裡面的參數是作者調出來的。



Fig. 2.1. Local response normalization, p. 4 [1].



Fig. 2.2. Local contrast normalization, p. 3 [11].



Fig. 2.3. Local input divisive normalization, p. 5 [12].

-----

Q5: ReLU

ReLU 是激活函數的一種,近年來比較多人使用。細節可以參考 [13]。

-----

Q6: Pooling

一般池化區不重疊 [14]。本論文使用重疊的方法,Top-1 錯誤率降低 0.4%,Top-5 降低 0.3%。

-----

Q7: Dropout

很快就講到 dropout 了。Dropout 是很簡單的概念,用在全連接層上。

參考圖3.1,在訓練時丟棄每層上面一些節點,每次丟棄的點是隨機選出,這樣等於每次是在不同的網路上訓練,可以有效避免 overfitting。

圖3.2講的是如果某個節點在訓練時出現的機率是p,由於測試時全部節點都保留,所以權重w要乘上p。

圖3.3是公式,這邊要先講解圖3.5比較清楚。

圖3.5左方是傳統的神經網路,如果我們先隨機產生一個值,1的機率是p,0的機率是(1-p),然後把值乘上權重再加總,這等於說有乘1被保留,乘0被丟棄,這樣就同時說明了圖3.3的公式與圖3.4的 Bernoulli distribution。

[16] 有深入的說明,[17], [18] 則是簡易的解說。Dropconnect [17] 則是 dropout 的特殊型 [16],雖然 [17] 聲稱它是一般化,我支持 [16] 的論點。

圖3.6是幾種激活函數搭配 drop 的比較,有意思的是 tanh 搭配 dropout 之後反而變差了!



Fig. 3.1. Dropout neural net model, p. 1930 [15].



Fig. 3.2. A unit at training time and at test time, p. 1931 [15].



Fig. 3.3. Formulae of a standard and dropout network, p. 1933 [15].



Fig. 3.4. Bernoulli distribution, p. 62 [16].



Fig. 3.5. Standard and dropout network, p. 1934 [15].



Fig. 3.6. Drop and activation function p. 5, [20].

-----

結論:

AlexNet 後續的 ZFNet、VGGNet 等都繼承了它的架構,即使是另一個路線的 GoogLeNet,也使用了它的技巧如 dropout 等 [8],可見 AlexNet 的成功!

-----

References

[1] 2012_Imagenet classification with deep convolutional neural networks

[2] 2014_Visualizing and understanding convolutional networks

[3] 1998_Gradient-Based Learning Applied to Document Recognition

[4] 2016_Mastering the game of Go with deep neural networks and tree search

[5] AI從頭學(一二):LeNet
http://hemingwang.blogspot.tw/2017/03/ailenet.html

[6] AI從頭學(一三):LeNet - F6
http://hemingwang.blogspot.tw/2017/03/ailenet-f6.html

[7] AI從頭學(二五):Kernel Visualizing
http://hemingwang.blogspot.tw/2017/05/aikernel-visualizing.html

[8] AI從頭學(一八):Convolutional Neural Network
http://hemingwang.blogspot.tw/2017/03/aiconvolutional-neural-network_23.html

[9] AI從頭學(一一):A Glance at Deep Learning
http://hemingwang.blogspot.tw/2017/02/aia-glance-at-deep-learning.html

[10] AI從頭學(二六):Aja Huang
http://hemingwang.blogspot.tw/2017/05/aiaja-huang.html

[11] 2009_What is the best multi-stage architecture for object recognition

[12] 2008_Why is real-world visual object recognition hard

[13] mAiLab_0005:Activation Function
http://hemingwang.blogspot.tw/2017/05/mailab0005activation-function.html

[14] mAiLab_0006:Pooling
http://hemingwang.blogspot.tw/2017/05/mailab0006pooling.html 

[15] 2014_Dropout, a simple way to prevent neural networks from overfitting

[16] 2016_Deep Learning

[17] 2013_Regularization of neural networks using dropconnect

[18] 2013_Maxout networks

Wednesday, May 24, 2017

AI 從頭學(二五):AlphaGo

AI 從頭學(二五):AlphaGo

2017/05/24

以四勝一負在2016年擊敗李世乭的 AlphaGo [1],在2017/5/23,再度以1/4目之差,小勝持黑的柯潔 [2]。AlphaGo 背後的靈魂人物,說是 Aja Huang 也不為過 [1]。

-----


Fig. 1. 黃士傑與AlphaGo對弈李世乭 [1]。



Fig. 2. 第 24 手「大飛」,第 54 手「斷」[2].

-----

Aja Huang 是台師大資工博士,碩士班跟博士班的題目都是圍棋 [1]。看完與柯潔的對奕之後,我特地找了他的博士論文來看 [3],參考文獻裡看似只有一篇跟深度學習有關 [4],其餘多屬強化學習的 MCTS [3]。

不過這篇 backpropagation [4] 並非我們熟悉的 BP 演算法 [5]。回過頭來再看 2016 年 DeepMind 發表的論文 [6],在 Huang 專門的 MCTS 之上,導入近年來最熱的深度學習 [7],Policy Network、Value Network、MCTS 三缺一不可,才是致勝的關鍵。

用 CNN 來下圍棋並非 DeepMind 首創 [8],早在1996年,即有學者提出用類神經網路下圍棋的概念 [9]。

[6]、[7]、[8]、[10] 一路追下去,[10] 這篇應該可以算是 AlphaGo alpha 版,裡面 CNN、TD、MCTS 都有。還不到十年,棋王就已不敵...

更早一點的研究,還有 [11]-[14]。

-----

References

[1] 創造AlphaGo的台灣「土博士」,他們眼中的黃士傑 _ 端傳媒 Initium Media
https://theinitium.com/article/20170116-taiwan-AlphaGo/

[2] 柯潔為何說「輸得沒脾氣」?8 個問題解讀人機大戰第一局 - INSIDE 硬塞的網路趨勢觀察
https://www.inside.com.tw/2017/05/23/analyzing-alphago-versus-ke-jie-round-1

[3] 應用於電腦圍棋之蒙地卡羅樹搜尋法的新啟發式演算法 SC Huang - 臺灣師範大學資訊工程研究所學位論文, 2011

[4] 2009_Backpropagation modification in Monte-Carlo game tree search

[5] AI從頭學(九):Back Propagation
http://hemingwang.blogspot.tw/2017/02/aiback-propagation.html

[6] 2016_Mastering the game of Go with deep neural networks and tree search

[7] 2015_Move evaluation in Go using deep convolutional neural networks

[8] 2014_Teaching deep convolutional neural networks to play Go

[9] 1996_The integration of a priori knowledge into a Go playing neural network

[10] 2008_Mimicking go experts with convolutional neural networks

[11] 2003_Local move prediction in Go

[12] 2003_Evaluation in Go by a neural network using soft segmentation

[13] 1996_The integration of a priori knowledge into a Go playing neural network

[14] 1994_Temporal difference learning of position evaluation in the game of Go

Tuesday, May 23, 2017

AI從頭學():Generative Adversarial Nets

AI從頭學():Generative Adversarial Nets

2017/05/23

前言:

施工中...

Summary:

Generative Adversarial Nets (GAN) [1] 自2014年推出以來,引 AI 界起很大的熱潮。GAN 的概念,是由 generative net (GN) 跟 discriminative net (DN) 相互對抗,最後 DN 不再能分辨 GN 生成的圖片是真是假,GN 就成功了(能產生以假亂真的圖片)。Adversarial 的觀念是新的,而 generative 跟 discriminative 的觀念則已超過十年 [2]。

有關 GAN 的簡單介紹,可以參考 [3], [4],較深入的討論,則可參考 [5]-[10]。[11], [12] 則有視覺化的訓練可以參考。

Log likelihood 是學習 GAN 的基礎 [13]-[16]。另外我們可以參考其他的論文來瞭解 GN [17]-[22]。最後則提供徹底掌握 GAN 所需的資料 [23]-[26]。

其實,以上資料並不足以徹底掌握 GAN。Wasserstein GAN [27]-[31] 才是完備的 GAN。而 Kullback–Leibler divergence [32] 與 Jensen–Shannon divergence [33] 算是基礎。

-----

Outline:

1. Formula
2. Generative Net
3. Deep Generative Models

本文重點有三:

1. GAN 公式
2. 生成網路構造
3. 瞭解 GAN 所需之相關資料

-----


Fig. 1.1a. Backpropagate derivatives through generative processes, p. 2 [1].



Fig. 1.1b. Random variable and probability distribution, p. 57 [23].



Fig. 1.1c. Expectation, p. 60 [23].



Fig. 1.1d. Normal distribution, also known as the Gaussian distribution, p. 63 [23].

-----


Fig. 1.2a. D and G play the following two-player minimax game with value function V (G;D), p. 3 [1].



Fig. 1.2b. The model can then be trained by maximizing the log likelihood, p. 2 [1].



Fig. 1.2c. Decomposition into the positive phase and negative phase of learning, p. 608 [23].




Fig. 1.3. Generative adversarial nets are trained by simultaneously updating the discriminative distribution, p. 4 [1].



Fig. 1.4. Minibatch stochastic gradient descent training of generative adversarial nets, p.4 [1].



Fig. 2.1a. DCGAN generator used for LSUN scene modeling, p. 4 [17].



Fig. 2.1b. A 100 dimensional uniform distribution Z, p. 4 [17].



Fig. 2.2. The architecture of the generator in Style-GAN, p. 324 [18].



Fig. 2.3. Text-conditional convolutional GAN architecture, p. 4 [19].



Fig. 2.4. A deconvnet layer (left) attached to a convnet layer (right), p. 822 [20].



Fig. 3.1. Deep generative models, p. vi [23].



Fig. 3.2. Deep learning taxonomy, p. 492 [24].



Fig. 3.3. Chapters 16-19, p. 671 [23].



Fig. 3.4. From section 3.14 to chapter 16, p. 560 [23].



Fig. 4.1. Fully-observed models [6].



Fig. 4.2. Transformation models [6].



Fig. 4.3. Latent bariable models [6].



Fig. 5.1. Probabilistic modeling of natural images, p. 563 [23], p. 8 [26].



Fig. 5.2. An illustration of the slow mixing problem in deep probabilistic models, p. 604 [23].



Fig. 5.3. Positive phase and negative phase, p. 611 [23].



Fig. 5.4. The KL divergence is asymmetric, p. 76 [23].



-----

References

1 GAN

[1] 2014_Generative adversarial nets

[2] 2007_Generative or discriminative, getting the best of both worlds

-----

2 GAN Internet

[3] 生成对抗式网络(Generative Adversarial Networks) – LHY's World
http://closure11.com/%E7%94%9F%E6%88%90%E5%AF%B9%E6%8A%97%E5%BC%8F%E7%BD%91%E7%BB%9Cgenerative-adversarial-networks/

[4] 能根據文字生成圖片的GAN,深度學習領域的又一新星 GigCasa 激趣網
http://www.gigcasa.com/articles/465963

[5] 深度学习与生成式模型 - Solomon1558的专栏 - 博客频道 - CSDN.NET
http://blog.csdn.net/solomon1558/article/details/52512459

[6] 生成式对抗网络GAN研究进展(一) - Solomon1558的专栏 - 博客频道 - CSDN.NET
http://blog.csdn.net/solomon1558/article/details/52537114

[7] 生成式对抗网络GAN研究进展(二)——原始GAN - Solomon1558的专栏 - 博客频道 - CSDN.NET
http://blog.csdn.net/solomon1558/article/details/52549409

[8] 生成式对抗网络GAN研究进展(三)——条件GAN - Solomon1558的专栏 - 博客频道 - CSDN.NET
http://blog.csdn.net/solomon1558/article/details/52555083

[9] 生成式对抗网络GAN研究进展(四)——Laplacian Pyramid of Adversarial Networks,LAPGAN - Solomon1558的专栏 - 博客频道 - CSDN.NET
http://blog.csdn.net/solomon1558/article/details/52562851

[10] 生成式对抗网络GAN研究进展(五)——Deep Convolutional Generative Adversarial Nerworks,DCGAN - Solomon1558的专栏 - 博客频道 - CSDN.NET
http://blog.csdn.net/solomon1558/article/details/52573596

[11] An introduction to Generative Adversarial Networks (with code in TensorFlow) – AYLIEN
http://blog.aylien.com/introduction-generative-adversarial-networks-code-tensorflow/

[12] Adverarial Nets
http://cs.stanford.edu/people/karpathy/gan/

-----

3 log likelihood

[13] 2009_Deep Boltzmann machines

[14] Likelihood function - Wikipedia
https://en.wikipedia.org/wiki/Likelihood_function

[15] 1.4 - Likelihood & LogLikelihood _ STAT 504
https://onlinecourses.science.psu.edu/stat504/node/27

[16] Chapter 18 Confronting the Partition Function
http://www.deeplearningbook.org/contents/partition.html

-----

4 Generator

[17] 2016_Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks

[18] 2016_Generative image modeling using style and structure adversarial networks

[19] 2016_Generative adversarial text to image synthesis

[20] 2014_Visualizing and understanding convolutional networks

[21] 2011_Adaptive deconvolutional networks for mid and high level feature learning

[22] 2016_A guide to convolution arithmetic for deep learning

-----

5 Goodfellow

[23] 2016_Deep Learning
https://github.com/HFTrader/DeepLearningBook/raw/master/DeepLearningBook.pdf

[24] 2016_Practical Machine Learning

[25] 2009_Learning multiple layers of features from tiny images

[26] 2011_Unsupervised models of images by spike-and-slab RBMs

-----

6 Goodfellow

[27] 2016_NIPS 2016 Tutorial, Generative Adversarial Networks

-----

7 Wasserstein GAN

[28] 令人拍案叫绝的Wasserstein GAN - 知乎专栏
https://zhuanlan.zhihu.com/p/25071913

[29] 生成式对抗网络GAN有哪些最新的发展,可以实际应用到哪些场景中? - 知乎
https://www.zhihu.com/question/52602529/answer/158727900

[30] 2017_Towards principled methods for training generative adversarial networks

[31] 2017_Wasserstein GAN

[32] 2017_ Improved training of Wasserstein GANs

[33] Kullback–Leibler divergence - Wikipedia
https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence

[34] Jensen–Shannon divergence - Wikipedia
https://en.wikipedia.org/wiki/Jensen%E2%80%93Shannon_divergence