Deep Learning
| Authors | Goodfellow, Ian Bengio, Yoshua Courville, Aaron |
| Series | Adaptive Computation and Machine Learning [0.0] |
| Tags | Artificial Intelligence, Deep Learning, Neural Networks, Machine Learning |
| Publisher | MIT Press |
| Published | 18 gen 2016 |
| Date | 18 giu 2021 |
| Languages | eng |
| Identifiers | uri: https://www.deeplearningbook.org/, oclc: 1183962587, Amazon.com, isbn: 9780262337373 |
| Formats | EPUB, PDF |
Description
ref:14.1 pt. 1 ch. 5 Machine Learning Basics is a good overview of all the types of problems that ML can solve.
§6.4.1 (ref:16.203 ff.) is on the universal approximation theorem.
§8.1.3 "Batch and Minibatch Algorithms" (ref:18.36-40) is a good explanation of terminilogy (batch vs. minibatch), batch size, etc., and the advantages of (mini)batch sizes that are the entire train set.
This book is a bit (or excessively?) formal.
1.2 Historical Trends in Deep Learning
01/24/251.2 Historical Trends in Deep Learning : 65
Unless new technologies enable faster scaling, artificial neural networks will not have the same number of neurons as the human brain until at least the 2050s.
01/24/251.2 Historical Trends in Deep Learning : 66
Even today’s networks, which we consider quite large from a computational systems point of view, are smaller than the nervous system of even relatively primitive vertebrate animals like frogs.
2.8 Singular Value Decomposition
01/25/252.8 Singular Value Decomposition : 100
Every real matrix has a singular value decomposition, but the same is not true of the eigenvalue decomposition.
2.11 The Determinant
01/26/252.11 The Determinant : 105
The determinant is equal to the product of all the eigenvalues of the matrix.
3.1 Why Probability?
01/28/253.1 Why Probability? : 117
When we say that an outcome has a probability p of occurring, it means that if we repeated the experiment (e.g., drawing a hand of cards) infinitely many times, then a proportion p of the repetitions would result in that outcome. This kind of reasoning does not seem immediately applicable to propositions that are not repeatable. If a doctor analyzes a patient and says that the patient has a 40 percent chance of having the flu, this means something very different—we cannot make infinitely many replicas of the patient, nor is there any reason to believe that different replicas of the patient would present with the same symptoms yet have varying underlying conditions. In the case of the doctor diagnosing the patient, we use probability to represent a degree of belief, with 1 indicating absolute certainty that the patient has the flu and 0 indicating absolute certainty that the patient does not have the flu. The former kind of probability, related directly to the rates at which events occur, is known as frequentist probability, while the latter, related to qualitative levels of certainty, is known as Bayesian probability.
5 Machine Learning Basics
01/29/255 Machine Learning Basics : 195
Machine learning is essentially a form of applied statistics with increased emphasis on the use of computers to statistically estimate complicated functions and a decreased emphasis on proving confidence intervals around these functions;
5.1 Learning Algorithms
01/29/255.1 Learning Algorithms : 195
Mitchell (1997) provides a succinct definition: “A computer program is said to learn from experience E with respect to some class of tasks T and performance measure P, if its performance at tasks in T, as measured by P, improves with experience E.”
5.2 Capacity, Overfitting and Underfitting
01/30/255.2 Capacity, Overfitting and Underfitting : 219
because we have more parameters than training examples.
01/30/255.2 Capacity, Overfitting and Underfitting : 221
Statistical learning theory provides various means of quantifying model capacity. Among these, the most well known is the Vapnik-Chervonenkis dimension, or VC dimension. The VC dimension measures the capacity of a binary classifier. The VC dimension is defined as being the largest possible value of m for which there exists a training set of m different x points that the classifier can label arbitrarily.
01/30/255.2 Capacity, Overfitting and Underfitting : 225
Inductive reasoning, or inferring general rules from a limited set of examples, is not logically valid.
!?
5.5 Maximum Likelihood Estimation
02/04/255.5 Maximum Likelihood Estimation : 254
mean squared error is the cross-entropy between the empirical distribution and a Gaussian model.
5.6 Bayesian Statistics
02/04/255.6 Bayesian Statistics : 261
making the Bayesian approach simple to justify, while the frequentist machinery for constructing an estimator is based on the rather ad hoc decision to summarize all knowledge contained in the dataset with a single point estimate.
5.7 Supervised Learning Algorithms
02/04/255.7 Supervised Learning Algorithms : 268
kernel trick. The kernel trick consists of observing that many machine learning algorithms can be written exclusively in terms of dot products between examples
02/04/255.7 Supervised Learning Algorithms : 271
The category of algorithms that employ the kernel trick is known as kernel machines, or kernel methods
5.8 Unsupervised Learning Algorithms
02/04/255.8 Unsupervised Learning Algorithms : 278
To achieve full independence, a representation learning algorithm must also remove the nonlinear relationships between variables.
5.10 Building a Machine Learning Algorithm
02/04/255.10 Building a Machine Learning Algorithm : 290
Some models, such as decision trees and k-means, require special-case optimizers because their cost functions have flat regions that make them inappropriate for minimization by gradient-based optimizers.
5.11 Challenges Motivating Deep Learning
02/04/255.11 Challenges Motivating Deep Learning : 291
curse of dimensionality. Of particular concern is that the number of possible distinct configurations of a set of variables increases exponentially as the number of variables increases
6 Deep Feedforward Networks
02/05/256 Deep Feedforward Networks : 314
activation functions that will be used to compute the hidden layer values
6.2 Gradient-Based Learning
02/05/256.2 Gradient-Based Learning : 329
From this point of view, we can view the cost function as being a functional rather than just a function.
6.4 Architecture Design
02/07/256.4 Architecture Design : 360
the universal approximation theorem (Hornik et al., 1989; Cybenko, 1989) states that a feedforward network with a linear output layer and at least one hidden layer with any “squashing” activation function (such as the logistic sigmoid activation function) can approximate any Borel measurable function from one finite-dimensional space to another with any desired nonzero amount of error, provided that the network is given enough hidden units.
02/07/256.4 Architecture Design : 361
There is no universal procedure for examining a training set of specific examples and choosing a function that will generalize to points not in the training set.
One cant give what one doesn't have.
6.5 Back-Propagation and Other Differentiation Algorithms
02/07/256.5 Back-Propagation and Other Differentiation Algorithms : 371
The idea of computing derivatives by propagating information through a network is very general and can be used to compute values such as the Jacobian of a function f with multiple outputs.
02/10/256.5 Back-Propagation and Other Differentiation Algorithms : 394
This table-filling strategy is sometimes called dynamic programming.
02/10/256.5 Back-Propagation and Other Differentiation Algorithms : 399
The back-propagation algorithm described here is only one approach to automatic differentiation. It is a special case of a broader class of techniques called reverse mode accumulation.
6.6 Historical Notes
02/11/256.6 Historical Notes : 404
Feedforward networks can be seen as efficient nonlinear function approximators based on using gradient descent to minimize the error in a function approximation.
02/11/256.6 Historical Notes : 405
gradient descent was not introduced as a technique for iteratively approximating the solution to optimization problems until the nineteenth century (Cauchy, 1847).
02/11/256.6 Historical Notes : 405
The book Parallel Distributed Processing presented the results of some of the first successful experiments with back-propagation in a chapter (Rumelhart et al., 1986b) that contributed greatly to the popularization of back-propagation and initiated a very active period of research in multilayer neural networks.
Rumelhart invented bprop?
02/11/256.6 Historical Notes : 408
The half-rectifying nonlinearity was intended to capture these properties of biological neurons: (1) For some inputs, biological neurons are completely inactive. (2) For some inputs, a biological neuron’s output is proportional to its input. (3) Most of the time, biological neurons operate in the regime where they are inactive (i.e., they should have sparse activations).
7 Regularization for Deep Learning
02/11/257 Regularization for Deep Learning : 411
constraints on a machine learning model, such as adding restrictions on the parameter values
7.4 Dataset Augmentation
02/11/257.4 Dataset Augmentation : 432
One way to improve the robustness of neural networks is simply to train them with random noise applied to their inputs.
02/11/257.4 Dataset Augmentation : 433
Dropout, a powerful regularization strategy that will be described in section 7.12, can be seen as a process of constructing new inputs by multiplying by noise.
8.1 How Learning Differs from Pure Optimization
02/14/258.1 How Learning Differs from Pure Optimization : 494
many useful loss functions, such as 0-1 loss, have no useful derivatives
02/14/258.1 How Learning Differs from Pure Optimization : 494
negative log-likelihood of the correct class is typically used as a surrogate for the 0-1 loss
8.2 Challenges in Neural Network Optimization
02/14/258.2 Challenges in Neural Network Optimization : 509
A test that can rule out local minima as the problem is plotting the norm of the gradient over time.
8.7 Optimization Strategies and Meta-Algorithms
02/18/258.7 Optimization Strategies and Meta-Algorithms : 564
Very deep models involve the composition of several functions, or layers. The gradient tells how to update each parameter, under the assumption that the other layers do not change. In practice, we update all the layers simultaneously. When we make the update, unexpected results can happen because many functions composed together are changed simultaneously, using updates that were computed under the assumption that the other functions remain constant.
02/18/258.7 Optimization Strategies and Meta-Algorithms : 575
The student network is much deeper and thinner
esprit de finesse?
02/19/258.7 Optimization Strategies and Meta-Algorithms : 578
In practice, it is more important to choose a model family that is easy to optimize than to use a powerful optimization algorithm.
02/19/258.7 Optimization Strategies and Meta-Algorithms : 583
Curriculum learning was also verified as being consistent with the way in which humans teach (Khan et al., 2011): teachers start by showing easier and more prototypical examples and then help the learner refine the decision surface with the less obvious cases. Curriculum-based strategies are more effective for teaching humans than strategies based on uniform sampling of examples and can also increase the effectiveness of other teaching strategies (Basu and Christensen, 2013).
Aristotelian: more known → lesser known
9.8 Efficient Convolution Algorithms
02/20/259.8 Efficient Convolution Algorithms : 635
Convolution is equivalent to converting both the input and the kernel to the frequency domain using a Fourier transform, performing point-wise multiplication of the two signals, and converting back to the time domain using an inverse Fourier transform. For some problem sizes, this can be faster than the naive implementation of discrete convolution.
9.10 The Neuroscientific Basis for Convolutional Networks
02/20/259.10 The Neuroscientific Basis for Convolutional Networks : 639
Convolutional networks are perhaps the greatest success story of biologically inspired artificial intelligence.
02/20/259.10 The Neuroscientific Basis for Convolutional Networks : 642
grandmother cells have been shown to actually exist in the human brain, in a region called the medial temporal lobe (Quiroga et al., 2005). Researchers tested whether individual neurons would respond to photos of famous individuals. They found what has come to be called the “Halle Berry neuron,” an individual neuron that is activated by the concept of Halle Berry
Did Feser or Jaki mention this?
10.2 Recurrent Neural Networks
02/21/2510.2 Recurrent Neural Networks : 663
The recurrent neural network of figure 10.3 and equation 10.8 is universal in the sense that any function computable by a Turing machine can be computed by such a recurrent network of a finite size.
10.12 Explicit Memory
02/26/2510.12 Explicit Memory : 727
neural networks lack the equivalent of the working memory system that enables human beings to explicitly hold and manipulate pieces of information that are relevant to achieving some goal
11.2 Default Baseline Models
02/26/2511.2 Default Baseline Models : 741
Early stopping should be used almost universally.
11.4 Selecting Hyperparameters
03/01/2511.4 Selecting Hyperparameters : 746
The learning rate is perhaps the most important hyperparameter. If you have time to tune only one hyperparameter, tune the learning rate.
03/01/2511.4 Selecting Hyperparameters : 747
When the learning rate is too small, training is not only slower but may become permanently stuck with a high training error. This effect is poorly understood (it would not happen for a convex loss function).
12.1 Large-Scale Deep Learning
03/04/2512.1 Large-Scale Deep Learning : 778
Distributed asynchronous gradient descent remains the primary strategy for training large deep networks and is used by most major deep learning groups in industry
03/04/2512.1 Large-Scale Deep Learning : 786
more precision is required during training than at inference time
12.5 Other Applications
03/27/2512.5 Other Applications : 836
12.5.1.1 Exploration versus Exploitation
left off here
Bibliography
02/11/25Bibliography : 1265
Cauchy, A. (1847). Méthode générale pour la résolution de systèmes d’équations simultanées. In Compte rendu des sciences de l’Académie des sciences , pages 536–538.
Cauchy invented gradient descent.
1 Introduction, at least, is worth reading. Figure 1.11 on p. 23 is interesting.
The first ¶ of §5.2.1 The No Free Lunch Theorem is sloppy, philosophically (cf. Feser, Immortal Souls):
Learning theory claims that a machine learning algorithm can generalize well from a finite training set of examples. This seems to contradict some basic principles of logic. Inductive reasoning, or inferring general rules from a limited set of examples, is not logically valid. To logically infer a rule describing every member of a set, one must have information about every member of that set.
His last sentence is describing deduction, not induction.
An introduction to a broad range of topics in deep learning, covering mathematical and conceptual background, deep learning techniques used in industry, and research perspectives.
“Written by three experts in the field, Deep Learning is the only comprehensive book on the subject.”
—Elon Musk, cochair of OpenAI; cofounder and CEO of Tesla and SpaceX
Deep learning is a form of machine learning that enables computers to learn from experience and understand the world in terms of a hierarchy of concepts. Because the computer gathers knowledge from experience, there is no need for a human computer operator to formally specify all the knowledge that the computer needs. The hierarchy of concepts allows the computer to learn complicated concepts by building them out of simpler ones; a graph of these hierarchies would be many layers deep. This book introduces a broad range of topics in deep learning.
The text offers mathematical and conceptual background, covering relevant concepts in linear algebra, probability theory and information theory, numerical computation, and machine learning. It describes deep learning techniques used by practitioners in industry, including deep feedforward networks, regularization, optimization algorithms, convolutional networks, sequence modeling, and practical methodology; and it surveys such applications as natural language processing, speech recognition, computer vision, online recommendation systems, bioinformatics, and videogames. Finally, the book offers research perspectives, covering such theoretical topics as linear factor models, autoencoders, representation learning, structured probabilistic models, Monte Carlo methods, the partition function, approximate inference, and deep generative models.
Deep Learning can be used by undergraduate or graduate students planning careers in either industry or research, and by software engineers who want to begin using deep learning in their products or platforms. A website offers supplementary material for both readers and instructors.
Tesseract uses LSTM. DeepL uses CNNs [correction: DeepL uses modified Transformer].
