← Back Deep Learning
Deep Learning

Description

ref:14.1 pt. 1 ch. 5 Machine Learning Basics is a good overview of all the types of problems that ML can solve.

§6.4.1 (ref:16.203 ff.) is on the universal approximation theorem.

§8.1.3 "Batch and Minibatch Algorithms" (ref:18.36-40) is a good explanation of terminilogy (batch vs. minibatch), batch size, etc., and the advantages of (mini)batch sizes that are the entire train set.

This book is a bit (or excessively?) formal.

1.2 Historical Trends in Deep Learning

01/24/251.2 Historical Trends in Deep Learning : 65

Unless new technologies enable faster scaling, artificial neural networks will not have the same number of neurons as the human brain until at least the 2050s.

01/24/251.2 Historical Trends in Deep Learning : 66

Even today’s networks, which we consider quite large from a computational systems point of view, are smaller than the nervous system of even relatively primitive vertebrate animals like frogs.

2.8 Singular Value Decomposition

01/25/252.8 Singular Value Decomposition : 100

Every real matrix has a singular value decomposition, but the same is not true of the eigenvalue decomposition.

2.11 The Determinant

01/26/252.11 The Determinant : 105

The determinant is equal to the product of all the eigenvalues of the matrix.

3.1 Why Probability?

01/28/253.1 Why Probability? : 117

When we say that an outcome has a probability p of occurring, it means that if we repeated the experiment (e.g., drawing a hand of cards) infinitely many times, then a proportion p of the repetitions would result in that outcome. This kind of reasoning does not seem immediately applicable to propositions that are not repeatable. If a doctor analyzes a patient and says that the patient has a 40 percent chance of having the flu, this means something very different—we cannot make infinitely many replicas of the patient, nor is there any reason to believe that different replicas of the patient would present with the same symptoms yet have varying underlying conditions. In the case of the doctor diagnosing the patient, we use probability to represent a degree of belief, with 1 indicating absolute certainty that the patient has the flu and 0 indicating absolute certainty that the patient does not have the flu. The former kind of probability, related directly to the rates at which events occur, is known as frequentist probability, while the latter, related to qualitative levels of certainty, is known as Bayesian probability.

5 Machine Learning Basics

01/29/255 Machine Learning Basics : 195

Machine learning is essentially a form of applied statistics with increased emphasis on the use of computers to statistically estimate complicated functions and a decreased emphasis on proving confidence intervals around these functions;

5.1 Learning Algorithms

01/29/255.1 Learning Algorithms : 195

Mitchell (1997) provides a succinct definition: “A computer program is said to learn from experience E with respect to some class of tasks T and performance measure P, if its performance at tasks in T, as measured by P, improves with experience E.”

5.2 Capacity, Overfitting and Underfitting

01/30/255.2 Capacity, Overfitting and Underfitting : 219

because we have more parameters than training examples.

01/30/255.2 Capacity, Overfitting and Underfitting : 221

Statistical learning theory provides various means of quantifying model capacity. Among these, the most well known is the Vapnik-Chervonenkis dimension, or VC dimension. The VC dimension measures the capacity of a binary classifier. The VC dimension is defined as being the largest possible value of m for which there exists a training set of m different x points that the classifier can label arbitrarily.

01/30/255.2 Capacity, Overfitting and Underfitting : 225

Inductive reasoning, or inferring general rules from a limited set of examples, is not logically valid.

!?

5.5 Maximum Likelihood Estimation

02/04/255.5 Maximum Likelihood Estimation : 254

mean squared error is the cross-entropy between the empirical distribution and a Gaussian model.

5.6 Bayesian Statistics

02/04/255.6 Bayesian Statistics : 261

making the Bayesian approach simple to justify, while the frequentist machinery for constructing an estimator is based on the rather ad hoc decision to summarize all knowledge contained in the dataset with a single point estimate.

5.7 Supervised Learning Algorithms

02/04/255.7 Supervised Learning Algorithms : 268

kernel trick. The kernel trick consists of observing that many machine learning algorithms can be written exclusively in terms of dot products between examples

02/04/255.7 Supervised Learning Algorithms : 271

The category of algorithms that employ the kernel trick is known as kernel machines, or kernel methods

5.8 Unsupervised Learning Algorithms

02/04/255.8 Unsupervised Learning Algorithms : 278

To achieve full independence, a representation learning algorithm must also remove the nonlinear relationships between variables.

5.10 Building a Machine Learning Algorithm

02/04/255.10 Building a Machine Learning Algorithm : 290

Some models, such as decision trees and k-means, require special-case optimizers because their cost functions have flat regions that make them inappropriate for minimization by gradient-based optimizers.

5.11 Challenges Motivating Deep Learning

02/04/255.11 Challenges Motivating Deep Learning : 291

curse of dimensionality. Of particular concern is that the number of possible distinct configurations of a set of variables increases exponentially as the number of variables increases

6 Deep Feedforward Networks

02/05/256 Deep Feedforward Networks : 314

activation functions that will be used to compute the hidden layer values

6.2 Gradient-Based Learning

02/05/256.2 Gradient-Based Learning : 329

From this point of view, we can view the cost function as being a functional rather than just a function.

6.4 Architecture Design

02/07/256.4 Architecture Design : 360

the universal approximation theorem (Hornik et al., 1989; Cybenko, 1989) states that a feedforward network with a linear output layer and at least one hidden layer with any “squashing” activation function (such as the logistic sigmoid activation function) can approximate any Borel measurable function from one finite-dimensional space to another with any desired nonzero amount of error, provided that the network is given enough hidden units.

02/07/256.4 Architecture Design : 361

There is no universal procedure for examining a training set of specific examples and choosing a function that will generalize to points not in the training set.

One cant give what one doesn't have.

6.5 Back-Propagation and Other Differentiation Algorithms

02/07/256.5 Back-Propagation and Other Differentiation Algorithms : 371

The idea of computing derivatives by propagating information through a network is very general and can be used to compute values such as the Jacobian of a function f with multiple outputs.

02/10/256.5 Back-Propagation and Other Differentiation Algorithms : 394

This table-filling strategy is sometimes called dynamic programming.

02/10/256.5 Back-Propagation and Other Differentiation Algorithms : 399

The back-propagation algorithm described here is only one approach to automatic differentiation. It is a special case of a broader class of techniques called reverse mode accumulation.

6.6 Historical Notes

02/11/256.6 Historical Notes : 404

Feedforward networks can be seen as efficient nonlinear function approximators based on using gradient descent to minimize the error in a function approximation.

02/11/256.6 Historical Notes : 405

gradient descent was not introduced as a technique for iteratively approximating the solution to optimization problems until the nineteenth century (Cauchy, 1847).

02/11/256.6 Historical Notes : 405

The book Parallel Distributed Processing presented the results of some of the first successful experiments with back-propagation in a chapter (Rumelhart et al., 1986b) that contributed greatly to the popularization of back-propagation and initiated a very active period of research in multilayer neural networks.

Rumelhart invented bprop?

02/11/256.6 Historical Notes : 408

The half-rectifying nonlinearity was intended to capture these properties of biological neurons: (1) For some inputs, biological neurons are completely inactive. (2) For some inputs, a biological neuron’s output is proportional to its input. (3) Most of the time, biological neurons operate in the regime where they are inactive (i.e., they should have sparse activations).

7 Regularization for Deep Learning

02/11/257 Regularization for Deep Learning : 411

constraints on a machine learning model, such as adding restrictions on the parameter values

7.4 Dataset Augmentation

02/11/257.4 Dataset Augmentation : 432

One way to improve the robustness of neural networks is simply to train them with random noise applied to their inputs.

02/11/257.4 Dataset Augmentation : 433

Dropout, a powerful regularization strategy that will be described in section 7.12, can be seen as a process of constructing new inputs by multiplying by noise.

8.1 How Learning Differs from Pure Optimization

02/14/258.1 How Learning Differs from Pure Optimization : 494

many useful loss functions, such as 0-1 loss, have no useful derivatives

02/14/258.1 How Learning Differs from Pure Optimization : 494

negative log-likelihood of the correct class is typically used as a surrogate for the 0-1 loss

8.2 Challenges in Neural Network Optimization

02/14/258.2 Challenges in Neural Network Optimization : 509

A test that can rule out local minima as the problem is plotting the norm of the gradient over time.

8.7 Optimization Strategies and Meta-Algorithms

02/18/258.7 Optimization Strategies and Meta-Algorithms : 564

Very deep models involve the composition of several functions, or layers. The gradient tells how to update each parameter, under the assumption that the other layers do not change. In practice, we update all the layers simultaneously. When we make the update, unexpected results can happen because many functions composed together are changed simultaneously, using updates that were computed under the assumption that the other functions remain constant.

02/18/258.7 Optimization Strategies and Meta-Algorithms : 575

The student network is much deeper and thinner

esprit de finesse?

02/19/258.7 Optimization Strategies and Meta-Algorithms : 578

In practice, it is more important to choose a model family that is easy to optimize than to use a powerful optimization algorithm.

02/19/258.7 Optimization Strategies and Meta-Algorithms : 583

Curriculum learning was also verified as being consistent with the way in which humans teach (Khan et al., 2011): teachers start by showing easier and more prototypical examples and then help the learner refine the decision surface with the less obvious cases. Curriculum-based strategies are more effective for teaching humans than strategies based on uniform sampling of examples and can also increase the effectiveness of other teaching strategies (Basu and Christensen, 2013).

Aristotelian: more known → lesser known

9.8 Efficient Convolution Algorithms

02/20/259.8 Efficient Convolution Algorithms : 635

Convolution is equivalent to converting both the input and the kernel to the frequency domain using a Fourier transform, performing point-wise multiplication of the two signals, and converting back to the time domain using an inverse Fourier transform. For some problem sizes, this can be faster than the naive implementation of discrete convolution.

9.10 The Neuroscientific Basis for Convolutional Networks

02/20/259.10 The Neuroscientific Basis for Convolutional Networks : 639

Convolutional networks are perhaps the greatest success story of biologically inspired artificial intelligence.

02/20/259.10 The Neuroscientific Basis for Convolutional Networks : 642

grandmother cells have been shown to actually exist in the human brain, in a region called the medial temporal lobe (Quiroga et al., 2005). Researchers tested whether individual neurons would respond to photos of famous individuals. They found what has come to be called the “Halle Berry neuron,” an individual neuron that is activated by the concept of Halle Berry

Did Feser or Jaki mention this?

10.2 Recurrent Neural Networks

02/21/2510.2 Recurrent Neural Networks : 663

The recurrent neural network of figure 10.3 and equation 10.8 is universal in the sense that any function computable by a Turing machine can be computed by such a recurrent network of a finite size.

10.12 Explicit Memory

02/26/2510.12 Explicit Memory : 727

neural networks lack the equivalent of the working memory system that enables human beings to explicitly hold and manipulate pieces of information that are relevant to achieving some goal

11.2 Default Baseline Models

02/26/2511.2 Default Baseline Models : 741

Early stopping should be used almost universally.

11.4 Selecting Hyperparameters

03/01/2511.4 Selecting Hyperparameters : 746

The learning rate is perhaps the most important hyperparameter. If you have time to tune only one hyperparameter, tune the learning rate.

03/01/2511.4 Selecting Hyperparameters : 747

When the learning rate is too small, training is not only slower but may become permanently stuck with a high training error. This effect is poorly understood (it would not happen for a convex loss function).

12.1 Large-Scale Deep Learning

03/04/2512.1 Large-Scale Deep Learning : 778

Distributed asynchronous gradient descent remains the primary strategy for training large deep networks and is used by most major deep learning groups in industry

03/04/2512.1 Large-Scale Deep Learning : 786

more precision is required during training than at inference time

12.5 Other Applications

03/27/2512.5 Other Applications : 836

12.5.1.1 Exploration versus Exploitation

left off here

Bibliography

02/11/25Bibliography : 1265

Cauchy, A. (1847). Méthode générale pour la résolution de systèmes d’équations simultanées. In Compte rendu des sciences de l’Académie des sciences , pages 536–538.

Cauchy invented gradient descent.


1 Introduction, at least, is worth reading. Figure 1.11 on p. 23 is interesting.

The first ¶ of §5.2.1 The No Free Lunch Theorem is sloppy, philosophically (cf. Feser, Immortal Souls):

Learning theory claims that a machine learning algorithm can generalize well from a finite training set of examples. This seems to contradict some basic principles of logic. Inductive reasoning, or inferring general rules from a limited set of examples, is not logically valid. To logically infer a rule describing every member of a set, one must have information about every member of that set.

His last sentence is describing deduction, not induction.


An introduction to a broad range of topics in deep learning, covering mathematical and conceptual background, deep learning techniques used in industry, and research perspectives.

“Written by three experts in the field, Deep Learning is the only comprehensive book on the subject.”
—Elon Musk, cochair of OpenAI; cofounder and CEO of Tesla and SpaceX

Deep learning is a form of machine learning that enables computers to learn from experience and understand the world in terms of a hierarchy of concepts. Because the computer gathers knowledge from experience, there is no need for a human computer operator to formally specify all the knowledge that the computer needs. The hierarchy of concepts allows the computer to learn complicated concepts by building them out of simpler ones; a graph of these hierarchies would be many layers deep. This book introduces a broad range of topics in deep learning.

The text offers mathematical and conceptual background, covering relevant concepts in linear algebra, probability theory and information theory, numerical computation, and machine learning. It describes deep learning techniques used by practitioners in industry, including deep feedforward networks, regularization, optimization algorithms, convolutional networks, sequence modeling, and practical methodology; and it surveys such applications as natural language processing, speech recognition, computer vision, online recommendation systems, bioinformatics, and videogames. Finally, the book offers research perspectives, covering such theoretical topics as linear factor models, autoencoders, representation learning, structured probabilistic models, Monte Carlo methods, the partition function, approximate inference, and deep generative models.

Deep Learning can be used by undergraduate or graduate students planning careers in either industry or research, and by software engineers who want to begin using deep learning in their products or platforms. A website offers supplementary material for both readers and instructors.


Tesseract uses LSTM. DeepL uses CNNs [correction: DeepL uses modified Transformer].

Unnamed image