AI Glossary

Deep Learning

Deep learning is machine learning using neural networks with multiple layers of computation that learn internal representations. The layers transform inputs into features useful for a task; “deep” refers to this layered structure.

· Updated · Chain of Thought

Model ArchitectureModel Training

A digit-recognition network takes an image’s pixel values, transforms them through hidden layers and produces scores for digit classes. Training changes weights so the output better matches labeled examples. The intermediate activations are learned representations, not a set of features a person must write individually.

Different architectures combine information differently: convolutional layers use local patterns, while attention lets positions exchange information according to their learned relevance. Depth lets later layers work with representations produced by earlier ones. There is no universal layer count at which every architecture suddenly becomes “deep.”

This matters because choosing a larger network is also choosing a training and evaluation burden. Depth alone does not establish useful performance, factual accuracy or robustness. Compare against simpler baselines and test on data outside training. For a small structured-data task, a linear model or tree ensemble may be worth trying before increasing network complexity.

Sources

Go deeper

  • 3Blue1Brown: But what is a neural network? video

    A visual introduction to weights, activations and layers, using handwritten digits.

From the conversation