Beyond CNN

2 minute read

Published:

A cat is a cat, but only if it doesn’t rotate 🙀

Representation learning powers deep learning. Instead of relying on hand-crafted features, deep learning models learn useful ones directly from data. A good representation learning should be meaningful, invariant to irrelevant changes, and disentangled into fundamental factors.

Convolutional Neural Networks (CNN) are a great example of representation learning. Instead of learning each part of an image separately, CNN reuse the same learned filters across all spatial locations, sharing weights throughout the image. This works because a cat is still a cat wherever it appears. A cat in the top-left of an image is the same cat in the bottom-right. This is translation equivariance by design, and it’s one of the core reasons CNNs work so well. Shifting the input image by a few pixels shifts the convolution output by the same amount! The network doesn’t have to re-learn anything.

However, conventional CNNs struggle with other geometric transformations, like rotations and reflections, often relying on brute-force data augmentation to learn them. When you rotate a handwritten “6” by 180°, it looks like a “9” to both humans and machines. A classical CNN has no built-in mechanism to link these two images. It must learn each orientation independently, which means:

  • Wasting model capacity on redundant rotated copies of every filter
  • Requiring heavy data augmentation (rotating training images) just to generalise
  • Still degrading significantly on unseen rotation angles

The figure below shows the same digit at 8 different orientations (rotation angles 0°, 45°, 90°, 135°, 180°, 225°, 270°, 315°). The pixel pattern changes completely at each angle, yet the semantic content (the identity of the digit) is identical. A classical CNN trained on upright digits must re-learn each orientation from scratch. Given this limitation, any model that bakes rotation invariance into its architecture rather than learning it from data has a fundamental advantage.

Digit at different orientations

Beyond the conventional CNN

Over the past decade, researchers have been trying to find a way to make neural networks naturally understand things like rotations and symmetry, instead of forcing them to learn these patterns from large amounts of extra data. In other words, they aim to mathematically “hard-bake” geometric symmetries directly into neural network architectures. I will explore methods in representation learning that embed built-in structural mechanisms to help models learn relationships between rotated versions of images. The works that I present here are only a small sample of a much larger body of research and are by no means exhaustive.