When machine learns not to panic every time something rotates

4 minute read

Published:

We begin the journey beyond conventional CNNs with a basic question: how can we teach a machine to separate the true core identity of an object (e.g. a cat) from superficial variations (like its rotation or position)? In 2014, in their paper “Learning the Irreducible Representations of Commutative Lie Groups”, researchers Taco Cohen and Max Welling tackled this question by looking outside of traditional computer science and borrowing ideas from theoretical physics and group theory. Their proposal is built around three complementary axes:

Disentanglement and invariance

Think of disentangling as taking a complex, blended fruit smoothie and teaching a machine to mathematically separate it back out into pure, individual piles of strawberries, bananas, and yogurt. In computer science, disentanglement means breaking down complex data (e.g. an image) into its individual ingredients (factors of variation), like separating an image into distinct elements such as object’s identity, shape, location, orientation, scale or even lighting condition. Once these factors are neatly separated, achieving invariance, which is the ability to ignore irrelevant changes (e.g. rotation of an object), becomes simple. If you want to identify an object regardless of where it is, you just focus on the identity factor and ignore the position factor, allowing the system to stay consistent even when unimportant details change.

Weyl’s principle

Physicists use this principle to figure out what the true, fundamental elementary particles of a system are, separating the real particles from the raw, messy numbers a measuring device might produce which has no physical meaning. This principle is based on the idea of symmetry. Even we change how we measure something (altering its surface appearance or measured values), the core reality of what we are observing stays exactly the same. Cohen and Welling realized that the pixels in a digital image are just like those raw physical measurements, and we need a way to find the “elementary particles” of the image. Although this principle is primarily used in physics, the concept is fully abstract and independent of the data type (such as images, optical flow, or audio), which makes it highly valuable for representation learning. More broadly, it can be applied to any context where a clear notion of symmetry exists.

Irreducibility and Lie group theory

Lie group is a mathematical framework used to describe continuous, smooth transformations. Think of the difference between a digital clock that jumps rigidly from 1:00 to 1:01, versus an old analog clock where the second hand sweeps in a perfectly smooth, continuous circle. Conventional neural networks only understood the rigid, jumping steps of a pixel grid. Lie group theory, however, provided the mathematical language to describe a smooth, unbroken sweeping motion—like rotating an image by any microscopic fraction of a degree, rather than just flipping it 90 degrees. In this context, irreducibility means breaking down these complex, continuous transformations into their absolute most fundamental, independent building blocks. Once a transformation is irreducible, it cannot be divided any further into smaller invariant parts, making it the purest elementary component of the system’s symmetry.


Applying these three ideas to computer vision, they built a smart probabilistic model called Toroidal Subgroup Analysis (TSA). Instead of relying on the brute-force method of showing a neural network thousands of manually rotated images, TSA learns the underlying mathematical rules of the transformation itself. The model was trained simply by looking at pairs of images: an original image and a transformed version of it. By observing how the image changes from the first to the second frame, the model learns to mathematically disentangle the complex, jumbled web of pixels into simple, independent pieces called irreducible representations.

Figure below shows model’s posterior distribution over a rotation angle (denoted as \(s\)) for three distinct pairs of images. The top row shows the initial and rotated images, while the corresponding blue graphs below plot the model’s calculated probability across all possible rotation angles from \(0\) to \(2\pi\). For the first pair (a plus sign rotated into an ‘X’), the distribution displays four distinct peaks due to the shape’s four-fold rotational symmetry, indicating four equally probable rotation angles. The second pair features a circular ‘O’ shape, which results in a completely flat, uniform distribution because its continuous symmetry makes every rotation angle equally valid. The third pair involves a ‘T’ shape lacking rotational symmetry, yielding a graph with a single prominent peak that represents high model certainty for a specific angle.


Souce: Cohen, T., & Welling, M. (2014). Learning the irreducible representations of commutative lie groups. In International Conference on Machine Learning

This work laid the foundation for building neural networks that can naturally handle continuous changes, such as cyclic rotation.

Continue to Chapter 1: How to learn representations under p4 and p4m symmetry groups