Protein Representation Learning
Published:
Protein representation learning: How we teach machines to understand proteins
Proteins
Proteins are fundamental to virtually all biological processes. They fold into intricate 3D shapes, interact with other proteins or molecules, and are behind nearly every cellular machinery. The ability to understand protein sequence, structure, dynamics, interactions, and functions at the molecular level has tremendous implications for therapeutic development, drug discovery, precision medicine, and the rational design of biomolecular interventions.

Proteins exhibit complex sequence context and intricate structural organization which govern their thermodynamic stability and functional specificity. Individual proteins may adopt multiple conformational states and carry out diverse functions depending on their structural state and cellular environment.
By analyzing protein sequences and structures we can answer multiple open questions in the protein science, including:
- The stability of the protein
- The functional capacity of the protein
- The impact of mutations on protein stability and activity
- The quality of the modeled protein structure, whether obtained experimentally or predicted using computational tools and enable:
- The rational design of proteins that are geometrically and physico-chemically compatible with a desired function or target
These analyses have a wide range of applications in explaining how proteins regulate cellular mechanisms and in enabling the rational design of selective and effective therapeutic agents. Therefore, a comprehensive understanding of the molecular mechanisms behind proteins is essential.
For a machine to reason about proteins, whether predicting their structure, function, or interactions, the first step is deciding how to represent them. Protein representation learning focuses on creating dense, rich and low-dimensional representations in a way that they reflect real biological meaning while still being practical and machine-learning-friendly for computational models. Here I give a short, practical overview of main ways proteins are represented and how data-driven and deep learning-based methods represent proteins as a compact vectors (i.e. embeddings).
Why representation matters
At the same time, a protein is:
- A sequence of amino acids
- A folded 3D object with physical constraints
- A dynamic system shaped by evolution
- An interconnected network of atoms linked by physical interactions
Any representation highlights some of these aspects and ignores others. The choice of representation often determines what a model can (and cannot) learn. A protein (or a protein complex) can be represented in different ways:
- Sequence
- Multiple sequence alignment
- Volumetric map
- Gaussian surface
- Molecular surface
- 3D molecular graph
- Point clouds
- Residue interaction network

Sequence-based representations
The most basic representation is the amino acid sequence itself.

Early protein sequence representations relied on simple, interpretable methods such as one-hot encoding of individual amino acids, but these approaches scale poorly and fail to model long-range dependencies.
Inspired by advances in natural language processing, modern methods instead learn dense sequence embeddings by treating proteins like language (language of life) using models trained on massive corpora of millions of sequences with self-supervised objectives like masked token prediction.

These learned embeddings capture rich evolutionary, structural, and functional signals directly from sequence data.

See my previous post on this topic:
Protein Language Models
Here is a hands-on session on protein sequence representation learning 
Evolutionary representations
Evolution leaves strong signals in protein families, and multiple sequence alignments (MSAs) make these signals explicit by revealing which residues are conserved or co-evolved across related sequences. Conserved positions often indicate functional or structural importance, while correlated mutations can suggest interacting residues that are spatially close in the folded protein or at the interface of protein-protein interaction. As a result, many structure prediction methods, including early versions of AlphaFold, have relied heavily on MSAs to extract rich evolutionary information, although this comes at the cost of expensive computation and limited applicability to orphan proteins or adaptive immune receptors such as antibodies, nanobodies, and T cell receptors (TCRs) that lack sufficient homologs.

Structure-based representations
When experimentally determined or computationally predicted 3D structures are available, proteins can be represented with a richer and more intuitive way, a geometric way. Instead of just looking at sequences, we can represent proteins using atomic- or residue-level coordinates, which capture how the protein is actually arranged in the 3D space.

3 dimensional convolutional neural networks (3D-CNN) can represent an intricate arangement of atoms in a space as a compact vector useful for classification or regression tasks.

3D volumetric maps
These coordinate-based representations are often either inherently invariant representations or processed with equivariant neural networks. An example of inherently invariant representation is 3D volumetric maps that are centered and oriented based on the common scaffold of the backbone of amino acids.

Here is a hands-on session on how to learn 3D volumetric maps of protein structures 
On the other hand equvariant neural networks naturally respect rotational and translational symmetry. It means the protein’s properties don’t change just because we rotate or shift it in space.
See my previous posts on this topic:
Geometric Deep Learning and Equivariance and Invariance
Graph representation learning
Another common approach is to model the structure as a graph or point cloud, where nodes represent residues or atoms and edges are spatial proximity, covalent or molecular bonds, or other types of interaction.

Graph neural networks (GNNs) are particularly well suited to this setting and have become increasingly popular for structural modeling. Using different variants of the message-passing algorithm, GNNs iteratively update node and edge features, as well as spatial coordinates in the case of SE(3)-Transformers equivariant architecture. At each iteration, nodes aggregate information from their neighbors, progressively incorporating information from increasingly distant nodes. After several message-passing steps, each node representation captures both its local neighborhood and the broader topology and features of the graph. These node representations can then be pooled into a single compact vector that summarizes the entire graph.

Here is a hands-on session on graph representation learning
The big advantage of these geometric approaches is that they explicitly encode physical and spatial information. However, obtaining reliable 3D structures can be expensive and time-consuming, and even predicted structures may be noisy or incomplete.
Conclusion
Protein representation learning is at the intersection of biology, physics, and machine learning, aiming to turn the complexity of proteins into forms that computer models can actually understand and use. One of its biggest strengths is transferability: once a model has learned a meaningful representation, it can be reused across many downstream tasks with little additional supervision, saving both time and data. That said, there is no single “best” representation! There are only ones that fit particular questions and constraints. For instance, if we study highly flexible regions of a protein, a static, structure-based representation might miss important biophysical realities, whereas a sequence-based approach, implicitly capturing aspects of dynamics and function, or an explicitly dynamics-based representation could be more appropriate. We have only begun to scratch the surface; as models become more multimodal and grounded in biophysical principles, representations are shifting from hand-crafted features to deeply learned ones, bringing us closer to truly capturing the rich and nuanced complexity of proteins.




