Graph ML: How Graph Neural Networks and Transformers Work

Graph machine learning (graph ML) is a family of techniques that train neural networks directly on data structured as graphs, where entities are nodes and their relationships are edges. Unlike standard deep learning, which expects neatly gridded inputs like images or sequential inputs like text, graph ML works with the messy, irregular connectivity that shows up in molecules, social networks, recommendation platforms, transportation grids, and protein structures. The field has grown rapidly since around 2017, and its core architecture, the graph neural network (GNN), now underpins applications from drug discovery to physics simulation. What makes graph ML interesting is also what makes it tricky: the same structural flexibility that lets it model almost anything also introduces unique theoretical ceilings and engineering headaches that researchers are still working through.

How Graph Neural Networks Actually Work

A GNN takes in a graph and produces a transformed graph as output, updating the information stored at each node, edge, and optionally a global summary, without changing the graph’s wiring. The dominant framework for doing this is called message passing. In each layer of the network, every node collects information from its neighbors, combines that information with its own features, and updates itself. Stack several of these layers, and each node gradually absorbs context from nodes farther and farther away in the graph.

The “graph-in, graph-out” design means GNNs progressively transform the feature vectors loaded into nodes, edges, and global context while preserving the connectivity of the input graph.1Distill. A Gentle Introduction to Graph Neural Networks – Section: Graph Neural Networks This is different from, say, flattening a graph into a table of features and feeding it to a standard neural network. By respecting the topology, GNNs can learn patterns that depend on who is connected to whom, not just what features each entity has.

Convolutions, Attention, and Other Flavors

Not all GNNs pass messages the same way. The two biggest families are graph convolutional networks and graph attention networks, and understanding the difference helps explain why certain models work better for certain tasks.

Graph convolutional networks (GCNs) treat all neighbors roughly equally, weighting each neighbor’s contribution based on the graph’s structure. You can think of this as a smoothing operation: each node becomes a weighted average of itself and its neighbors. Interestingly, researchers have shown that whether you design these convolutions in the “spectral” domain (thinking about the graph’s mathematical frequencies) or the “spatial” domain (thinking about literal neighborhoods), the two approaches turn out to be theoretically equivalent under a general framework.2arXiv. Bridging the Gap Between Spectral and Spatial Domains in Graph Neural Networks – Section: Abstract That equivalence matters because it means insights from one camp apply to the other.

Graph attention networks (GATs) take a different approach: instead of weighting neighbors uniformly, they let each node learn to pay more attention to some neighbors than others. By stacking layers where nodes attend over their neighbors’ features, GATs can implicitly assign different importance weights to different nodes in a neighborhood without expensive matrix operations or needing to know the full graph structure in advance.3arXiv. Graph Attention Networks – Section: Abstract This selective focus often helps on tasks where not all connections are equally informative, such as citation networks where some references are more relevant than others.

More recent work has pushed attention-based models deeper. One challenge is that naively stacking many GAT layers causes problems (more on that shortly). Researchers have proposed techniques like layer-wise scaling scores that regulate attention coefficients based on how much overlap exists between neighborhoods at each layer, acting as a “soft” regulation to preserve information from nodes with large overlap.4Knowledge-Based Systems. Simple and deep graph attention networks – Section: Layer-wise self-adaptive GAT These engineering tricks make it practical to build deeper attention networks without the usual degradation.

The Expressiveness Ceiling

GNNs have a known theoretical limit on what they can distinguish. The standard message-passing GNN is provably no more powerful than a classical algorithm called the Weisfeiler-Lehman (WL) graph isomorphism test, which is a procedure for checking whether two graphs have the same structure. In practice, this means there are pairs of graphs with genuinely different structures that a standard GNN will treat as identical because their neighborhoods look the same at every step of message passing.5arXiv. How Powerful are Graph Neural Networks? – Section: Abstract

This ceiling has been confirmed from multiple angles. Popular variants like GCNs and GraphSAGE fall strictly below this upper bound, meaning they miss distinctions that even the WL test would catch. A model called the Graph Isomorphism Network (GIN) was designed to be provably as powerful as the WL test, making it the most expressive architecture within the standard message-passing class.6arXiv. How Powerful are Graph Neural Networks? – Section: Abstract Extensions to dynamic and attributed graphs have further characterized when this ceiling holds and when it can be pushed higher.7PubMed. Weisfeiler-Lehman goes dynamic: An analysis of the expressive power of Graph Neural Networks for attributed and dynamic graphs – Section: Abstract

This matters in practice because if your task requires telling apart certain symmetric substructures, a vanilla GNN will fail silently. It will give you a confident answer that happens to be wrong because it literally cannot see the difference. Higher-order GNNs and graph transformers are two strategies being explored to push past this limit, though both come with higher computational costs.

Over-Smoothing and Over-Squashing

Even below the expressiveness ceiling, GNNs face two problems that get worse as you make them deeper. Over-smoothing happens when stacking too many message-passing layers causes every node’s representation to converge toward the same value, washing out the local features that distinguish one node from another. Over-squashing is the opposite bottleneck: information from distant nodes gets compressed through narrow graph pathways, losing detail before it reaches where it is needed.

These two problems are not independent. Research has revealed that both are intrinsically tied to a property of the graph’s structure (specifically its spectral gap), and they sit on opposite ends of a trade-off. Sharpening node features to fight over-smoothing makes over-squashing worse, and relaxing the graph to ease over-squashing invites more smoothing.8arXiv. On the Trade-off between Over-smoothing and Over-squashing in Deep Graph Neural Networks – Section: Abstract This trade-off has been independently confirmed through a physics-inspired lens, where both issues can be characterized through the properties of filtering functions applied during message passing.9arXiv. Unifying over-smoothing and over-squashing in graph neural networks: A physics informed approach and beyond – Section: Abstract

One promising mitigation strategy is graph rewiring: selectively adding or removing edges to reshape the graph’s connectivity before or during training. For instance, algorithms based on discrete curvature measures can identify bottleneck edges and rewire them in a computationally efficient way while preserving important graph properties.10arXiv. On the Trade-off between Over-smoothing and Over-squashing in Deep Graph Neural Networks – Section: Abstract Another direction involves graph transformers, which sidestep local message passing entirely by letting every node attend to every other node, though this introduces its own scaling challenges.

Graph Transformers

Transformers reshaped natural language processing and computer vision, and researchers have been adapting them to graphs. The core idea is to replace (or augment) neighborhood-based message passing with global self-attention, letting each node look at every other node in the graph. This directly attacks over-squashing because information no longer has to hop through intermediate nodes to travel long distances.

The main challenge is encoding the graph’s structure into the transformer. In text, position in a sequence is straightforward. In a graph, there is no single natural ordering of nodes. Early approaches linearized the graph into a sequence and encoded absolute position, but this lost the precise relative relationships between nodes. Others encoded relative position using bias terms but missed the interplay between node features, edge features, and topology. More recent methods like Graph Relative Positional Encoding avoid linearization entirely, capturing both node-topology and node-edge interactions in the attention mechanism.11arXiv. GRPE: Relative Positional Encoding for Graph Transformer – Section: Abstract The field is still actively debating how much structural inductive bias to bake into graph transformers versus letting the attention mechanism figure it out from data.

Beyond Simple Graphs

Real-world data rarely fits neatly into a single type of node connected by a single type of edge. Graph ML has expanded to handle several richer structures.

Heterogeneous graphs contain multiple node types and edge types. A biomedical knowledge graph, for example, might have drug nodes, disease nodes, gene nodes, and protein nodes, connected by edges representing interactions like “treats,” “causes,” or “binds to.” These graphs can be formally represented with mapping functions that assign each node and edge to its specific type.12PubMed Central. Heterogeneous graph neural networks for link prediction in biomedical networks – Section: 2 Problem formulation Handling heterogeneity usually means learning separate transformation weights for each edge type or using type-aware attention, so the model does not conflate a “treats” edge with a “causes” edge.

Temporal graphs add a time dimension. Social networks gain and lose connections, financial transactions unfold over time, and communication patterns shift hour by hour. Continuous dynamic graph neural networks have emerged to learn fine-grained temporal representations from such data. One persistent challenge is that most methods only aggregate local neighborhood information and ignore global structural changes as the network evolves, which loses context about how broader patterns shift over time.13Information Sciences. Learning continuous dynamic network representation with transformer-based temporal graph neural network – Section: Abstract Newer architectures model these continuous-time dynamics rather than treating the graph as a series of discrete snapshots, capturing the fact that real-world graphs vary continuously.14arXiv. Continuous Temporal Graph Networks for Event-Based Graph Data – Section: Abstract

Knowledge graphs represent factual information as triples (entity-relation-entity) and are used in search engines, question-answering systems, and biomedical databases. A major application of graph ML here is link prediction: figuring out which connections are likely missing from an incomplete knowledge graph. Embedding-based methods that learn low-dimensional representations of entities and relations have shown strong performance on standard benchmarks for this task.15ACM Transactions on Knowledge Discovery from Data. Knowledge Graph Embedding for Link Prediction: A Comparative Analysis – Section: Abstract

Where Graph ML Is Already Deployed

The most visible industrial deployment of graph ML is in recommendation systems. Pinterest developed PinSage, a graph convolutional network trained on a graph with 3 billion nodes and 18 billion edges representing pins, boards, and their relationships. PinSage uses random walks and graph convolutions to generate item embeddings that encode both content features and graph structure. In A/B tests, it produced higher-quality recommendations than comparable deep learning and graph-based alternatives.16arXiv. Graph Convolutional Neural Networks for Web-Scale Recommender Systems – Section: Abstract Pinterest later extended this with MultiBiSage, which models diverse entity types (users, idea pins, creators) and heterogeneous interactions (add-to-cart, follow, long-click) across multiple bipartite graphs, learning higher-quality embeddings than the original single-graph approach.17arXiv. MultiBiSage: A Web-Scale Recommendation System Using Multiple Bipartite Graphs at Pinterest – Section: Abstract

In drug discovery, molecules are natural graphs: atoms are nodes, bonds are edges. Graph-based and sequence-based deep learning methods have been developed that achieve top-ranking performance on molecular property prediction benchmarks, including an AI Cures open challenge for COVID-19-related drug discovery where graph-based methods achieved the number one ranking in both major evaluation metrics.18Bioinformatics. Advanced graph and sequence neural networks for molecular property prediction and drug discovery – Section: Results The appeal of GNNs for chemistry is that the model operates on a representation that mirrors how chemists actually think about molecules, as connected structures rather than strings of text or flat fingerprints.19PubMed Central. Advanced deep learning methods for molecular property prediction – Section: GNN-based methods

Physics Simulation and Engineering

One of graph ML’s most impressive applications is learning to simulate physical systems. MeshGraphNets, for example, treat the computational mesh used in physics simulations as a graph, with mesh nodes as graph nodes and mesh edges as graph edges. The model passes messages along the mesh to predict how the system evolves at each time step. It can accurately predict dynamics across aerodynamics, structural mechanics, and cloth simulation, and can even adapt the mesh resolution during the simulation.20arXiv. Learning Mesh-Based Simulation with Graph Networks – Section: Abstract

This approach has been extended to urban-scale engineering problems. PIGNN-CFD, for example, uses a physics-informed graph neural network to predict wind fields around buildings, working directly on the irregular unstructured meshes that computational fluid dynamics simulations produce. Once trained, it runs one to two orders of magnitude faster than the traditional simulation it was trained on while maintaining consistent accuracy, and it can scale to predict wind fields of arbitrarily large urban scenes.21Building and Environment. PIGNN-CFD: A physics-informed graph neural network for rapid predicting urban wind field defined on unstructured mesh – Section: Abstract The speed gain matters because traditional fluid dynamics simulations of city blocks can take hours or days; a trained GNN can produce comparable results in seconds.

Geometric Deep Learning and Protein Structure

Standard GNNs treat graphs as abstract topological objects, ignoring the physical coordinates of nodes. In molecular biology, the 3D positions of atoms and residues matter enormously. Geometric deep learning extends graph ML with networks that respect the symmetries of 3D space: if you rotate or translate a protein, the network’s output should transform accordingly rather than changing unpredictably.

Equivariant graph neural networks have been applied to protein structure tasks with strong results. One architecture based on geometric vector perceptrons outperformed reference methods on three out of eight tasks in a structural biology benchmark and tied for first on two others, proving competitive with more computationally expensive approaches that use higher-order mathematical representations.22arXiv. Equivariant Graph Neural Networks for 3D Macromolecular Structure – Section: Abstract Another system, EnQA, leverages 3D-equivariant GNNs to estimate the accuracy of protein structural models by working with the structural features extracted from AlphaFold2 predictions.23PubMed Central. 3D-equivariant graph neural networks for protein model quality assessment – Section: Abstract These methods represent a shift in how the field thinks about structural prediction: instead of engineering handcrafted geometric features, the network learns to extract them directly from the 3D graph.

Generative Models for Molecules and Proteins

Graph ML is not limited to analyzing existing graphs. Generative graph models learn to create new graphs from scratch, which is transformative for scientific design tasks. Diffusion models, which became the state of the art for image generation, have been adapted to operate on graphs for generating novel molecules and protein structures.24arXiv. A Survey on Graph Diffusion Models: Generative AI in Science for Molecule, Protein and Material – Section: Abstract The idea is to start with noise and iteratively denoise it into a valid molecular graph, guided by desired properties like binding affinity or synthesizability. This falls under the broader category of AI-generated content in science, where the goal is not just prediction but design of new materials, drugs, and biological structures.

Self-supervised learning has also gained traction for graph ML, especially in domains where labeled data is scarce. In medical imaging, for instance, brain functional connectivity networks derived from fMRI scans can be augmented and used for contrastive learning without any diagnostic labels, with the self-supervised pretrained model then fine-tuned for detecting brain disorders.25PubMed Central. Self-supervised graph contrastive learning with diffusion augmentation for functional MRI analysis and brain disorder detection – Section: Abstract This approach is valuable because clinical datasets are often small and expensive to label, and graph contrastive learning allows the model to learn useful structural representations from unlabeled data first.

Scaling Challenges

Running GNNs on web-scale graphs with billions of nodes and edges is a serious engineering problem. The core issue is that message passing creates data dependencies between nodes: to compute the embedding for one node, you need embeddings from its neighbors, who in turn need embeddings from their neighbors, and so on. This neighborhood explosion makes even a two-layer GNN touch a potentially enormous fraction of the graph for each target node.

Two broad strategies have emerged. Full-graph training loads the entire graph and processes it at once, which works for smaller graphs but hits memory walls fast. Sample-based training, by contrast, samples a subgraph or a batch of neighborhoods for each training step, trading exact computation for scalability. Researchers have argued that sample-based training is the more promising path forward for scaling to very large graphs.26ACM SIGOPS Operating Systems Review. Scalable Graph Neural Network Training – Section: Abstract Distributed training across multiple machines introduces additional challenges around partitioning the graph, generating mini-batches, and managing the massive volume of feature data that must be communicated between machines.27arXiv. Distributed Graph Neural Network Training: A Survey – Section: Abstract Pinterest’s deployment of PinSage on a 3-billion-node graph remains one of the benchmarks for what production-grade graph ML engineering looks like.

Explainability and Adversarial Vulnerability

As graph ML moves into consequential domains like healthcare and finance, two concerns grow louder: can we understand why a GNN made a specific prediction, and can we trust that it will not be easily fooled?

Explainability for GNNs means identifying which nodes, edges, or subgraph patterns mattered most for a given prediction. Methods like PGExplainer have been adapted for subgraph-enhanced GNNs, producing human-interpretable explanations of graph classification decisions. Experiments on both real and synthetic datasets suggest these frameworks can successfully explain the decision process of a GNN on classification tasks.28arXiv. Explainability in subgraphs-enhanced Graph Neural Networks – Section: Abstract This is harder than explainability for images (where you can just highlight regions of the input), because graph explanations involve both which nodes and which connectivity patterns contributed to the output.

On the robustness side, GNNs have been shown to be vulnerable to adversarial attacks: small, carefully chosen modifications to a graph’s edges or node features can significantly change the model’s predictions. This poses real challenges for deploying GNNs in safety-critical scenarios like fraud detection or clinical decision support.29arXiv. Understanding the Robustness of Graph Neural Networks against Adversarial Attacks – Section: Abstract The vulnerability is partly structural: because each node’s representation depends on its neighbors, perturbing a single edge can ripple through the message-passing process and corrupt representations far from the point of attack. Defending against these attacks while preserving model performance is an active research area with no clean solution yet.