GINE experiment¶
This stage asks whether vector-valued edge state needs another sparse backend or can complete the existing message-passing algebra.
Composition¶
x_u [N, F] -- source gather --+
+-- ReLU -- target sum --+-- MLP --> y_v
e_uv [E, D] -- linear to F ----+ |
+-- (1 + eps) x_v
Graph.edge_values already gathers node state into original COO edge order.
Graph.sum_edges exposes the CSR segment sum already used privately by target
softmax. The layer adds only ordinary tinygrad linear maps and ReLU:
This follows GINE from Hu et al.. The pinned PyG implementation exposes the same edge projection, ReLU message, sum, self term, and caller-supplied update network. Tinymesh owns a fixed two-linear update instead of a module protocol. Graph-JEPA's official implementation uses GINE as its patch encoder, making MUTAG bond state the first live caller.
Decision¶
At revision
5a72bd7,
CPU and Metal produced identical float32 evidence:
| Measurement | Result |
|---|---|
| Initial MSE | 5.0000 |
| Update-weight gradient | [-11, -1] |
| Aligned-edge MSE after one step | 0.9225 |
| Reversed-edge MSE | 1.9225 |
| Erased-edge MSE | 1.8100 |
| Learned update weight | [0.55, 0.05] |
One SGD step reduced aligned loss by 81.6%. Reversing the two edge types made
loss 2.08x worse; erasing them made it 1.96x worse. The destination nodes
and their source node fields are otherwise identical, so the learned update
uses aligned edge identity rather than a node-only shortcut.
Focused tests independently match a host edge sum, return each destination gradient to its original COO edges, preserve leading axes through one sparse call, and cover empty edges. Correct edge reduction and GINE therefore enter the public API. Patch extraction, self-supervision, and representation quality remain experiments.
Limits¶
The public layer accepts one homogeneous [N, F] node tensor and one aligned
[E, D] edge tensor. Its epsilon is fixed, its edge projection is always
learned, and its update is a two-linear ReLU MLP. Heterogeneous graphs,
trainable epsilon, leading model axes, edge-conditioned attention, and model
quality are not claimed.