Recent trends in scientific machine learning have hinted that physics-grounded training objectives may represent atomistic machine learning’s closest analog to the expressiveness of image reconstruction in computer vision. In this spirit, with Zatom-2, researchers at the University of Cambridge and beyond have studied a new way of pretraining atomistic generative models (disclosure: conflict of interest).
Zatom-2’s key idea is to unify atomistic tasks and data with a single tokenization scheme, consisting of: (1) the new atom1 tokenizer for chemical (molecule and material) data, which infers (atom) element types and lattice geometry parameters from 3D coordinate representations alone; and (2) atom14 for biological (protein) data, which infers amino acid (token) sequences from such geometric representations as well. In this setting, multitask pretraining, in particular with energy and force prediction tasks, empirically provides the best generative modeling configuration for molecule and material data, and also yields the best transfer learning results for protein generation tasks. If you’re curious to learn more about this work, its source code, documentation, and model weights are freely available.
Multiple sequence alignments are a key idea in AI for biology these days, thanks primarily to AlphaFold. What is not as well understood is how such multi-sequence alignments might be useful for learning rich representations in other domains of science. In CellMSA, the authors propose to refactor such alignments into the context of single-cell biology, to learn cellular representations for downstream tasks.
With an AlphaFold-inspired, multi-sequence-equipped architecture in hand, CellMSA delivers strong performance across a range of cellular prediction tasks, suggesting that the pairwise gene representations learned from its cell-gene sequence alignments are broadly useful for cellular biology. For those interested in learning more, CellMSA’s source code is available on GitHub.
How may physics inform the representations of generative models going forward? How might cellular representation learning benefit from enhanced physical (biological) supervision or transfer? In the future, do you think scaling laws might hold between low-level physical domains and higher-order biological domains?