Sunday, August 30, 2026

Rapid Generation of Transition-State Conformer Ensembles via Constrained Distance Geometry

Stefan P. Schmid, Henrik Seng, Thibault Kläy, and Kjell Jorner (2026)
Highlighted by Jan Jensen

Here's another example of a tool that has made research life in the group significantly easier. racerTS takes a single TS structure and generates an ensemble of TS conformers. While this could already be done with CREST or GOAT, those methods were too expensive for the workflows we are developing in the group, so we basically ignored the problem. For minima, we used RDKit for conformational sampling, and we needed something similar for TSs—and that is exactly what racerTS provides.

racerTS is an RDKit-based pipeline for rapidly generating transition-state conformer ensembles using constrained distance geometry. It requires an XYZ structure of a TS guess, the total molecular charge, and the atom indices defining the reaction center; reactant or product SMILES may optionally be supplied. The coordinates and charge are converted into an RDKit Mol object. The reaction-center atoms and their automatically identified nearest neighbors collectively define the “frozen atoms.” ETKDG then generates conformers by repeatedly sampling a distance-bounds matrix while preserving the reaction-center geometry through positional and distance constraints. The resulting structures are refined with MMFF94, with UFF used as a fallback for unsupported atom types. Duplicate conformers are removed when their heavy-atom RMSD relative to a lower-energy conformer is below 0.125 Å. Structures more than 20 kcal/mol above the force-field minimum are discarded, and the surviving ensemble is written to a multi-structure XYZ file. For improved energetic ranking, the full workflow can additionally optimize the ensemble with GFN-FF while fixing the frozen atoms, rank the conformers using GFN2-xTB single-point energies, and repeat the pruning with a 6 kcal/mol window.

Across 20 diverse reactions, racerTS was approximately four times faster than CREST and 500 times faster than GOAT on a single CPU core. This comparison used GFN2-xTB//GFN-FF for racerTS and CREST, whereas GOAT was run at the more expensive GFN2-xTB level. racerTS, CREST, and GOAT all gave median activation-energy errors below 0.18 kcal/mol, with racerTS yielding an error of 0.17 kcal/mol. The xTB post-optimization step was important for racerTS’s accuracy; omitting it increased the speed but degraded the energy ranking. Overall, racerTS offers accuracy and conformational coverage comparable to those of CREST—though it is somewhat less exhaustive than GOAT—at a substantially lower computational cost, with occasional larger energy outliers for individual reactions.




This work is licensed under a Creative Commons Attribution 4.0 International License.



Friday, July 31, 2026

From SMILES Codes for Reactants and Products to Transition States With VeloxChem

Bastiaan van Hoorn, Patrick Norman and Mårten S. G. Ahlquist (2026)
Highlighted by Jan Jensen



Imagine if you could just draw a reaction in, say, ChemDraw, press a button, and get the TS (like shown above). Well now you pretty much can.**

Many algorithms construct TS guess structures by interpolating between reactant and product structures and therefore require that the atoms are in the same order. This can be very tedious to set up by hand and makes inclusion into automated workflows difficult.

In addition, I've found that the success of interpolation methods for biomolecular reactions depends heavily on how the reactant and product structures are aligned. Some methods try to solve this by rigid alignment, but that tends not to be sufficient for large, flexible molecules.

This paper offers automatic atom mapping and alignment, which is a massive step forward, in addition to a novel FF-based interpolation method that even includes some conformational sampling. And it works for "real-world" systems as you can see above.

Although VeloxChem can also optimise TSs, you can also interface it with other packages though a Python API. Here you can extract the interpolation path from VeloxChem (see this example notebook) and use that info (e.g. the TS guess or the aligned, atom-mapped reactant and product coordinates) as input to other packages if you want. 

This is a really great tool that we are already using quite a bit in my group.

**You have to provide SMILES strings, which you can easily get from ChemDraw. For transition metal complexes (TMCs) you have to give (non-mapped) coordinates instead, because VeloxChem uses RDKit. ChemDraw can't produce RDKit-compatible SMILES for TMCs (but Schrödinger Sketcher can). Also, the 3D structures that RDKit produces for TMCs are not very good. However, my group and others are working on fixing this, so that should be available soon. 



This work is licensed under a Creative Commons Attribution 4.0 International License.

Saturday, June 27, 2026

Developing Pharmaceutically Relevant Pd-Catalyzed C−N Coupling Reactivity Models Leveraging High-Throughput Experimentation

Seung Kyun Ha, Dipannita Kalyani, Michael S. West, Jessica Xu, Yu-hong Lam, Thomas Struble, Spencer Dreher, Shane W. Krska, Stephen L. Buchwald, and Klavs F. Jensen (2025)
Highlighted by Jan Jensen

Yield prediction is one of the most difficult and important challenges for machine learning applied to chemistry. This paper is a useful contribution because it provides a relatively large and systematic high-throughput dataset of ca. 4000 Pd-catalyzed C−N coupling reactions, spanning a wide variety of secondary amines and aryl bromides relevant to medicinal chemistry.

One important caveat is that the study does not address full reaction-condition optimization. All scope reactions are run using a single set of reaction conditions: one catalyst, one base, and one solvent system. The task is therefore better described as substrate-scope prediction under fixed conditions, rather than general prediction of reaction yield across arbitrary reaction conditions.

Significantly, the authors provide a useful reality check on the quality of yield data. For 32 repeated reactions, the measured product Liquid Chromatography Area Percent (LCAP) values correlate poorly, with (R^2 = 0.35). This experimental variability motivates their decision to treat the problem as binary classification rather than regression. A threshold of 20% product LCAP is chosen to define a “successful” reaction, and the repeated reactions are then consistent under this classification scheme in 27 cases. This supports a broader cautionary point: if yields from carefully controlled HTE experiments are already noisy at the level of absolute values, then predicting precise yield values from heterogeneous literature or web-scraped data is likely to be extremely difficult, and perhaps unrealistic in many settings.

The authors construct four different test sets to ascertain whether ML models can be used to extrapolate to unseen amines (amine OSS), aryl bromides (ArX OSS), or both (Both OSS) in addition to standard interpolation (DRS(n) where n is the percentage of the dataset used for training).

The authors compare several model classes and molecular representations, including random forests, decision trees, AdaBoost, fully connected neural networks, and MPNNs using Chemprop. Input features include one-hot encodings, Morgan fingerprints, quantum-mechanical fingerprints, molecular graphs, and combinations of these. Overall, the best models are usually either random forests with fingerprint-based descriptors, sometimes augmented with QM descriptors for the reacting components, or MPNNs. However, the optimal model and representation depend on the data split, which is itself an important result: there is no single universally best model for all generalization tasks.

The best model for each split is then used to design a corresponding prospective validation library of 96 reactions. The models are first retrained on the full experimental dataset using the best architecture, input features, and hyperparameters identified from the retrospective modeling. For the DRS validation library, the DRS25 settings are used. Each validation library is constructed so that approximately half of the reactions are predicted to give >20% LCAP and half are predicted to give <20% LCAP. The confidence threshold is >0.9 for the Amine OOS, ArX OOS, and DRS libraries, and >0.8 for the Both OOS library. For OOS amines or aryl halides, the selected substrates must also have a maximum Tanimoto similarity <0.7 to the corresponding substrates used in the model-building dataset. Thus, the validation libraries are not random samples of chemical space; they are enriched for reactions where the model is sufficiently confident.

The prospective validation results are impressive. For the Amine OOS library, the RF model gives 11 false positives, and no false negatives. For the ArX OOS library, the MPNN gives 3 false positives, and 2 false negatives. For the Both OOS library, the RF model performs less well but still gives useful enrichment, with most errors arising from false positives rather than false negatives. For the DRS25 library, the RF model performs extremely well, with essentially perfect precision and only one false negative. Overall, the models are especially good at avoiding false negatives, which is important in a medicinal chemistry setting because false negatives could cause chemists to discard reactions that would actually work.

Having said that, this study represents something close to a best-case scenario for reaction-outcome prediction. The dataset is large by the standards of synthetic chemistry, with around 4000 systematically generated reactions. The reactions are all run under the same conditions, reducing experimental heterogeneity. The positive rate is also relatively high: about 35% of the reactions exceed the 20% LCAP threshold. This makes the classification task easier than many realistic discovery settings where successful reactions are much rarer. Finally, because the dataset is large and the hit rate is high, the models can make a substantial number of high-confidence predictions, which enables the construction of balanced validation libraries with 50% predicted successes and 50% predicted failures. In smaller, noisier, or more imbalanced datasets, this level of prospective performance would likely be much harder to achieve.


This work is licensed under a Creative Commons Attribution 4.0 International License.

Sunday, May 31, 2026

Harnessing AtomisticSkills for Agentic Atomistic Research

Bowen Deng, Bohan Li, Matthew Cox, Hoje Chun, Juno Nam, Artur Lyssenko, Sathya Edamadaka, Jurgis Ruza, Xiaochen Du, Nofit Segal, Jesus Diaz Sanchez, Mingrou Xie, Ty Perez, Yu Yao, Miguel Steiner, Sauradeep Majumdar, Charles B. Musgrave III, Anirban Chandra, Abhirup Patra, Detlef Hohl, Connor W. Coley, Ju Li, Rafael Gómez-Bombarelli (2026)
Highlighted by Jan Jensen




AtomisticSkills is a hierarchical research framework in which skills encode reusable, mid-level scientific workflows, while tools provide low-level, type-checked computational operations that agents can reliably call to execute those workflows. While LLMs can in principle assemble such workflows from package documentation and first principles, in practice their performance degrades as context length grows. Put another way, trying to keep the manuals and execution details for RDKit, ORCA, and related tools in context at the same time is likely to increase hallucinations in the proposed workflow. Instead, complicated workflows are distilled by experts into SKILL.md files that outline how the tools are to be used, and these can be loaded into general-purpose coding agents such as Claude Code and Codex.

I especially like this last point. AtomisticSkills lets researchers use a tool they may already be familiar with, but apply it to new scientific problems. It looks like an interesting way to share robust workflows with non-experts. Take, for example, the installation of AtomisticSkills  itself: it basically amounts to downloading the repository and telling Codex to “Install AtomisticSkills according to its docs/setup.md guide,” after which the agent interactively guides the user through comparatively complicated steps such as creating environments, configuring API keys, and registering MCP servers.

For example, while we have made the xTB version of our EsNuEl workflow available through a web server, making the DFT versions available there was not practical. Installing it locally from the repo is of course possible, but perhaps a little intimidating for the target group of synthetic chemists. An approach like this might be more palatable: package the workflow as a skill, provide tested scripts and examples, and let a general-purpose coding agent guide the user through local setup and execution.


Thursday, April 30, 2026

Density Functional Theory Surrogate Enables Fast and Broad Computational Evaluation of Homogeneous Transition Metal Catalytic Energy Landscapes

Kevin P. Quirion, Wang-Yeuk Kong, Britton Stanley, Jyothish Joy, and Daniel H. Ess (2026)
Highlighted by Jan Jensen


It has been about 10 months since Meta FAIR released the Universal Models for Atoms, or UMA, machine-learning interatomic potentials. Since then, the first independent benchmarking studies have begun to appear, and this paper by Quirion and co-workers asks a very practical question: can UMA be used as a fast surrogate for DFT in homogeneous organometallic catalysis?

The authors examine seven catalytic/organometallic case studies taken from the literature, including Ir pincer alkane dehydrogenation, Rh hydroformylation, Ru olefin metathesis, Pd Buchwald–Hartwig amination, Cu-catalyzed difluorocarbene insertion, Ni asymmetric radical capture/reductive elimination, and a dinuclear Ni–Ni naphthyridine-diimine cycloaddition.

For literature geometries, they recompute reaction energies using ωB97M-V/def2-TZVPD single points, which is close to the level of theory that UMA is trained to reproduce. They then compare these values to UMA-S and UMA-M single-point energies, and in many cases also to UMA-optimized structures and energies. 

The headline result is encouraging: in most cases, UMA tracks ωB97M-V very well, often within a few kcal/mol and with good agreement in relative barriers and reaction-profile shapes. This is particularly impressive because the systems include different metals, oxidation-state changes, large ligands, charged species, and transition states. For routine conformer screening, preliminary mechanism mapping, or fast evaluation of many candidate catalysts, this suggests UMA could be genuinely useful.

There are, however, two important problem cases.

The first is the Cu-catalyzed difluorocarbene insertion, where the key issue is an open-shell singlet intermediate. UMA could not locate the TS1e transition state during optimization or NEB, gave unphysical conformational changes when optimizing the singlet 3e, and predicted the triplet state of 3e to be much lower than the singlet. At first glance this looks like a UMA failure, but ωB97M-V itself has similar problems with the singlet–triplet energetics. So this is not simply a machine-learning-potential problem. UMA is trained to reproduce ωB97M-V-like energies and forces; it should not be expected to magically repair failures of the underlying DFT reference method. The more specific concern is that UMA also has practical difficulties optimizing the open-shell singlet surface and locating the associated transition state. It was not tested whether ωB97M-V had the same problem.

The second problem case is the dinuclear Ni–Ni naphthyridine-diimine diene cycloaddition. Here UMA struggles with the relative spin states and barriers. In particular, it does not reproduce the same doublet/quartet ordering as ωB97M-V, and it overstabilizes some parts of the profile. This is perhaps less surprising because OMol25 did not include multinuclear transition-metal complexes, and the authors note that the naphthyridine-diimine ligand is not represented in the training set. Interestingly, the optimized geometries are not disastrous: UMA-S gives heavy-atom RMSDs of roughly 0.22 Å for the doublet and 0.36 Å for the quartet relative to the reported M06-L structures. So the failure is more severe for relative energetics and spin-state ordering than for generating plausible structures.

Overall, the study is a strong endorsement of UMA as a practical tool for organometallic mechanism work, provided it is used with the same caution one would apply to DFT. UMA appears especially promising for rapid conformer screening, approximate reaction-profile generation, and preoptimization before higher-level single-point calculations.

One unresolved issue is training-set overlap. The authors write that the OMol25 training database is so large that it “cannot be easily queried,” and that UMA does not provide an intrinsic nearest-neighbor or structure-comparison analysis for new inputs. That is a real limitation: if a benchmark system, or something very close to it, is already in the training data, the benchmark is much less informative about out-of-distribution generalization.

At the same time, the paper also states that the authors queried the dataset for the naphthyridine-diimine ligand and provide code in the Supporting Information. So the situation is somewhat unclear. The database may be inconvenient to search, but it does not seem impossible to search. For future UMA benchmark studies, it would be very useful to include at least a basic training-set check: for example, filtering OMol25 by metal, composition, charge, spin state, ligand identity, and local coordination environment. This would help distinguish cases where UMA is genuinely extrapolating from cases where it is interpolating within a familiar chemical neighborhood.

Wednesday, March 25, 2026

Stochastic tensor contraction for quantum chemistry

Jiace Suna and Garnet Kin-Lic Chan (2026)
Highlighted by Jan Jensen


What this paper lacks in terms of punchy title, it makes up for in content. I guess I would have gone with something like "Monte Carlo Meets Coupled Cluster: Slashing the Cost of CCSD(T)" or "Stochastic Tensor Contraction Pushes CCSD(T) Toward Mean-Field Cost". 

Anyway, tensor contraction is the algebraic core of much of quantum chemistry: large multidimensional arrays representing amplitudes and integrals are multiplied and summed over shared indices to produce energies and intermediates. It matters because these contractions set the scaling wall for methods like CCSD(T), where the formal cost rises far faster than Hartree–Fock. 

This study uses importance samplling to evaluate the tensor contraction, Importance sampling means drawing the most important terms in a sum more often than the unimportant ones, while reweighting so the final estimator stays unbiased. Here, Sun and Chan use it to evaluate high-order tensor contractions stochastically.

The headline result is that stochastic tensor contraction (STC) drives the scaling of CCSD(T) down dramatically: from the usual O(N^6) and O(N^7) down to O(N^4). In practice, water-cluster tests show very large FLOP reductions and wall-time crossovers at surprisingly small sizes. 

Figure 7 in the paper is the real selling point, because it compares against the incumbent approximate workhorse, DLPNO-CCSD(T), on 20 realistic molecules. STC is faster than DLPNO for every system in the set, with speedups ranging from 2.5× to 32×, while also delivering smaller errors than all DLPNO/Normal results and 15 of 20 DLPNO/Tight results. Just as importantly, the STC errors stay tightly clustered around the chosen target of 0.2 kcal/mol, whereas DLPNO errors vary much more from system to system. That makes STC look not just fast, but controllable. 

Table 3 sharpens that message. Averaged over the benchmark set, STC has a mean absolute error of 0.2 kcal/mol at a geometric mean runtime of 10.7 min, compared with 3.00 kcal/mol / 58 min for DLPNO/Normal, 0.70 kcal/mol / 159 min for DLPNO/Tight, and 773 min for exact CCSD(T). So the paper’s central claim is not merely better asymptotic scaling, but a roughly order-of-magnitude win in both time and error relative to state-of-the-art local correlation in this benchmark. 

One caveat: while the speed-up is undeniably impressive, another likely limiting factor is memory. The paper notes the use of density fitting “to reduce memory requirements,” but does not really quantify memory use or memory scaling in the same systematic way as FLOPs and wall time. Given that modern CC implementations are often limited as much by storage and movement of intermediates as by raw arithmetic, that omission stands out. 

Overall, this is prototype code, but very exciting prototype code. It will be very interesting to see whether this stochastic route can mature into something that genuinely displaces DLPNO-CCSD(T) as the default reduced-cost gold-standard method. Code: GitHub repository



This work is licensed under a Creative Commons Attribution 4.0 International License.



Saturday, February 28, 2026

Classical solution of the FeMo-cofactor model to chemical accuracy and its implications

Huanchen Zhai, Chenghan Li, Xing Zhang, Zhendong Li, Seunghoon Lee, and Garnet Kin-Lic Chan (2026)
Highlighted by Jan Jensen



The FeMo cofactor in nitrogenase enzymes is often mentioned as the killer application of quantum computing (QC) in chemistry. That is due to its complex electronic structure, which has made is difficult to model accurately. However, Chan and co-workers now claim to have computed the electronic energy to, by their estimate, chemical accuracy by conventional means.

They have done so by a series of calculations as indicated in the figure above. The CPU requirements are not given in detail, but the authors point out that no supercomputer was needed. 

Interestingly, the authors found that the ground state wavefunction is not inherently strongly multireference. Rather the main challenge is to identify the correct (mostly) single-reference state.

Where does that leave chemical applications of QC? For one thing, it moves the goalpost further back. The active space is the one typically used to estimate QC requirements, but it may have to be expanded to include MOs from the surrounding protein to accurately capture the chemistry, which would require even larger quantum computer. But that will be even further into the future with plenty of time for conventional approached to get there first. 

In my opinion, the case for QC-based quantum chemistry was never very strong, and this study is just another blow.

Wednesday, January 28, 2026

Predicting Enantioselectivity via Kinetic Simulations on Gigantic Reaction Path Networks

Yu Harabuchi, Ruben Staub, Min Gao, Nobuya Tsuji, Benjamin List, Alexandre  Varnek, and Satoshi Maeda (2026)
Highlighted by Jan Jensen



The automated predict of chemical reaction networks have thus far been limited to relatively small systems, typically with less than 50 atoms (including Hs) due to computational expense. This study goes significantly beyond this by studying a system with 228 atoms.

This is made possible by three things: 

1. While the system is big, the reaction is relatively simple, so the reaction network is relatively small. 

The reaction is an acid-catalysed cyclisation reaction involving a relatively small and chemically simple molecules. It is the (chiral) acid catalyst that contributes most of the atoms. The reaction itself has three steps: protonation of alkene group, intramolecular C-O bond formation on the activated alkene, deprotonation of the O to regenerate the catalyst. Most of the atoms are chemically inert, and there are 12 chemically active atoms (defined by the user). In all, the study identified 74 possible intermediates/products and only about half of those are chemically distinct if you ignore chirality. 

2. Cheap surrogate energy function

They use a Δ-ML approach that corrects the xTB energy and gradient to obtain better accuracy. The ML model is trained on-the-fly against DFT calculations. 

3. Massive computational resources 

In spite of 1 and 2 they this study required massive computational resources. They don't address this point specifically, other than to mention that it requires millions of gradient evaluations, but Maeda stressed this point during his talk at the WATOC last year. 

So this is not exactly a routine application. 

Wednesday, December 31, 2025

One step retrosynthesis of drugs from commercially available chemical building blocks and conceivable coupling reactions

Babak Mahjour, Felix Katzenburg, Emil Lammi, and Tim Cernak (2025)
Highlighted by Jan Jensen

What are important reactions that we currently can't perform? I asked myself this a few years ago and found that there were very few papers in the literature that addressed this. It turns out that I possessed the skills to figure it out for myself if I had only had the idea. The idea being that "the most valuable couplings would utilize the most abundant building blocks to form the most common types of bonds found in [a] target dataset."

As an example, the authors took a list of 9028 known drugs and asked how many could potentially be made in a single step from molecules in the MilliporeSigma catalog by hypothetical coupling reactions. The answer turns out to be 2573 (28%), which is a surprisingly large number. The most common reaction was the coupling of alkyl alcohols and alkyl amines, followed by alkyl acid-alkyl amine and alkyl acid-alkyl alcohols. All reaction for which there's no robust and generally applicable synthetic protocol, although AFAIK, although Zhang and Cernak took a stab at the alkyl acid-alkyl amine coupling. 

I really wish there were more papers like this. Identifying important questions to work on is just as important as solving them, and the latter is almost always a communal effort.


This work is licensed under a Creative Commons Attribution 4.0 International License.



Thursday, November 27, 2025

From Random Determinants to the Ground State

Hao Zhang and Matthew Otten (2025)
Highlighted by Jan Jensen


The paper introduces a method they call TrimCI that very efficiently finds a relatively small set of determinants that accurately describes strongly correlated systems. (Well, it actually works for any system, but the main advantage is for strongly correlated systems). 

Unlike most new correlation methods, this one is actually simple enough to describe in a few sentences. TrimCI starts by constructing a set of orthogonal (non-optimised!) MOs (e.g. by diagonalising the AO overlap matrix). From these MOs you construct a small number of random determinants (e.g.100), construct the wavefunction (i.e. construct the Hamiltonian matrix and diagonalise, as per usual). Then you compute all the Hamiltonian elements between this wavefunction ($H_{ij}$) and the remaining determinants and add determinants with sufficiently large |$H_{ij}$| to the wavefunction. Finally, there is the trimming step "which removes negligible basis states by first diagonalising randomised blocks of the core and then performing a global diagonalising step on the surviving set." And repeat.


The authors find that this approach converges much quicker than other similar methods, using many fewer determinants. Another big advantage is that the method does not require a single-determinant ground state as a starting point and is thus not sensitive to how much such a single-determinant deviates from the actual wavefunction.

So, what's the catch here? In order to be practically useful, we need to compute energy differences with mHa accuracy, and I did not see any TrimCI results for chemical systems where the energy had converged to that kind of accuracy. It's possible that error cancellation can help here, but that needs to be investigated. The authors do look at extrapolation, which looks promising, but needs to be systematically investigated. Yet another option is to use the (compact) TrimCI wavefunction as an ansatz for dynamic-correlation methods.

It's also not clear what AO basis set it used for some of these calculations (including the one shown above). I suspect small basis sets are used and even FCI energies with very small basis sets are of limited practical use. Are the TrimCI calculations on large systems still practical with more realistic basis sets?

Nevertheless, this seems like a very promising step in the right direction.


This work is licensed under a Creative Commons Attribution 4.0 International License.



Friday, October 31, 2025

Electron flow matching for generative reaction mechanism prediction

Joonyoung F. Joung, Mun Hong Fong, Nicholas Casetti, Jordan P. Liles,  Ne S. Dassanayake & Connor W. Coley (2025)
Highlighted by Jan Jensen



While the title says reaction mechanism prediction, it's really reaction mechanism-based reaction outcome prediction. The approach uses glow matching (a generalization of diffusion-based  approaches) to predict changes to the bond-electron (BE) matrix (basically connectivity matrix with the lone pair electron count on the diagonal), thus ensuring mass and charge conservation because changes in the BE matrix are constrained to sum to 0. The method is trained 1.4 million elementary reaction steps derived primarily from the USPTO dataset.  

Recursive predictions yield a complete reaction mechanism step by step, starting from the reactants. (I assume the products are defined as the state where no more changes are predicted.) The method is probabilistic so several different reaction outcomes are possible if the process is repeated, and ranked according to frequency. Another option is to use DFT calculation to rank the different mechanisms.
 
Like any ML method its applicability is tied to the training set. For example, of 22,000 reactions from patents reported in 2024 that were not assigned  a specific reaction class in the Pistachio dataset, the approach successfully recovered products in only 351 cases. However, the authors show that a new reaction class can be added with as few as 32 examples.

Thursday, September 25, 2025

Fundamental Study of Density Functional Theory Applied to Triplet State Reactivity: Introduction of the TRIP50 Dataset

William B. Hughes, Mihai V. Popescu, Robert S. Paton (2025)
Highlighted by Jan Jensen


While this paper presents an interesting and useful benchmark and dataset of barrier heights involving organic molecules in the triplet state, I am highlighting this paper for a different reason. 

While compiling the data set the authors "observed a  common tendency for triplet SCF calculations to converge non-Aufbau solutions, resulting in catastrophic predictions in both thermochemistry and activation energy  barriers and leading to errors as high as 26.4 kcal/mol." They go on to note that "Since such  errors  cannot be predicted a priori, manual inspection of spin densities for triplet-state calculations can be helpful to ensure the lowest triplet state has been converged with KS-DFT."

I remember the days when the SCF would routinely fail to converge even for simple singlet ground state molecules. So when they did converge and you got odd results, one of the first things you checked was the orbitals. But those days are long gone and I don't think it would occur to me now. I'd be much more likely to ascribe it to some deficiency in the functional. 

I now wonder how many of such wrong conclusions are scattered throughout the literature, especially for molecules with "funky" electronic structure, such as transition metal complexes.  Manual inspection of the MOs is not going to be a practical option for many of these studies, and SCF stability checks did not identify all problems! 

However, most QM packages have several options for the MO guess and it might be a good idea to use more than one of them and check whether they all converge to the same SCF solution. It'll be just like the old days.


This work is licensed under a Creative Commons Attribution 4.0 International License.



Friday, August 29, 2025

UMA: A Family of Universal Models for Atoms

Brandon M. Wood, Misko Dzamba1, Xiang Fu, Meng Gao, Muhammed Shuaibi, Luis Barroso-Luque, Kareem Abdelmaqsoud, Vahe Gharakhanyan, John R. Kitchin, Daniel S. Levine, Kyle Michel, Anuroop Sriram, Taco Cohen, Abhishek Das, Ammar Rizvi, Sushree Jagriti Sahoo, Zachary W. Ulissi, C. Lawrence Zitnick (2025)
Highlighted by Jan Jensen

I use xTB extensively in my research and I am often asked why don't switch to machine learning potentials (MLPs) instead. My answer has always been that they have too many limitations: limited atom types, no charged molecules, can't handle reactions, efficiency on CPUs, solvent effects, etc. I know these can be overcome by making bespoke MLPs, then it is not really a simple replacement for xTB, but a whole new workflow. 

However, the new UMA MLP from Meta seems to address all but one of my concerns (more on that below). UMA is trained to reproduce DFT energies and gradients calculated for nearly half a billion 3D structures, spanning molecules, surfaces, reactions, etc, containing atoms from virtually all of the periodic table. It is also possible to specify the charge and multiplicity and the cost seems to be comparable to xTB when running on CPUs, when interfaced with ORCA. So this is all very encouraging.

Two main questions remain. One is the accuracy, and by that I mean how faithfully it reproduces ωB97M-V/def2-TZVPD results (in the case of molecules) for molecules outside it's training set. AFAIK nothing is published yet, but encouraging results are being shared online.

The other main question is how to include implicit solvent effects. In cases where it is OK to optimise i the gas phase, one option is to compute the solvation energy with some other method and add it to the gas phase UMA results. Even if you do that at the DFT level, UMA has still saved you a lot if time. However, if the problem requires optimisation in solvent, then you have to use a faster method like xTB to compute the solvent effects on the gradient in order to get any time-savings. Depending on how well xTB does on the system of interest, this could "contaminate" the UMA results. Alternatively, a purely ML approach would basically amount to redoing UMA for molecules with continuum solvation included. Explicit solvation is fine in principle, but impractical for routine applications.

Anyway, before this is resolved there could be some fairly routine applications that still cannot be address satisfactorily with MLPs.



This work is licensed under a Creative Commons Attribution 4.0 International License.

Thursday, July 31, 2025

Computing solvation free energies of small molecules with first principles accuracy

 J. Harry Moore, Daniel  J. Cole, and Gábor Csányi (2025)
Highlighted by Jan Jensen


Local CCSD(T) has made it possible to reach near chemical accuracy for many real life applications. However, most real life applications happen in solution,  where the only realistic option is still continuum solvation methods, which, in general, do offer chemical accuracy (especially for charged systems). 

In principle, this can be fixed explicit solvation at the CCSD(T) level but of course the need for sampling makes this practically impossible at present. However, this paper by Moore and co-workers is a step in that direction. 

The idea is that ML potentials now approach chemical accuracy and are fast enough for proper sampling. The authors developed a MLP that is compatible with alchemical transformation (needed for sampling efficiency) and showed that experimental logP values (the difference in solvation energy between water and octanol) of drug-like compounds can be predicted within 0.65 kcal/mol (0.45 log units), i.e chemical accuracy.

Furthermore, the calculations took "only" about 4-7 days per molecule (octanol simulations converge slower than water) on a single node, containing either 8 NVIDIA A100, or 8 NVIDIA L40S, GPUs. While this is too slow for routine applications it is fast enough to create benchmark sets. This is great news since experimental solvation energies only are measured for relatively small molecules.

However, there are some caveats. The main one is that only neutral molecules where tested, since the MLP only was trained on neutral compounds, and it is not clear whether the same accuracy can be obtained for charged systems. 



This work is licensed under a Creative Commons Attribution 4.0 International License.

Friday, June 27, 2025

g-xTB: A General-Purpose Extended Tight-Binding Electronic Structure Method For the Elements H to Lr (Z=1–103)

Thomas Froitzheim, Marcel Müller, Andreas Hansen, and Stefan Grimme (2025)
Highlighted by Jan Jensen

This highlight is coming to you live from Oslo where Grimme presented the release of (a preliminary version of) g-xTB at WATOC yesterday. Grimme has been talking about this method for a few years now and many in the community (not least me) have been waiting for this moment with some anticipation.

To cut a long story (and paper!) short g-xTB offers mid-level DFT accuracy at semi-empirical cost (30-50% slower than GFN2-xTVB) including reaction energies and barrier heights. That's quite a statement! Structures of TM-complexes also seems to be improved.

This comes about a month after the release of the general ML-potential UMA and the paper offers some tantalising preliminary comparisons. For example, MAEs for reaction energies and forward barrier heights for the BH9 data shown above are 1.6 and 1.9 kcal/mol, respectively! However, UMA also appears to have some problems with some large extended systems, some non-covalent interactions, and some TM complexes. 

Given that UMA and g-xTB are roughly the same costs (on CPUs) it will be interesting to see how these methods will co-evolve over the coming years. 

Note that you can apply both methods through ORCA 



This work is licensed under a Creative Commons Attribution 4.0 International License.



Friday, May 30, 2025

Repurposing quantum chemical descriptor datasets for on-the-fly generation of informative reaction representations: application to hydrogen atom transfer reactions

Javier E. Alfonso-Ramos, Rebecca M. Neeser, Thijs Stuyver (2024)
Highlighted by Jan Jensen



If you have very little data, the single most useful thing you can do is find good descriptors. Sigman, Doyle, and others have shown this very nicely for reactivity predictions of transition metal containing catalysts, but there's less systematic work for other types of reactions. 

In this paper, Stuyver and co-workers suggest a descriptor set for barriers of hydrogen atom transfer (HAT) reactions that are based valence bond (VB) theory. In practise this translates to computing the bond dissociation energies (BDEs) without relaxing the geometry, and combining them with the BD free energies (BDFE, where ΔBDFE corresponds to ΔGrp). In addition, atomic Mulliken charges, spin densities, and buried volume are also added. All descriptors are predicted by surrogate models to avoid QM-based calculations.

Using these descriptors they get significantly better barrier predictions compared to fingerprint or graph convolution representation, even using simple models such as linear regression. Even the simple Bell-Evans-Polanyi model (a linear model based solely on ΔGrp) outperforms the models using fingerprints and graph convolution, with an R2 of 0.71 compared to 0.65 for graph convolution. For, comparison the R2s for the VB-based descriptors are 0.80-0.85, depending on the ML-model.

I wonder what other approximate chemical methods contain inspirations for new descriptors?


This work is licensed under a Creative Commons Attribution 4.0 International License.



Wednesday, April 30, 2025

Known Unknowns: Out-of-Distribution Property Prediction in Materials and Molecules

Nofit Segal, Aviv Netanyahu, Kevin P. Greenman, Pulkit Agrawal, Rafael Gómez-Bombarelli (2025)
Highlighted by Jan Jensen

Pat Walters recently wrote a blog post called Why Don’t Machine Learning Models Extrapolate? where he showed that ML models trained on a low-MW molecules cannot accurately predict the MW of larger molecules, while linear regression appears to do OK. Here I want to highlight a recently proposed ML method by Segal et al. that does appear to be able to extrapolate with decent accuracy. 

Before diving into the paper I let me try to explain why some traditional ML methods such as tree-based methods and NNs might struggle with extrapolation, using MW as the property of interest

Linear Regression. Let's start with linear regression (I'll skip the bias for simplicity),  

$$y_{pred} = X_1w_1+X_2w_2+... +X_nw_n$$

If $X_i$ is the number of atom $i$ in the molecule and $w_i$ is the atomic weight of atom $i$, then you have a perfect model that is able to extrapolate (assuming all atom types are represented in the training data). However, if X is a binary fingerprint then you obviously won't get good extrapolated values, so the molecular representation can also impacts whether a model can extrapolate accurately (I'll return to this point below). Pat uses count fingerprints, which contains the heavy-atom count, so the model "just" has to learn to infer the H-atom count from the remaining fragments and it does a pretty good job of that.

Random Forest. In RF each tree (i) predicts one of the $y$ values in the training set and the prediction is simply the average of $N$ trees:  

$$y_{pred} = (y_1 + y_2 + ... +y_N)/N $$

So the largest possible value for $y_{pred}$ is the maximum $y$ values in the training set $\max(y)$. Thus, RF is fundamentally incapable of any extrapolation.

LGBM. In gradient-boosted tree methods each tree $i$ tries to predicts the deviation from the mean of the training set ($lr$ is the learning rate, which is typically a small number):

$$ y_{pred} = \langle y \rangle +lr(\Delta y_1 + \Delta y_2 + ... + \Delta y_N) $$

The largest possible value of $\Delta y_i$ is the largest deviation from the mean found in the training set, but one can image a combination of $lr$, $N$, and $\Delta y_i$'s where test set-$y_{pred}$ could be larger than any $y$ value in the training set, but it is also unlikely that is it going to be significantly bigger, since it is tied to the mean of the training set. This is indeed what Pat finds.

Feed Forward Neural Networks. For an NN the prediction is a linear combination of the outputs from the activation functions from the last hidden layer:

$$y_{pred} = a_1 w_1+a_2 w_2+... +a_n w_n$$

If the activation function is something like sigmoid or tanh, then the maximum value of $a_i$ is 1, no matter the molecule. So the model is clearly going to have a very hard time extrapolating at all, in analogy with linear regression using binary fingerprints.

With the ReLU activation function the model has some theoretical chance of extrapolating. In fact, if the input is the simple atom-count vector discussed above one could arrive at a perfect model by setting most of the weights to 0, a few to 1, and some of the weights in the last layer to the atomic weights. However, the odds of finding that global error-minimum by traditional training techniques is very small. But note that it is a practical rather than a fundamental limitation, for this particular property and molecular representation. As before, if you use a binary fingerprint it will be impossible to make an method capable of extrapolating, no matter what activation function you use. However, a-ReLU NN can predict values larger than $\max(y)$, given the right molecular representation. Which brings us to ...

Graph Neural Networks. In GNNs (e.g. ChemProp) the molecular representation that is fed to the NN is some combination of atomic descriptor-vectors. The most popular way to combine the atomic vectors is to average them, which will result in molecular descriptor vectors that is roughly independent of molecular size. This will make accurate extrapolation harder, even when using ReLU.  

The approach by Segal et al. Segal et al. have developed a method called bilinear transduction, which optimises a bilinear function, with contributions from one model that depends on $X$ and one that depends on $\Delta X$, i.e. the difference between molecules in the test set.

$$ y_{pred} = f_1(\Delta X) g_1(X) + f_2(\Delta X) g_2(X) + ... + f_n(\Delta X) g_n(X)$$

The basic idea is that while the training set may not contain, say, molecules with 15 C atoms, it contains molecules with 10 C atoms, and molecule pairs that differ by 5 C atoms. So if you combine that knowledge you should be able to make a reasonable extrapolation to 15 C atoms. 

This turns out to work quite well as you can see from the rightmost panel here (for the FreeSolv data set)


This work is licensed under a Creative Commons Attribution 4.0 International License.

Monday, March 31, 2025

Designing Target-specific Data Sets for Regioselectivity Predictions on Complex Substrates

Jules  Schleinitz,  Alba  Carretero-Cerdán, Anjali  Gurajapu, Yonatan  Harnik, Gina  Lee, Amitesh  Pandey, Anat  Milo, and Sarah Reisman (2025)
Highlighted by Jan Jensen



Have you ever looked at a poorly performing ML model and thought: "Hmm, maybe I should make my training set smaller"? Me neither. 

Well, this paper shows an example where that actually works. In particular, they show examples where regioselectivity predictions for some molecules are improved by using only some of the available data. They show that if you start with a very small training set you get the wrong prediction of where in a molecule the reaction occurs and when you add more data the model often makes the right prediction eventually. However, if you keep adding data the model starts making a wrong prediction again! In other words, they get a much better classification model if the make a bespoke training set for each molecule.

This raises two important questions: 1) Does this apply to all datasets and properties? and 2) How do you figure out which data points to include in your training set for a particular molecule when you don't know the right answer?

If the answer to the first question is "no" (and it probably is), then how do we figure out when it this is a good strategy? (other than trial-an-error).  I suspect that the predictions of local properties (such as the reactivity of an atom) are more likely to benefit from bespoke training sets, compared to global properties such as solubility. But that is just a guess.

Another guess, it that this will apply mostly to small, inhomogeneous datasets. If so, we could easily generate bespoke models for each individual prediction on-the-fly if that would lead to better predictions. But we need to figure out the answers to question 2 first. 

I also think that if we can understand how additional data can hurt a models performance, it would give us some valuable insights into how ML models learn.



This work is licensed under a Creative Commons Attribution 4.0 International License.



Friday, February 28, 2025

GOAT: A Global Optimization Algorithm for Molecules and Atomic Clusters

Bernardo de Souza (2025)
Highlighted by Jan Jensen


If you want to predict accurate reaction energies and barrier heights of typical organic molecules then you are spending a significant portion of CPU time on the conformational search. While generating a large number of random starting structures often works OK for smaller molecules (with less than, say, 15 rotatable bonds) it fails for larger molecules where the odds of randomly generating the global minimum quickly approach zero. You thus need methods that focus the search arounds low energy regions of the PES.

In essence, the the algorithm walks up in some random direction, detects when a conformation barrier has been crossed, minimises the energy, and decides whether a new conformer has been found. New conformers are then included in the ensemble using a Monte-Carlo criteria with simulated annealing. The process is repeated until no new low energy conformers are found.

GOAT is better or very similar to CREST for all but one organic molecule tested. For organometallic complexes GOAT is better or similar to CREST, except for three cases where CREST fails in some way. For small molecules GOAT is a bit slower than CREST, but for large molecules GOAT is usually considerably faster.

GOAT is thus a valuable addition to the computation chemistry toolbox.


This work is licensed under a Creative Commons Attribution 4.0 International License.



Wednesday, January 29, 2025

Applying statistical modeling strategies to sparse datasets in synthetic chemistry

Brittany C. Haas, Dipannita Kalyani, Matthew S. Sigman (2025)
Highlighted by Jan Jensen



A few weeks ago I listened to an online talk by Matt Sigman and one thing that really surprised me was the remarkable success he had with very, very simple decision trees (sometimes just with a single node!). I wanted to learn more about it and luckily for me he has just published this excellent perspective.

First of all, these methods are used out of necessity because the data sets are relatively small (typically <1000 and often <100). So why do they work so well? Each application is on one particular organometallic catalyst and reaction type, and the experimental data usually comes from the same lab. The descriptors are usually obtained by DFT calculations and carry a lot of high quality chemical information. In fact, the authors make the point that if the approach fails they look for new descriptors rather than new ML methods. In fact you can view the single-node decision tree as an automated way of finding the single best descriptor in a collection.

You could of course ask: if you are doing DFT calculations, why not simply compute the barriers of interest rather then using ML. The problem is that problem like yield optimisation translate to very small changes (on the order of 0.5 - 2 kcal/mol) in barrier height, and even CCSD(T) is simply not up to the task for the systems of this size, where conformational sampling/Boltzmann averaging, explicit solvent effects, and anharmonic effects become important. 

Could this approach be applied to other problems, like drug discovery? The application would probably have to be something like IC50 prediction for lead optimisation where the molecules share a common core. One main difference to organometallic catalysis is that the activity of the catalyst is a function of a relatively localised molecular region compared to ligand-protein binding, which has contributions from most of the molecule. It thus seems difficult to find one or a few descriptors that capture the binding, and that can be computed reliably by DFT calculations - especially if no experimental protein-ligand structures are available. 

However, this paper suggest that it may be better to focus on developing such descriptors instead of new ML methods.



This work is licensed under a Creative Commons Attribution 4.0 International License.