Bastiaan van Hoorn, Patrick Norman and Mårten S. G. Ahlquist (2026)
Highlighted by Jan Jensen

This work is licensed under a Creative Commons Attribution 4.0 International License.
Important recent papers in computational and theoretical chemistry
A free resource for scientists run by scientists
Bastiaan van Hoorn, Patrick Norman and Mårten S. G. Ahlquist (2026)
Highlighted by Jan Jensen

Joonyoung F. Joung, Mun Hong Fong, Jihye Roh, Zhengkai Tu, John Bradshaw, and Connor Wilson Coley (2024)
Highlighted by Jan Jensen
Figure 1 from the paper. (c) the authors 2024
If you don't follow this particular subject, you might be surprised to learn that there isn't a large database of elementary reactions relevant to organic synthesis. Until now.
While datasets such as Reaxys contain millions of reactions, they are typically multistep reactions. That's mostly fine for training retrosynthesis algorithms (although the authors present discuss some disadvantages), but presents a challenge if you want to use more physically based methods such as QM to predict reactivity. For example, while there are some databases of transition states (TSs) they are typically for synthetically irrelevant reactions. So, for example, while very promising methods have been developed for TS prediction, they have been trained on these datasets and are thus have limited practical applicability to synthesis.
This paper is an important step towards fixing this:
"We identified the most popular 86 reaction types in Pistachio and curated elementary reaction templates (Figure 1c) for each of these 86 reaction types with 175 different reaction conditions (e.g., types of mechanisms). ... By applying these expert elementary reaction templates to the reactants in Pistachio, we obtained the recorded products as well as unreported byproducts and side products. We systematically selected and preserved the mechanistic pathways leading to the formation of the recorded product for each entry, resulting in a comprehensive dataset comprising 1.3 million overall reactions and 5.8 million elementary reactions."
The next step is now to use this data to obtain TSs for these elementary reactions - a difficult but important challenge to the CompChem community.

This work is licensed under a Creative Commons Attribution 4.0 International License.
Thijs Stuyver (2024)
Highlighted by Jan Jensen


Chenru Duan, Yuanqi Du, Haojun Jia, and Heather J. Kulik (2023)
Highlighted by Jan Jensen
Part of Figure 1 from the paper.
As anyone who has tried it will know, finding TSs is one of the most difficult, fiddly, and frustrating tasks in computational chemistry. While there are several methods aimed at automating the process, they tend to have a mixed success rate or be computationally expensive and, often, both.
This paper looks to be an important first step in the right direction. The method produces a guess at a TS structure based on the coordinates of the reactants and products. Notably, the input structures need not be aligned or atom mapped!
The method achieves a median RMSD of 0.08 Å compared to the true TSs and it often so good that single point energy evaluation gives a reliable barrier. The method also predicts a confidence scoring model for uncertainty quantification, which allows you to a priori judge whether such a single point is sufficient or whether a TS search is warranted. The approach allows for accurate reaction barrier estimation (2.6 kcal/mol) with DFT optimizations needed for only 14% of the most challenging reactions.
So, the method's not going to do away with manual TS searches entirely, but it is going to be invaluable for large scale screening studies. As the authors note, the method can likely also be adapted to the prediction of barrier heights, which could potentially be used to pre-screen reactions on a much, much bigger scale.
The paper is an important proof-of-concept study, but needs to be trained on much larger data sets (note that it is only trained on C, N, and O containing molecules), which are non-trivial to obtain. But the method could likely be used to obtain these data sets in an iterative fashion.
This work is licensed under a Creative Commons Attribution 4.0 International License.



![]() TS1 | |
![]() TS2 | ![]() TS3 |
![]() |
| Fig 1: Schematic flow chart of the AI-assisted MD simulation algorithm. |






3
|
4
|
TS [6+4]
|
TS Cope
|


1TS
|
3TS1F
|
3TS1Cl
|



Concerted TS
|
Stepwise TS
|

Concerted TS
|
Stepwise TS
|