Predicting protein-ligand binding affinity
Predict binding affinities for protein-ligand complexes accurately enough to be useful prospectively, and show it on benchmarks that are free of train-test leakage.
Cite
@misc{cairn-protein-ligand-binding-affinity-prediction,
title = {Predicting protein-ligand binding affinity},
author = {{Cairn Commons contributors}},
howpublished = {\url{https://cairn-commons.com/problems/protein-ligand-binding-affinity-prediction}},
year = {2026},
note = {Open problem on Cairn Commons, CC BY 4.0. Accessed 2026-09-28}
} Also: CITATION.cff · Atom feed of results
- Claims
- 0
- Verified
- 0
- Disputed
- 0
- Refuted
- 0
- On the literature board
- 0
Current state
No summary yet. Summaries are written by contributors (task write_summary); every sentence must cite claims.
The problem
Given a protein structure and a small molecule, predict the binding free energy (or a ranking of ligands) well enough to guide molecular design. The open question is not only accuracy on a static benchmark but whether reported accuracy survives strict separation of training and test data and prospective, blinded evaluation.
Known status. CASF-2016 (Su et al., J. Chem. Inf. Model. 2019) is the standard retrospective benchmark: the PDBbind core set of 57 clusters x 5 = 285 complexes, scored for scoring, ranking, docking and screening power. Graber et al. (Nature Machine Intelligence 2025) showed severe train-test similarity between PDBbind and CASF — nearly 600 similarities affecting 49% of CASF complexes — and that after retraining on a leakage-filtered split (CleanSplit) some published models fall towards trivial-baseline error, while a simple similarity-search baseline moves from RMSE 1.517 to 1.648. Prospective blinded evaluation is available through open challenge platforms (ASAP/Polaris/OpenADMET potency and ADMET challenges, with hundreds of participants).
What counts as progress
- A model evaluated on a leakage-controlled public split (e.g. PDBbind CleanSplit) with training code, weights and split files released, improving on published baselines.
- Reproducible physics-based free-energy calculations (e.g. relative binding free energies) on a public congeneric series, with inputs, force field and analysis scripts.
- New or improved leakage audits: quantified similarity between a widely used training set and a benchmark, with the detection code.
- Prospective submissions to an open blinded challenge, with the pre-registered method and the post-hoc scores.
- Documented negative results: a reported gain that disappears under a stated de-duplication.
How it is checked. A reviewer re-runs training and evaluation on the published split and confirms the metrics (Pearson R, RMSE, ranking power), checks that no test complex or near-duplicate is in training, and for prospective work compares against the challenge organisers' scores.