ECCB (2026)
Sungjoon Park, Soorin Yim, Dongyun Kim, Kiwoong Yoo, Doyeong Hwang, Kyungwook Lee, Jongseong Jang, Kiyoung Kim
Abstract
Binding affinity governs how proteins interact and underlies essential biological processes. Computational approaches have been developed to simulate and predict protein binding, but the scarcity of high-quality data has imposed significant constraints. One consequence is that most methods focus on predicting mutational changes in binding affinity (ΔΔG), rather than binding affinity (ΔG) itself. This practice risks overfitting to skewed data distributions, limiting the generalizability of predictions. Recent advances in protein structure prediction have enabled computational modeling of protein conformations in mass, providing rich structural information from which binding interactions can be largely explained. However, leveraging these advances for effective prediction of binding affinity has yet to translate into reliable predictions.
Results:We present PreFold-dG, a model that estimates binding affinities of protein complexes utilizing intermediate embeddings from Boltz-2, an open-source foundation model for protein structure prediction. Our approach aggregates residue-level information weighted by inter-residue distance, and predicts ΔG directly rather than its derivative, ΔΔG. PreFold-dG achieved state-of-the-art performance on common binding affinity benchmarks and demonstrated robustness on independent test sets. Ablation studies suggest that predicting ΔΔG without grounding in ΔG is prone to overfitting to specific datasets. We further validated our model through case studies on real-world antibody evolution data.