Literature-Augmented Knowledge Graph Embedding for Repurposing Candidate Molecules in Hepatic FibrosisYAN Yiyun*
-
Abstract
Objective This study aims to construct a diseasespecific knowledge graph and apply multiple knowledge graph embedding models to prioritize candidate molecules with experimental validation potential for hepatic fibrosis drug repurposing. Methods The DRKG hepatic fibrosis subgraph was integrated with PubMed literature abstracts, and a large language model (LLM) was used to extract subjectpredicateobject (SPO) triples. After entity normalization, relation standardization, and nodetype reclassification, a literatureaugmented knowledge graph was built, containing 122 775 nodes, 292 711 edges, 7 entity types, and 19 standardized relation types. Four KGE models—TransE, RotatE, RGCN, and CompGCN—were trained as internal baselines for link prediction. Combined with multimodel normalized scoring, entitytype constraints, nonmolecular entity exclusion, canonical name deduplication, and finegrained classification strategies, a final ranked list of candidates was generated from an initial pool of 24 211 candidates. Results A total of 5 041 experimentally verifiable candidate molecules were identified. Top candidates include curcumin, pirfenidone, resveratrol, quercetin, berberine, salvianolic acid B, EGCG, taurine, astragaloside IV, and apigenin, all supported by L1_direct evidence in the knowledge graph. Knownpositive recovery analysis detected all 14 positive controls, with 13 in the main list, 12 in Top30, and 14 in Top200, indicating strong internal positive enrichment. Molecular docking of apigenin (ranked #10) against FAK (PDB: 2J0L) yielded a best binding energy of −7.51 kcal/mol, with all 9 conformations below the −5.0 kcal/mol threshold. Conclusions This study provides interpretable candidate prioritization and mechanistic moleculartargetpathway clues for hepatic fibrosis drug repurposing, preliminarily validates the physical binding feasibility of knowledgegraphprioritized candidates, and offers computational strategies and data support for subsequent candidate screening and experimental validation.
-
-