Paper deep dive
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
Thomas MacDougall, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Maxim Malkov, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/21/2026, 6:21:46 AM
Summary
This paper introduces 3D-Fit, a benchmarking strategy to evaluate Large Language Models (LLMs) on 3D spatial molecule generation under complex constraints. It compares general-purpose LLMs against specialized diffusion models using conditions like protein pockets, anchor fragments, pharmacophore points, and mandatory interactions. The study finds that while LLMs lag behind state-of-the-art diffusion models, they demonstrate promising capabilities in handling multiple simultaneous spatial constraints.
Entities (9)
Relation Signals (8)
3D-Fit → usesdataset → CrossDocked2020
confidence 95% · comprehensive synthetic benchmark datasets of 3D conditions based on the test complexes from CrossDocked2020
3D-Fit → usesdataset → PLINDER
confidence 95% · comprehensive synthetic benchmark datasets of 3D conditions based on the test complexes from CrossDocked2020 and PLINDER
3D-Fit → supportsconstraint → Anchor Fragments
confidence 93% · We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments
3D-Fit → supportsconstraint → Pharmacophore Points
confidence 93% · spatial constraints, including anchor fragments, pharmacophore points
3D-Fit → supportsconstraint → Mandatory Pocket-Ligand Interactions
confidence 93% · spatial constraints, including ... mandatory pocket-ligand interactions
3D-Fit → evaluates → LLMs
confidence 92% · assessing LLM performance on multi-conditioned spatial molecule generation
3D-Fit → comparesagainst → Diffusion Models
confidence 90% · compare state-of-the-art diffusion methods with recent proprietary and open-weight foundation LLMs
3D-Fit → usesoutputformat → Simplified SDF
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.
Tags
Links
- Source: https://arxiv.org/abs/2607.18144v1
- Canonical: https://arxiv.org/abs/2607.18144v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
71,567 characters extracted from source content.
Expand or collapse full text
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints Thomas MacDougall 1 Maksim Kuznetsov 1 Roman Schutski 3 Rim Shayakhmetov 2 Maxim Malkov 2 Vladimir Aladinskiy 2 Alex Aliper 2 Alex Zhavoronkov 1,2,3 1 Insilico Medicine Canada Inc. 2 Insilico Medicine AI Limited 3 Insilico Medicine Hong Kong Ltd. Abstract Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit – a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups. 1 Introduction Designing molecules that satisfy strict spatial constraints is a challenging task in computational chemistry. In realistic drug discovery scenarios, 3D generative models should produce molecular structures simultaneously satisfying multiple heterogeneous spatial requirements. While specialized diffusion models [1–7] have demonstrated leading performance in standard pocket- conditioned molecular generation, adapting them to simultaneously handle multiple heterogeneous spatial constraints, although feasible [8–11], remains non-trivial. Balancing these diverse condition types within a single generative process is challenging, as constraints may differ in scale, rigidity, and importance, and may even conflict with one another. As a result, effective multi-constraint generation therefore requires careful model design, accurate condition encoding, and robust training or inference. Alternatively, Large Language Models (LLMs) have achieved remarkable success across various computational chemistry and drug discovery tasks [12–14], largely owing to their innate capacity to seamlessly process complex task instructions and handle multiple constraints simultaneously. Despite these advances, their ability to explicitly reason about physics and 3D environments remains largely underexplored. Within the specific domain of 3D molecular design, a handful of recent approaches—such as XYZTransformer [15], BindGPT [16], and nach0-pc [17]—have trained special- ized language models using textual formulations of 3D structures, relying on a textual description of the spatial environment and/or generated structure. Nevertheless, it remains systematically unverified whether general-purpose LLMs possess the capability to navigate complex 3D molecular design. Preprint. arXiv:2607.18144v1 [cs.LG] 20 Jul 2026 The main goal of this work is to assess 3D capabilities of the current state-of-the-art LLMs, whether these models can understand, reason about, and accurately generate ligands that satisfy wide sets of geometrically grounded structural constraints. This benchmark is designed to move beyond pocket- conditioned ligand generation by introducing a richer set of structurally explicit generation conditions (e.g., specific interaction patterns, pharmacophore features, and fragment anchoring), reflecting the criteria medicinal chemists routinely consider when forming molecular hypotheses. Here, we broaden the conditioning space to multiple properties and build comprehensive synthetic benchmark datasets of 3D conditions based on the test complexes from CrossDocked2020 [18] and PLINDER [19]. In this work, we propose the 3D-Fit benchmark to evaluate 3D molecular generation under a variety of spatial conditions. The benchmark evaluates protein pocket-only and single-constraint protein pocket-conditioned generation (with anchor fragments or pharmacophore points) to compare state-of- the-art diffusion methods with recent proprietary and open-weight foundation LLMs; also it covers multi-constraint protein pocket-conditioned generation to analyze LLMs success rates and failure modes. To make such evaluation possible, we design spatial condition representations, a structured molecular output format, and a prompting protocol tailored to 3D molecular generation. Finally, we provide a detailed analysis of model performance, including success cases, common failure modes, and the challenges faced by LLMs under increasingly complex spatial constraints. 2 Related Work Structure-based Molecular BenchmarksModern structure-based generative models are based on generated datasets such as CrossDocked2020[18], and experimental repositories of structures such as PDBbind[20] and Binding MOAD[21], which provide curated binding protein-ligand complexes. Building on these sources, several benchmarks were proposed: CBGBench[22] uses CrossDocked- style data to evaluate de novo generation, linker design, and scaffold hopping through metrics focusing on interaction geometry and substructure validity (mostly 2D); POKMOL-3D[23] presents 32protein targets to benchmark pocket-conditioned 3D generation using active molecule confor- mations; Durian[24] uses experimental affinity and structures to evaluate generative performance; and MolGenBench provides a large-scale evaluation of target-aware lead optimization and de novo design across120targets using220,005validated active molecules with some of the tasks having 3D. Our benchmark and dataset differ from the mentioned projects by providing a comprehensive set of multiple spatial conditions for every structure, sheer size of the data and a collection of robust computational metrics. Pocket-conditioned diffusion models Pocket-conditioned diffusion models are the leading ap- proaches in structure-based drug discovery at the time of writing. Foundational works such as TargetDiff[2] and MolDiff[25] introduced SE(3)-equivariant diffusion processes to jointly generate molecular coordinates and atom types. Building on these, several works proposed the methods to enhance binding accuracy, like PMDM[4] which generates the joint distribution of protein and ligand atoms or IPDiff[9] which introduces explicit interaction priors to guide the sampling process. DiffSBDD[3] introduced scalar property conditioning to pocket-constrained generation, followed by DiffBP[26] that achieved SOTA accuracy by targeting the 2D properties of druglikeness and QED. Other notable enhancements in pocket-conditioned generation include incorporation of binding-aware features (BindDM[27]). Lastly, SeFMol[7] employed reinforcement learning to guide the diffusion process toward high-affinity candidates. Pocket-conditioned diffusion with auxiliary conditioningFollowing the success of the diffusion- based models in pocket-conditioned generation, a series of works introduced auxiliary geometric conditions to enhance their practical applicability in drug design. The simplest form of such conditions are rigid anchor fragments. DiffDec[8] was one of the first models to introduce a preservation mechanism to generate new molecular components around fixed scaffolds. The recent FDC-Diff[28] achieved SOTA results in fragment-to-lead optimization. Pharmacophore points provide a more general description of spatial patterns matching a range of atoms or atomic groups. Recently, MolSnapper[10] achieved precise 3D pharmacophore matching during pocket-conditioned generation. Another way to control spatial generation is to specify mandatory pocket-ligand interaction priors. IPDiff [9] was one of the first diffusion models introducing this condition to pocket-aware generative process. The recent DiffPharma[11] combines pharmacophore and mandatory pocket- ligand interaction conditioning for even better control of molecular generation. 2 GPT 5.5 Opus 4.8 Gem. 3.1 Grok 4.3 Qwen 3.5 GLM 5 MolSnapper PocketXMol DiffSBDD DiffPharma -12 -10 -8 -6 -4 -2 0 Docking Score (kcal/mol, 0-clipped) LLMsDiffusion Models DockingClashesConformation StrainRawOptimized 0 20 40 60 80 100 Clash / Strain Success Rate (%) Figure 1: 3D-Fit benchmark selected results for pocket-only conditioned generation on CrossDocked. 3 Spatial Condition Descriptions Protein Pockets are localized regions of a target protein structure that form cavities capable of accommodating a ligand structure. Pockets are defined by the subset of residues whose atoms line the cavity and determine its geometry and physicochemical environment. In practice, a protein pocket is usually determined by the residues within a given distance cutoff point from a ligand. Mandatory Protein–Ligand Interactions are a set of key binding contacts that a candidate ligand is expected to reproduce with certain residues of the target protein pocket. Interaction points are valuable because they provide an interpretable, target-specific description of the binding mode that complements general pocket geometry and can guide molecule generation or evaluation toward ligands that preserve critical contacts with the protein. Anchor Fragments are chemically significant ligand substructures that are expected to be preserved or approximately reproduced in a generated molecule. These fragments are typically extracted from a reference ligand in its bound conformation and correspond to substructures that contribute to the ligand binding mode. Anchor fragments are useful because they provide a direct way to constrain molecular generation around experimentally or structurally observed ligand geometry while still allowing modifications in the remaining parts of the molecule. Pharmacophore Points describe the essential spatial arrangement of interaction features responsible for the recognition and binding within a target pocket. They provide an interpretable, target-aware representation of a ligand’s binding mode. While there are protein-based and ligand-based pharma- cophore points, we utilize the ligand-based pharmacophores that abstract the molecular structure into key functional features derived primarily from the ligand. By requiring candidate molecules to satisfy these pharmacophore points, the search can be guided toward compounds that maintain critical binding interactions while still allowing the discovery of chemically diverse and novel molecules. 4 3D-Fit Benchmark Our 3D-Fit benchmark is a framework for applying and evaluating multiple 3D condition satisfaction for molecular generation. It can use any test set of 3D pocket-ligand complexes as condition sources, but in this work we focus on two popular datasets for generative chemistry,CrossDocked2020[18] andPLINDER[19]. Both datasets are well-established sources of test protein-ligand complexes with varying test split strategies and comprehensive specialist model baselines. For each test example, we consider the Protein Pocket as the main 3D condition and up to three additional conditions: Mandatory Pocket-Ligand Interaction Points, Anchor Fragments and Pharmacophore Points. We also propose concise, token-efficient textual descriptions of 3D conditions in the input and a robust textual output format for small molecules in 3D. 3 CA at (-31.063,−9.930, -15.675) Carbon (C) within 0.5 Angstroms of (−22.711,−5.663, -21.678) Hydrophobic group within 0.5 Angstroms of (-25.266, -2.904,−24.220) Hydrogen-bond donor interaction with Residue 17, THR, Chain B Pocket description Mandatory interactions description Anchor fragments description Pharmacophore points description Textual description of spatial conditions Figure 2: Visualization of spatial conditions and their corresponding textual descriptions. 4.1 Textual Representation of Spatial Conditions In our benchmark, we describe the four spatial conditions using compact textual descriptions that can be provided directly as text to an LLM. The general prompt for LLMs benchmarking and examples of the following textual descriptions of spatial conditions are provided in Appendix A. Protein PocketSince pockets typically contain hundreds or thousands of atoms, a compact descrip- tion is crucial for comprehensive yet efficient LLM benchmarking. Although the PDB format [29] is a versatile container for protein structures, it includes substantial redundancy (headers, repeated fields, occupancy/B-factors, element annotations, alternate locations, etc.), which results in unnecessary large token consumption for LLM-based pipelines. In our work, we propose a textual representation designed to optimize the length of textual protein description without the loss of critical information. To preserve the sequential nature of proteins in this representation, amino acids are described sequentially, beginning with a block that specifies the index and three-letter code of amino acid. For example, the third tyrosine residue is represented as‘Residue 3, TYR’. Then, each amino acid’s atom is described by its name and coordinates, such as‘CB at (3.255,−1.106, 8.438)’. We limit atom details to just the atom name and coordinates with three digits after the decimal point, because other properties (atom type, valence, connectivity, charge) can be inferred with standard amino acids. All heavy atoms within an amino acid are listed starting from the backbone and proceeding to the side-chain atoms of the residue (N→CA→C→...→NH1→NH2). In case if the target protein contains more than one chain, we specify the chain before the block of amino acids’ descriptions corresponding to this chain, e.g. ‘Chain A:’. Since the overall geometry is largely determined by heavy-atom backbones and side chains, we work with hydrogen-depleted proteins to further reduce the number of input tokens. Mandatory Pocket-Ligand Interactions Each requirement specifies an interaction type and the corresponding pocket residue (chain ID, residue index, and residue name). For example, an interaction constraint can be expressed as: ‘Hydrophobic interaction with Residue 13, GLU, Chain A’. Anchor Fragments We describe anchor fragments in a compact form as a set of anchor atoms. Each anchor atom is defined by (i) its atom type and (i) a spherical tolerance region that indicates where this atom may be placed, represented by the sphere center and the radius. This representation provides a concise way to constrain ligand placement while allowing small positional variability. For example, to specify a carbon anchor atom within 0.5 Å of a given point, we describe it as: ‘Carbon (C) within 0.5 Angstroms of (−16.191,−11.325, 8.531)’. Pharmacophore PointsAs a textual representation, we specify pharmacophore points by type and their 3D positions. If the pharmacophore point has an associated directionality, such as a donor point that should be oriented toward a protein-pocket residue, we do not specify this direction explicitly. Instead, we allow the model to infer the appropriate orientation from the protein pocket description. For example, to specify a hydrogen-bond acceptor that must be positioned within0.5Å of a given coordinate, we provide the LLM with the following instruction: ‘Hydrogen bond acceptor within 0.5 Angstroms of (22.710, 32.862,−24.262)’. 4 RDKit3D 7 6 0 0 0 0 0 0 0 0999 V2000 -1.1021 0.9459 -1.1727 N 0 0 0 0 0 0 0 0 0 0 0 0 -0.3654 0.7773 0.0624 C 0 0 1 0 0 0 0 0 0 0 0 0 -1.1177 -0.0777 1.0602 C 0 0 0 0 0 0 0 0 0 0 0 0 -1.3757 -1.3658 0.5635 O 0 0 00 0 0 0 0 0 0 0 0 0.9449 0.0945 -0.1579 C 0 0 00 0 0 0 0 0 0 0 0 1.2854 -0.2577 -1.3167 O0 0 00 0 0 0 0 0 0 0 0 1.7305 -0.1166 0.9612 O 0 0 00 0 0 0 0 0 0 0 0 2 1 1 6 2 3 1 0 3 4 1 0 2 5 1 0 5 6 2 0 5 7 1 0 M CHG 2 1 1 7 -1 M END SDF 1-1.102 0.946 -1.173 N +1 2-0.365 0.777 0.062 C 0 3-1.118 -0.078 1.060 C 0 4-1.376 -1.366 0.564 O 0 50.945 0.095 -0.158 C 0 61.285 -0.258 -1.317 O 0 71.731 -0.117 0.961 O -1 2 1 1 6 2 3 1 0 3 4 1 0 2 5 1 0 5 6 2 0 5 7 1 0 Keep essential Discard redundant Set explicit indices Simplified SDF Figure 3: Zwitterionic form of Serine amino acid as SDF and Simplified SDF. 4.2 Textual Representation of 3D Molecule Output The choice of the output format for the generated 3D ligand is essential for our benchmark. The format should be familiar to the LLM’s internal knowledge base, and similar to condition representations. It should also be non-redundant and robust, since unnecessary components and format complexity slow down generation and increase the chance of errors that may render the entire output invalid. In our work, we propose Simplified SDF (see Fig. 3), a modified version of the standard SDF [30] designed to reduce the number of textual tokens and format fragility by retaining only key information needed for accurate molecule reconstruction. This format first describes the index, symbol, coordinates and charges of each atom, and then specifies molecular connectivity by listing bonded atom pairs together with their bond types. As an ablation study, we also evaluated the SMILES+XYZ format used in [16][31] as another token- efficient representation (see Appendix C). We found that the Simplified SDF consistently outperforms this alternative on nearly all metrics for general-purpose LLMs. We hypothesize that this is because the benchmarked LLMs were trained on corpora containing data similar to the SDF/Simplified SDF family, as well as the ability to specify atom coordinate before specifying their connectivity. 4.3 Test Datasets and Conditions Preparation We compare the performance of each model on two test datasets,CrossDocked2020v1.3 and PLINDER 2024-06/v2. Each dataset was processed according to five principles: (i) Literature-consistent test split: ForCrossDocked2020, we used the test split0from the downsampled version of the dataset. For PLINDER, we also used the provided test split. (i) Valid molecule loading: For both datasets, complexes were included if both the protein and the ligand are valid, i.e. they could be loaded with ProDy [32] and RDKit [33] without errors. (i) Binding quality: ForCrossDocked2020, we used complexes with the distance of RMSD < 0.5 Å and Vina [34] score <−6. For PLINDER, we used all test set complexes as they are. (iv) Drug-like molecule size: In both datasets, we filtered ligands to retain only those with more than 25heavy atoms to better reflect drug-like molecules in the final test sets. We did not apply additional drug-likeness filters to avoid further biasing the data distribution given our focus on 3D conditions. (v) Protein pocket cutoff: All pockets were identified from full protein structures by selecting all atoms for every residue with at least one atom within 10 Å of the reference ligand center. The final test sets consist of948complexes forCrossDocked2020and505complexes forPLINDER. For each test example in both datasets, 3D spatial conditions are sampled according to the following processes. The concept of the 3D-Fit framework allows for additional sampling or runtime sampling of other protein-ligand complexes for future work in applications such as Reinforcement Learning, as long as the test complexes are not used for any training. For our reported experiments, the sampling was pre-computed and fixed for each example to maintain consistent comparison across models. Protein Pocket forms the core conditional identity for each test example. For each test object, the pocket-ligand complex is randomly shifted by up to 50 Å in each axis, to prevent any potential coordinate memorization by LLMs in the case of the test set leakage. 5 Model PocketStructure UniDockPB Inter.ParsedRO5N HvyPB Intra. CD PLCD PL CD PLCD PLCD PL CD PL R O R OR O R OR O R O GPT 5.519. -6.5 20. -6.16 999 1009910099 10023. 23.23 30 31 37 GPT 5.4127. -6.8 125. -6.40 65 0 6666 6749 4329. 28.3 4 2 3 Opus 4.856. -6.5 53. -6.21 90 4 9590 9689 9521. 21.47 55 53 58 Opus 4.7 52. -6.5 52. -6.12 96 3 9596 9595 9421. 21.55 66 57 62 Opus 4.638. -6.9 33. -6.52 58 5 6258 6258 6222. 22.27 31 29 35 Sonnet 4.696. -6.8 96. -6.41 66 0 6466 6465 6423. 23.11 12 9 11 Gem. 3.157. -6.4 51. -6.01 93 4 9494 9585 9023. 23.41 49 41 48 Grok 4.1122. -6.7 110. -6.20 21 0 2521 2621 2522. 22.10 10 10 10 Grok 4.3101. -6.8 114. -6.20 66 0 6666 6664 6421. 21.45 45 46 46 Qwen 3.589. -7.0 88. -6.61 82 1 8183 8279 7725. 25.1 2 2 3 GLM 5108. -6.9 106. -6.30 46 0 4346 4344 4022. 23.3 3 3 4 PocketXMol-6.3-8.2-4.9-7.19193 93 9493 9456 5529. 27.89 8891 90 DiffSBDD-4.2 -5.8 -3.5 -4.990 98 9199999988 9217. 15.70 71 73 74 PDMD2.1 -7.7 -1.4 -7.31 87 1 8688 8758 4931. 33.7273 70 71 DiffPharma-4.0 -5.6 -3.2 -4.684 93 89 9594 9582 8617. 15.64 66 68 69 MolSnapper-6.8 -8.6 -6.2 -8.292 999198100 10086 6627. 33.89 90 8586 SeFMol-0.2 -6.2 -3.3 -5.40 100 1 100100 100989917. 16.51 52 55 55 IPDiff---2.9 -5.5--1 100- 100-78- 17.-- 54 57 BindDM -2.9 -5.3 -1.9 -4.40 991 99100 10073 7219. 17.43 50 44 53 Table 1: Benchmarking Results for Pocket-Only Conditioned Generation. Mandatory Pocket-Ligand Interactions are extracted from the reference protein-ligand complex with the ProLIF library [35]. Using the Prolif package, each protein/ligand complex is annotated with all important binding interactions according to their type and residue. One random interaction from the annotated list is sampled and used as a necessary 3D condition. We consider the following interaction types: hydrophobic contacts; hydrogen bonds;π–πstacking; van der Waals contacts; ionic interactions involving anionic or cationic groups; and cation–π / π–cation interactions. Anchor Fragments are extracted from the ground-truth ligand pose in the holo complex. The scope of potentially chemically important fragments is large, ranging from small functional groups to large scaffold structure. We focus on relatively small anchor fragments, to evaluate fragments as a 3D condition but to allow diversity in the molecule output. For each reference ligand in its binding conformation, we split the molecule according to the BRICS fragmentation algorithm [36] and randomly select a fragment with between3and8heavy atoms. If no fragment meets this criterion, we instead fragment the molecule at all rotatable bonds and select within the same size range. If none are found, we choose a random fragment from the original BRICS decomposition. Pharmacophore Point is extracted using the Pmapper [37] package. We annotate each ligand with all possible pharmacophore points independently of the pocket. While it is possible to have multiple points, we sample a single random point from the annotated list to use as a condition. We consider the following pharmacophore types: hydrogen-bond donors, hydrogen-bond acceptors, hydrophobic groups, positive ionizable groups, negative ionizable groups, aromatic groups, and exclusion regions. 6 Model PocketAnchorStructure UniDockPB Inter.SRParsedRO5N HvyPB Intra. CD PLCD PLCD PL CD PLCD PLCD PL CD PL R O R OR O R OR O R OR O R O GPT 5.57.1 -6.4 2.9 -5.921 99 32 9999 0 99 199 9999 9922. 22.46 55 53 59 GPT 5.482. -6.6 61. -5.92 66 6 6667 0 66 067 6654 5726. 26.4 7 4 8 Opus 4.828. -6.2 23. -5.77 89 13 9590 1 95090 95909521. 21.50 54 63 67 Opus 4.7 23. -6.1 15. -5.78 91 13 9591 1 95091 95909320. 20.54 59 65 69 Opus 4.611. -6.4 5.7 -5.911 74 18 8474 1 84 274 8474 8421. 21.43 50 56 62 Sonnet 4.633. -6.4 23. -5.75 70 9 7470 0 74 070 7469 7221. 21.23 26 24 29 Gem. 3.139. -5.9 21. -5.47 9411 9694 1 95095 9682 8822. 22.57 64 57 66 Grok 4.145. -6.9 29. -6.14 17 4 1517 0 16 017 1616 1520. 20.3 3 2 3 Grok 4.355. -6.9 34. -6.32 37 4 3337 0 32 038 3433 2922. 23.6 6 3 3 Qwen 3.559. -6.9 35. -6.31 49 4 4750 0 47 050 4842 4127. 27.0 0 0 0 GLM 543. -6.2 32. -5.52 42 2 4743 0 47 043 4741 4621. 21.3 4 4 5 PocketXMol-6.3 -7.9-4.9 -7.087 90 76 8590 1586 990 8659 3027. 34.83 80 7270 DiffSBDD-2.3 -4.9 -1.6 -4.26888 70939818 99 10989979 7818. 16.787473 77 PDMD-5.4-8.5 -4.0-7.336 58 32 5259 9 51 460 5336 3229. 27.48 48 44 44 Table 2: Benchmarking Results for Pocket+Anchor Fragment Conditioned Generation. 5 Experiments In this section, we benchmark recent closed- and open-weight large language models, together with diffusion models, and analyze their success and failure modes. The modular design of 3D-Fit allows conditions to be added or removed from the evaluation depending on the model’s capabilities. 5.1 LLM Models In this work, we use the following recently released LLMs as baselines: Proprietary Closed-Weight LLMs: GPT 5.5 [38]; GPT 5.4 [39]; Claude 4.7 Opus [40]; Claude 4.6 Opus [41]; Claude 4.6 Sonnet [42]; Gemini 3.1 Pro [43]; Grok 4.1 Fast Reasoning [44]. Open-Weight LLMs: Qwen-3.5 (397b-a17b) [45]; DeepSeek v3.2 [46]; GLM-5 [47]. During benchmarking, LLMs were not allowed to perform web searches or use any external tools, ensuring that we evaluated only the knowledge and reasoning capabilities of the models themselves. All LLM results were sampled with a temperature of1.0, a top p of1.0, maximum new tokens of 8192without counting the reasoning tokens, and the default API parameters beyond this. Where it was supported by the model and provider, reasoning effort/reasoning level was set to "high". 5.2 Specialist 3D Diffusion Models We also compare against specialist 3D molecular generative models conditioned on various combina- tions of 3D constraints. We indicate them with purple color. We list the condition sets considered in our benchmark and the corresponding diffusion models supporting them: Pocket-only: PocketXMol [48], SeFMol [7], DiffPharma [11], MolSnapper [10], IPDiff [9], BindDM [27], DiffSBDD [3] and PMDM [4]. Pocket+Anchor fragments: PocketXMol [48], DiffSBDD [3] and PMDM [4]. Pocket+Pharmacophore points: DiffPharma [11] and MolSnapper [10]. 5.3 Metrics and Evaluation We evaluated all generated molecular structures against sampled conditions for each protein-ligand complex. Evaluation of the generated molecules is done on the basis of two key pillars of success; 7 Model PocketPhar.Structure UniDockPB Inter.SRParsedRO5N HvyPB Intra. CD PLCD PLCD PL CD PLCD PLCD PL CD PL R O R OR O R OR O R OR O R O GPT 5.512. -6.4 11. -5.910 99 16 9999 99 100 999910099 9922. 22.32 41 29 39 GPT 5.4109. -6.6 108. -6.21 67 2 6568 68 66 6669 6648 4628. 28.5 9 7 9 Opus 4.852. -6.5 41. -6.12 84 5 9085 85 90 9085 9085 9021. 21.42 48 44 50 Opus 4.7 43. -6.4 43. -6.04 92 6 9493 93 94 9493 94919320. 20.51 6250 61 Opus 4.622. -6.8 23. -6.24 60 6 6560 60 65 6560 6559 6422. 22.26 31 29 35 Sonnet 4.665. -6.7 55. -6.32 61 3 6561 61 65 6561 6560 6422. 22.8 11 8 11 Gem. 3.148. -6.1 37. -5.74 949 949494959495 9585 8623. 23.41 51 43 52 Grok 4.1102. -7.490. -6.60 15 0 1515 15 16 1616 1613 1424. 24.1 1 2 2 Grok 4.374. -6.7 70. -6.31 52 3 5451 51 53 5353 5550 5122. 22.31 31 30 31 Qwen 3.589. -7.478. -6.81 42 2 4542 42 45 4543 4535 3726. 26.0 0 0 0 GLM 564. -6.9 55. -6.31 27 2 3327 27 33 3327 3324 3022. 22.1 1 1 1 DiffPharma-4.2-5.7 -3.3-4.988 9489 9573 73 78 7795 9584 8717. 16.5355 6162 MolSnapper-6.1 -8.1 -4.4 -7.38499 829991 91 94 94100 10086 6427. 32.90 91 86 89 Table 3: Benchmarking Results for Pocket+Pharmacophore Point Conditioned Generation. 3D Molecular Validity and 3D Condition Success. For binary pass/fail metrics, we report the Success Rate (SR, %) on all outputs, to not bias results by validity. 3D Molecular Validity metrics evaluate the quality of the generated 3D molecules independently of any external conditions. We first report whether the input molecule can be successfully reconstructed and sanitized (Parsed). We also show the average number of heavy atoms (N Hvy) to assess the size of generated molecules. Along with these metrics, we compute the percentage of molecules that passes Lipinski’s rule of five (RO5) to roughly estimate the druglikeness of molecules. We then use PoseBusters [49] to assess 3D plausibility. Specifically, we group the conformation-related PoseBusters checks into a single metric, PoseBusters Intra-Molecular (PB Intra), and count a molecule as successful only if it passes all of the following filters: bond lengths, bond angles, internal steric clashes, aromatic ring flatness, non-aromatic ring non-flatness, and double bond flatness. 3D Condition Success metrics assess the molecule’s satisfaction of the generation conditions tied to each test example. We report this metrics for ligand position before (R=Raw) and after (O=Optimized) local ligand pose optimization with UniDock. For Pharmacophores, Anchor Fragments, and Pocket- Ligand Interaction point conditions, we evaluate success from the 3D positioning of atoms/features. If the generated molecule contains the conditioned feature, atom, or interaction in the desired location or with the desired residue, we consider it as a success. We do not evaluate the 2D-connectivity of Anchor Fragments or other features, only 3D placement. For Protein Pockets, evaluating success is more nuanced as there is no obvious pass condition universal for all proteins, so we report two metrics. We calculate and report Unidock [50] Score (Median, kcal/mol, Lower is better). We highlight with red if the median model Unidock Score higher than−6threshold. We also group the pocket-related PoseBusters checks into a single metric, PoseBusters Inter-Molecular (PB Inter), and count a molecule as successful only if it passes all of the following filters: protein-ligand maximum distance, minimum distance to the protein and volume overlap with the protein. 5.4 Results We organize the results according to the following conditioning settings: (i) Pocket-only Conditioning is the most widely supported setting, and diffusion models that support additional constraints typically also support pocket-only conditioning. Table 1 shows these results. (i) Pocket+Fragment Conditioning and Pocket+Pharmacophore Conditioning are supported by only a small number of diffusion models. Tables 2 and 3 show these results, respectively. (i) Pocket+Mandatory Interaction+Pharmacophore+Fragment Conditioning is not natively supported by any diffusion or other model we are aware of. Still LLMs are easy to prompt in this setting, so we compare only the LLMs models. Table 4 shows these results. 8 Model PocketAnchorPhar.M.Interact.Structure UniDockPB Inter.SRSRSRParsedRO5N HvyPB Intra. CD PLCD PLCD PLCD PLCD PL CD PLCD PLCD PL CD PL R O R OR O R OR O R OR O R OR O R OR O R O GPT 5.53.0 -6.3 1.7 -5.726 99 37 9999 199 099 99 99 9973 28 82 2899 9998 9821. 21.41 48 42 48 GPT 5.468. -6.7 56. -6.11 56 2 5961 2 63 260 60 61 6127 17 31 1961 6347 5026. 26.2 2 1 2 Opus 4.811. -6.1 6.4 -5.59 70 15 7670 0 76 070 70 76 7644 20 52 2270 7669 7420. 21.24 26 32 35 Opus 4.7 10. -6.0 5.5 -5.41176 208476 184 076 76 84 8337 19 46 2176 8473 8220. 20.28 32 36 39 Opus 4.67.2-6.2 5.1-5.610 68 17 7968 179 068 68 79 7935 18 41 2268 7967 7821. 21.31 34 3841 Sonnet 4.615. -5.8 9.6 -5.16 60 10 6761 0 67 060 60 66 6633 16 37 1861 6758 6420. 20.17 18 18 21 Gem. 3.129. -6.1 17. -5.55 929 9393194193939494682773269395808423. 23.353932 38 Grok 4.148. -7.149. -6.63 18 2 1319 0 13 017 17 13 127 4 6 419 1315 1220. 20.3 3 2 2 Grok 4.343. -7.0 44. -6.52 43 2 3543 0 35 043 43 34 3421 11 18 1044 3634 3223. 23.1 1 1 2 Qwen 3.544. -7.3 41. -6.31 47 2 4849 149 048 47 48 4722 11 26 1150 4936 3627. 27.0 0 0 0 GLM 523. -6.1 18. -5.53 39 5 3939 0 39 039 39 38 3828 9 31 1039 3933 3621. 21.1 1 1 2 Table 4: Results for Pocket+Interaction+Anchor+Pharmacophore Conditioned Generation. 6 Discussion In this section, we discuss the key observations from the benchmarking results. Increasing Spatial Conditions for LLMs Across all experiments, our results demonstrate that LLMs show emerging ability to follow spatial constraints, especially when the number of conditions increases. The successful parsing metrics indicate that the models generally understand the required output format. However, the PoseBusters intra-molecular filters (PB Intra) reveal variability in the quality of the generated conformations. LLMs perform particularly well on Anchor Fragment and Pharmacophore Point conditioning, possibly because these conditions are described in a form that closely resembles the corresponding output entries needed to satisfy them. In contrast, the models struggle more with Mandatory Interaction Point condition, which may be due to their more abstract nature and poor understanding of specific protein-ligand interactions. When additional conditions are provided, the models achieve substantially better Pocket metrics compared to the Pocket-only setting. This suggests that more seed-ligand information helps the models generate better molecules, which begin to fit inside the protein pocket. The results also indicate that the models may perform only limited exploration beyond molecules that directly satisfy the given conditions; therefore, adding more constraints may encourage more accurate exploration of chemical space in order to satisfy all specified requirements. Docking Scores of Generated MoleculesAlthough LLMs can often produce molecules that satisfy some structural constraints after optimization, their raw poses are poor binders (see Fig. 4). Before local UniDock optimization (R), all LLM UniDock scores are above the−6threshold, and are often strongly positive, indicating severe steric clashes with the pocket or highly unfavorable placement. Local UniDock optimization (O) substantially improves these poses, bringing most LLM scores to the moderate range of roughly−6.0to−7.0(see Fig. 6). However, they still remain behind the best pocket-specific diffusion models, which achieve substantially lower optimized scores, e.g. MolSnapper, PocketXMol and PDMD. Diffusion models also produce poses with roughly the same binding mode before and after optimization. See Appendix B for the details. These results suggest that while LLMs generally follow the desired spatial conditions, they often do so by producing poses with non-optimal geometry and fail to produce poses that can be optimized without substantial ligand repositioning. In particular, the generated molecules frequently exhibit external steric clashes with the pocket, as reflected by poor raw UniDock scores and low inter-molecular PoseBusters pass rates, as well as internal structural problems indicated by weaker intra-molecular validity. Thus, satisfying the requested spatial constraints alone is not sufficient for generating strong binders: the molecules must also adopt physically plausible geometries, both with respect to the protein environment and their own internal structure. Diffusion Models and Intermolecular Filters The PoseBusters inter-molecular filters show a different failure mode. For LLMs, raw inter-molecular pass rates are generally very low, consistent with their poor raw UniDock scores, but local optimization can substantially improve both docking scores and inter-molecular validity. Diffusion models are more mixed: several methods obtain high 9 inter-molecular pass rates after optimization, but this does not always correspond to strong UniDock scores. For example, DiffSBDD, DiffPharma, SeFMol, IPDiff, and BindDM can pass many inter- molecular filters while still having relatively weak optimized UniDock scores, whereas MolSnapper combines high inter-molecular validity with the best docking scores. Thus, inter-molecular filter success is useful for detecting physically implausible poses, but it is not by itself a sufficient proxy for poses quality. UniDock scores remain the more direct metric for ranking molecular poses, while PoseBusters filters should be interpreted as complementary validity checks rather than definitive measures of binder quality. 7 Limitations and Impact Limitation The lack of statistical significance in our experimental results is a limitation. This decision was made to instead conduct a better analysis of more available models. Another limitation is choosing random fragments as anchor fragments. Other works involving fixed fragments consider important substructures such as scaffolds. We consider fragments as a 3D constraint independent of chemical validity and the protein enviornment. Another limitation is the usage of one pharmacophore point and one mandatory interaction point when our developed sampling strategy supports more. Broader ImpactsPotential positive impacts of using off-the-shelf LLMs for drug discovery include reducing the costs required to develop new medicines. At the same time, these capabilities raise dual-use concerns, as widely available models could potentially be misused to generate harmful or toxic compounds. Although our work focuses on evaluation rather than deployment, it highlights the need for misuse-aware filtering, access controls in high-risk settings, and chemical-structure safeguards in future LLM-based molecular design systems. 8 Conclusion In this work, we introduce 3D-Fit, a benchmark for evaluating 3D molecular generation under increasingly complex spatial conditions. We also define textual representations for these conditions and propose a structured output format for generated 3D molecules. Our results show that current frontier LLMs exhibit a meaningful ability to parse spatial constraints instructions and generate molecules that satisfy explicit local 3D constraints. In particular, LLMs perform well when conditions are expressed in a form that can be directly copied, mirrored, or locally reconstructed in the output, such as anchor fragments and pharmacophore points. Moreover, adding more spatial conditions often improves pocket-related metrics, suggesting that additional seed-ligand information helps LLMs place generated molecules more accurately in 3D space. These findings indicate that LLMs possess emerging capabilities for instruction-following in 3D environments. However, our benchmark also reveals important limitations. Although LLMs often follow the desired spatial conditions, they frequently produce poses with non-optimal geometry. Before local UniDock optimization, their docking scores are consistently poor, often indicating severe steric clashes with the protein pocket. After optimization, the scores improve substantially, but still remain weaker than those of the best diffusion-based generators. LLM-generated molecules also show internal structural issues, as reflected by weaker intra-molecular validity compared to the strongest specialized models. Thus, local constraint satisfaction alone is not sufficient for effective structure-based design: generated molecules must also be physically plausible both internally and in their placement relative to the pocket. Overall, 3D-Fit highlights both the promise and the current shortcomings of off-the-shelf LLMs for 3D molecular design. LLMs are increasingly capable of handling multi-condition generation prompts, but they do not yet match diffusion models in physical plausibility and binding quality. We hope this benchmark will support future work on domain-specific training, improved 3D molecular representations, and more reliable evaluation for multi-constraint structure-based drug design. 9 Code and Data Availability The benchmark’s code data are accessible via the following link:https://github.com/ insilicomedicine/bench-3d-fit. 10 References [1]Xingang Peng, Shitong Luo, Jiaqi Guan, Qi Xie, Jian Peng, and Jianzhu Ma. Pocket2Mol: Efficient molecular sampling based on 3D protein pockets. In Kamalika Chaudhuri, Ste- fanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceed- ings of the 39th International Conference on Machine Learning, volume 162 of Proceed- ings of Machine Learning Research, pages 17644–17655. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/peng22b.html. [2] Jiaqi Guan, Wesley Wei Qian, Xingang Peng, Yufeng Su, Jian Peng, and Jianzhu Ma. 3D equivariant diffusion for target-aware molecule generation and affinity prediction. In The Eleventh International Conference on Learning Representations, 2023. URLhttps: //openreview.net/forum?id=kJqXEPXMsE0. [3]Arne Schneuing, Charles Harris, Yuanqi Du, Kieran Didi, Arian Jamasb, Ilia Igashov, Weitao Du, Carla Gomes, Tom L. Blundell, Pietro Lio, Max Welling, Michael Bronstein, and Bruno Correia. Structure-based drug design with equivariant diffusion models. Nature Computational Science, 4(12):899–909, Dec 2024. ISSN 2662-8457. doi: 10.1038/s43588-024-00737-x. URL https://doi.org/10.1038/s43588-024-00737-x. [4]Lei Huang, Tingyang Xu, Yang Yu, Peilin Zhao, Xingjian Chen, Jing Han, Zhi Xie, Hailong Li, Wenge Zhong, Ka-Chun Wong, and Hengtong Zhang. A dual diffusion model enables 3d molecule generation and lead optimization based on target pockets. Nature Communications, 15(1):2657, Mar 2024. ISSN 2041-1723. doi: 10.1038/s41467-024-46569-1. URLhttps: //doi.org/10.1038/s41467-024-46569-1. [5] Siyi Gu, Minkai Xu, Alexander S Powers, Weili Nie, Tomas Geffner, Karsten Kreis, Jure Leskovec, Arash Vahdat, and Stefano Ermon. Aligning target-aware molecule diffusion models with exact energy optimization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=EWcvxXtzNu. [6]Changda Gong, Jiaojiao Fang, Yan Tang, Guixia Liu, Yun Tang, and Weihua Li. SGEDiff: a subgraph-enriched diffusion model for structure-based 3d molecular generation. Journal of Cheminformatics, 17(1):175, Dec 2025. ISSN 1758-2946. doi: 10.1186/s13321-025-01123-z. URL https://doi.org/10.1186/s13321-025-01123-z. [7] Xudong Zhang, Sanqing Qu, Fan Lu, Jianmin Wang, Zhixin Tian, Shangding Gu, Yan- ping Zhang, Alois Knoll, Shaorong Gao, Guang Chen, and Changjun Jiang.Steering semi-flexible molecular diffusion model for structure-based drug design with reinforcement learning. Science Advances, 12(16):eady9955, 2026. doi: 10.1126/sciadv.ady9955. URL https://w.science.org/doi/abs/10.1126/sciadv.ady9955. [8]Junjie Xie, Sheng Chen, Jinping Lei, and Yuedong Yang. DiffDec: Structure-aware scaffold decoration with an end-to-end diffusion model. Journal of Chemical Information and Modeling, 64(7):2554–2564, 2024. doi: 10.1021/acs.jcim.3c01466. URLhttps://doi.org/10.1021/ acs.jcim.3c01466. PMID: 38267393. [9] Zhilin Huang, Ling Yang, Xiangxin Zhou, Zhilong Zhang, Wentao Zhang, Xiawu Zheng, Jie Chen, Yu Wang, Bin CUI, and Wenming Yang. Protein-ligand interaction prior for binding- aware 3d molecule diffusion models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=qH9nrMNTIW. [10] Yael Ziv, Fergus Imrie, Brian Marsden, and Charlotte M. Deane. MolSnapper: Conditioning diffusion for structure-based drug design. Journal of Chemical Information and Modeling, 65 (9):4263–4273, 2025. doi: 10.1021/acs.jcim.4c02008. URLhttps://doi.org/10.1021/ acs.jcim.4c02008. PMID: 40248896. [11] Masami Sako, Nobuaki Yasuo, and Masakazu Sekijima. Interaction-constrained 3d molec- ular generation using a diffusion model enables structure-based pharmacophore model- ing for drug design.npj Drug Discovery, 3(1):8, Mar 2026.ISSN 3005-1452.doi: 10.1038/s44386-026-00040-x. URL https://doi.org/10.1038/s44386-026-00040-x. 11 [12]Debjyoti Bhattacharya, Harrison J. Cassady, Michael A. Hickner, and Wesley F. Reinhart. Large language models as molecular design engines. Journal of Chemical Information and Modeling, 64(18):7086–7096, 2024. doi: 10.1021/acs.jcim.4c01396. URLhttps://doi.org/10.1021/ acs.jcim.4c01396. PMID: 39231030. [13]Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, and Chaowei Xiao. Conversational drug editing using retrieval and domain feedback. In The Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview. net/forum?id=yRrPfKyJQ2. [14]Botao Yu, Frazier N. Baker, Ziqi Chen, Xia Ning, and Huan Sun. LlaSMol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tuning dataset. In First Conference on Language Modeling, 2024. URLhttps://openreview. net/forum?id=lY6XTF9tPv. [15]Daniel Flam-Shepherd and Alán Aspuru-Guzik. Language models can generate molecules, materials, and protein binding sites directly in three dimensions as XYZ, CIF, and PDB files, 2023. URL https://arxiv.org/abs/2305.05708. [16] Artem Zholus, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Daniil Polykovskiy, Sarath Chandar, and Alex Zhavoronkov. BindGPT: A scalable framework for 3d molecular design via language modeling and reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence, 39(24):26083–26091, Apr. 2025. doi: 10.1609/aaai.v39i24.34804. URL https://ojs.aaai.org/index.php/AAAI/article/view/34804. [17]Maksim Kuznetsov, Airat Valiev, Alex Aliper, Daniil Polykovskiy, Elena Tutubalina, Rim Shayakhmetov, and Zulfat Miftahutdinov. nach0-pc: Multi-task language model with molecular point cloud encoder. Proceedings of the AAAI Conference on Artificial Intelligence, 39(23): 24357–24365, Apr. 2025. doi: 10.1609/aaai.v39i23.34613. URLhttps://ojs.aaai.org/ index.php/AAAI/article/view/34613. [18]Paul G. Francoeur, Tomohide Masuda, Jocelyn Sunseri, Andrew Jia, Richard B. Iovanisci, Ian Snyder, and David R. Koes. Three-dimensional convolutional neural networks and a cross- docked data set for structure-based drug design. Journal of Chemical Information and Modeling, 60(9):4200–4215, 2020. doi: 10.1021/acs.jcim.0c00411. URLhttps://doi.org/10.1021/ acs.jcim.0c00411. PMID: 32865404. [19]Janani Durairaj, Yusuf Adeshina, Zhonglin Cao, Xuejin Zhang, Vladas Oleinikovas, Thomas Duignan, Zachary McClure, Xavier Robin, Gabriel Studer, Daniel Kovtun, Emanuele Rossi, Guoqing Zhou, Srimukh Veccham, Clemens Isert, Yuxing Peng, Prabindh Sundareson, Mehmet Akdel, Gabriele Corso, Hannes Stärk, Gerardo Tauriello, Zachary Carpenter, Michael Bron- stein, Emine Kucukbenli, Torsten Schwede, and Luca Naef. PLINDER: The protein-ligand interactions dataset and resource. In ICML’24 Workshop ML for Life and Material Science: From Theory to Industry Applications, 2024. URLhttps://openreview.net/forum?id= 7UvbaTrNbP. [20] Renxiao Wang, Xueliang Fang, Yipin Lu, Chao-Yie Yang, and Shaomeng Wang. The PDBbind database: Methodologies and updates. Journal of Medicinal Chemistry, 48(12):4111–4119, 2005. doi: 10.1021/jm048957q. URLhttps://doi.org/10.1021/jm048957q. PMID: 15943484. [21]Swapnil Wagle, Richard D. Smith, Anthony J. Dominic, Debarati DasGupta, Sunil Kumar Tripathi, and Heather A. Carlson. Sunsetting Binding MOAD with its last data update and the addition of 3d-ligand polypharmacology tools. Scientific Reports, 13(1):3008, Feb 2023. ISSN 2045-2322. doi: 10.1038/s41598-023-29996-w. URLhttps://doi.org/10.1038/ s41598-023-29996-w. [22]Haitao Lin, Guojiang Zhao, Odin Zhang, Yufei Huang, Lirong Wu, Cheng Tan, Zicheng Liu, Zhifeng Gao, and Stan Z. Li. CBGBench: Fill in the blank of protein-molecule complex binding graph. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=mOpNrrV2zH. 12 [23]Haoyang Liu, Yifei Qin, Zhangming Niu, Mingyuan Xu, Jiaqiang Wu, Xianglu Xiao, Jinping Lei, Ting Ran, and Hongming Chen. How good are current pocket-based 3d generative models?: The benchmark set and evaluation of protein pocket-based 3d molecular generative models. Journal of Chemical Information and Modeling, 64(24):9260–9275, 2024. doi: 10.1021/acs.jcim.4c01598. URLhttps://doi.org/10.1021/acs.jcim.4c01598. PMID: 39629985. [24]Dou Nie, Huifeng Zhao, Odin Zhang, Gaoqi Weng, Hui Zhang, Jieyu Jin, Haitao Lin, Yufei Huang, Liwei Liu, Dan Li, Tingjun Hou, and Yu Kang. Durian: A comprehensive benchmark for structure-based 3d molecular generation. Journal of Chemical Information and Modeling, 65(1):173–186, 2025. doi: 10.1021/acs.jcim.4c02232. URLhttps://doi.org/10.1021/ acs.jcim.4c02232. PMID: 39681323. [25]Xingang Peng, Jiaqi Guan, Qiang Liu, and Jianzhu Ma. MolDiff: Addressing the atom- bond inconsistency problem in 3D molecule diffusion generation. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 of Pro- ceedings of Machine Learning Research, pages 27611–27629. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/peng23b.html. [26] Haitao Lin, Yufei Huang, Odin Zhang, Siqi Ma, Meng Liu, Xuanjing Li, Lirong Wu, Jishui Wang, Tingjun Hou, and Stan Z. Li. DiffBP: generative diffusion of 3d molecules for target protein binding. Chem. Sci., 16:1417–1431, 2025. doi: 10.1039/D4SC05894A. URLhttp: //dx.doi.org/10.1039/D4SC05894A. [27]Zhilin Huang, Ling Yang, Zaixi Zhang, Xiangxin Zhou, Yu Bao, Xiawu Zheng, Yuwei Yang, Yu Wang, and Wenming Yang. Binding-adaptive diffusion models for structure-based drug design. Proceedings of the AAAI Conference on Artificial Intelligence, 38(11):12671–12679, Mar. 2024. doi: 10.1609/aaai.v38i11.29162. URLhttps://ojs.aaai.org/index.php/ AAAI/article/view/29162. [28]Haotian Chen, Yiting Shen, Jichun Li, and Weizhong Zhao. An effective fragment-based dual conditional diffusion framework for molecular generation. Briefings in Bioinformatics, 27(1): bbaf727, 01 2026. ISSN 1477-4054. doi: 10.1093/bib/bbaf727. URLhttps://doi.org/10. 1093/bib/bbaf727. [29]Helen M. Berman, John Westbrook, Zukang Feng, Gary Gilliland, T. N. Bhat, Helge Weissig, Ilya N. Shindyalov, and Philip E. Bourne. The protein data bank. Nucleic Acids Research, 28 (1):235–242, 01 2000. ISSN 0305-1048. doi: 10.1093/nar/28.1.235. URLhttps://doi.org/ 10.1093/nar/28.1.235. [30]Arthur Dalby, James G. Nourse, W. Douglas Hounshell, Ann K. I. Gushurst, David L. Grier, Burton A. Leland, and John Laufer. Description of several chemical structure file formats used by computer programs developed at molecular design limited. Journal of Chemical Information and Computer Sciences, 32(3):244–255, May 1992. ISSN 0095-2338. doi: 10.1021/ci00007a012. URL https://doi.org/10.1021/ci00007a012. [31]Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov, Daniil Polykovskiy, Annika Brundyn, Aastha Jhunjhunwala, Anthony Costa, Alex Aliper, Alán Aspuru-Guzik, and Alex Zhavoronkov. nach0: multimodal natural and chemical languages foundation model. Chem. Sci., 15:8380–8389, 2024. doi: 10.1039/D4SC00966E. URLhttp://dx.doi.org/10. 1039/D4SC00966E. [32] Ahmet Bakan, Lidio M. Meireles, and Ivet Bahar. ProDy: Protein dynamics inferred from theory and experiments. Bioinformatics, 27(11):1575–1577, 06 2011. ISSN 1367-4803. doi: 10. 1093/bioinformatics/btr168. URL https://doi.org/10.1093/bioinformatics/btr168. [33]Greg Landrum, Paolo Tosco, Brian Kelley, Ric, David Cosgrove, sriniker, gedeck, Riccardo Vianello, NadineSchneider, Eisuke Kawashima, Gareth Jones, Dan N, Andrew Dalke, Brian Cole, Matt Swain, Samo Turk, AlexanderSavelyev, Alain Vaucher, Maciej Wójcikowski, Ichiru Take, Vincent F. Scalfani, Daniel Probst, Kazuya Ujihara, guillaume godin, Rachel Walker, Juuso Lehtivarjo, Axel Pahl, Francois Berenger, jasondbiggs, and strets123. rdkit/rdkit: 2023_09_3 (q3 2023) release, December 2023. URL https://doi.org/10.5281/zenodo.10275225. 13 [34]Oleg Trott and Arthur J. Olson. AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of Computational Chemistry, 31(2):455–461, 2010. doi: https://doi.org/10.1002/jcc.21334. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/jcc.21334. [35] Cédric Bouysset and Sébastien Fiorucci. ProLIF: a library to encode molecular interactions as fingerprints. Journal of Cheminformatics, 13(1):72, Sep 2021. ISSN 1758-2946. doi: 10.1186/s13321-021-00548-6. URL https://doi.org/10.1186/s13321-021-00548-6. [36]Jörg Degen, Christof Wegscheid-Gerlach, Andrea Zaliani, and Matthias Rarey. On the art of compiling and using ’drug-like’ chemical fragment spaces. ChemMedChem, 3(10):1503–1507, 2008. doi: https://doi.org/10.1002/cmdc.200800178. URLhttps://chemistry-europe. onlinelibrary.wiley.com/doi/abs/10.1002/cmdc.200800178. [37] Alina Kutlushina, Aigul Khakimova, Timur Madzhidov, and Pavel Polishchuk. Ligand- based pharmacophore modeling using novel 3d pharmacophore signatures. Molecules, 23 (12), 2018. ISSN 1420-3049. doi: 10.3390/molecules23123094. URLhttps://w.mdpi. com/1420-3049/23/12/3094. [38]OpenAI. GPT-5.5 system card, April 2026. URLhttps://deploymentsafety.openai. com/gpt-5-5/gpt-5-5.pdf. [39] OpenAI. GPT-5.4 thinking system card, March 2026. URLhttps://deploymentsafety. openai.com/gpt-5-4-thinking/gpt-5-4-thinking.pdf. [40]Anthropic. System card: Claude opus 4.7, April 2026. URLhttps://cdn.sanity.io/ files/4zrzovbb/website/037f06850df7fbe871e206dad004c3db5fd50340.pdf. [41]Anthropic.System card: Claude opus 4.6, February 2026.URLhttps://w-cdn. anthropic.com/6a5fa276ac68b9aeb0c8b6af5fa36326e0e166d.pdf. [42]Anthropic. System card: Claude sonnet 4.6, February 2026. URLhttps://w-cdn. anthropic.com/bbd8ef16d70b7a1665f14f306e88b53f686a75.pdf. [43]Gemini Team. Gemini 3.1 pro model card, February 2026. URLhttps://storage. googleapis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card. pdf. [44]xAI.Grok 4.1 model card, November 2025.URLhttps://data.x.ai/ 2025-11-17-grok-4-1-model-card.pdf. [45]Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URLhttps: //qwen.ai/blog?id=qwen3.5. [46]DeepSeek-AI, Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Fucong Dai, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoran Wei, Haowei Zhang, Haowen Luo, Haozhe Ji, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, Jialiang Huang, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jingchang Chen, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jinhua Zhu, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexin Huang, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Liang Zhao, Liangsheng Yin, Lihua Guo, Lingxiao Luo, Linwang Ma, Litong Wang, Liyue Zhang, M. S. Di, M. Y Xu, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingxu Zhou, Panpan Huang, Peixin Cong, Peiyi Wang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, S. H. Liu, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Songyang Zhou, Tao Ni, Tao Yun, Tian Pei, Tian Ye, Tianyuan Yue, Wangding Zeng, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjun Gao, Wentao Zhang, Xi Gao, Xiangwen Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, 14 Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingkai Yu, Xingyou Li, Xinyu Yang, Xinyuan Li, Xu Chen, Xuecheng Su, Xuehai Pan, Xuheng Lin, Xuwei Fu, Y. Q. Wang, Yang Zhang, Yanhong Xu, Yanru Ma, Yao Li, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, Yi Yu, Yichao Zhang, Yifan Ding, Yifan Shi, Yiliang Xiong, Ying He, Ying Zhou, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuduan Wang, Yue Gong, Yuhan Wu, Yuheng Zou, Yukun Li, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Z. F. Wu, Z. Z. Ren, Zehua Zhao, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhiyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Zizheng Pan, Zongqing Yao, Bei Feng, Hui Li, J. L. Cai, Jiaqi Ni, Lei Xu, Meng Li, Ning Tian, R. J. Chen, R. L. Jin, S. S. Li, Shuang Zhou, Tianyu Sun, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xinnan Song, Xinyi Zhou, Y. X. Zhu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, Dongjie Ji, Jian Liang, Jianzhong Guo, Jin Chen, Leyi Xia, Miaojun Wang, Mingming Li, Peng Zhang, Ruyi Chen, Shangmian Sun, Shaoqing Wu, Shengfeng Ye, T. Wang, W. L. Xiao, Wei An, Xianzu Wang, Xiaowen Sun, Xiaoxiang Wang, Ying Tang, Yukun Zha, Zekai Zhang, Zhe Ju, Zhen Zhang, and Zihua Qu. DeepSeek-v3.2: Pushing the frontier of open large language models, 2025. URL https://arxiv.org/abs/2512.02556. [47]GLM-5-Team, :, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, Chenzheng Zhu, Congfeng Yin, Cunx- iang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Haoran Wang, Huilong Chen, Jiajie Zhang, Jian Jiao, Jiaqi Guo, Jingsen Wang, Jingzhao Du, Jinzhu Wu, Kedong Wang, Lei Li, Lin Fan, Lucen Zhong, Mingdao Liu, Mingming Zhao, Pengfan Du, Qian Dong, Rui Lu, Shuang-Li, Shulin Cao, Song Liu, Ting Jiang, Xiaodong Chen, Xiaohan Zhang, Xuancheng Huang, Xuezhen Dong, Yabo Xu, Yao Wei, Yifan An, Yilin Niu, Yitong Zhu, Yuanhao Wen, Yukuo Cen, Yushi Bai, Zhongpei Qiao, Zihan Wang, Zikang Wang, Zilin Zhu, Ziqiang Liu, Zixuan Li, Bojie Wang, Bosi Wen, Can Huang, Changpeng Cai, Chao Yu, Chen Li, Chengwei Hu, Chenhui Zhang, Dan Zhang, Daoyan Lin, Dayong Yang, Di Wang, Ding Ai, Erle Zhu, Fangzhou Yi, Feiyu Chen, Guohong Wen, Hailong Sun, Haisha Zhao, Haiyi Hu, Hanchen Zhang, Hanrui Liu, Hanyu Zhang, Hao Peng, Hao Tai, Haobo Zhang, He Liu, Hongwei Wang, Hongxi Yan, Hongyu Ge, Huan Liu, Huanpeng Chu, Jia’ni Zhao, Jiachen Wang, Jiajing Zhao, Jiamin Ren, Jiapeng Wang, Jiaxin Zhang, Jiayi Gui, Jiayue Zhao, Jijie Li, Jing An, Jing Li, Jingwei Yuan, Jinhua Du, Jinxin Liu, Junkai Zhi, Junwen Duan, Kaiyue Zhou, Kangjian Wei, Ke Wang, Keyun Luo, Laiqiang Zhang, Leigang Sha, Liang Xu, Lindong Wu, Lintao Ding, Lu Chen, Minghao Li, Nianyi Lin, Pan Ta, Qiang Zou, Rongjun Song, Ruiqi Yang, Shangqing Tu, Shangtong Yang, Shaoxiang Wu, Shengyan Zhang, Shijie Li, Shuang Li, Shuyi Fan, Wei Qin, Wei Tian, Weining Zhang, Wenbo Yu, Wenjie Liang, Xiang Kuang, Xiangmeng Cheng, Xiangyang Li, Xiaoquan Yan, Xiaowei Hu, Xiaoying Ling, Xing Fan, Xingye Xia, Xinyuan Zhang, Xinze Zhang, Xirui Pan, Xu Zou, Xunkai Zhang, Yadi Liu, Yandong Wu, Yanfu Li, Yidong Wang, Yifan Zhu, Yijun Tan, Yilin Zhou, Yiming Pan, Ying Zhang, Yinpei Su, Yipeng Geng, Yong Yan, Yonglin Tan, Yuean Bi, Yuhan Shen, Yuhao Yang, Yujiang Li, Yunan Liu, Yunqing Wang, Yuntao Li, Yurong Wu, Yutao Zhang, Yuxi Duan, Yuxuan Zhang, Zezhen Liu, Zhengtao Jiang, Zhenhe Yan, Zheyu Zhang, Zhixiang Wei, Zhuo Chen, Zhuoer Feng, Zijun Yao, Ziwei Chai, Ziyuan Wang, Zuzhou Zhang, Bin Xu, Minlie Huang, Hongning Wang, Juanzi Li, Yuxiao Dong, and Jie Tang. GLM-5: from vibe coding to agentic engineering, 2026. URL https://arxiv.org/abs/2602.15763. [48]Xingang Peng, Ruihan Guo, Fenglin Guo, Ziyi Wang, Jiayu Sun, Jiaqi Guan, Yinjun Jia, Yan Xu, Yanwen Huang, Muhan Zhang, Jian Peng, Xinquan Wang, Chuanhui Han, Zihua Wang, and Jianzhu Ma. Unified modeling of 3d molecular generation via atomic interactions with pocketXMol. Cell, 189(7):1904–1922.e28, Apr 2026. ISSN 0092-8674. doi: 10.1016/j.cell. 2026.01.003. URL https://doi.org/10.1016/j.cell.2026.01.003. [49]Martin Buttenschoen, Garrett M. Morris, and Charlotte M. Deane. PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences. Chem. Sci., 15:3130–3139, 2024. doi: 10.1039/D3SC04185A. URLhttp://dx.doi.org/10.1039/ D3SC04185A. 15 [50]Yuejiang Yu, Chun Cai, Jiayue Wang, Zonghua Bo, Zhengdan Zhu, and Hang Zheng. Uni- Dock: GPU-accelerated docking enables ultralarge virtual screening. Journal of Chemical Theory and Computation, 19(11):3336–3345, 2023. doi: 10.1021/acs.jctc.2c01145. URL https://doi.org/10.1021/acs.jctc.2c01145. PMID: 37125970. 16 A Benchmarking templates General template for LLM benchmarking Generate the three-dimensional structure of a ligand molecule given the following 3D generation conditions: The ligand molecule must bind to the following protein pocket, meaning good interactions with the protein pocket atoms and no clashes (atoms with very close coordinates) with the protein pocket atoms: Atoms are grouped first by chain, then by residue, and are referenced using atom names from PDB-style protein notation. Pocket description The ligand molecule must contain the following atoms and their 3D coordinates within 0.5 Angstroms of the given coordinates: Anchor fragments description The ligand molecule must have the following pharmacophore points: Pharmacophore points description The ligand molecule must have the following interactions with specific residues in the protein pocket: Mandatory interactions description The most important condition is generating a valid molecule with plausible 2D and 3D structure. This means reasonable bond lengths, angles, and no overlapping with protein pocket. If all conditions are not satisfiable simultaneously, prioritize the validity of the generated ligand molecule. The output molecule must have a valid 2D and 3D structure enclosed in <sdf> and </sdf> tags. In this format, the molecule is provided as a simplified SDF/MolBlock string. Which has two parts, an atom block and a bond block. The atom block contains a line for each heavy atom (non-hydrogen atom), 6 values per line. First is the atom index, starting from 1, followed 3 values for the x y z coordinates in Angstroms with 3 decimal places each, fifth is the atomic symbol and sixth is the charge on the atom (0 for no charge, -1 for negative charge, +1 for positive charge) The bond block contains a line for each bond, 4 values per line. First is the start atom index from the atom block, the second is the end atom index from the atom block, the third is the bond order (1 for single, 2 for double, 3 for triple) and fourth is the stereoscopy of the bond (if applicable, 0 for none, 1 for up, 6 for down). Bonds do not have a unique identifier index, the start and end atom indices are used to identify the bond. Hydrogen atoms are not included in this format and are instead determined implicitly from the available valencies of the listed heavy atoms Simplified SDF representation example Important: Before finalizing, explicitly verify the output molecule: format validity, 3d plausibility, all heavy atoms have associated coordinates, and all conditions are satisfied. If uncertain or unable to satisfy all conditions, prefer a ligand molecule that satistifes the original validity conditions: The ligand molecule must be a valid molecule, meaning it is a single connected structure with correct valencies for all atoms. The ligand molecule must have a 3D structure, and the bond lengths and angles must obey chemical rules such that the strain energy of the molecule is minimal. The ligand molecule should be druglike and fill the pocket of the protein, forming good interactions with the protein pocket atoms without overlapping. To be druglike, the ligand molecule should satisfy Lipinski’s rule of five. The ligand molecule should have a number of heavy atoms between 20 and 40. The ligand molecule should have a number of hydrogen bond donors less than 5. The ligand molecule should have a number of hydrogen bond acceptors less than 10. The ligand molecule should have a logP value less than 5. Do not output any molecule unless all checks pass. If you see the problems with unrealistic atom coordinates, stop generating the current molecule, close the tags and generate a new molecule from scratch in a new set of mol tags. Repeat this process as necessary to output a valid molecule. 17 Simplified SDF representation example Forexample,amoleculewiththeSMILESstring ’Cc1c(C2C(F)(F)C2)n(Cc2cscn2)c1’ with (<x>, <y>, <z>) coordinates for each atom can be output as: <sdf> 1 30.987 6.655 24.653 C 0 2 31.351 7.912 24.029 C 0 3 30.624 8.775 23.180 C 0 4 31.510 9.726 22.702 C 0 5 31.368 10.813 21.681 C 0 6 30.007 11.245 21.155 C 0 7 28.976 11.750 22.162 C 0 8 27.672 11.818 21.343 C 0 9 27.928 11.007 20.084 C 0 10 26.905 10.209 19.755 F 0 11 28.130 11.857 19.044 F 0 12 29.209 10.235 20.333 C 0 13 32.716 9.446 23.320 N 0 14 34.031 10.014 23.112 C 0 15 35.271 9.181 23.283 C 0 16 35.399 7.840 22.977 C 0 17 36.876 7.257 23.558 S 0 18 37.373 8.844 23.863 C 0 19 36.438 9.775 23.782 N 0 20 32.587 8.443 24.204 C 0 1 2 1 0 2 3 1 0 3 4 2 0 4 5 1 0 5 6 1 0 6 7 1 0 7 8 1 0 8 9 1 0 9 10 1 0 9 11 1 0 9 12 1 0 4 13 1 0 13 14 1 0 14 15 1 0 15 16 2 0 16 17 1 0 17 18 1 0 18 19 2 0 13 20 1 0 20 2 2 0 12 6 1 0 19 15 1 0 </sdf> 18 Examples of conditions Pocket description example Format: <pdb_atom_name> at (<x>, <y>, <z>) Chain A: Residue 109, VAL: N at (68.271, 2.050, 1.527) CA at (67.422, 1.308, 2.444) C at (67.471, -0.161, 2.090) O at (67.515, -0.532, 0.919) CB at (65.933, 1.702, 2.355) CG1 at (65.410, 2.021, 3.730) CG2 at (65.733, 2.843, 1.411) ... Residue 228, PRO: N at (58.328, 5.229, 16.806) CA at (57.922, 4.667, 15.514) C at (56.907, 5.548, 14.808) O at (56.674, 5.326, 13.600) CB at (57.330, 3.320, 15.891) CG at (56.746, 3.565, 17.244) CD at (57.729, 4.482, 17.928) OXT at (56.356, 6.445, 15.480) Chain B: Residue 57, ALA: N at (53.770, -1.069, -1.125) CA at (55.007, -0.739, -0.422) C at (55.902, 0.060, -1.360) O at (56.089, -0.311, -2.525) CB at (55.704, -2.004, 0.009) ... Residue 186, TRP: N at (49.481, 14.438, 11.606) CA at (49.181, 13.268, 10.788) C at (48.290, 13.692, 9.624) O at (47.515, 14.636, 9.741) CB at (48.472, 12.187, 11.610) CG at (48.364, 10.864, 10.875) CD1 at (49.224, 9.798, 10.960) CD2 at (47.354, 10.487, 9.929) NE1 at (48.811, 8.787, 10.125) CE2 at (47.667, 9.181, 9.479) CE3 at (46.216, 11.126, 9.416) CZ2 at (46.883, 8.503, 8.540) CZ3 at (45.437, 10.451, 8.482) CH2 at (45.777, 9.150, 8.055) 19 Anchor fragments description example Format: Element Name (<atomic_symbol>) within 0.5 Angstroms of (<x>, <y>, <z>) Carbon (C) within 0.5 Angstroms of (57.658, 3.414, 8.907) Carbon (C) within 0.5 Angstroms of (58.645, 4.259, 8.129) Oxygen (O) within 0.5 Angstroms of (59.242, 3.567, 7.049) Carbon (C) within 0.5 Angstroms of (59.795, 4.701, 9.013) Oxygen (O) within 0.5 Angstroms of (59.550, 4.399, 10.428) Carbon (C) within 0.5 Angstroms of (60.108, 6.147, 8.872) Oxygen (O) within 0.5 Angstroms of (61.217, 6.406, 9.748) Carbon (C) within 0.5 Angstroms of (58.999, 7.101, 9.267) Pharmacophore points description example Format: <pharmacophore_type> within 0.5 Angstroms of (<x>, <y>, <z>) Hydrogen bond acceptor within 0.5 Angstroms of (54.373, 6.889, 9.107) Mandatory interactions description example Van der Waals interaction with any atom of residue 74, ILE from chain B 20 Opus 4.8 PocketXMol GPT 5.5 MolSnapper Figure 4: Generated structures for5TBOpocket: raw poses in magenta, UniDock-optimized in cyan. B Generated Structures and Unidock Score Distributions Figure 4 shows several generated examples from different models for the same target. The raw poses from LLMs are significantly different than the final optimized poses, whereas for diffusion models the poses are similar. Figure 6 shows the distributions of all Unidock scores from each model for both raw and optimized poses. This figure shows how significantly the LLM-generated poses are optimized compared to diffusion poses, as well as showing the difference in final optimized score between LLMs and Diffusion Models. C Alternative Format: SMILES+XYZ We additionally benchmarked another 3D ligand output representation, referred to as Enumerated SMILES+XYZ (see Fig. 5) Enumerated SMILES+XYZ For this format, we extend the compact SMILES+XYZ textual representation of 3D molecular structures introduced in BindGPT [16] and nach0-pc [17]. In these works, the molecular graph is first specified using a SMILES string, after which the 3D coordinates of the heavy atoms are provided in the same order as the corresponding atoms appear in the SMILES representation. We extend this format in a similar to Simplified SDF way by applying an index to each atom in SMILES format, and to its corresponding 3D coordinates. We relax the format fragility by supporting an arbitrary ordering of atomic coordinates as long as coordinate entries cover all heavy atoms in the original SMILES string. Comparison with Simplified SDFWe benchmarked LLMs on the full-conditioning task (pocket, anchor fragments, pharmacophore points, and mandatory interactions) and found that Simplified SDF consistently outperforms Enumerated SMILES+XYZ (see Table 5). Enumerated SMILES+XYZ [c:1]1[c:2][c:3][n:4][c:5][c:6]1| 4 N 1.092 -0.768 0.027|2 C 0.120 1.403 0.009|3 C 1.213 0.555 0.036|1 C -1.123 0.785 -0.028| 6 C -1.216 -0.591 -0.036|5 C -0.087 -1.384 -0.008 Figure 5: Example of Enumerated SMILES+XYZ representation 21 DS ModelFMT Valid Mol PB Intra Internal Energy PB Inter Pocket Uni- Dock PH4. 3D Anch. Mand. Int. CrossDocked2020 GPT-5.4 SDF 0.979 0.4200.530 0.3080.608 -5.320 0.730 0.977 0.242 SMI 0.808 0.0910.230 0.3420.665 -5.183 0.742 0.660 0.271 Sonnet 4.6 SDF 0.678 0.4050.530 0.4390.441 -4.844 0.519 0.678 0.175 SMI 0.467 0.2820.344 0.3620.290 -4.594 0.430 0.457 0.124 Opus 4.6 SDF 0.588 0.4000.504 0.2300.466 -5.465 0.436 0.583 0.228 SMI 0.440 0.2510.341 0.2100.363 -5.368 0.432 0.431 0.175 Gemini 3.1 Pro SDF 0.965 0.7200.885 0.5430.827 -5.246 0.733 0.965 0.487 SMI 0.969 0.7180.891 0.5420.804 -5.129 0.893 0.969 0.480 Qwen 3.5 SDF 0.898 0.3390.496 0.6400.574 -4.639 0.670 0.891 0.255 SMI 0.823 0.2820.398 0.6670.442 -4.278 0.709 0.787 0.191 GLM 5 SDF 0.400 0.2700.274 0.3440.226 -4.275 0.350 0.399 0.119 SMI– PLINDER GPT-5.4 SDF 0.990 0.3860.448 0.3270.479 -4.732 0.790 0.988 0.317 SMI 0.812 0.0570.133 0.3540.507 -4.609 0.768 0.614 0.240 Sonnet 4.6 SDF 0.685 0.4400.554 0.3960.317 -4.246 0.578 0.683 0.230 SMI 0.434 0.2340.299 0.2790.176 -3.913 0.404 0.416 0.158 Opus 4.6 SDF– SMI 0.489 0.3150.388 0.2200.277 -4.463 0.473 0.483 0.218 Gemini 3.1 Pro SDF 0.978 0.6790.895 0.5170.707 -4.771 0.786 0.978 0.594 SMI 0.960 0.6590.877 0.4950.640 -4.629 0.909 0.960 0.584 Qwen 3.5 SDF 0.869 0.3190.450 0.5980.400 -4.138 0.685 0.863 0.309 SMI 0.770 0.2080.299 0.5840.250 -3.741 0.677 0.713 0.204 GLM 5 SDF 0.420 0.2670.281 0.2890.152 -3.753 0.382 0.414 0.172 SMI– Table 5: 3D Quality Metrics and Condition Successes for Simplified SDF and Enumerated SMILES+XYZ formats. 22 0 50 100 150 200 UniDock Score (kcal/mol) Crossdocked (Raw) Proprietary LLMsOpen-weight LLMsSpecialized Models 12 10 8 6 4 UniDock Score (kcal/mol) Crossdocked (Optimized) 0 50 100 150 200 UniDock Score (kcal/mol) Plinder (Raw) GPT 5.5GPT 5.4 Opus 4.8Opus 4.7Opus 4.6 Sonnet 4.6 Gemini 3.1-Pro Grok 4.1Grok 4.3 Qwen 3.5 397B GLM 5 PocketXMol DiffSBDD PMDM SeFMol IPDiff DiffPharma MolSnapper BindDM 10 8 6 4 2 UniDock Score (kcal/mol) Plinder (Optimized) Figure 6: Distributions of Unidock scores for all targets and condition sets, compared between models. 1% and 99% outliers are removed. 23