OpenPLBind 2021+ v2
Released: 2026-07-25

Numeric-ready release
---------------------
Entries: 4,039 (one structure and one selected Kd, Ki, or IC50 label per PDB)
Classic: 3,024
Extended: 1,015
Validated ligand SDF: 3,116
Point labels: 4,031
Interval labels: 8
Refined Set v2: 727

Label normalization
-------------------
The published v2 files contain the corrected normalization. Sixty-six source
rows changed during validation (scientific notation, comma-formatted values,
one p-metric qualifier, and two replicate aggregates). Post-fix validation
found zero raw-to-nM, pActivity, qualifier-direction, interval, or replicate
issues.

Coordinate scope
----------------
The release uses a PDBbind-style per-entry layout and task semantics. It is not
the same coordinate preparation as the original PDBbind biological-unit
workflow. Protein coordinates retain all non-focal polypeptide chains from the
first deposited model. Assembly expansion, structure repair, protonation, water
retention, and retention of other nonpolymers are not applied. Pockets contain
complete protein residues with any heavy atom within 10 A of the selected
ligand. The publishable structure directory contains physical copies, not hard
links.

Structure package
-----------------
Source directory:
C:\Users\ustc\Documents\Multi-agent Self-evolving Crowdsourceing\output\openplbind_training_set_numeric_ready_20260725\structures

Uncompressed files: 39,467
Uncompressed bytes: 9,246,446,302
Uncompressed GiB: 8.611

Public files
------------
INDEX_openplbind_PL_data.4039
    Compact PDBbind-style labels with corrected pActivity values.
training_manifest.tsv
    Complete machine-readable labels, provenance, assay context, and paths.
structure_index.tsv
    Per-entry structure file index.
numeric_ready_summary.json
    Release counts and normalization validation summary.
OpenPLBind_refined_set_v2.tsv
    Final 727-entry Refined Set.
OpenPLBind_refined_selection_audit_v2.zip
    Refined selection candidates, audit tables, rules, and summary.
