Loading…
PDBx/mmCIF Ecosystem: Foundational Semantic Tools for Structural Biology
[Display omitted] •Community Driven Data Representation for structural biology data.•Provides common data representation and tools that accelerates scientific discovery.•Provides data and software infrastructure for the Protein Data Bank (PDB) Core Archive.•Provides essential data infrastructure sup...
Saved in:
Published in: | Journal of molecular biology 2022-06, Vol.434 (11), p.167599-167599, Article 167599 |
---|---|
Main Authors: | , , , , , , , , , , , , , , , , , , , , , |
Format: | Article |
Language: | English |
Subjects: | |
Citations: | Items that this one cites Items that cite this one |
Online Access: | Get full text |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Summary: | [Display omitted]
•Community Driven Data Representation for structural biology data.•Provides common data representation and tools that accelerates scientific discovery.•Provides data and software infrastructure for the Protein Data Bank (PDB) Core Archive.•Provides essential data infrastructure supporting FAIR data delivery.
PDBx/mmCIF, Protein Data Bank Exchange (PDBx) macromolecular Crystallographic Information Framework (mmCIF), has become the data standard for structural biology. With its early roots in the domain of small-molecule crystallography, PDBx/mmCIF provides an extensible data representation that is used for deposition, archiving, remediation, and public dissemination of experimentally determined three-dimensional (3D) structures of biological macromolecules by the Worldwide Protein Data Bank (wwPDB, wwpdb.org). Extensions of PDBx/mmCIF are similarly used for computed structure models by ModelArchive (modelarchive.org), integrative/hybrid structures by PDB-Dev (pdb-dev.wwpdb.org), small angle scattering data by Small Angle Scattering Biological Data Bank SASBDB (sasbdb.org), and for models computed generated with the AlphaFold 2.0 deep learning software suite (alphafold.ebi.ac.uk). Community-driven development of PDBx/mmCIF spans three decades, involving contributions from researchers, software and methods developers in structural sciences, data repository providers, scientific publishers, and professional societies. Having a semantically rich and extensible data framework for representing a wide range of structural biology experimental and computational results, combined with expertly curated 3D biostructure data sets in public repositories, accelerates the pace of scientific discovery. Herein, we describe the architecture of the PDBx/mmCIF data standard, tools used to maintain representations of the data standard, governance, and processes by which data content standards are extended, plus community tools/software libraries available for processing and checking the integrity of PDBx/mmCIF data. Use cases exemplify how the members of the Worldwide Protein Data Bank have used PDBx/mmCIF as the foundation for its pipeline for delivering Findable, Accessible, Interoperable, and Reusable (FAIR) data to many millions of users worldwide. |
---|---|
ISSN: | 0022-2836 1089-8638 |
DOI: | 10.1016/j.jmb.2022.167599 |