Skip to main navigation Skip to search Skip to main content

SARS-CoV2 Docking Dataset for MLMol Language Model (50M)

Dataset

Description

This is a processed molecular dataset from this https://doi.ccs.ornl.gov/ui/doi/348 adding up to 50M molecules for the training and 486K molecules for the validation. Instructions on how to use/run/train this dataset can be found here: https://code.ornl.gov/candle/mlmol
Date made availableMay 20 2022
PublisherOak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States). Oak Ridge Leadership Computing Facility (OLCF); Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)

Cite this