Dataset / Documentation

Understand the data.

Dataset card, file contracts and evaluation guidance.

Start with the metadata.

  1. Search or filter the trajectory catalog.
  2. Inspect sequence, chains, partition and QC information on a detail page.
  3. Download individual files or copy the whole-sample command when Hugging Face hosting is available. Metadata CSV export is available now.
  4. When data hosting opens, retrieve matched XTC/GRO/PDB files and verify their SHA-256 values.

A runnable trajectory-reading example and evaluator will follow with the data release. They are Coming soon in this preview.

trajectories/{md_id}/rna.xtc
trajectories/{md_id}/rna.gro
trajectories/{md_id}/rna.pdb
metadata/sample_metadata.csv
metadata/residue_mapping.csv
metadata/qc_summary.csv
splits/{train,val,test_struct,test_flex}.csv

Source & curation

RNA-Solo provides experimental RNA structures. Original PDB/mmCIF records restore source molecular context. RNA-Solo target chains are retained; additional RNA chains require reliable assignments, heavy-atom contact within 5 Å after removal of non-RNA components, and the partner-size limit. When partners are retained, combined RNA length is limited to 512 nt.

Proteins, DNA, organic ligands and RNA chains outside the retention policy are removed. Unresolved backbone gaps, unsupported chemistry and untraceable topology changes are excluded.

Simulation protocol

ParameterSetting
EngineGROMACS
RNA force fieldAmber RNA OL3
WaterOPC, explicit
SaltNeutralization + 150 mM NaCl
Production100 ns · 300 K · 1 bar
Thermostat / barostatNosé–Hoover / Parrinello–Rahman
Saved coordinatesEvery 100 ps · RNA only
ExcludedEquilibration, water and ions

A final 5-ns NPT equilibration is excluded from released production. Detailed configuration files: Coming soon.

Files, mappings & units

md_id is the trajectory join key. rnasolo_id identifies the RNA-level source; split_identity_id records the corrected grouping identity used for partitioning.

Use matching release structures with each XTC. GRO and XTC store distances in nm; PDB uses Å. Readers may convert units. Website motion summaries are in Å. Sequence concatenation follows recorded chain boundaries, which are zero-based, end-exclusive. Released PDB chain labels can differ from canonical mmCIF chain IDs; detail pages show that mapping.

Coordinates follow the release coordinate policy. Fitted boxes are not physical simulation boxes; do not apply a new minimum-image convention based on them.

UseFramesIntervalTime range
Released trajectories1001100 ps0–100 ns
Benchmark references1011 ns0–100 ns

Motion summaries & quality control

The catalog’s median RMSD is the median, over all 1,001 frames, of heavy-atom RMSD to released frame 0 without an additional alignment. Heavy atoms exclude names beginning with H. It is a descriptive release-coordinate statistic, not the pairwise C1′ RMSD used during flexibility split construction and not a residue RMSF.

Mean radius of gyration is averaged over frames using heavy-atom coordinates relative to their unweighted centroid.

Coordinate integrity, backbone continuity, periodic-boundary behavior and multi-chain separation are checked. Sustained chain separation triggers exclusion and re-simulation of separated components. Current and inherited diagnostic warnings are shown separately. A pass does not establish convergence.

RMSF describes per-residue fluctuation magnitude; covariance includes directionality; normalized motion coupling (NMC) describes coupling between residue motions. Per-residue descriptors are Coming soon.

Evaluation protocol

View manuscript benchmark results →

PartitionPurpose
trainModel training
valModel and hyperparameter selection
test_structStructural generalization
test_flexHigh-flexibility generalization
guardAuditing only; excluded from training, model selection and formal evaluation

Partition counts & proportions → · Download split files →

Related trajectories are grouped by identity, source, lineage, sequence and structural similarity before partitioning. Partitions preserve related groups; never randomly split trajectory frames. Select checkpoints using validation, then evaluate the same selected checkpoint on both test panels. Guards do not participate in training, selection or formal evaluation.

Generation uses an initial conformer and a 101-frame, 1-ns grid. Geometry uses the future 100 frames; RMSF uses all 101. Metrics aggregate within each RNA and then across RNAs. Most manuscript methods use K=4 generated samples; ConfRover uses K=1.

Single-conformer prediction evaluates covariance, RMSF and NMC. Direct predictors and frozen-representation probes remain separate in the results. Evaluator artifact/version and runnable reproduction: Coming soon.

Intended use & limitations

RNA dynamics learning, trajectory generation and dynamics descriptor prediction within this dataset’s simulation conditions. Each production trajectory covers 100 ns. The release does not represent full protein–RNA or ligand environments, chemical reactions, or comprehensive sampling of slower processes.

Sampled conformations and motion summaries alone do not establish equilibrium, free energies or kinetic rates. RNA–protein comparisons are conditional on the matched cohorts, atom choices and simulation protocols.

Version, license & citation

RNADynBench-v0.1-revision-20260920-x0

The 20260920 revision updated split membership and grouping identity. The x0 revision regenerated initial PDB/GRO coordinates from saved XTC frame 0. Trajectories, metadata and splits were unchanged by that serialization repair; structure hashes changed.

Paper / arXiv
Coming soon
Authors & contact
Coming soon
Dataset DOI
Coming soon
License
Coming soon
BibTeX
Coming soon
Code & weights
Coming soon