How to Import Monomers in Bulk

Prev Next

This guide shows how to register many monomers at once from a CSV or Excel file, follow the import, and fix the rows that failed. The worked example is a small library a peptide project typically needs: two hydrophobic replacements and two PEG linkers, one of them with a deliberate mistake.

Prerequisites: the Manager or Administrator role in Biotoolkit Monomers, and the SMILES of every monomer with its attachment points written as [*:n].

Step 1: Download the template

On the Home page, click Import Monomer, then Need Help. The Import Monomer Help dialog summarises the rules and offers a Template File button, which downloads monomer_import_template.xlsx. Start from it: the column headers must match exactly, and their order does not matter.

The Import Monomer Help dialog listing what the import does and the important notes, with the Template File button

Step 2: Fill in one row per monomer

The columns:

Column Required Content
SYMBOL yes The symbol HELM sequences will use. Unique within your organisation for the polymer type, and not held by a public monomer.
NAME yes Full name.
SMILES yes Structure with [*:n] attachment points.
POLYMER_TYPE yes PEPTIDE, RNA or CHEM.
MONOMER_TYPE yes Backbone or Branch; Undefined for CHEM.
NATURAL_ANALOG for PEPTIDE and RNA Single-letter code of the natural residue, for example L. - for CHEM.
AUTHOR no Free text.
RGROUPS no Cap group per attachment point, as R1:H;R2:OH.

The full column reference is in Import File Format.

The example library:

SYMBOL NAME SMILES POLYMER_TYPE MONOMER_TYPE NATURAL_ANALOG RGROUPS
Ahp 2-Aminoheptanoic acid CCCCC[C@H](N[*:1])C([*:2])=O PEPTIDE Backbone L R1:H;R2:OH
MeNva N-methyl-norvaline CCC[C@H](N(C)[*:1])C([*:2])=O PEPTIDE Backbone V R1:H;R2:OH
PEG2 Diethylene glycol linker [*:1]OCCOCCO[*:2] CHEM Undefined - R1:H;R2:OMe
PEG3 Triethylene glycol linker [*:1]OCCOCCOCCO[*:2] CHEM Undefined - R1:H;R2:H

Three things to check in your own rows:

  • The symbol is free. The public library already holds the common non-natural residues: norleucine is Nle, norvaline is Nva, sarcosine is meG, 2-aminoisobutyric acid is Aib. A row that reuses a public symbol fails, and a row whose structure is already public fails too. Browse the Home report, or search the name, before inventing a symbol. See Public Monomers.
  • Attachment points are numbered. [*] without a number is not an attachment point.
  • Cap groups come from the vocabulary. H, OH, NH2, Azide and Ethynyl are shipped. The PEG2 row above uses OMe, which is not one of them, on purpose: it is the row that will fail. See Attachment Points and CAP Groups.

Save the file as .xlsx or .csv.

Step 3: Upload the file

On the Monomer Import page, drop the file on the Import Monomer File zone or click it to browse, then click Import Monomer. The import runs in the background: the page returns immediately, and the File History report below refreshes every few seconds until the import finishes.

The Monomer Import page: the drop zone and Import Monomer button, and the File History report with two finished files, both with a Download button in the Error File column

Each row of File History is one file. Status shows where it stands, Summary gives the row counts, and Duration Seconds how long it took. Show more columns from the report's Actions menu if you want the parsed, succeeded and failed counts side by side.

A short delay between the upload and the appearance of the monomers on the Home page is normal: each structure is validated and drawn by the Biotoolkit service before it is stored.

Step 4: Read the outcome

The file above ends with Status Error and the summary Processed 4 rows: 3 succeeded, 1 failed. The status is about the file: a single failed row is enough for Error, and the three other monomers are in the library all the same.

A row of your file has one of three outcomes:

  • Succeeded, created. The symbol did not exist for that polymer type in your organisation.
  • Succeeded, updated. A private monomer with the same symbol and polymer type already existed, and the file's values replaced it. If the structure or the classification changed, the previous state is archived as a version. See Versioning and Deletion.
  • Failed. The row was rejected. It does not stop the other rows.

Public monomers are never touched by an import, and a row that carries the symbol or the structure of a public monomer fails.

On the Home page, monomers created by an import show Unknown in Created By. The File History row is where the uploader is recorded.

Step 5: Fix the failed rows

When at least one row failed, the Error File column of File History shows a Download button. The error file holds the failed rows of your file with one extra column, ERROR_MESSAGE.

For the PEG2 row it reads:

Invalid cap group "OME" for R2. Valid values: Azide, Ethynyl, H, NH2, OH

Correct the value to R2:H in your file, keep only that row, and upload it again. Re-importing a row that already succeeded is harmless: it is reported as updated, and no version is written when nothing material changed.

The messages you are most likely to meet:

Message Cause
Symbol is required, Name is required, SMILES is required, Polymer type is required, Monomer type is required A mandatory column is empty on that row.
Unknown polymer type: … / Unknown monomer type: … A value outside the lists. The message ends with the expected values.
Invalid cap group "…" for R2. Valid values: … A cap outside the vocabulary.
Invalid R-group format "…". Expected format: R1:H or R2:OH A pair in RGROUPS is not label:cap.
Cannot modify protected public monomer 'Nva' (public seed data) A public monomer holds that symbol. Use it, or choose another symbol.
Monomer with this structure is already registered: … The same structure is already registered under another symbol, yours or public.

If the message is An error has occurred, check your file format … at upload time, the file itself could not be read: check the extension, the header row, and that the file is not empty.

Next steps