Attachment Points and CAP Groups

Prev Next

A monomer is not a molecule but a fragment of one: it has open bonds where it joins its neighbours. This article explains how those bonds, the attachment points or R-groups, are written in a SMILES, what a cap group is, and why the service insists on both.

A monomer is a residue, not a free molecule

Alanine as a free amino acid is C[C@H](N)C(=O)O. Inside a peptide it has lost a hydrogen from its amine and a hydroxyl from its carboxylic acid, and it is bonded to two neighbours instead. The monomer registered in the library is that residue, with the two open bonds marked:

C[C@H](N[*:1])C([*:2])=O

Each [*:n] is a dummy atom standing for a bond to something else. The number is the R-group label: R1 and R2 here. HELM connects monomers by naming the R-groups on each side, so R2 of one residue bonds to R1 of the next.

The conventions behind the numbers

The numbers are a convention, not a chemical fact, and the tools that build sequences rely on it.

Polymer type Monomer type R1 R2 R3
PEPTIDE Backbone Amine nitrogen, N-terminal side Carbonyl carbon, C-terminal side Side chain, when the residue can branch, such as the epsilon amine of lysine or the thiol of cysteine
PEPTIDE Branch Bond to the backbone residue's R3
RNA Backbone sugar 5' oxygen, to the previous phosphate 3' oxygen, to the next phosphate 1' carbon, to the base
RNA Backbone phosphate To the 3' of the previous sugar To the 5' of the next sugar
RNA Branch base Bond to the sugar's R3
CHEM Undefined Free choice, but numbered from R1 without gaps

Labels run from R1 to R25. They must start at R1 and increase without a gap: a monomer with R1 and R3 but no R2 is rejected.

Notations the service accepts

The service normalises every structure to the [*:n] form, and that is the form it shows when you reopen a monomer. It also reads:

Notation Example Where you meet it
Atom map on a dummy atom, [*:n] C[C@H](N[*:1])C([*:2])=O The recommended form, used by the import file and the form.
CXSMILES R-labels C[C@H](N[*])C([*])=O \|$;;;_R1;;_R2$\| Exported by chemistry toolkits; the label block after the \| names each dummy atom.
Folded cap atom, [H:1], [OH:2] C[C@H](N[H:1])C([OH:2])=O The original Pistoia HELM monomer library. The cap atom itself carries the map number.

[R1] as a bracket atom, and an unnumbered [*], are not attachment points. The first is not valid SMILES for the service; the second is a dummy atom with no label and is rejected with Number of R-Groups in Attachmentlist is NOT equal to number of R-Groups in Smiles.

Cap groups: what sits on an unused attachment point

An attachment point that is not bonded to a neighbour still has to be something. The cap group is the small group that completes the valence there: for a peptide backbone, a hydrogen on R1 gives the free amine, and a hydroxyl on R2 gives the free carboxylic acid. The cap is what the toolkit adds when it draws or computes a monomer at the end of a chain, or on its own.

The service stores the cap group as a name from a controlled vocabulary, and turns it into chemistry with a SMILES template where {n} is the R-group number:

Cap group Template Chemistry it completes
H [*:{n}][H] An amine, a thiol, an alcohol, an alkyl chain end. The default.
OH O[*:{n}] A carboxylic acid on R2 of an amino acid, a phosphate oxygen.
NH2 N[*:{n}] A C-terminal amide, a guanidine or amine attached through its carbon.
Azide [*:{n}]N=[N+]=[N-] An azido handle for click chemistry.
Ethynyl [*:{n}]C#C A terminal alkyne, the other partner of a click reaction.

Cap groups are written in the import file as R1:H;R2:OH and chosen from a selector per attachment point on the form. A cap outside the vocabulary is rejected with the list of valid values, never silently replaced: substituting a hydroxyl for an unknown cap would store a different molecule without anyone noticing. Your Discngine contact can add a cap group to your organisation's vocabulary when your chemistry needs one.

Why a wrong cap matters

The cap decides the molecule that tools compute from your monomer when the attachment point is free:

  • R2:H on an amino acid makes its free form an aldehyde, and the molecular weight and logP of every peptide ending on it are wrong;
  • R1:OH on an amino acid makes the N-terminus a hydroxylamine;
  • a cap chosen on the wrong side of a multi-atom template, such as writing the azide as if it were a single atom, gives a pentavalent carbon.

The cap does not change how the monomer bonds inside a chain. Two monomers with the same skeleton and different caps are the same residue in the middle of a sequence, and different molecules at its ends.

Related