Canonical SMILES and Duplicate Detection

Prev Next

The same molecule can be written in countless ways. This article explains how the service decides that two monomers are the same structure, what that decision does and does not take into account, and why a missing stereocentre is the most common cause of an unexpected duplicate error.

One molecule, many SMILES

C[C@H](N[*:1])C([*:2])=O, [*:1]N[C@@H](C)C(=O)[*:2] and O=C([*:2])[C@@H](C)N[*:1] are three spellings of L-alanine as a residue. A library that compared text would register three monomers. A library that compared nothing would let a curator register alanine under a new symbol by mistake, and every alignment would then see two different residues where there is one.

The service takes a middle path. On every write, it converts the structure to a canonical isomeric SMILES: a single spelling that a given molecule always produces, whatever the input order, coordinates or notation. That string is the monomer's canonical key. It is shown in the Version Log as Canonical key, and it is what duplicate detection compares.

What the canonical key keeps and what it drops

The key keeps everything that makes a molecule a different molecule:

  • the connectivity and bond orders;
  • the elements, charges and isotopes;
  • tetrahedral and double-bond stereochemistry written with @, @@, / and \.

It normalises everything that is only presentation:

  • atom order and ring-closure numbering;
  • 2D coordinates;
  • the notation of attachment points: [*:1], a CXSMILES _R1 label and a folded [H:1] all become the same dummy atom;
  • CXSMILES extension blocks, including enhanced stereo groups such as &1 or or1. Only the absolute configuration written on the atoms survives.

That last point has a consequence worth knowing. Two monomers drawn as enantiomers, but whose stereocentre is declared in an and group rather than as absolute, produce the same key and collide. The service is right, the input is ambiguous: an and group says "either configuration", so both drawings describe the same racemate. Declare the centre as absolute when you mean one enantiomer.

The duplicate rules

When a monomer is created or updated, three checks run against the live monomers of the same polymer type. A match with a deleted monomer is handled separately: the registration is held until an administrator revives the deleted monomer or the correspondence is judged. See Versioning and Deletion.

Check Scope Error
Same canonical key as another live monomer Your organisation and the public library Monomer with this structure is already registered: meG. If you expected a distinct structure, verify your input carries full stereochemistry …
Same symbol as a public monomer Public layer Cannot modify protected public monomer 'Nva' (public seed data)
Same symbol as one of your monomers, on creation Your organisation Monomer with this symbol is already registered: MeA

The symbol rule within your own organisation is handled differently by each entry point. The form refuses to create a second monomer with an existing symbol. A bulk import treats a matching symbol as the same monomer being updated, so a file can be re-imported to correct a library. See How to Import Monomers in Bulk.

The name is also unique within a polymer type of your organisation, so two monomers cannot share a full name even with distinct symbols.

Why "duplicate" usually means "missing stereo"

The error most curators meet is a duplicate structure they did not expect. In almost every case, the new SMILES has no stereo tag on a centre that the existing monomer defines, or the reverse.

Suppose Nle is registered as CCCC[C@H](N[*:1])C([*:2])=O and a colleague tries to register D-norleucine as CCCCC(N[*:1])C([*:2])=O. The second SMILES has no @, so it describes the racemate, and the racemate is a different key from the L form: no collision, and a wrong monomer is now in the library. Later, someone registers the D form properly as CCCC[C@@H](N[*:1])C([*:2])=O: still no collision. The library holds L, D and racemic norleucine under three symbols, and only one of them was intended.

Now the reverse. A residue was registered without its stereo tag, and the curator re-registers the L form under a new symbol. The keys differ, so it goes through, and the racemic monomer stays behind as a trap. The right move is to edit the existing monomer, not to register a second one.

The duplicate message says as much: If you expected a distinct structure, verify your input carries full stereochemistry (some molfiles omit stereocenters, e.g. sulfoxide R/S). Sulfoxides are the classic case: many drawing tools do not write the sulfur stereocentre at all, so the R and S sulfoxides of a residue produce one key.

Attachment points and caps in the key

The dummy atoms are part of the key, so a residue with two attachment points and the same residue with three are different monomers. The cap groups are not in the key: they live with the attachment point definition, and a change of cap is a material change tracked by versioning rather than a duplicate matter. See Attachment Points and CAP Groups.

Practical rules

  • Write full stereochemistry on every centre you know. Use X-style unknowns only when the centre is genuinely undefined.
  • Declare a single enantiomer as absolute, not as an and or or group.
  • When you get a duplicate error, open the existing monomer named in the message and compare the two drawings before doing anything else. Nine times out of ten the fix is to edit it, not to register another.
  • Prefer the public monomer when it exists. Your sequences stay comparable with everyone else's.

Related