Sequence Notation and Residue Numbering
Sequence notation is the convention for writing a peptide as one-letter or three-letter residue codes from N-terminus to C-terminus, together with the numbering used to identify positions.
A peptide sequence is written from the N-terminus on the left to the C-terminus on the right, using either three-letter codes such as His-Ala-Glu or the compact one-letter alphabet in which those become HAE. The one-letter code is not systematic: some letters are the residue's initial, others were assigned for the sound of the name or because the obvious letter was taken, which is why tryptophan is W, glutamine Q, phenylalanine F and asparagine N. Anything outside the standard twenty has no one-letter equivalent and has to be spelled out in brackets, so a modified analog is usually written in mixed notation with the non-standard residues named in full.
Numbering runs from the N-terminal residue as position one, and modifications are conventionally written as the residue with its position, as in a bracketed D-alanine at position two. Disulfide bonds are given as pairs of cysteine positions. Terminal modifications are appended, most commonly a trailing amide notation for a C-terminal amide.
Reading this correctly is what lets you tell two analogs apart. A single bracketed substitution is often the entire difference between a native hormone with a two-minute half-life and a weekly drug, and between two research compounds sold under similar names.
The reliable source of confusion is which numbering scheme is in force. GLP-1 is quoted both in proglucagon-derived numbering, where the active hormone begins at residue seven and the DPP-4 site is described at position eight, and in a renumbered form where the same hormone starts at one and that site is position two. Both are correct and they refer to the same atom. Listings that mix the two, or that quote a substitution position without saying which convention it belongs to, produce sequences that cannot be checked against anything.