The twenty building blocks
Most peptides are built from a set of twenty standard amino acids. Each has its own chemical character, and each is represented by a three-letter abbreviation (for example, Gly for glycine) and a single letter (G). A peptide sequence is simply these codes written in order, from one defined end of the molecule to the other.
Reading a sequence
Sequences are conventionally written from the amino end (the N-terminus) to the carboxyl end (the C-terminus). Because the order carries the identity of the molecule, sequences are written precisely and are referenced in a peptide’s documentation. Modifications to the ends or to individual residues are also noted, as they change the molecule.
Why researchers care about notation
Consistent notation lets a research material be described unambiguously. When a specification sheet lists a sequence, it is stating exactly which molecule is intended, independent of any trade or common name.
Frequently asked questions
What is the N-terminus?
It is the amino end of a peptide, the conventional starting point when a sequence is written out.
Why use letter codes instead of full names?
Codes make long sequences compact and unambiguous, which is helpful for labels, specifications and records.
