Proteomics & Bioanalytik

Prof. Dr. Andreas Tholey

The Systematic Proteomics & Bioanalytics group develops and applies advanced mass spectrometry-based proteomics methods to address biological and biomedical research questions, with a particular focus on protein and proteoform characterization.


The research of the “Systematische Proteomics & Bioanalytik” group is based on two pillars:

Method development for MS-based proteomics and proteoform-centric proteomics (bottom-up proteomics (BUP), middle-down proteomics (MDP), top-down proteomics (TDP), N-/C-terminomics, low-input proteomics (LIP)). Method development encompasses the entire analytical pipeline, starting from sample preparation to LC- and CE-separation schemes, up to the mass spectrometric detection and characterization of peptides, proteins, and proteomes.

Application of these methodologies onto biomedical, biological and biotechnological research questions.

Proteomics and Bioanalytics two pillars

The two pillars are strongly linked to each other; thus, analytical developments are stimulated by the needs of the application projects, and newly developed analytical approaches are transferred back, allowing to provide state-of-the-art methods to the projects.


Our group is not organized as a core facility or service unit, but operates as a research-oriented bioanalytical platform, establishing novel analytical technologies and applying these, together with established methods, in cooperation projects.

We are always open for cooperations! In case you aim to perform proteomics experiments, please contact Prof. A. Tholey (via e-mail).

What can we offer for cooperation projects?

  • Protein identification in cell lysates, supernatants, body fluids, tissues, tissue sections.
  • Quantitative analyses: identification of abundance changes of proteins under different biological conditions.
  • Identification of posttranslational modifications (e.g., phosphorylation, ac(et)ylation/lipidation, oxidation (including disulphide bridges), glycosylation, modification by (reactive) metabolites, and many more).
  • N- and C-terminomics
  • Low-input proteomics: analysis of samples derived, e.g., from low cell numbers (below 50 cells).
  • Identification of protein-protein interactions.
  • Proteome-wide identification of proteolytic processes.
  • Single-protein characterization, e.g., antibodies, antibody drug conjugates, etc.

(Much) more on request – we are happy to discuss your analytical problem with you!


Protein biosynthesis begins with the process of transcription, which produces mRNA transcripts from the genomic information, followed by translation of the polynucleotide to the protein level. In addition, alternative start and stop codons have recently been described to expand the space of genetic information, ultimately leading to the formation of alternative proteins (altProt).

At both stages of protein biosynthesis, numerous biological processes can occur that alter the final protein products from the genetically encoded information. While at the transcript level alternative splicing can lead to different mRNAs, finally leading to the formation of protein isoforms (Figure 1) at the protein level chemical modifications can occur both during (co-translational) and after translation (post-translational modifications, PTMs). The latter events include both covalent modification of amino acid side chains, of which more than 300 have been described to date, and proteolytic processing (e.g., truncation) of the peptide chain. As a result, an exponentially larger number of different protein molecules (species) are produced from a given number of genes; these protein species are called proteoforms.

Formation of proteoforms

Figure 1: Formation of proteoforms. Out of a single gene, via several routes at transcription (e.g., alternative splicing), different isoforms can be formed. Further processing during or after the translation (co- or posttranslational modifications) leads to the formation of protein species (= proteoforms) which have different molecular compositions, physicochemical properties and structures. Proteoforms of a given gene therefore can have completely different biological functions or may even be involved in different biological processes. This makes proteoforms the real active currency of life! The number of proteoforms in humans can only be estimated (in the above example it is five different PrF from one gene), with current estimates reaching from ca. 6 Mio. to more than 1 billion different proteoforms formed out of only ca. 22,500 genes. As the number of coding sequences is presently steeply increasing due to the identification of short open reading frames (e.g., encoding for microproteins) and alternative open reading frames, with their gene products also undergoing posttranslational modifications, the number of proteoforms in cells will even more increase.

By altering amino acid side chains or peptide backbone length, different proteoforms formed from a single encoded protein sequence will have different physicochemical properties such as molecular mass, net charge, or hydrophobicity. These properties influence protein folding and ultimately determine the structure of the proteoforms. As the structure and physicochemical properties define the protein function, different proteoforms of a given protein may have different molecular functions and can be involved in entirely different biological processes. Proteoforms of a given protein may differ in their subcellular localization, their interaction partners, and may have modulated (e.g., enzyme kinetics) or completely different biological activities.

Therefore, the knowledge of the proteoforms within a biological system is an absolute need to understand the molecular processes driving the biology. As this information can neither be inferred from genomics nor from transcriptomics information, proteomics technologies are required to identify, quantify and characterize proteoforms in complex biological environments to unravel their role in biological processes.

Currently, bottom-up proteomics (BUP) is the dominant approach in proteome analysis. The bottom-up workflow involves proteolytic digestion of proteins into peptides using a working protease, followed by separation of the complex peptide mixture, e.g., by liquid chromatography (LC), and mass spectrometry (MS)-based analysis to determine peptide masses and derive peptide sequence information (MS/MS). A critical step in BUP is the derivation of protein level information from the peptide level. This process, protein inference, allows the identification and quantification of protein groups, but as shown in Figure 2, information about the native proteoforms is lost due to proteolytic digestion.

BUP (left) vs. TDP (right).

Figure 2: BUP (left) vs. TDP (right). In the BUP approach, all proteins/proteoforms in a sample are digested, forming a large number of peptides (peptides 1-7; their position in the proteoforms PrF 1-4 is shown in the middle box). After analysis, mainly by LC-MS/MS, these peptides can be used for the inference of the encoded protein. However, it is not possible to unambiguously assign the exact PrF from which a given peptide originates from. In contrast, TDP aims the direct identification of the PrF, thus all information about a given proteoform (e.g., side chain modifications, truncations, and especially the combination of these modifications in a given PrF) is retained. LC-MS/MS is also the dominating approach in TDP, however, using different analytical parameters both for the separation and the MS analysis.

While BUP is an established and sensitive technology, the loss of proteoform-level information creates a need for alternative analytical approaches. Top-down proteomics (TDP) fills this gap. In TDP, analysis is performed at the intact protein level, which a priori preserves all information about a proteoform within a single molecule. While the direct analysis of intact proteins rather than digested peptides is intuitive, numerous analytical challenges have prevented the widespread use of this technology in the past. Major challenges, of which only a selection is listed here, occur at all levels of the analytical pipeline: (1) sample preparation: reduced solubility of intact proteins when released from their cellular environment; (2) separation of intact proteins by LC is less straightforward than that of peptides; (3) MS and MS/MS: lower sensitivities, e.g., due to the broad charge state distribution in electron spray ionization (ESI) MS; the need for efficient but still error-prone deconvolution algorithms; the need for high-resolution MS instrumentation; very complex MS/MS spectra, often with incomplete ion series; (4) data interpretation: lack of data integration algorithms; lack of defined data curation criteria (e.g., false discovery rates). To overcome these challenges in TDP is a major aim of our research.

Proteolytic processing or alternative start/stop codons can lead to the formation of truncated proteoforms or proteoforms with non-canonical N- or C-termini. For the proteome-wide identification of N- and C-termini, N- and C-terminomics approaches have been developed, mainly based on BUP. On the other hand, TDP a priori identifies the intact proteoform and thus both termini in a given proteoform. Our group works on the improvement of BUP-based terminomics approaches and on the integration with TDP (integrative multilevel proteoformics).

Application-driven projects (examples)

  • Short open reading frame encoded peptides (SEP), microproteins, and proteins encoded by alternative ORFs (altProt).
  • Antimicrobial peptides.
  • Amyloid proteomics.
  • Mode of action studies for new drugs: identification of targets and off-targets at the proteoform level.
  • Spatial proteo(for)mics.
  • Immuno-peptidomics.
  • Paleo-proteomics.

  • Molecular & Cellular Proteomics

    ·

    Properties, Origin, and Consistency of Truncated Proteoforms Across Top-Down Proteomic Studies

    Kaulich PT, Fulcher JM, Tholey A

  • Proteomics

    ·

    Top-Down Proteomics: Why and When?

    Kaulich PT, Tholey A

  • Nature Methods

    ·

    Influence of different sample preparation approaches on proteoform identification by top-down proteomics

    Kaulich PT, Jeong K, Kohlbacher O, Tholey A

  • Angewandte Chemie International Edition

    ·

    Digital Microfluidics and Magnetic Bead-Based Intact Proteoform Elution for Quantitative Top-down Nanoproteomics of Single C. elegans Nematodes

    Leipert J, Kaulich PT, Steinbach MK, Steer B, Winkels K, Blurton C, Leippe M, Tholey A

  • iScience

    ·

    Proteoforms expand the world of microproteins and short open reading frame-encoded peptides

    Cassidy L, Kaulich PT, Tholey A