In 1998, Ortho Biotech changed the formulation of Eprex, a recombinant erythropoietin product sold across Europe, replacing human serum albumin with polysorbate 80 as a stabilizer. Within several years, clinicians began reporting a rare and severe complication in patients who had been receiving the drug for years without incident: pure red cell aplasia, a condition in which the body stops producing red blood cells altogether. Nicole Casadevall and colleagues documented the pattern in the New England Journal of Medicine in 2002, tracing dozens of cases to neutralizing antibodies that cross-reacted with patients' own endogenous erythropoietin.
The counterintuitive part is not that a drug provoked an immune response. It is that the molecule itself had not changed. Patients had been safely exposed to the same amino acid sequence for years. Something about the manufacturing change, not the peptide's structure, appeared to have broken a tolerance that had held for decades of clinical use.
That case, along with subsequent guidance from the FDA in 2014 on immunogenicity assessment for therapeutic protein products, helped establish immunogenicity testing as a mandatory, parallel track in peptide and protein drug development, not an afterthought bolted on after efficacy data comes in. This article walks through the mechanics: how HLA presentation works, how researchers use in silico prediction tools to flag risk before a peptide ever reaches a clinic, what intrinsic and extrinsic factors tip the balance toward an immune response, and why predicting peptide immunogenicity through HLA and MHC epitope analysis has become one of the central questions in the field rather than a footnote to it.
What Immunogenicity Means for a Therapeutic Peptide
The FDA's 2014 Guidance for Industry defines immunogenicity as the capacity of a substance to provoke an immune response. That definition is narrower than it sounds. Immunogenicity is not the same as toxicity, and it is not the same as an allergic reaction, though the three can overlap in outcome.
For peptide drugs specifically, immunogenicity is primarily a T-cell-dependent process. It begins with antigen-presenting cells, immune cells that capture, process, and display fragments of a peptide on their surface for inspection by T-cells. This distinguishes peptide immunogenicity from the simpler antibody reactions associated with, say, a bee sting or a food allergen.
Because this process runs on a separate biological track from a drug's intended pharmacology, immunogenicity risk assessment has to run parallel to efficacy testing throughout development, not sequentially after it. A peptide can be extremely effective at its intended biological target and still fail in the clinic because it triggers anti-drug antibodies. The remainder of this piece follows that risk assessment workflow in order: the mechanism of HLA presentation, the prediction tools built to anticipate it, the risk factors that modify it, and the clinical consequences when prediction falls short.
The HLA System: How Peptides Get Presented to T-Cells
In humans, the Major Histocompatibility Complex goes by a more specific name: the Human Leukocyte Antigen, or HLA, system. According to research published in the Journal of Immunology Research on MHC-I and MHC-II in T-cell-mediated immunity, these cell surface proteins are the gatekeepers of adaptive immunity. Nothing gets flagged as foreign to a T-cell unless an HLA molecule has already presented it.
There are two classes at work, and they handle different cargo. MHC Class I molecules present endogenous peptides, fragments generated from proteins made inside the cell, to CD8+ cytotoxic T-cells. MHC Class II molecules present exogenous peptides, meaning material from outside the cell, such as an injected peptide drug, to CD4+ helper T-cells. A therapeutic peptide administered by injection is far more likely to be processed through the Class II pathway, since it enters the body from the outside rather than being synthesized within a cell.
Whether a given peptide fragment actually binds an HLA molecule depends on anchor residues, specific amino acid positions along the peptide chain that fit into pockets within the HLA binding groove. The fit has to be precise. A peptide with the right length and the right anchor residues at the right positions will sit stably in the groove long enough to be displayed to a passing T-cell. A peptide slightly off in sequence may not bind at all, and simply passes through unnoticed.
HLA MHC Polymorphism and Epitope Prediction in Immunogenicity
HLA genes are, according to research summarized in Trends in Pharmacological Sciences on HLA polymorphism in pharmacogenomics, the most polymorphic genes in the human genome. Thousands of distinct HLA allele variants exist across the population, each with a differently shaped binding groove.
This is the reason a single peptide sequence can be immunologically invisible in one person and provoke a full T-cell response in another. The peptide has to fit the groove of the HLA allele a particular person happens to carry. A sequence that binds tightly to HLA-DRB1*15:01, a common allele in some European populations, may bind weakly or not at all to HLA-DRB1*07:01, carried by a different segment of the population.
The practical consequence for clinical trial design is significant. A drug's immunogenic risk is not a single number that applies uniformly across a study population. It is a distribution, shaped by which HLA alleles happen to be represented among enrolled subjects. A peptide that shows a clean safety profile in a trial population enriched for one set of alleles could behave very differently once it reaches a broader, more genetically diverse patient population after approval. This is part of why post-market surveillance for anti-drug antibodies remains standard practice long after a peptide drug clears clinical trials.
In Silico Epitope Prediction: The Tools of the Trade
Given the scale of HLA diversity, checking every peptide sequence against every allele in a wet lab would be impractical. This is where computational prediction has become indispensable. Three tools dominate the field: NetMHCpan, the IEDB Analysis Resource, and MHCflurry.
All three are machine learning systems trained on experimentally validated peptide-HLA binding data collected in the Immune Epitope Database, a large public repository of documented binding measurements. The basic logic is consistent across tools: feed in a peptide sequence, and the algorithm returns a predicted binding affinity across potentially thousands of HLA alleles, flagging which combinations are likely to result in stable presentation.
NetMHCpan uses an artificial neural network architecture and has been trained to generalize across HLA alleles even when direct experimental data for a specific allele is sparse. MHCflurry takes a similar machine learning approach with a different underlying model architecture and training pipeline. The IEDB Analysis Resource functions somewhat differently, offering a consensus method that draws on multiple prediction algorithms rather than relying on a single model. Comparing the three by training dataset size and allele coverage gives researchers a sense of which tool suits which application, since coverage gaps for rarer HLA alleles remain an issue across all three systems. That comparison also illustrates a broader point: prediction accuracy is only as good as the experimental data underlying the training set, and some HLA alleles are simply better characterized than others.
None of these tools confirm immunogenicity. They estimate binding potential. A peptide predicted to bind an HLA molecule strongly may or may not actually trigger a T-cell response once presented, since binding is a necessary but not sufficient condition for immune activation. This gap between predicted binding and confirmed immune response is exactly why in silico screening functions as a filter, narrowing thousands of candidate sequences down to a manageable shortlist for experimental follow-up, rather than as a final verdict.
From Prediction to Confirmation: In Vitro and Regulatory Steps
The European Medicines Agency's 2017 Guideline on Immunogenicity Assessment of Therapeutic Proteins lays out a risk-based approach that treats in silico prediction as the first of several checkpoints, not the last. Peptides flagged as high-risk binders move into in vitro testing before any human exposure occurs.
Two assay types dominate this stage. MHC binding assays directly measure whether a candidate peptide physically binds purified HLA proteins in a controlled lab setting, confirming or refuting the computational prediction. T-cell proliferation assays go a step further, exposing human T-cells to the peptide-HLA complex and measuring whether the T-cells actually activate and multiply in response, which is a closer proxy for what would happen in a living patient.
Even after a peptide clears both computational and in vitro screening and enters clinical trials, monitoring does not stop. Anti-drug antibody testing continues throughout clinical development and frequently into post-market surveillance, because predictive tools, however refined, cannot eliminate uncertainty about how a diverse human population will respond. This layered approach, moving from algorithm to test tube to clinic to ongoing surveillance, reflects an industry-wide acknowledgment that no single step in the pipeline can be trusted in isolation.
Intrinsic and Extrinsic Risk Factors Beyond Sequence
Sequence and HLA fit are not the only variables at play. Intrinsic factors, meaning properties inherent to the peptide molecule itself, include the amino acid sequence, any post-translational modifications, and aggregation propensity, the tendency of individual peptide molecules to clump together into larger particles. Research published in the Journal of Pharmaceutical Sciences on the immunogenicity of protein aggregates found that aggregates are handled very differently by the immune system than monomeric, individually dissolved peptide molecules.
Aggregates are more readily taken up by antigen-presenting cells, in part because their larger particulate structure resembles the kind of material those cells evolved to capture, such as bacterial debris. This more efficient uptake can translate into a stronger downstream T-cell response compared to the same peptide sequence in monomeric form.
Extrinsic factors, meaning conditions introduced during manufacturing and delivery rather than inherent to the molecule, add another layer. Synthesis impurities left over from peptide production, formulation excipients such as stabilizers or preservatives, and even route of administration, whether a peptide is injected subcutaneously, intravenously, or delivered another way, can all influence how the immune system encounters and processes the drug. Weighing the relative contribution of sequence, aggregation state, and manufacturing impurities to overall immunogenicity risk, as several of the cited studies attempt to do, underscores that no single factor tells the whole story. It is the combination that determines outcome.
Breaking Immune Tolerance: Why Some Peptides Trigger a Response and Others Don't
The immune system is, under normal circumstances, tolerant of self-peptides, meaning it has learned not to attack proteins that are naturally present in the body. A therapeutic peptide that closely mimics an endogenous sequence benefits from this tolerance. A peptide that diverges from any naturally occurring human sequence does not.
One trigger for breaking tolerance is the presence of a novel T-cell epitope, a fragment not represented anywhere in the human proteome. The immune system has no prior exposure to that fragment and treats it as foreign by default, regardless of how therapeutically beneficial the parent molecule might be.
Epitope novelty is not the only factor. Co-stimulatory signals, sometimes described as danger signals, appear to play a role alongside epitope presentation in determining whether a full adaptive immune response gets underway. This may help explain why
Sources
- explorationpub.com — explorationpub.com
- immudex.com — immudex.com
- frontiersin.org — frontiersin.org
- annualreviews.org — annualreviews.org
- creative-biostructure.com — creative-biostructure.com
- nih.gov — pmc.ncbi.nlm.nih.gov
- nih.gov — pmc.ncbi.nlm.nih.gov
- biorxiv.org — biorxiv.org
- oup.com — academic.oup.com
- nih.gov — pubmed.ncbi.nlm.nih.gov
- acs.org — pubs.acs.org
- nih.gov — pmc.ncbi.nlm.nih.gov
- nih.gov — pmc.ncbi.nlm.nih.gov
- frontiersin.org — frontiersin.org
- nih.gov — pmc.ncbi.nlm.nih.gov
- ludwigcancerresearch.org — ludwigcancerresearch.org
- ascopubs.org — ascopubs.org
- aacrjournals.org — aacrjournals.org
- frontiersin.org — frontiersin.org

