5 Engineering More Stable Proteins
© 2023-2026 Romas Kazlauskas. All rights reserved. Last revised: March 2026.
Summary. Protein function requires the folded protein form, but this form is unstable mainly because it readily unfolds into a flexible, unstructured form. Protein folding is favored by burying of hydrophobic side chains and hydrogen bonding between the amino acids. Protein unfolding is favored by the increase in flexibility of the main chain of amino acids upon unfolding. Protein stability is usually measured by the reversible unfolding of the protein with either heat or chemical additives like urea. Protein stabilization involves making substitutions that shift the folding-unfolding balance toward the folded form. Stabilizing substitutions can either stabilize the folded conformation or destabilize the unfolded ensemble. This tutorial emphasizes web-based tools to identify substitutions that stabilize proteins. Besides unfolding, other sources of protein instability are chemical modifications like oxidations or cleavage by proteases and aggregation of partly unfolded proteins into insoluble particles.
Key learning goals
- Proteins are dynamic, equilibrating collections of structures. Most of the structures are folded, but a small fraction are unfolded. This protein unfolding is the primary source protein instability. Stabilization shifts the folding-unfolding balance toward folding, but does not prevent unfolding.
- To measure protein stability, one measures the ratio of folded to unfolded protein. Under normal conditions, the amount of unfolded protein is too small to detect. Adding denaturants such as urea or heating the sample increases the amount of unfolded protein to measurable amounts. Extrapolation back to normal conditions reveals the stability of the protein.
- One way to stabilize proteins is to restore amino acid residues that are conserved in homologs. The rationale for this approach is evolution conserves residues that provide some benefit; one benefit is a contribution to stability.
- Another way to stabilize proteins is to destabilize the unfolded form. The main driving force for unfolding is the increased flexibility of the unfolded form. Decreasing the flexibility of unfolded form by adding disulfide crosslinks or introducing proline residues often stabilizes proteins.
- The most direct way to stabilize proteins is to create or strengthen attractive interactions between amino acids in the folded conformation. Although proteins will remain dynamic and still unfold, the stronger interactions either slow down unfolding or speed up refolding.
5.1 Introduction
Protein stability usually refers to resistance to unfolding. Stresses like high temperatures, organic cosolvents, high substrate or product concentrations, extremes of pH or high ionic strength can all cause unfolding. In many cases, stabilizing a protein to one stress also stabilizes it to other stresses. For example, proteins that tolerate high temperatures often also tolerate organic cosolvents.[1]
One advantage of more stable proteins is the ability to use them as biocatalysts in artificial environments where naturally-occurring proteins are unstable. For example, using biocatalysis at higher than physiological temperatures yields faster reaction rates, higher substrate solubility, and lower solution viscosities. Reactions at high temperature can also avoid microbial contamination since most contaminating microbes cannot grow at high temperature. Biocatalysts that tolerates organic cosolvents and high concentrations of substrates and products simplify the scale-up of industrial processes by requiring less solvent, smaller equipment, and less complicated product isolation.
The second advantage of more stable proteins is extended lifetime or storage at ordinary temperatures. For example, manufacture of β-lactam antibiotics with penicillin G acylase recycles the immobilized enzyme many times to lower the overall cost. More stable therapeutic proteins last longer during storage and are more resistant to proteases in serum.
A third advantage of more stable proteins is higher yields of soluble protein during manufacture, especially for small, single-domain proteins. Over-expression of recombinant proteins in microbes creates high concentrations of protein. The unfolded forms can aggregate into insoluble particles called inclusion bodies. More stable single-domain proteins aggregate less and yield higher amounts of soluble protein.[2]
A fourth advantage is that more stable proteins are also more ‘evolvable’ or able to acquire beneficial traits. Substitutions that change protein function may destabilize proteins. If the destabilized protein fails to fold, then the potential improvement is lost. More stable proteins can tolerate greater numbers of destabilizing mutations and are thus more evolvable than their marginally stable variants.[3] One source of stable proteins is thermophiles and other extremophiles.[4] For example, the Pfu DNA polymerase used for the polymerase chain reaction (PCR) comes from the hyperthermophile Pyrococcus furiosus, Figure Figure 5.1. This polymerase catalyzes DNA synthesis at 72 °C and tolerates the high temperatures (95 °C) needed to dissociate complementary DNA strands.
One disadvantage of enzymes from thermophiles is the typically low catalytic activity of these enzymes at room temperatures.[5] Since thermophiles do not live at room temperature, there is no selective pressure for high activity at room temperature. If the application requires activity at room temperature, enzymes from thermophiles may not be suitable. Pfu DNA polymerase has negligible activity at room temperature. Other limitation of enzymes from thermophiles is the lack of other needed protein characteristics. For example, applications may require unusual substrate specificity unavailable in thermophiles or human proteins to minimize immune response.
Many human disease-causing mutations are associated with amino acid substitutions that decrease protein stability.[6] For example, mutations that decrease the stability of fructose bisphosphate aldolase cause hereditary fructose intolerance and mutations of the tumor suppressor protein p53 cause some cancers.[7] The ability to predict substitutions that alter protein stability may help identify mutations that cause genetic diseases and treat them with genetic editing technologies. For example, sickle cell disease, caused by the lower stability of the Glu6Val variant of \(\beta\)-globin, has been treated by expressing the unmutated fetal \(\beta\)-globin gene.[8]
5.2 Native and denatured conformations equilibrate
The native protein state, N, is the folded, functional form. It is compact with the hydrophobic side chains mainly buried and the polar side chains mainly exposed to solvent. It has a specific structure (or similar set of structures), typically with α-helices, β-sheets, and turns folded in a particular arrangement, Figure Figure 5.2. While the structure is dynamic and moves, it has an overall stable structure. Most of the amino acids interact with each other; only amino acids on the protein surface interact with solvent water. In contrast, the denatured protein state, D, is not functional and is not a single state, but a collection or ensemble of more or less folded states. The main chain makes large, random motions and the amino acids interact mainly with solvent water, not with each other.
The native form of a typical protein is more stable than the denatured state by \(\sim \! 10\) kcal/mol, Figure Figure 5.3. A Gibbs energy of unfolding, \(\Delta G^{unfold}\), of +10 kcal/mol corresponds to an equilibrium constant \(e^{-(10,000/RT)}\) or \(4.6 \times 10^{-8}\) at room temperature indicating that unfolding is rare, eq. Equation 5.1. Folding and unfolding are fast for single domain proteins. Native forms continuously unfold to denatured states and these continuously refold to native forms.
\[\Delta G^{unfold} = -RTln\left( K^{unfold} \right) \text{ or } K^{unfold} = e^{-\Delta G^{unfold}/RT} \tag{5.1}\]
Protein solutions contain unfolded protein molecules even for stably folded proteins. In the case where the native form is \(\sim \! 10\) kcal/mol more stable than the unfolded form, one in twenty-two million protein molecules unfolds at any given time in solution at room temperature, eq. Equation 5.2. This tiny fraction is too small to detect by spectroscopic methods, but it is nevertheless a large number of molecules. A solution containing 1 mg of a 30 kDa protein contains \(4 \times 10^{16}\) molecules; of these, nearly a billion (\(10^9\)) are in the denatured form at any time.
\[K^{unfold} = \frac{[D]}{[N]}=4.6 \times 10^{-8} \text{ or } [N] = 22 \times 10^6 \cdot [D] \tag{5.2}\]
The main driving force for protein folding is the hydrophobic effect, which is the tendency of non-polar solutes to cluster in water. Burying a –CH\(_2\) group contributes 1.1 ± 0.5 kcal/mol to protein stability. The hydrophobic effect provides \(\sim \! 60\)% of the driving force to collapse the amino acid chain into a compact structure.[9] The folded form of a protein is its lowest energy conformation in dilute solutions. (In concentrated protein solutions, oligomeric or aggregated states may be lower in energy.) The chains orient to maximize the hydrophobic effect and also to make attractive interactions between amino acids: hydrogen bonds and other electrostatic interactions.
Hydrogen bonds contribute the remaining 40% of the driving forces for protein folding. The hydrogen bonds between protein atoms each contribute on average 1.1 ± 0.8 kcal/mol each to protein stability. The unfolded protein makes hydrogen bonds between protein and water, so this energy contribution is the net gain in hydrogen bond strength. Hydrogen bonds also define how the protein will fold into helices and sheets.
The dominant driving force for unfolding of the main protein chain is the increase in flexibility. The denatured ensemble is a collection of highly flexible forms. This flexibility creates many conformations or micro-states and the existence of these states is the main driving force for unfolding. The micro-states create a favorable entropy contribution, \(-T\Delta S\), to the Gibbs energy. To find the Gibbs energy contribution of entropy, one multiplies the entropy change by temperature and a negative sign since an increase in entropy lowers Gibbs energy, eq. Equation 5.3, where \(W_{folded}\) and \(W_2\) are the different numbers of micro-states in the states being compared.
\[ \Delta G_{unfolded-folded} = -T\Delta S_{2-1} = -RTln\left( \frac{W_2}{W_1} \right) \tag{5.3}\]
To find the effect of a molecule’s conformational flexibility on Gibbs energy, one compares the number of conformations in the flexible state to the non-flexible state. For example, one can estimate the difference in Gibbs energy between the folded and unfolded states for one amino acid in a protein at 300 °K only due to differences in backbone flexibility while ignoring any differences in side-chain flexibility. Assume that the backbone has a single conformation in the folded state, N, but can adopt three staggered conformations along each of the two rotatable bonds (ψ, φ) in the unfolded state, D, Figure Figure 5.4.
The approach is to estimate the difference in the number of conformations available to the folded and unfolded states and then convert this difference to Gibbs energy using eq. Equation 5.3 above. The amino acid in the folded state has one available conformation while in the unfolded state, an amino acid can adopt \(3\cdot3 = 9\) possible conformations. Converting this difference in possible conformations into Gibbs energy yields:
\[\Delta G_{N-D}= -RTln\left( \frac{W_2}{W_1} \right) = 1.987 \text{ cal/mol}\cdot ^{\circ}\text{K} \cdot 300 \:^{\circ}\text{K} \cdot ln\left( \frac{1}{9} \right) = +1.3 \text{ kcal/mol}\]
Thus, entropy due to differences in the backbone flexibility favors the unfolded state by 1.3 kcal/mol for each amino acid residue. More accurate estimates that account for the unequal likelihood of the nine conformations and differences in side chain flexibility yield a similar number: 1.4 kcal/mol per amino acid residue.[10] For a typical protein of 300 amino acids, this entropy effect of backbone flexibility contributes 400 kcal/mol, which is large as compared to the balance of \(\sim \! 10\) kcal/mol in favor of the folded state. Hydrophobic interactions and hydrogen bonds in the native state counterbalance this flexibility advantage of the denatured state with a slightly larger Gibbs energy contribution. Protein stability is a balance between one set of large forces favoring the folded state and another set favoring the unfolded state.
5.3 Measuring the folding-unfolding equilibrium
Measuring the equilibrium constant requires measuring the ratio of folded and unfolded protein at equilibrium. Under normal conditions, the equilibrium amount of unfolded protein is too small to measure. Instead, researchers use destabilizing conditions where easily detectable amounts of both the folded and unfolded protein exist, Figure Figure 5.5. Typical destabilizing conditions are additives like urea or guanidinium chloride, changes in the solution pH, or increases in temperature. After measuring the equilibrium constant at these destabilizing conditions, researchers extrapolate back to normal conditions to get the desired equilibrium constant under normal conditions.
Most of this book treats protein folding and unfolding as a cooperative, two-state process (Figure Figure 5.2 above). The entire protein is either folded or denatured, and the process is entirely reversible. There are no partially folded proteins or folding intermediates. The cooperative folding might start with one amino acid adopting an α-helix conformation, quickly followed by neighboring amino acids adopting the α-helix structure to create intrachain hydrogen bonds. This assumption of a simple, reversible two-state process is reasonable for single domain monomeric proteins. In these cases, adding stabilizing substitutions anywhere in the protein can contribute to stabilization since the entire protein unfolds simultaneously.
Changing the environment does not gradually change protein structure from folded to unfolded. Proteins remain fully folded or fully unfolded, but their ratio changes with the environment. Unfolding is cooperative because the interactions that stabilize the folded protein are stronger in combination than their individual contributions. The protein remains folded even when some stabilizing interactions break, but breaking too many destabilizes the others leading to complete unfolding of the protein.[11]
Section Section 5.5.1 below will briefly consider multiple domain and oligomeric proteins. Multiple domain proteins usually fold and unfold stepwise via intermediates. The same stabilization strategies apply, but the substitutions must be in unfolded regions, not anywhere in the protein.
5.3.1 Denaturation with urea, \(\Delta G_{H_2O}\)
Urea unfolds proteins because it 1) solvates exposed hydrophobic groups to reduce the hydrophobic effect and 2) forms hydrogen bonds to the backbone to disrupt secondary structures. Typical proteins unfold at 3-5 M urea; the solubility limit of urea is 18 M at room temperature.
The urea-induced denaturation experiment involves preparing solutions of protein in various concentrations of urea, waiting until the folding-unfolding reaches equilibrium, and measuring the amounts of native and denatured protein. The most common method to detect the folded and unfolded proteins is measuring the intensity of protein absorbance or fluorescence, Figure Figure 5.6, but any method that distinguishes between the native and denatured states is suitable. Some examples are measuring shifts in the wavelength of the fluorescence maximum, changes in the circular dichroism spectra, changes in the NMR spectra, decreases in catalytic activity, or changes in the dye fluorescence upon binding to hydrophobic regions of the unfolded protein.
The reasoning above predicts that the fluorescence emission wavelength increases for the denatured protein, but it is not easy to predict the relative intensity of this emission. The intensity change also depends wavelength chosen to monitor the fluorescence. For the example in Figure Figure 5.6\(\text{a}\), the fluorescence intensity at 320 nm (near the emission maximum of the native protein) decreases as the protein unfolds, but the fluorescence intensity at 360 nm (near the emission maximum of the denatured protein) increases (not shown).
To extract the constants for unfolding equilibrium at different urea concentrations, \(N \rightleftharpoons D\) from the fluorescence changes, one needs to know the fractions of native, \(F_N\), and denatured protein, \(F_D\). The ratio of these fractions is the equilibrium constant, eq. Equation 5.4. One assigns the fluorescence in low or no denaturant, \(Y_N\), to the native protein and the fluorescence at high denaturant concentrations, \(Y_D\), to the denatured protein. The supporting information includes a derivation of this equation. Intermediate values of fluorescence, \(Y_{obs}\), correspond to a mixture of native and denatured states. Only data in the transition region of the denaturation experiment contribute to the analysis because when the protein is nearly completely folded or completely unfolded, the data do not precisely define the ratio of the two forms.
\[K^{unfold}=\frac{F_N}{F_D}=\frac{Y_{obs}-Y_N}{Y_D-Y_{obs}} \tag{5.4}\]
Applying the a linear extrapolation approach[12] to these measured equilibrium constants reveals the Gibbs energy of unfolding in water. This linear extrapolation approach starts by converting these equilibrium constants to Gibbs energy values according to eq. Equation 5.1 above and plotting them versus urea concentration, Figure Figure 5.7. The plot reveals a straight line with a negative slope since adding more urea favors protein unfolding. The line crosses the x-axis at \(\Delta G^{unfold} = 0\). The goal of this experiment is to measure \(\Delta G^{unfold}\) without any urea, that is, \([\text{urea}] = 0\). Extrapolation of the line to y-axis intercept reveals the Gibbs energy of unfolding in pure water, \(\Delta G^{unfold}_{H_2O}\).
The equation for this line has a slope of -m and and a y-intercept at \(\Delta G_{H_2O}^{unfold}\). The m-value is the absolute value of the slope and indicates the sensitivity of the protein to denaturant, which depends on the solvent-accessible surface area exposed upon unfolding.[12] A protein that unfolds completely has a higher m-value than a similar protein that unfolds only partially. Similar proteins with different m-values indicate different degrees of unfolding in their denatured states. Typical m-values for proteins undergoing urea-induced unfolding are \(\sim \! 1000\) cal/(mol·M).This linear extrapolation approach is a convenient way to measure protein stability and the supporting information includes an example calculation.
\[\Delta G^{unfold} = \Delta G_{H_2O}^{unfold}-m\cdot[\text{urea}] \tag{5.5}\]
Positive Gibbs energies of unfolding in pure water indicate that unfolding of the protein is unfavorable. A higher Gibbs energy of unfolding indicates a more stable folded protein. The molecular basis for the linear relationship between unfolding Gibbs energy and urea concentration may be a binding interaction between urea and protein. The binding of urea destabilizes the protein and the fraction of urea bound increases linearly with concentration.
5.3.2 Heat denaturation, \(\Delta G^{unfold}(T)\)
Heat also unfolds proteins. Increases in temperature increase the entropy contribution to Gibbs energy of unfolding, eq. Equation 5.6 where \(\Delta H^{unfold}\) and \(\Delta S^{unfold}\) are, respectively, the changes in enthalpy and entropy upon unfolding. The entropy contribution due to the increase in flexibility of the unfolded protein increases at higher temperature and eventually leads to unfolding.
\[\Delta G^{unfold} = \Delta H^{unfold} - T\Delta S^{unfold} \tag{5.6}\]
Here we consider reversible heat-induced unfolding. Section Section 5.6.1 below considers irreversible unfolding caused by aggregation of the unfolded protein. As with urea-induced denaturation, this unfolding can be monitored spectroscopically, Figure Figure 5.8. For reversible unfolding, the equilibrium constant yields the Gibbs energy of unfolding at these high temperatures. The melting temperature, \(T_m\), is the temperature where the amounts of folded and unfolded protein are equal; \(\Delta G^{unfold} = 0\).
These measured Gibbs energies of unfolding at high temperatures can be extrapolated to room temperature, but this extrapolation is more complex than extrapolating the Gibbs energy of unfolding to pure water from a urea denaturation curve.[13] One might expect a plot of ΔGunfold versus temperature to yield a straight line according to eq. 3.7 above, where ΔSunfold is the slope and ΔHunfold is the y-intercept. However, the Gibbs energy of protein unfolding varies with temperature according to a more complex equation, eq. Equation 5.7.[14] Extrapolating the melting curve data to lower temperatures also requires measuring \(\Delta C_p\) which is the change in heat capacity upon unfolding, Figure Figure 5.8.
\[\Delta G^{unfold} = \Delta H^{unfold}_0 - T\Delta S^{unfold}_0 + \Delta C_p\left[(T - T_0) - T\ln(T/T_0)\right] \tag{5.7}\]
Here \(T_0\) is the reference temperature and \(\Delta H^{unfold}_0\) and \(\Delta S^{unfold}_0\) are the enthalpy and entropy of unfolding at that temperature. For reactions involving small molecules, the heat capacity of starting materials and products are similar, and differences may be ignored. Eq. Equation 5.7 simplifies to eq. Equation 5.6 when \(\Delta C_p = 0\).
However, the heat capacities of the folded and unfolded proteins differ significantly. Another way to state the reason for the more complex equation is that for protein unfolding \(\Delta H^{unfold}\) and \(\Delta S^{unfold}\) vary with temperature while eq. Equation 5.6 assumes they are constant. The extra terms in eq. Equation 5.7 account for the variation of \(\Delta H^{unfold}\) and \(\Delta S^{unfold}\) with temperature.
The unfolded protein has a higher heat capacity than the folded protein because unfolding exposes hydrophobic groups to water. Water surrounds these hydrophobic groups with an ice-like structure, which has slightly higher energy than bulk water. The fluctuation of water between bulk and ice-like structure increases the heat capacity.
The change in heat capacity for protein unfolding, \(\Delta C_p\), is always positive with a typical value of 14 cal/mol·deg per residue or 4.2 kcal/mol·deg for a 300 aa protein.[15] Differential scanning calorimetry can measure this change in heat capacity. Differential scanning calorimetry measures the heat required to increase the temperature of a sample. The baseline value before unfolding reveals the heat capacity of the folded protein, while the baseline value after unfolding reveals the heat capacity of the unfolded protein. The difference is the change in heat capacity. In the absence of experimental data, the web tool SCooP[16] predicts the parabolic curve for temperature-dependent unfolding of a protein based on its structure.
The parabolic curve in Figure Figure 5.8\(\text{b}\) predicts that proteins will also denature at low temperatures. For most proteins the predicted cold denaturation temperature, Tl, lies well below the freezing point of water as shown in the sketch, but for some proteins, the cold denaturation temperature lies in the liquid range of water. For example, the yeast protein frataxin denatures at both 7 °C and at 30 °C.[17] At high temperatures the increase in conformational entropy (flexibility) of the protein atoms drives unfolding, while at low temperatures a decrease of the hydrophobic effect drives unfolding. At lower temperatures, the cost of ordering water molecules around a hydrophobic group decreases. As the tendency of hydrophobic groups to aggregate in water decreases, the folded form loses a stabilizing force.
Increases in melting temperature usually correspond to increases in protein stability. Comparing the unfolding or melting temperatures of protein variants is a common, approximate measure of their stability. One assumes that the unfolding process is similar for protein variants so the parabolic curve of Gibbs energy of unfolding versus temperature (Figure Figure 5.8\(\text{b}\) above) remains similar. One heats a protein sample slowly while monitoring the protein tryptophan fluorescence (Figure Figure 5.8\(\text{a}\)). The temperature corresponding to the midpoint of the change in fluorescence indicates the melting temperature. Usually one is on the right side of the parabola describing protein heat stability so an increase in melting temperature corresponds to a increase in protein stability.
A comparison of pairs of homologous proteins from thermophiles and mesophiles showed that the melting temperatures, \(T_m\), were 31.5 °C higher and the Gibbs energies of unfolding 8.7 kcal/mol higher for the thermophiles.[18] From this comparison, we can derive a rule of thumb that stabilizing a protein by 1 kcal/mol will increase the melting temperature by 3.6 °C, eq. Equation 5.8, see dotted line in Figure Figure 5.8\(\text{b}\).
\[\Delta \Delta G^{unfold} (\text{kcal/mol}) \sim \Delta T_m (\text{°C})/3.6 \text{ °C/kcal/mol} \tag{5.8}\]
In many cases, the Gibbs energy of heat-induced unfolding,\(\Delta G^{unfold}(T)\), is comparable to that obtained by urea unfolding. When these values differ, it is likely because the unfolded states differ in the two experiments.
5.4 Stabilization to cooperative unfolding
Stabilizing a dynamic object like a protein differs from stabilizing a rigid object like a chair. To stabilize a chair one can add a supporting brace, which keeps the chair in the desired structure. In contrast, a protein folds and unfolds continuously; it is dynamic. Adding a stabilizing interaction to a protein does not prevent unfolding. Instead, the stabilizing interaction shifts the folding-unfolding equilibrium toward folding so the protein spends more time in the folded form. Since the equilibrium constant is the ratio of the forward and reverse rate constants, eq. Equation 5.9, shifting the equilibrium constant also changes the rates of folding and unfolding. A stabilizing interaction may slow down unfolding, speed up folding, or both; but, unlike stabilizing a chair, stabilizing a protein does not prevent unfolding.
\[ N \xrightleftharpoons[k_{fold}]{\,k_{unfold}\,} D \tag{5.9}\]
Stabilizing a protein also differs from stabilizing a rigid object because one can stabilize a protein by destabilizing the unfolded form.[19] Everyday objects like chairs are not dynamic and cannot be stabilized this way. This book uses the expression ‘stabilize a protein’ or ‘increase protein stability’ to mean either stabilizing the folded form or destabilizing the unfolded form.
One reason that it is hard to predict which substitutions will stabilize a protein is that the structure of the denatured state (an ensemble of rapidly equilibrating structures) is unknown. Stabilizing substitution must affect the folded and unfolded forms differently. A substitution that equally stabilizes (or destabilizes) the folded and unfolded forms does not change stability of the protein. X-ray crystal structures reveal changes to the folded form, but changes to the denatured state are invisible because its structure is unknown. Stabilization can occur without strengthening the folded structure. For example, seven amino acid substitutions dramatically stabilized xylanase (\(T_m\) increased by 25 °C), but the structures of the wild-type and stabilized proteins showed similar interactions between amino acids in both structures.[20] In this case, the stabilizing substitutions likely destabilized the unfolded state and changed its structure, but nature of these changes is unknown because the structure of the unfolded state is unknown.
A second reason that it is hard to predict stabilizing substitutions is that the substitutions must not hinder catalytic or binding ability of the target proteins. Substitutions that increase stability often also decrease activity and vice versa.[21] One reason for this trade-off is that catalysis and binding rely on interactions of the substrate or transition state with unsatisfied hydrogen bonds, exposed hydrophobic groups and unpaired charges in the protein. Satisfying these interactions with amino acid substitutions will stabilize the protein, but also disrupt binding and catalysis. Researchers typically avoid substitutions near the active site when searching for stabilizing substitutions to prevent disrupting binding and catalysis.
Some protein stabilization strategies seek to stabilize the native protein, others destabilize the denatured conformation, and, for a few, the molecular mechanism is unknown, Table Table 5.1. The design strategies that seek to stabilize the native protein require a protein structure or at least a homology model to start the design. A homology model is an extrapolated 3D-structure of a protein based on the known structure of a similar protein. Swiss Model and AlphaFold are web tools to generate homology models from protein sequences. Strategies that seek to destabilize the unfolded ensemble also require a protein structure to ensure that the changes do not distort the folded form. Strategies with unknown mechanism do not require a protein structure. For example, restoring residues that are conserved in homologs (consensus sequence approach, see below) requires only a protein sequence. Random mutagenesis followed by screening to find stabilized variants also does not require a protein structure. Chapters 9 and 10 discuss random mutagenesis methods.
| Strategy | Rationale | Example |
|---|---|---|
| restore residues conserved in homlogs | not specified | Consensus Finder identifies conserved amino acids in homologs that are missing from target protein |
| add disulfide crosslinks | destabilize unfolded protein | Disulfide by Design identifies suitable locations, additional analysis needed to narrow choices |
| add Pro residues | destabilize unfolded protein | Analysis of structure to identify locations suitable for proline WHAT-IF |
| substitutions in or near flexible regions | stabilize folded protein | Random mutagenesis in or near flexible regions |
| random mutagenesis and screening | not specified | Random mutagenesis anywhere followed by screening for stabilized variants |
| optimize electrostatic interactions | stabilize folded protein | TKSA-MC identifies destabilizing electrostatic interactions |
| optimize hydrophobic and other interactions | stabilize folded protein | Modeling with Rosetta and FoldX to identify stabilizing substitutions |
To engineer a change, one needs to choose both the location of the substitution and the replacement amino acid. The ‘restore residues conserved in homologs’ approach predicts both, so it is easy to implement. The ‘substitutions in or near flexible regions’ approach defines the location, but not the replacements. Additional modeling or experiments are needed to identify the replacements. The ‘add disulfide links’ approach specifies that cysteine is the replacement amino acid, but additional modeling is needed to find suitable locations. Design strategies like ‘improve hydrophobic and other interactions’ are the least specific and require computer modeling to predict both the location and the replacement amino acid. Each of the approaches in Table Table 5.1 will be described in more detail below. The effect of each stabilizing mutation is typically small (0.5 kcal/mol), so substantial stabilization of the protein (2-4 kcal/mol) requires multiple substitutions and often more than one design strategy.
5.4.1 Restore conserved amino acids
Evolution conserves amino acids that contribute to protein function. This contribution may be to structure, catalysis, stability or other protein property. The consensus approach hypothesizes that conserved amino acids residues are more likely to contribute to protein stability than a non-conserved amino acid. To use the consensus approach to stabilize a protein, one identifies amino acid residues conserved within homologous proteins, but which are missing in the target protein. Restoring these conserved amino acids is more likely to stabilize the protein than random substitutions.
Consensus design works because natural proteins rarely follow the consensus sequence at all structurally important positions. Natural proteins only have to be stable enough to fulfill their biological function; proteins with stabilities above a certain threshold will have no further selection advantage. There are many stabilizing residues, but individual proteins do not need all of them to reach the threshold of being stable enough. Thus, replacing a residue with the corresponding consensus amino acid may improve the stability or folding efficiency of a protein of interest. The same reasoning leads to the conclusion that the stability of a protein designed by consensus can be higher than that of the proteins in the alignment.
The consensus sequences approach does not hypothesize any molecular basis for the stabilization. It relies only on amino acid sequence comparisons from bioinformatics and does not require any structural information.[22,23] It is easy to implement, and several web-based tool to generate a consensus sequence for a protein are available: Consensus Finder[24] or FireProt.[25] Typically researchers restore one or a few conserved amino acids. Usually at least half of these predicted substitutions prove to be stabilizing. Pantoliano and coworkers first suggested that substitutions to the consensus amino acid may stabilize proteins and showed that the consensus-type mutation Met50Phe increased the unfolding temperature of subtilisin BPN' by 1.8 °C[26] and many others have used this approach. For example, a sequence alignment of 351 homologs of Staphylococcus nuclease A identified the His124Glu substitutions as potentially stabilizing, Figure Figure 5.9.
One can also restore all the conserved amino acids simultaneously. Lehmann and coworkers predicted the consensus sequence of fungal phytases based on the sequence alignment of thirteen sequences, Figure Figure 5.10.[22] The consensus phytase differed by at least 80 amino acids from the parent sequences. A synthetic gene encoded the consensus phytase containing all these changes. The consensus protein was 15-26°C more thermostable than any of its parents.
At some positions, counting the numbers of occurrences identifies the consensus amino acid. For example, all thirteen sequences start with QVL, and most have S in the next position (boxed), so the consensus sequence begins with QVLS, Figure Figure 5.10. At other positions, identifying the consensus amino acid requires care to ensure unbiased sampling of sequences. Identifying the most common residues requires a uniform sampling of amino acid sequences throughout the evolutionary tree, but the available amino acid sequences may not be evenly distributed. For the phytase example above, researchers reduced the bias of multiple sequences from different strains of the same species by weighting the amino acid sequences so that each Aspergillus species had equal weight. For example, the list included five sequences from Aspergillus fumigatus strains, Figure Figure 5.10. Each of these sequences was weighted by 0.2 in deriving the consensus to avoid bias toward this species. For example, consider the last amino acid in Figure Figure 5.10 (position 73, boxed). The first two sequences (A. terreus) have A at this position, the next three (A. niger) and two others (Emericella and Talaromyces) have S, the five A. fumigatus sequences have K, and the last sequence (Myceliophthora) has V. Although S and K both occur five times, K occurs only in the five A. fumigatus sequences and receives a total weight of 1.0, while S occurs in three different species and gets a total weight of 3.0. The consensus amino acid at this position is S. The consensus sequence is not unique, but depends on the set of comparison sequences. In later work, as more phytase sequences became available, researchers tested a revised consensus sequence, which further increased stability.
In other cases, a database search may yield thousands of similar sequences and it is impractical to assign different weights to the sequences. To minimize phylogenetic bias in these cases, one clusters similar sequences (typically >90% identity) together and uses only one sequence from each cluster for the consensus calculation. The web-based tools to find consensus sequences mentioned above use this clustering to minimize phylogenetic bias.
One disadvantage of the consensus sequence approach is that its conservative nature predicts small numbers of substitutions, often less than a handful or even none. Although experimental screening reveals that many stabilizing substitutions exist, the consensus sequence approach, by focusing on those conserved among homologs, misses most of them.
Ancestral proteins are extinct proteins that correspond to branch points in a phylogenetic tree. Gene synthesis and expression of these proteins in microbial hosts resurrects these proteins. In many cases, ancestral proteins are more stable than modern proteins. Some ancestral proteins may be more stable because the earth was warmer in prehistoric times. However, ancestral proteins from times when the earth's temperature was only slightly warmer are also more stable, so other effects may contribute.[28] One contribution is that reconstructed ancestral proteins are often similar to the consensus sequence for that group since many descendants retain the ancestral residues. However, ancestral proteins differ from consensus proteins is two ways. First, ancestral proteins identify amino acids conserved within a related group, while consensus proteins include data from all homologous proteins, including those outside a group. Second, the consensus approach weights all sequences equally, but ancestral sequence reconstruction takes branch lengths into account, so more recent changes count less in an ancestral sequence prediction.
5.4.2 Destabilize the unfolded ensemble
The main property favoring the unfolded ensemble is its flexibility. Reducing this flexibility destabilizes the unfolded ensemble, thus shifting the folding-unfolding equilibrium toward folding. Two ways to reduce this flexibility are to add disulfide cross-links and to replace amino acid residues with proline. In both cases the replacement amino acid is defined; the challenge is to find suitable locations for the replacement. The ideal site in the folded structure would fit the replacement perfectly, so the only effect of the substitution is to reduce the flexibility of the unfolded ensemble.
Add disulfide cross links. Replacing a nearby pair of amino acid residues with cysteines followed by spontaneous oxidation creates a disulfide cross link, Figure Figure 5.11. Such disulfide links often stabilize proteins to unfolding. For example, the replacement of Ala43 and Ser80 in Bacillus RNAse with cysteines followed by spontaneous oxidation yielded a more stable protein that unfolded in 5.77 M urea as compared to 4.58 M urea for the wild type.[29] This engineering involved both amino acid replacements and the formation of a cross link, both of which could contribute to changes in stability. To measure the effect of the cross link alone, the researchers compared the stabilities of the oxidized disulfide form with the reduced dithiol form of the protein. These two forms differ only in the presence or absence of a cross link. For the example above, the Gibbs energy of unfolding of the oxidized form was 3.2 kcal/mol higher than for the reduced dithiol form demonstrating that the cross link stabilized the protein.
The reason that disulfide cross links stabilize proteins is that cross links reduce the flexibility of the unfolded protein. Upon unfolding, the cross link remains and limits the movement of the unfolded protein. Flexibility creates disorder (entropy), which is stabilizing. Reducing flexibility of the unfolded protein destabilizes it. The cross link constrains the distance between the linked cysteines. In the folded form, the cysteines are already close, so this constraint has little effect. In contrast, this constraint dramatically limits the movement and flexibility of the unfolded protein. Since increased flexibility and the associated increased entropy is the main reason that proteins unfold, reducing the flexibility of the unfolded form shifts the folding-unfolding equilibrium toward folding. This shift stabilizes the protein to unfolding.
A common misconception is that disulfide cross links prevent unfolding; they do not. Proteins remain flexible, dynamic structures even after adding disulfide cross links. They still unfold in urea solutions and upon heating. The RNAse protein in the example above still unfolded in urea solutions, although unfolding required higher urea concentrations that for wild-type. A similar misconception is that protein cross links strengthen and stabilize the folded protein state. They do not. The locations chosen for a disulfide cross link are already nearby in the folded protein. The introduction of the cysteines and the cross link may slightly change the folded protein stability, but this change is most often a destabilization, see below.
The chain entropy model predicts the amount of expected stabilization from a disulfide cross link. This model calculates the change in entropy due to reduced flexibility of the unfolded protein state caused by the disulfide cross link. This model assumes that the introduction of cysteines and formation of the disulfide cross link has no effect on the stability of the folded protein state. The amount of protein stabilization expected from disulfide cross link depends on the size of the ring created by the cross-link. Larger rings limit the flexibility of more amino acids and are therefore more stabilizing. Equation eq. Equation 5.10 below estimates the chain entropy effect of disulfide cross link on the unfolded state.[30]
\[\Delta \Delta S^{unfold} = -2.1 \text{cal/mol} \cdot \ ^{\circ}\!K - \frac{3}{2}R\ln n \tag{5.10}\]
The \(\Delta \Delta S^{unfold}\) refers to change due to the crosslink for the entropy change associated with unfolding; the \(n\) is the number of residues in the ring created by the crosslink and \(R\) is the gas constant. A link between amino acid residues 43 and 80 creates a ring of 38 residues; one more than the difference between 80 and 43. This equation comes from from polymer theory estimates of the probability that the two ends of the chain will coincide in the same volume element. If linked, the two ends must remain in the same volume element, while if unlinked they may move apart and allow greater numbers of conformations. To estimate the Gibbs energy change of the disulfide cross link using the chain entropy model, one multiplies by \(-T\) where \(T\) is the temperature (in °K) yields the anticipated Gibbs energy contribution, eq. Equation 5.11.
\[\Delta \Delta G^{unfold} \text{ (cal/mol)} = -T \cdot \Delta \Delta S^{unfold} = T \left[ 2.1+\frac{3}{2} \cdot 1.987 \cdot \ln n \right] \tag{5.11}\]
For example, the chain entropy model predicts that at 25 °C a cross link between amino acids 43 and 80 will stabilize a protein by 3.9 kcal/mol, which is slightly higher than the measured value of 3.2 kcal/mol in Bacillus RNase, Figure Figure 5.11. The stabilization predicted by eq. Equation 5.11 increases non-linearly with the number of amino acid residues in the ring, Figure Figure 5.12. The estimated stabilization for the smallest possible ring (two amino acids) is 1.2 kcal/mol. The stabilization increases to 3.0 kcal/mol at \(n\) = 15, which is the average separation of disulfide links in natural proteins.[31] With larger numbers of amino acid residues in the ring, the increase slows because the cross link restricts a smaller proportion of the motions in larger rings as compared to smaller rings.
The measured values for stabilization often do not match that predicted by the chain entropy model, likely because the chain entropy model ignores any effect on the folded state. The x-ray structure of Bacillus RNase 43-80 disulfide link mentioned above shows a small reorganization of backbone structure that may account for the slightly lower than predicted stabilization (3.2 vs. 3.9 kcal/mol).[29] Most of the examples in Figure Figure 5.12 seem to fit in this category. Two other disulfide links introduced in Bacillus RNase also are unusual cases. The disulfide at positions 85-102 stabilized the protein more than predicted by the chain entropy model, while the one at 70-92 was so much lower that it destabilizes the protein. The x-ray structures of the more-stable-than-expected protein showed no structural changes as compared to wild type, while the destabilized cross-linked protein showed substantial reorganization that apparently destabilized the folded protein.
To stabilize proteins by introducing a disulfide link, one must identify locations where the disulfide cross link fits in the folded protein without disrupting it. This step requires a structure of the target protein, although a homology model may also be used. Most prediction programs use geometry-only calculations where the distances and angles required for a disulfide crosslink are fit to the current fixed locations of possible amino acid pairs. Three geometry-only programs available as web tools are SSBOND (currently offline),[36] Disulfide by Design,[37] and YOSSHI.[38] The web interface requires only a pdb file of the protein structure for the calculation. YOSSHI also identifies disulfide cross links that occur in homologs. Since these cross links have been tested by natural selection, they may be prioritized for testing. Figure Figure 5.13 shows part of the results of an SSBOND prediction.
Typically the programs first identify amino acid pairs with suitable Cβ-Cβ distances (2.9-4.6 Å) and second, estimate the energy cost of distorting a disulfide bond to fit. Links with high predicted strain energies should be avoided. For example, conformation 1 in Figure Figure 5.13 between Ala4 and Asp89 is predicted to be destabilizing. The expected strain energy (total energy, TOTNRG column) is 6.2 kcal/mol, which is larger than the predicted stabilization energy for this link using eq. Equation 5.11 (4.6 kcal/mol). Since disulfide cross links that create larger rings are predicted to be more stabilizing by the chain entropy model, one normal favors larger rings over smaller rings.
Since geometry-only calculations ignore interactions between amino acids, one must consider these before choosing sites for mutagenesis. First, most researchers also avoid cross links near the active site to prevent possible disruption of catalysis. Second, geometry-only calculations do not consider possible loss of favorable interactions due to removal of existing amino acids, nor do they consider potential bumping interactions between the cysteines and adjacent amino acids. Examination of the protein structure can identify these potential problems. Some researchers also include a molecular dynamics simulations to generate alternative protein conformations.[39] These simulations could identify additional locations where a disulfide cross link can fit or identify locations that should be avoided since the cross link causes large shifts in the backbone that may destabilize the folded protein state.
Disulfide bonds usually form spontaneously upon air oxidation of the cysteines. If the application requires the protein to remain in the cytoplasm, which is a reducing environment, the disulfide bonds may not form spontaneously making this stabilization approach unsuitable. Mutant strains of E. coli where the enzyme thioredoxin reductase is inactivated can form disulfide cross links within the cytoplasm.[40]
While disulfides are the most common way to cross link the amino acid chain for protein stabilization, there are other possibilities.[41] For example, connecting the N- and C-termini into a ring also restricts the main chain motion of the unfolded protein and stabilizes proteins to unfolding. Formation of this ring requires special methods as well as a protein fold where the N- and C-termini are near one another.
Introduce proline residues. One can also reduce the backbone flexibility of the unfolded ensemble without creating crosslinks by selective amino acid substitutions. Proline is the least flexible amino acid because its ring limits its backbone conformations. Similarly, glycine is the most flexible amino acid since no side chain limits the possible backbone conformations. Replacing any amino acid with proline is expected to reduce the flexibility of the denatured ensemble and stabilize the folded protein. Replacing glycine with any other amino acid should have a similar effect. To achieve a net stabilization, these substitutions should not otherwise stabilize the denatured ensemble and must not destabilize the folded form with unfavorable interactions. As in other cases, one should avoid changes in the active site of an enzyme to prevent disrupting catalysis.
\[\text{flexibility: Gly > Ala \& 17 other aa > Pro}\]
Stabilization starts by finding a location where introducing a proline does not disrupt the native structure or interactions. Substitution of an amino acid with proline at such a location should cause no change in the Gibbs energy of the native state, Figure Figure 5.14. The Gibbs energy of the denatured state for the proline variant will be higher than for the wild type due to reduced flexibility of the denatured state.
Introducing proline is a reliable approach to stabilizing proteins. Replacing a typical amino acid with proline reduces the number of conformation in the denatured ensemble by an estimated factor of 7.5 or a \(\Delta S\) of -4 cal/mol· deg,[42] eq. Equation 5.12, which corresponds to a Gibbs energy contribution of 1.2 kcal/mol at 298 °K.
\[\Delta S = R \ln \left[\frac{\text{\# of conformations for typical aa}}{\text{\# of conformations for proline}}\right]= R \ln \left( 7.5 \right) = 4 \text{ cal/mol·deg} \tag{5.12}\]
This approach requires identifying a suitable location for the proline. The main chain angles for proline fit well in three places: near the start of an α-helix, the i+1 position in a type I or II β-turn[42] or the i position of a type II β-turn.[43] These choices of proline location consider only main chain angles; the success of any substitution also assumes minimal changes to side chain interactions. The WHAT-IF web tool at can identify these locations in a protein structure (Click on ‘mutation prediction’, then ‘Suggest proline mutations’). The first example of protein stabilization by proline substitution was an Ala82Pro substitution in T4 lysozyme, which occurs near the start of an α-helix and stabilized the protein by 0.8 kcal/mol, Table Table 5.2.[44] The restricted backbone angles of proline fit well near the start of an α-helix and the rest of the structure also fit a proline residue. The first four residues of a helix lack a partner to which the main chain N-H can donate a hydrogen bond, so the lack of an N-H in proline is not a disadvantage at the start of a helix. The typical success rate for stabilization upon proline substitution is \(\sim \! 50\)%.
| Protein | Location of substitution |
\(T_m\), \(^{\circ}\)C | \(\Delta \Delta G^{unfold}\) (kcal/mol) |
|---|---|---|---|
| wild type | none | 64.7 | 0 |
| Ala82Pro | near start of \(\alpha\)-helix | 66.8 | 0.8 |
| Gly77Ala | near start of \(\alpha\)-helix | 65.6 | 0.4 |
Although replacing glycine with alanine also reduces the flexibility of the denatured state, other considerations limit the effectiveness of this substitution. Replacing glycine residues with alanine reduces the number of conformations available to the unfolded ensemble by approximately a factor of three, which corresponds to an Gibbs energy contribution of 0.4 kcal/mol at 298 °K, which is smaller than that of proline substitution, 1.2 kcal/mol. As with the addition of proline, one must ensure that the glycine to alanine substitution does not strain the folded form. Since glycine can adopt conformations that are not accessible to alanine (e.g., a left-handed helix conformation which occurs in some β-turns), replacements at these locations would destabilize proteins. Second, alanine can create stabilizing interactions within the denatured state that counterbalance the reduced flexibility so the net effect may be zero.[45] The Gly77Ala example in Table Table 5.2 above stabilized the protein, but the x-ray structure suggests that it stabilized the folded form due to the burial of the hydrophobic methyl group and not due to changes in flexibility of the denatured state.
5.4.3 Stabilize the folded form
Strengthening interactions between amino acids in the folded protein is the most obvious approach to stabilizing proteins. Starting from the three-dimensional structure of the protein, one designs various improved interactions. One difficulty with this approach is that proteins already create stabilizing interactions, so this design searches for additional interactions. One must predict both the location for the changes and the replacement amino acids.
Flexible regions identify weak spots in proteins. The effect of flexible regions in proteins can be confusing. Flexibility is a weakly stabilizing feature because it increases entropy.[46] This flexibility may contribute a factor of two, or 0.4 kcal/mol, to stability. While this flexibility is stabilizing, it also indicates that interactions with other amino acids are weak. Creating strong interactions between amino acids such as hydrogen bonds and hydrophobic interactions (several kcal/mol) in place of the weak stabilization of flexibility can be a net gain for stability, Figure Figure 5.15. Thus, it is not removing flexibility that stabilizes the protein, but the addition of new interactions between the amino acids. Loops on the surface and the N- and C-termini of the protein are often the most flexible.
There are two steps to this stabilization approach: first to identify the flexible regions of the protein and second, to identify stabilizing substitutions. One way to identify flexible areas is molecular dynamics simulation, which directly models protein motion.[39] Another way to identify flexible regions is the high B-factors in the x-ray crystal structure.[47] B-factors describe the spreading of electron density assigned to that atom. This spread may be due to movement of the atom during the x-ray analysis (temperature-dependent atomic vibrations), or may be due to the atom occupying several fixed positions in the structure (static disorder), which suggests motion in solution. Errors in model-building can also increase B-factors. The B-FITTER program identifies the twenty amino acids in a protein that have the highest B-factors. B-FITTER calculates the average of the B-factors for all atoms (excluding hydrogen) for each amino acid and displays a list of the twenty amino acid residues with the highest B-factors.
The second step is to identify stabilizing substitutions. One can make changes within the flexible region or next to flexible region. Interactions between amino acids require partners so either approach can yield improvements, but one comparison found larger stabilization with the replacements in the neighboring areas.[48] Some researchers targeted the flexible regions with random mutations,[47] while others used modeling to predict stabilizing substitutions.[48]
Copy features from thermotolerant homolog. Proteins from microorganisms that live in extreme environments (such as high temperatures >80 °C; extremes of pH, high salt, high pressure) tolerate their extreme environment better than homologous proteins from mesophiles (organisms that grow best at moderate temperatures). One protein stabilization strategy is to transfer the amino acids responsible for this extra stability in proteins from extremophiles to the corresponding proteins from mesophiles. This difficulty of this strategy is identifying which amino acids to transfer since some sequence differences impart increased stability, but most of the differences reflect random genetic drift. One approach relies on amino acid sequence comparison within the extremophiles. The expectation is that stabilizing amino acids are conserved in the proteins from extremophiles, but missing from the target protein from mesophiles. This approach differs from the consensus sequence approach, identifies residues conserved among all homologs, including both those from extremophiles and mesophiles. One pitfall of this approach is that residues may be conserved within extremophiles because they are closely related, not because they contribute to stability. For example, comparison of a pectate lyase with four homologs from thermophiles identified nine substitutions common to the enzymes from thermophiles, but absent in the target enzyme.[49] However, only one of these nine substitutions (Arg236Phe) significantly increased the thermostability of the target lyase (12-fold), likely by improving hydrophobic packing.
A second approach relies on the structures of proteins from extremophiles and the identification of stabilizing interactions within them. The stabilizing interactions are not apparent from the structure, so identifying them requires modeling. For example, researchers identified which ion pairs in adenylate kinase from a thermophile contribute to stability using molecular dynamics simulations.[50] During the simulated movement of the protein some ion pairs broke apart, while three of them remained tightly paired. Transfer of these tightly-paired ion pairs into a less stable homolog increased its melting temperature by 0.8 to 3.3 °C.
Stabilizing interactions are not apparent from a comparison of structures because no single factor dominates.[51] Each protein uses a unique mixture of interactions to increase protein stability. For example, many proteins from thermophiles contain more extensive hydrogen bond networks to strengthen electrostatic interactions, but some proteins from thermophiles do not include more extensive hydrogen bond networks but are nevertheless stable at high temperature. Similarly, some, but not all, proteins from thermophiles contain increased atomic packing to strengthen hydrophobic interactions, increased numbers of ion pairs, shortened loops to minimize interactions with the solvent, increased occurrence of Ala in helices and increased oligomerization to enhance interactions between amino acids in the folded state. Other stabilizing features are not apparent in the folded structure because their effect is to destabilize the denatured state. This wide range of features strengthens the argument that there are multiple paths to protein stabilization.
Remove destabilizing electrostatic interactions on the protein surface. Charged residues create favorable electrostatic interactions between residues with opposite charge and unfavorable interactions between residues with the same charge. Optimizing these interactions in the folded protein, especially by removing destabilizing interactions on the surface of the protein, stabilizes the protein. For example, five substitutions on the surface of an acylphosphatase increased the Gibbs energy of unfolding by 2.2 kcal/mol without affecting catalytic activity.[52] Three of these substitutions reversed the charge of the residues and two substitutions introduced new charges.
The two reasons to choose substitutions on the protein surface are that they are likely to be far from the active site and that they are likely to maintain good solvation. Avoiding mutations near the active site increases the probability that the modified protein will maintain catalytic activity (or binding in the case of a binding protein). Choosing residues with good solvation for substitution makes the prediction of stabilizing or destabilizing more accurate. The effect of a charged amino acid residue on protein stability depends on the differences in both electrostatics and solvation in the folded versus unfolded states. Since all residues are assumed to be well-solvated in the unfolded state, choosing residues that are well solvated in the folded state allows the prediction to ignore solvation in the comparison. One reason to avoid substitutions on the protein surface is that they may have a smaller effect on stability than substitutions in the interior of the protein. The environment of residues on the protein surface changes less drastically upon unfolding than the environment of residues in the interior of the protein. Nevertheless, residues on the surface move apart when the protein unfolds, thereby changing the electrostatic interactions.
TKSA-MC is a web-tool to identify residues for replacement due to unfavorable electrostatic interactions.[53] It calculates the electrostatic contribution of all the charged residues in the protein and suggests replacing them if they make an unfavorable electrostatic contribution and lie on the protein surface. It does not suggest what the replacement amino acid should be, but uncharged polar residues or oppositely charged residues are the obvious choices. The electrostatic calculation requires a protein structure and considers the pH of the solution, the distance between the charges, different dielectric constants for the protein and solvent, and the fraction of the residue exposed to solvent. For example, TKSA-MC predicts that replacement of residues Asp36, Asp40, and Glu42 in Staphyloccocal binding protein would stabilize it, Figure Figure 5.16.
Most electrostatic interactions are stabilizing and only a few are destabilizing, so this approach identifies only a few locations for mutagenesis. The ability to predict new stabilizing interactions would yield additional locations for mutagenesis, but this approach has not been tested with this web tool.
Modeling to identify stabilizing substitutions. Rosetta[54] and FoldX[55] are modeling programs widely used to predict protein stability. Both are available online. The programs model the physical interactions of the atoms in the folded form, but Rosetta adds statistical analysis of different properties extracted from protein databases.
FoldX is force field specifically for predicting protein stability. FoldX requires an experimental structure of a protein as the starting point. It does not do geometry optimization or conformational searching; it only calculates unfolding Gibbs energy for the protein and variants of the protein. For protein engineering, it predicts substitutions that stabilize or destabilize a protein. The force field equation contains terms for steric clash and electrostatic interaction like the physics-based force fields, but it also contains terms for solvation and entropy, Equation 5.13.
\[ \begin{split} \Delta G &= a \cdot \Delta G_{vdw} + b \cdot \Delta G_{solvH} + c \cdot \Delta G_{solvP} + d \cdot \Delta G_{wb} \Rightarrow \text{solvation terms}\\ & + e \cdot \Delta G_{hbond} + f \cdot \Delta G_{el} + g \cdot \Delta G_{kon} \quad \quad \quad \quad \Rightarrow \text{electrostatic terms}\\ & + h \cdot T \Delta S_{mc} + k \cdot T \Delta S_{sc} \qquad \qquad \quad \quad \Rightarrow \text{entropy terms}\\ & + l \cdot \Delta G_{clash} \qquad \qquad \quad \quad \Rightarrow \text{steric clash term}\\ \end{split} \tag{5.13}\]
These added terms estimate the Gibbs energies of various interactions of the amino acids with each other in the folded form as compared to interaction with solvent, which mimics the unfolded form. For example, the four solvation terms estimate the attractive van der Waals interaction between water and the protein (from the energy of removal of a protein atom from water to vapor phase), the desolvation of hydrophobic and polar groups upon folding (from the energy of removal of a protein atom from water to an organic solvent), and specific interactions with water when the water molecules that make more than two hydrogen bonds with the protein. In these cases the water molecules are included as atoms in the calculation. The terms in the FoldX force field were scaled with respect to each other to match the predicted Gibbs energy to experimental measurements of >1000 stabilizing and destabilizing amino acid substitutions. A FoldX calculation is part of the HotSpot Wizard web tool (Sumbalova et al., 2018) at: http://loschmidt.chemi.muni.cz/hotspotwizard.
Rosetta’s equation for energy is a hybrid of simple terms that model physical interactions combined with knowledge-based terms to improve accuracy. These knowledge-based terms come from statistical analysis of known protein structures. For example, while FoldX estimates electrostatic interactions using Coulomb’s law and hydrogen bonding terms, Rosetta also weights the estimate by the probability that the two charged atoms occur nearby in PDB structures. Molecular modeling methods typically identify substitutions to increase hydrophobic interactions and various electrostatic interactions. For example, Rosetta predicted improved hydrophobic interactions in cytosine deaminase with three substitutions (A23L, I140L, V108I), which increased the apparent melting temperature by 10 °C.[56] The supporting information contains a detailed example of using Rosetta to identify a stabilizing substitution.
The energies from a molecular mechanics calculation like Rosetta are energy relative to a hypothetical unstrained molecule. A geometry optimization gives you a stable, reasonable structure, but the energy value needs a comparison value to make sense. For example, if you calculate the energy of the wt and a mutant, you can conclude which one of those folded structures is more stable. (The folded state is a collection of many conformations, so you are assuming that this single geometry-optimized structure is a good representative all folded structures that contribute to the folded state.) If you further assume that the energies of the unfolded states are identical, then the difference between the energies of the wt and mutant is the difference in unfolding energies. The units in Rosetta are ‘Rosetta energy units’ which are similar to kcal/mol.
The disadvantage of computational predictions of stabilizing substitutions is their low accuracy, often below 10%. Computer predictions are often incorrect either because the calculated energy of the molecular structure is incorrect or because the structure itself is incorrect and the protein favors a different conformation.
The first type of error is due to approximations used to speed up calculations (e.g., fixing the protein backbone, omitting explicit water molecules, ignoring entropic effects) and due to training of computers using small proteins; in short, the force field is inaccurate. To minimize this first type of errot, Floor and coworkers[48] suggested checking for the following in computer-predicted substitutions: 1) steric clashes, 2) internal cavities, 3) solvent-exposed aromatic residues, 4) solvent-exposed methionine, 5) a hydrogen bond donor or acceptor that has fewer hydrogen bonding interactions than in wild type, 6) a large hydrophobic patch on the protein surface, 7) uncompensated removal of a salt bridge, and 8) destabilization of an \(\alpha\)-helix by removal of \(\alpha\)-helix capping. The supporting information contains PyMOL scripts to check for some of the problems. Even with these checks the success rate for predicting substitutions to stabilize a haloalkane dehalogenase was only 9.2%.
The second type of error, an incorrect structure, is an error of conformational sampling. Computer modeling usually adjusts the residues near the substitution, but in reality a substitution may also affect more distant regions or cause more dramatic rearrangements in the main chain positions. Simulating alternative conformations using molecular dynamics with explicit water molecules adds these interactions and improves the success rate to 13-19%.[39,57] Finding these changes is computationally expensive since it requires long molecular dynamics simulations of protein motions. Machine learning approaches may be a faster alternative by using statistical patterns to accelerate conformational searching.[58]
Even with a perfect structure of the folded protein, errors in predicting the effect of a substitution remain because the computations ignore the effect of the substitution on the unfolded protein state. Protein stability depends on the difference between the folded and unfolded states; ignoring the unfolded state is a major approximation.
Another approach to increasing success is to average the predictions of numerous methods. The assumption is that errors arising from different simplifications will cancel out. For example, the web server PROSS (Protein Repair One-Stop Shop[59]) combines the homology search used by the consensus sequence approach with computational design. First, PROSS searches for homologs to identify which substitutions are allowed at each position. The rationale for this evolution-based constraint is to favor variants that maintain their original molecular function. While the consensus sequence approach predicts stabilizing substitutions from the frequency of occurrence, PROSS adds an energy calculation to predict which substitutions are best. This calculation requires a three-dimensional structure of target protein. The rationale for adding energy calculations is that individual proteins differ so that an amino acid residue that is suitable in most homologs may not be suitable in the target protein. For example, stabilization of acetylcholine esterase involved a replacement for glycine at position 416. There was no clear choice for a replacement based on the frequency of occurrence because nine amino acids, including glycine, appeared at this position with similar frequency. Computational modeling of the replacements using Rosetta eliminated the most commonly-occurring histidine because it fit poorly and suggested third-most-frequently occurring glutamine because it formed a hydrogen bond with a nearby tyrosine.
5.5 Stabilization to stepwise unfolding
So far the discussion assumed that the protein is a single domain protein that unfolds cooperatively, reversibly, and in a single step. Multiple domain proteins likely unfold stepwise with one domain unfolding first followed by other domains. For these proteins, stabilizing substitution must stabilize the domain that unfolds first. Proteins may also associate into oligomers. Oligomerization of folded proteins creates stabilizing interactions between amino acids at the protein-protein interface. (If the interactions were destabilizing, the proteins would not associate.) Creating additional interactions at the protein-protein interface is a new stabilization strategy available for oligomeric proteins.
5.5.1 Multiple domain proteins
Multiple domain proteins may unfold in a single step like single domain proteins, but more often they unfold stepwise with each domain unfolding separately. Stepwise unfolding creates an metastable intermediate with part of the protein folded and part unfolded. The stabilization must focus on the domain that unfolds first.
Bacterial cocaine esterase consists of three domains: an α/β-hydrolase domain, a cap domain, and a jelly roll domain. Molecular dynamics simulations indicated that the cap domain unfolded first. Two substitutions in the cap domain (T172R/G173Q) increased the half-life from \(\sim \! 0.2\) h to 6 h at 37 °C,[60] which corresponds to stabilization of 2.1 kcal/mol. Modeling suggested that the T172R substitution strengthened interactions within the cap domain, while the G173Q substitution strengthened interactions between the cap and α/β-hydrolase domains by creating a new hydrogen bond.
Eijsink and coworkers dramatically stabilized a neutral protease from Bacillus, NprT, to self-degradation by stabilizing it to unfolding, Table Table 5.3.[61]. Proteolytic degradation proceeds from the unfolded form and does not depend strongly on amino acid sequence for this broad specificity protease, so stabilization to unfolding slows the proteolytic degradation. The substitutions increased the apparent melting temperature of the protein. Upon melting, the protease self-degrades rapidly. The stabilizing substitutions were either copied from a heat tolerant homolog or rationally designed to stabilize the protein. The half-life at 70 °C increased dramatically from <0.5 to 170 min, which corresponds to 4.0 kcal/mol. The researchers named this stabilized enzyme ‘boilysin’ because of its ability to remain active at 100 °C.
| \(\Delta T_m\), \(^{\circ}\)C |
Substitution | How identified | Molecular basis for stabilization |
|---|---|---|---|
| +12.3 | Ala69Pro Thr63Phe |
heat tolerant homolog heat tolerant homolog |
less flexible unfolded domain hydrophobic interaction |
| +7.6 | Gly58Ala Thr56Ala Ala4Thr |
heat tolerant homolog heat tolerant homolog heat tolerant homolog |
less flexible unfolded domain not determined not determined |
| +6.9 | Ser65Pro 8-60Cys link |
rational design rational design |
less flexible unfolded domain less flexible unfolded domain |
The stabilizing mutations all cluster in one region of the protein, Figure Figure 5.17. This multiple domain protein unfolds stepwise, and this region of the protein is the part that unfolds first and is then cleaved by other protease molecules. Stabilizing this region to unfolding was the key to preventing proteolytic degradation.
5.5.2 Oligomeric proteins
Dimerization or oligomerization of monomeric native state proteins creates new protein-protein interactions while decreasing protein-water interactions and therefore stabilizes the folded native form, Figure Figure 5.18. In contrast, if unfolded or partially unfolded forms dimerize and oligomerize, then these interactions stabilize the unfolded structures, which eventually leads to aggregation.
Substitutions at the oligomer interface that strengthen the interaction between monomers will stabilize proteins. In dramatic example, a replacement to remove a negative charge at the dimer interface of malate dehydrogenase (Glu165Gln or Glu165Lys) increased the melting temperature by 24 °C.[62] Another approach is to covalently link the monomers into dimers by introducing a disulfide link. Comparison of the structures of a peptidase and a thermostable homolog identified a cross-link between monomers in the thermostable homolog. Adding the this link to the peptidase increased its denaturation temperature by 30 °C.[63]
5.6 Weaknesses other than unfolding
While unfolding is the weak point of most proteins, in some cases, other weaknesses dominate protein instability. Cooking irreversibly aggregates egg white proteins so that they do not refold and restore a liquid upon cooling. Other examples are chemical modification of proteins such oxidation or degradation by proteases. These other weaknesses may be related to unfolding. Proteins unfold at least partially before they aggregate or are degraded by proteases. Chemical modification can promote unfolding. A different type of protein instability is short serum half-life, which depends on biochemical processes that clear proteins from the bloodstream.
5.6.1 Protein aggregation
Protein aggregation is the irreversible association of partially or fully unfolded proteins into insoluble particles. Aggregation-prone regions assemble by intermolecular β-structured interactions to form the core of the aggregate. Aggregation can occur when unfolded proteins encounter one another. For example, cooking an egg first unfolds the egg white proteins to expose hydrophobic regions to the solvent. These regions associate with similar exposed hydrophobic regions of other egg white protein molecules leading to oligomerization and eventually insoluble aggregates, eq. Equation 5.14, Figure Figure 5.19. Here \(k_{agg}\) is the aggregation rate constant.
\[N \xrightleftharpoons[]{\,K^{unfold}\,} D \xrightarrow{k_{agg}} \text{aggregate} \tag{5.14}\]
Aggregation can also occur under mild conditions. For example, overexpressing protein in bacteria creates high concentrations of unfolded protein. If protein folding is slow, the unfolded proteins can aggregate into insoluble particles called inclusion bodies. Biotherapeutic proteins can also aggregate during storage, which creates a danger of an immune response to the aggregates.
Measuring a loss of function after heating identifies aggregation as a contribution to protein instability. For example, one measures the enzyme activity at room temperature, incubates the sample at an elevated temperature, then cools it to room temperature and measures enzyme activity again. The decrease in activity reveals the fraction of enzyme irreversibly unfolded due to heating. The values reported are typically the half-life at a specific temperature. For non-catalytic proteins, the change in intrinsic fluorescence can measure the amount of natively folded protein remaining. The activity loss depends on both how much of the protein unfolds upon heating (its inherent stability) and on its ability to refold upon cooling (which may be prevented by aggregation), eq. 3.x above.
To convert the measured loss of activity to half-life, researchers assume the inactivation follows first-order kinetics, eq. Equation 5.15, where \(A\) is the activity at time \(t\), \(A_0\) is the initial enzyme activity and \(k\) is the inactivation rate constant. The natural logarithm of enzyme activity (\(\ln A\)) decreases linearly with time. A plot of several measurements of the natural logarithm of the remaining activity at different times yields a straight line, with a slope of \(-k\) and a y-intercept of \(\ln A_0\).
\[\ln A = -k \cdot t + \ln A_0 \tag{5.15}\]
By measuring the inactivation rate constants for both the wild-type enzyme and the engineered variant, eq. Equation 5.16, yields the change in Gibbs energy of activation, \(\Delta \Delta G^‡\), for the rate of enzyme inactivation. If the mutant is more stable, then \(\frac{k_{mut}}{k_{wt}}\) is less than one, so the value of \(\Delta \Delta G^{\ddagger}\) is positive, indicating an increase in the activation energy to unfold the protein.
\[\Delta \Delta G^‡ = -RT\ln \left[ \frac{k_{mut}}{k_{wt}} \right] \tag{5.16}\]
In most cases, the rate determining step in enzyme inactivation is the unimolecular unfolding step, so the assumption of first-order kinetics is justified. However, refolding of the enzyme may also depend on aggregation of the protein into an insoluble precipitate, so the rate determining step may involve several enzyme molecules. In these cases, the inactivation will not fit eq. Equation 5.15, which assumes first order kinetics. Another way to measure irreversible denaturation by heat is to measure an apparent melting temperature, \(T_{m, app}\). It is not a true melting temperature, only an apparent one, because it is an irreversible process.
Two strategies to reduce aggregation are 1) to reduce unfolding (\(K_{eq}\) in eq. Equation 5.14) or 2) slow down aggregation (\(k_{agg}\) in eq. Equation 5.14). Minimizing unfolding of the protein will use strategies above in Section Section 5.4. For example, a combination of three substitutions designed to stabilize the native cytosine deaminase slowed its aggregation 30-fold.[56]
Since aggregation occurs mainly through hydrophobic interactions, one method to slow aggregation is to minimize the hydrophobicity of solvent-exposed regions. To reduce the aggregation of antibodies, Chennamsetty and coworkers[64] first identified hydrophobic areas on the surface including hydrophobic patches that become exposed as the protein moves. Replacing hydrophobic residues with hydrophilic ones within these patches reduced aggregation. Several web tools predict substitutions needed to reduce protein aggregation: Aggrescan3D[65] and Solubis.[66] After predicting the aggregation-prone regions (hydrophobic, β-sheet forming stretches) the web tools predict substitutions to 1) reduce their unfolding and exposure to solvent by strengthening interactions between the aggregation-prone area and the rest of the protein or 2) eliminate the aggregation propensity of the region. These substitutions are charged residues (R, K, D, or E) which reduce hydrophobic interactions needed for aggregation or proline, which hinders the formation of β-sheet structures in aggregates. The tools use FoldX to choose the substitution that is expected to best stabilize the folded form.
Hindering aggregation can even overcome a decrease in stability. Substitution A281E in Candida antarctica lipase B decreased its melting temperature from 58 to 51 °C indicating a decline in stability. However, this substitution increased the half-life of this enzyme at 70 °C 22-fold demonstrating a reduced propensity to aggregate.[67] This substitution makes a hydrophobic region of the enzyme less hydrophobic and less likely to aggregate.
A more drastic approach to reducing the association between protein molecules is supercharging. Supercharging is extensive substitutions on the protein surface to increase the net charge as high as +48, so the protein molecules repel one another and cannot aggregate.[68] Introducing large numbers of charge residues on the protein surface also reduce its hydrophobicity as charged amino acids replace uncharged amino acids. For example, green fluorescent protein (GFP, net charge = -7) loses its fluorescence at 100 °C due to unfolding, Equation 5.17. Cooling to room temperature does not restore fluorescence because the unfolded protein has aggregated and precipitated. In contrast, green fluorescent protein variants engineered to have a +36 or -30 net charge also lost their fluorescence upon heating to 100 °C due to unfolding, but upon cooling, they did not aggregate and regained 62 and 28%, respectively, of their original fluorescence. This supercharging did not make the GFP variants more stable; in contrast, the GFP variants unfolded more readily than wild-type in urea, but supercharging did reduce their propensity to aggregate in the unfolded state.
\[ \begin{split} \color{green} GFP^{7-}_{folded} &\xrightleftharpoons[]{\,\text{heat}\,} GFP^{7-}_{unfolded} \xrightarrow{k_{agg}} \text{aggregate}\\ \color{green} GFP^{30-}_{folded} &\xrightleftharpoons[]{\,\text{heat}\,} GFP^{30-}_{unfolded}\\ \end{split} \tag{5.17}\]
Two disadvantages of supercharging are the potential to destabilize the protein and to disrupt its function. Residues with the same charge create unfavorable electrostatic interactions, which destabilize proteins. A change in electrostatic environment can also shift the p\(K_a\) of residues in the active site, which may affect binding or catalysis. For example, catalysis (\(k_{cat}\)) by a supercharged glutathione-S-transferase (net charge -40 for the dimer) was three-fold slower than wild-type (net charge +2 for the dimer).
5.6.2 Chemical modification
The side chains of three amino acids readily undergo spontaneous chemical modification: 1) asparagine deamidate to aspartate, 2) methionine oxidizes to a sulfoxide and 3) cysteine oxidizes to disulfide protein oligomers as well as sulfur oxides such as sulfenic acids (RS-OH). Replacement of problematic residues with a non-reactive residue eliminates the possibility of modification, but replacement of every asparagine, methionine, and cysteine in a protein is rarely needed. Many of the residues may not undergo modification, and even when they do, some modified residues will have little effect on protein properties. Spontaneous deamidation of asparagine to aspartate is a common chemical modification of proteins. The most reactive are Asn-Gly sequences in sterically unhindered regions which have half-lives as short as 6 d in physiological conditions.[69] The least reactive asparagines are \(10^5\)-fold more stable. Deamidation converts the neutral Asn residue to a negatively charged Asp residue. This change can impair function, destabilize the protein or have no effect. The web tool at www.deamidation.org predicts asparagine deamidation rates based on the 3-D structures of proteins. Glutamine can also deaminate to glutamate, but the rate is about a thousand-fold slower, except for the case of an N-terminal Gln, which deaminates at a rate similar to Asn. Oxidation of sulfur atoms in methionine and cysteine alters a proteins properties and may destabilize or inactivate it. The first protein engineering of an industrial enzyme was the removal of an oxidation sensitive methionine from the detergent protease subtilisin.[70] The existing subtilisin could tolerate most of the harsh conditions of laundry (heat, high pH, and detergent), but it could not tolerate bleach. Bleach oxidized a methionine near the active site to the methionine sulfoxide (R-S(O)-CH3), which hindered binding of the substrate proteins to inactivate the protease. Replacement of this problematic methionine with alanine created a bleach-tolerant protease. A cysteine-to-serine replacement stabilizes several cytokine drugs. This change does not affect biological activity, but avoids oligomerization by the formation of non-native intermolecular disulfide links between proteins. The specific activity of interferon-β, when expressed in E. coli, was about 10-fold lower than the native protein. The researchers hypothesized that oligomerization through intermolecular disulfide bonds caused the lower activity.[71] Replacement of Cys17 with Ser eliminated oligomerization and restored the specific activity to that of the native protein. The commercial drug, Betaseron\(^®\), contains this substitution. A similar Cys125Ser substitution stabilizes human interleukin 2, marketed as Proleukin\(^®\).
Some applications require proteins to work under conditions where other chemical modifications are possible. For example, glucose, an aldehyde, can react with lysine to form an imine. The enzymatic conversion of corn syrup (glucose) to high fructose corn syrup to increase its sweetness requires the enzyme xylose isomerase to tolerate high concentrations of glucose. Stabilization involved replacing a surface lysine that reacted with the glucose.[72] At very high or low pH, irreversible chemical reactions can degrade the protein. For example, base-catalyzed β-elimination at pH > 8 alters cystine (disulfide) residues’ side chains. Peptide backbone links next to aspartic acid residues can hydrolyze at pH < 4.
5.6.3 Serum half-life
Extending the serum half-life of a therapeutic protein enhances its efficacy because it remains active for a longer time. It also lowers cost because less therapeutic protein is required and improves delivery because injections are less frequent.[73,74] Three biochemical processes limit the serum half lives of proteins: (i) degradation by proteolytic enzymes, (ii) rapid filtration by the kidneys and (iii) receptor-mediated endocytosis. Small proteins (<50 kDa) typically have short serum half lives (5–50 min).
To stabilize proteins to proteolysis researchers identify the cleavage sites and modify them to slow hydrolysis. For example, human glucagon-like peptide-1 (GLP-1) has a serum half-life of only 2-3 minutes, in part due to rapid proteolytic cleavage by dipeptidyl peptidase 4 between the 8Ala-9Glu linkage, Figure Figure 5.20\(\text{a}\). To slow down proteolysis researchers replaced the alanine residue with glycine (removal of the methyl group at C\(\alpha\)) or with 2-aminoisobutyric acid, Aib (addition of another methyl group at C\(\alpha\)).[75] Researchers replaced the alanine, not the glutamate because proteases are more selective for the amino acid on the carbonyl side of the amide bond rather than the amino acid on the amine side of the amide bond. Other protease that occur in serum besides dipeptidyl peptidase 4 are neutral endopeptidase, plasma kallikrein, and plasmin.
To slow filtration by the kidneys, researchers increase the molecular weight of the therapeutic protein to above ~50 kDa. In the case of GLP-1, researchers increased the molecular weight by adding an octadecanoic diacid group, which binds to serum albumin, which has a molecular weight of 66.5 kDa, Figure Figure 5.20\(\text{b}\) above. The engineering of additional glycosylation sites or the chemical addition of poly(ethylene glycol) chain are other ways to increase the molecular weight of a therapeutic protein.
The choice of endogenous albumin as a binding protein also minimizes the third process that lowers serum half-life, receptor-mediated endocytosis. Endothelial cells lining the bloodstream internalize proteins via receptor-mediated endocytosis, which transfers the proteins from the blood to adjacent cells. However, some proteins, such as serum albumin and immunoglobulin G (IgG) are recycled back to the bloodstream. Upon endocytosis proteins transfer to the endosome, then to the lysosome for degradation. However, the endosome is acidic and contains FcRn receptor proteins, which tightly bind albumin and the Fc region of IgG proteins. Instead of degradation, this complex moves back to the cell surface where the higher pH of the blood weakens the binding between albumin or IgG and FcRn, causing release of the albumin or IgG back to the bloodstream. This recycling mechanism accounts for the long half-life of albumin and IgG in the blood. The modified form of GLP-1 in Figure Figure 5.20\(\text{b}\) above has a half-life of ~160 hours (3,200-fold longer than GLP-1) is an treatment for type 2 diabetes (Ozempic\(^®\)) and obesity (Wegovy\(^®\)).
Etanercept is another example where receptor-mediated recycling extends the serum half-life, but in this case using the Fc fragment of IgG, Figure Figure 5.21.[76]. The active protein is the extracellular domain of p75 tumor necrosis factor receptor, which binds the tumor necrosis factor cytokine to reduce inflammation in rheumatoid arthritis. Two copies of the active protein increase its affinity 100-4000-fold over monomeric counterparts, likely by better mimicking the dimeric structure of the receptor. These two active proteins are fused to the Fc fragment of human IgG, which increases serum half-life by enabling recycling of filtered protein back to the bloodstream.
Another approach to increasing the serum half-life of therapeutic proteins is to create insoluble depot of the protein that slowly releases into the bloodstream. One example is the noncovalent entrapment of a GLP-1 homolog in poly(, -lactide-co-glycolide) to form 60 \(\mu\)m microspheres. Diffusion and hydrolysis of the polymer release the GLP-1 homolog with a half-life of 5-6 days. Another example is a long-acting insulin called insulin glargine. The protein sequence has been modified to alter the isoelectric point so that it precipitates as microcrystals at physiological pH and then slowly dissolves over twenty-four hours. The modified insulin is formulated in pH 4 solution where it is soluble, but then precipitates upon injection in the muscle.
5.7 Concluding remarks
Most single substitutions increase the stability of a protein by ≤1 kcalmol (3–4 °C increase in melting temperature), so large stabilizations require multiple stabilizing mutations. In many cases, especially when the substitutions are far from one another, their stabilization effects will be approximately additive. For example, a heat-stabilized α-amylase contains at least ten amino acid substitutions, Table Table 5.4.[77] This enzyme catalyzes the hydrolysis of starch to glucose oligomers in the manufacture of corn syrup. High temperatures (90 °C) increase the solubility of the starch and speed up hydrolysis, but require a heat-tolerant α-amylase. Stabilizing substitutions include replacement of amino acid residues that can undergo chemical modification, substitutions to stabilize the native form, substitutions to destabilize the unfolded form and substitutions discovered by random mutagenesis where the stabilization mechanism is unknown.
| Approach | Specifics |
|---|---|
| prevent chemical modification |
remove oxidation (M197) and deamidation sites (N188, Q284) |
| stabilize N | bury Ca\(^{2+}\) ion (A181T), minimize electrostatic destabilization (H156Y) |
| destabilize D | introduce proline (R124P) |
| random mutagenesis |
M15T, H133I, N199S, A209V |
In another example, twelve substitutions in a halohydrin dehalogenase increased its apparent melting temperature by 28 °C and increased its ability to tolerate organic solvents.[57] The researchers first identified stabilizing single substitutions and then combined the best twelve into a single variant. The increased stability was mainly due to redistributed surface charges and improved interactions between subunits in this homotetrameric enzyme.
References
Supporting Information
SI 5.1. Derivation of equation relating spectral changes and equilibrium constant, eq. Equation 5.4.
Deriving the relationship between spectral changes and equilibrium constant is straightforward. The protein is either in the native state or in the denatured state, so the sum of the two fractions is one: \(F_N + F_D = 1\). The observed fluorescence at each urea concentration, \(Y_{obs}\), is the sum of the fluorescence contributions from the native state, \(Y_N · F_N\), where \(Y_N\) is the fluorescence intensity from the pure native state, and the denatured state, \(Y_D · F_D\), where \(Y_D\) is the fluorescence intensity from the pure denatured state.
\[Y_{obs} = Y_N · F_N + Y_D · F_D\]
Substituting \(F_N = 1 - F_D\) yields:
\[F_D = (Y_{obs} - Y_N)/ (Y_D - Y_N)\]
Similarly substituting \(F_D = 1 - F_N\) yields:
\[F_N = (Y_D -Y_{obs})/ (Y_D - Y_N)\]
Dividing these two equations yields the equation below, which is the same as eq. Equation 5.4.
\[K^{unfold}=\frac{F_N}{F_D}=\frac{Y_{obs}-Y_N}{Y_D-Y_{obs}}\]
SI 5.2. Example of linear extrapolation of protein unfolding data in urea Protein A was dissolved in solutions of different concentrations of urea and allowed to reach equilibrium. The fluorescence spectra of these solutions showed an increase in fluorescence at 2-5 M urea, Figure Figure 5.22. Using this fluorescence data, calculate the Gibbs energy of unfolding in pure water for protein A.
To solve this problem, first estimate the fluorescence of the native and denatured states. The fluorescence of the native state is the flat part of the curve before unfolding begins. The average of the values at 0 and 1 M urea yields \(Y_N\) = 71. The fluorescence of the denatured state is the flat part of the curve after unfolding is complete. The average of the values at 6-9 M urea yields \(Y_D\) = 155.
Next, estimate the equilibrium constant between the native and denatured forms at each urea concentration in the unfolding region. For this example: \(K^{unfold} = (Y_{obs} - 71)/ (155 - Y_{obs})\), so one can calculate \(K_{unfold}\) at different urea concentrations by substituting the experimental values of \(Y_{obs}\). This procedure yields the values in Table Table 5.5. Convert the equilibrium constants to Gibbs energy using R = 1.987 cal/ mol °K and T = 298 °K, which yields the \(\Delta G^{unfold}\) values in the table.
| [urea] (M) |
\(K^{unfold}\) | \(\Delta G^{unfold}\) (kcal/mol) |
|---|---|---|
| 2.0 | 0.14 | 1.2 |
| 3.0 | 0.40 | 0.54 |
| 4.0 | 4.6 | -0.90 |
| 5.0 | 16 | -1.6 |
Finally, plot the \(\Delta G^{unfold}\) values as a function of urea concentration, fit the data to a straight line and extrapolate to pure water (\([\text{urea}]\) = 0), Figure Figure 5.23.
The extrapolation yields a y-intercept of +3.3 kcal/mol, which is the Gibbs energy of unfolding of this protein in pure water. This protein is less stable than typical proteins, but even for this one, the equilibrium amount of denatured protein in pure water is only 0.38%, which would be difficult to measure directly.
Problems
5.1. The following are examples of interactions that can stabilize proteins: a) improved core packing, b) higher backbone rigidity, c) increased surface polarity, d) increased salt bridges, e) increased hydrogen bonds between amino acids, f) replace residues with proline, g) increase number of charged amino acids. For each interaction, account for the stabilizing effect by describing the effect that it has on the folded protein and on the unfolded protein.
5.2. a) Floor and coworkers* used the ‘Dynamic Disulfide Discovery algorithm’, which they developed previously,† to predict disulfide cross link locations. That algorithm yielded seven residues pairs. (You do not need to look up the details of this algorithm.) The Dynamic Disulfide Discovery algorithm excludes residue pairs that are within 15 positions of each other. Explain why this is a good idea. If you apply this criterion to the results of SSBOND, how many residues pairs are still suitable for engineering?
Besides excluding residues pairs that are within 15 positions of each other, suggest another criterion that you could use to exclude a suggested pair. Provide a rationale for your criterion.
The apparent melting temperatures for the wild type and two of the variants are listed below. In which variant did the disulfide bond stabilize the protein? Explain how you came to your conclusion. For the variant that did not stabilize the protein, suggest a reason why it did not work based on the data in the table. Explain your reasoning. Hint: Consider the structure of amino acids R and D.
| Substitutions | \(T_m\)(oxidized) | \(T_m\)(reduced) |
|---|---|---|
| none (wt) | 51 \(^{\circ}\)C | 49 \(^{\circ}\)C |
| A5C/A185C | 56 \(^{\circ}\)C | 49 \(^{\circ}\)C |
| R20C/D70C | 50 \(^{\circ}\)C | 45.5 \(^{\circ}\)C |
- Compare a prediction to a measured value. For the prediction, use equations 5.10 and 5.11 in the text to predict the expected Gibbs energy of stabilization for the two variants in part d above. For the experimental values, estimate the measured Gibbs energy of stabilization from the change in \(T_m\) using equation 5.8 (page 104) of the text. Comment on the match or mismatch between predictions and measured values. For the temperature, use the melting temperature for the variant that you are comparing.
5.3. Predicting stabilizing single substitutions with Rosetta. Jones and coworkers identified several substitutions that stabilize the esterase SABP2.[78] Table Table 5.6 lists four of the best substitutions. The increase in thermostability was measured two ways. The first method measured the time for irreversible inactivation at 60 °C. The sample was incubated at 60 °C for 15 min, then cooled to room temperature and the remaining catalytic activity was measured. The loss in catalytic activity was assumed to be first order in enzyme concentration and the expected half-life was calculated. Wild-type SABP2 had a half-life of 3.5 min. The single substitution variants in Table Table 5.6 all showed a longer half-life. The second two variants showed a larger increase in the half-life than the first two variants. The second measure of thermostability involved unfolding of the protein in urea and compared the concentration of urea where the protein was half unfolded, [urea]\(_{1/2}\). Wild-type SABP2 was half-unfolded at 2.23 M urea. The variants unfolded at higher concentrations of urea. The increase in urea concentration needed to half-unfold the protein was higher for the first two variants than for the second two.
| Substitution | Increase in half-life at 60 °C |
Increase in [urea]\(_{1/2}\) |
|---|---|---|
| L65F | 1.2-fold | +1.5 M |
| E215P | 1.5-fold | +1.0 M |
| Q221M | 6.6-fold | +0.4 M |
| A230V | 3.9-fold | +0.4 M |
This question tests how well Rosetta predicts these four stabilizing substitutions. A webtool for Rosetta is available at ROSIE2 (Rosetta Online Server that Includes Everyone). In particular, we will use the tool stabilize-PM (point mutations), which predicts stabilizing single substitutions.[79] Using the tool is simple; one specifies the pdb code for the protein (1y7i for SABP2) and which positions you would like to mutate (A[65,215,221,230]), which means the four positions listed in chain A). By default, Rosetta tests 19 amino acids (all except cysteine) as replacements. During the calculations, Rosetta allows the two amino acids on either side of the mutated postion to move freely, all amino acids within 10 Å of the mutated position to move under weak constraints and all remaining amino acids to move only slightly (strong constraints). The results of the calculations (heat map in Figure Figure 5.24) show only one predicted stabilizing substitution, Leu65Phe, which is one of the experimentally identified stabilizing substitutions.
- Thermostability was measured two different ways, which gave differing results about which substitution was the most stabilizing. Which of these experimental values should be compared to the Rosetta calculations? Explain your choice.
- Rosetta failed to predict the Glu215Pro stabilizing substitution. Explain what was omitted from the Rosetta calculation that could have caused this error.
You may explore the Rosetta web tools for your own protein engineering project. The webtool requires login with a Github account, but is otherwise simple to use.
Answers
Click to show answers
5.2. a) The stabilization increases with the size of the ring formed. If the residues are within 15 positions, the expected stabilization is only 3.1 kcal/mol (using eq 3.13 in the text, n = 16, T = 300 \(^\circ\)K) and the actual stabilization is often less, so focusing on the larger rings increases the likelihood of success.
Avoid pairs near the active site. Avoid pairs like 2 where you replace an ion pair (Arg, Asp) which stabilizes the folded state. The replacement would destabilize both the folded and unfolded states leading to no net stabilization.
The A5C/A185C cross link stabilizes the protein because the melting temperature of the oxidized form increases by 5 \(^\circ\)C as compared to wt. The R20C/D70C cross link destabilizes the protein because the melting temperature of the oxidized form decreases by 1 \(^\circ\)C as compared to wt. Replacing the R/D pair with a cystine cross link removes an ion pair from the folded protein. The replacement destabilized the folded form (mp of reduced form is 3.5 \(^\circ\)C lower than for wt).
A5C/A185C, n = 181, T = 327.5 \(^\circ\)K, predicted \(\Delta \Delta G = 5.8\) kcal/mol; observed 5/3.6 = 1.4 kcal/mol; also poor agreement. The cysteines seem to fit because the mp of the reduced form is unchanged from wt. The crosslink may not restrict the flexibility as much as expected.
R20C/D70C, n = 51, T = 323.5 \(^\circ\)K, predicted \(\Delta \Delta G = 4.5\) kcal/mol; observed -1/3.6 = -0.28 kcal/mol. The poor agreement may be due to removal of a ion pair from the folded form. Removal destabilizes folded form, but has no effect on the unfolded form because unfolding breaks the ion pair.
* Floor, R. J., Wijma, H. J., Colpa, D. I., Ramos-Silva, A., Jekel, P. A., Szymanski, W. et al. (2014) Computational library design for increasing haloalkane dehalogenase stability. ChemBioChem 15, 1660-72.↩︎
† Wijma et al. (2014) Computationally designed libraries for rapid enzyme stabilization. Prot. Eng. Des. Sel. 27, 49–58; http://doi.org/10.1093/protein/gzt061↩︎