Peptalyzer™ supports noncanonical amino acids across multiple physicochemical features used in peptide analysis. These include hydropathy, hydrophilicity, membrane partitioning, and composition-based polarity classification. Each of these models was originally developed for the 20 canonical amino acids. Extending them to noncanonical residues requires explicit modeling assumptions, not direct transfer.
This article explains how noncanonical amino acids are handled in each feature, what is calculated directly, what is approximated, and where the limits are. The goal is to keep all calculations transparent, so it is always clear which parts of the result are grounded in canonical data and which rely on curated or proxy values.
Across all features, Peptalyzer™ follows a consistent approach:
- Use canonical constants where they exist
- Introduce curated values when a close chemical analog is defensible
- Use proxy or model-based values when no direct mapping exists
- Disable calculations when no reliable approximation can be justified
Each section describes how a given feature is computed, how noncanonical residues are integrated, and what level of confidence can be expected from the result.
Analyze Modified Peptides with Peptalyzer™
Use Peptalyzer™ to include noncanonical amino acids in your sequence and assess their impact on physicochemical properties and synthesis behavior.
📘 What will you learn here?
Peptide Charge Properties
Peptalyzer™ computes peptide charge using a fractional Henderson–Hasselbalch model, not discrete charge states. Canonical residues are handled directly through established pKa scales: Asp (D), Glu (E), Cys (C), and Tyr (Y) contribute acidic terms, while His (H), Lys (K), and Arg (R) contribute basic terms. All other canonical residues are treated as side-chain neutral. Free N- and C-termini are always included as independent ionizable groups.
For noncanonical amino acids, charge behavior is fully driven by residue-library metadata. Most supported noncanonical residues (Aib, AmPhe, beta-Ala, Cit, Hyp, Nle, Nva) are defined with ionization_role = none, meaning they contribute no side-chain charge but remain fully included in the sequence. The only exception is Orn, which is treated as a basic residue through analog mapping to Lysine, enabling standard ionization modeling.
Net charge is evaluated at the selected pH using fractional protonation, and pI is determined by bisection on the same charge function, ensuring internal consistency within each selected pKa scale (IPC2_peptide, Bjellqvist, EMBOSS, Lehninger). Note that the noncanonical residues are mapped onto this same ionizable-group framework through predefined metadata rules rather than independent pKa parameterization.
Canonical Handling
- Direct use of pKa scales with fractional charge contributions
- Acidic: D, E, C, Y
- Basic: H, K, R
- All others: neutral side chains
- Termini handled explicitly
Noncanonical Handling
- Metadata-driven ionization behavior
- Neutral residues (Aib, AmPhe, beta-Ala, Cit, Hyp, Nle, Nva) contribute zero side-chain charge
- Orn uses Lys-like ionization (analog transfer)
Limitations
- No residue-specific pKa values for noncanonical amino acids
- Orn may overestimate charge due to Lys approximation
- beta-Ala treated within α-peptide framework despite structural differences
- No environment-dependent pKa correction (intrinsic model only)
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct | neutral | low | No side-chain charge |
| R (Arg) | yes | direct | basic | low | Strongly basic |
| N (Asn) | yes | direct | neutral | low | Non-ionizable side chain |
| D (Asp) | yes | direct | acidic | low | Fully modeled |
| C (Cys) | yes | direct | acidic | low | Weak acid |
| Q (Gln) | yes | direct | neutral | low | Non-ionizable |
| E (Glu) | yes | direct | acidic | low | Fully modeled |
| G (Gly) | yes | direct | neutral | low | Non-ionizable |
| H (His) | yes | direct | basic | low | Partially protonated near neutral pH |
| I (Ile) | yes | direct | neutral | low | Non-ionizable |
| L (Leu) | yes | direct | neutral | low | Non-ionizable |
| K (Lys) | yes | direct | basic | low | Fully modeled |
| M (Met) | yes | direct | neutral | low | Non-ionizable |
| F (Phe) | yes | direct | neutral | low | Non-ionizable |
| P (Pro) | yes | direct | neutral | low | Non-ionizable |
| S (Ser) | yes | direct | neutral | low | Non-ionizable |
| T (Thr) | yes | direct | neutral | low | Non-ionizable |
| W (Trp) | yes | direct | neutral | low | Non-ionizable |
| Y (Tyr) | yes | direct | acidic | low | Weak acid |
| V (Val) | yes | direct | neutral | low | Non-ionizable |
| Aib | partial | metadata | neutral | low | No ionization |
| AEEA | partial | metadata | neutral | low | Neutral non-ionizable PEG-like linker (ionization_role = none); no side-chain charge |
| Ahx | partial | metadata | neutral | low | Neutral non-ionizable linker (ionization_role = none); no side-chain charge |
| AmPhe | partial | metadata | neutral | low | No ionization |
| beta-Ala | partial | metadata | neutral | medium | β-backbone outside calibration |
| Cit | partial | metadata | neutral | low | Neutral urea side chain |
| Hyp | partial | metadata | neutral | low | No ionization |
| Nle | partial | metadata | neutral | low | No ionization |
| Nva | partial | metadata | neutral | low | No ionization |
| Orn | partial | analog (→ Lys) | basic | medium | May overestimate charge |
GRAVY and Kyte–Doolittle Hydropathy
Peptalyzer™ computes hydropathy using a single Kyte–Doolittle (KD) value per residue, shared across both outputs: GRAVY (global average hydropathy) and the Kyte–Doolittle Membrane/Core profile (position-resolved hydropathy along the sequence). Canonical amino acids use the original KD constants directly. Noncanonical residues are included only when a KD value is defined in the residue library (physchem_scales.hydropathy_kd). In the current implementation, Cit, Nle, Nva, and Orn have explicit curated KD values and are treated as supported, while Aib, AmPhe, beta-Ala, and Hyp are included through curated proxy or model assumptions (partial support). GRAVY is calculated as the arithmetic mean of all residue KD values, and the KD profile reuses the same values along the sequence. Smoothing does not change per-residue KD values but modifies the visual profile by averaging local segments. If any residue lacks a KD value, both outputs are disabled to avoid silent fallback.
Canonical Handling
- Direct use of Kyte–Doolittle constants for all 20 amino acids
- Same per-residue values used for GRAVY and KD profile
Noncanonical handling
- Uses residue-library hydropathy_kd values
- Cit, Nle, Nva, Orn: explicit curated values (treated as supported)
- Aib, AmPhe, beta-Ala, Hyp, Ahx, AEEA: curated proxy/model KD values (partial support; excluded in strict mode, included with caveats in exploratory mode)
Limitations
- KD is a residue-level hydropathy scale and does not account for backbone type, conformation, or sequence context
- Proxy values for noncanonical residues represent model assumptions, not experimental KD measurements
- beta-Ala and Aib fall outside the original KD calibration domain. It is assigned a neutral proxy value (≈0) due to lack of direct analog; this is a pragmatic approximation rather than a chemically derived KD value.
- Aib alters conformational behavior (helix promotion), which is not captured by residue-level KD values.
- In strict mode, partial-support residues are excluded; in exploratory mode, they are included with explicit caveats
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | KD 1.8 | low | Used directly in GRAVY and KD profile |
| R (Arg) | yes | direct canonical | KD -4.5 | low | Used directly in GRAVY and KD profile |
| N (Asn) | yes | direct canonical | KD -3.5 | low | Used directly in GRAVY and KD profile |
| D (Asp) | yes | direct canonical | KD -3.5 | low | Used directly in GRAVY and KD profile |
| C (Cys) | yes | direct canonical | KD 2.5 | low | Used directly in GRAVY and KD profile |
| Q (Gln) | yes | direct canonical | KD -3.5 | low | Used directly in GRAVY and KD profile |
| E (Glu) | yes | direct canonical | KD -3.5 | low | Used directly in GRAVY and KD profile |
| G (Gly) | yes | direct canonical | KD -0.4 | low | Used directly in GRAVY and KD profile |
| H (His) | yes | direct canonical | KD -3.2 | low | Used directly in GRAVY and KD profile |
| I (Ile) | yes | direct canonical | KD 4.5 | low | Hydrophobic reference |
| L (Leu) | yes | direct canonical | KD 3.8 | low | Hydrophobic reference |
| K (Lys) | yes | direct canonical | KD -3.9 | low | Used directly in GRAVY and KD profile |
| M (Met) | yes | direct canonical | KD 1.9 | low | Used directly in GRAVY and KD profile |
| F (Phe) | yes | direct canonical | KD 2.8 | low | Used directly in GRAVY and KD profile |
| P (Pro) | yes | direct canonical | KD -1.6 | low | Used directly in GRAVY and KD profile |
| S (Ser) | yes | direct canonical | KD -0.8 | low | Used directly in GRAVY and KD profile |
| T (Thr) | yes | direct canonical | KD -0.7 | low | Used directly in GRAVY and KD profile |
| W (Trp) | yes | direct canonical | KD -0.9 | low | Used directly in GRAVY and KD profile |
| Y (Tyr) | yes | direct canonical | KD -1.3 | low | Used directly in GRAVY and KD profile |
| V (Val) | yes | direct canonical | KD 4.2 | low | Hydrophobic reference |
| Ahx | partial | Nle/Leu-like linker proxy | 3.8 | high | Omega-amino-acid linker; conservative KD-compatible approximation |
| AEEA | partial | PEG-like Gln proxy | −3.5 | high | Neutral polar PEG-like linker proxy; non-standard KD-compatible approximation |
| Aib | partial | curated model proxy | KD 1.5 | high | Outside KD calibration (α,α-disubstitution) |
| AmPhe | partial | Phe-like proxy | KD 2.8 | medium | α-methyl effect not encoded |
| beta-Ala | partial | neutral proxy | KD 0.0 | high | β-backbone outside KD model |
| Cit | yes | curated value (Gln-like) | KD -3.5 | medium | Ureido vs amide difference |
| Hyp | partial | polar Pro-like proxy | KD -2.5 | medium | Hydroxylation approximated |
| Nle | yes | curated value (Leu-like) | KD 3.8 | low–medium | Close aliphatic analog |
| Nva | yes | curated value | KD 4.0 | medium | Linear vs branched side chain |
| Orn | yes | curated value (Lys-like) | KD -3.8 | medium | Shorter side chain than Lys |
Computational Note
GRAVY is computed as Σ(KDᵢ) / n. The Kyte–Doolittle profile uses the same per-residue KD values and applies smoothing for visualization without modifying the underlying scale.
Hopp–Woods Hydrophilicity Profile
Peptalyzer™ computes the Hopp–Woods hydrophilicity profile as a position-by-position residue property, not as a sequence average. Each residue contributes a single Hopp–Woods value along the sequence, enabling visualization of local hydrophilic and hydrophobic regions. Canonical amino acids use the original Hopp–Woods constants directly. Noncanonical residues are included only when a curated hydropathy_hw value is available in the residue library. In the current implementation, Nle and Nva are treated as having explicit curated values, while AmPhe, Cit, Hyp, and Orn are included through analog or proxy-based values (partial support). Aib and beta-Ala remain unsupported and block the feature if present. The feature operates in two modes: strict mode, which excludes all partial-support residues, and exploratory mode, which includes them with explicit caution. The Hopp–Woods scale is historically derived for epitope and surface exposure prediction and is typically interpreted over short sequence windows rather than as an absolute physicochemical property.
Canonical Handling
- Direct use of Hopp–Woods constants for all 20 amino acids
- One value per residue, used directly in the sequence profile
Noncanonical Handling
- Uses residue-library hydropathy_hw values when available
- Nle and Nva are treated as having curated values, with Nva carrying a higher expected deviation due to side-chain linearity (treated as supported)
- Hyp: polarity-adjusted Pro-like proxy (shifted toward a Ser-like hydrophilicity value)
- AmPhe, Cit, Orn: proxy/model values (partial support)
- Ahx, AEEA, Aib, and beta-Ala: unsupported (feature blocked)
Proxy values reflect polarity similarity but do not capture backbone or conformational effects.
Limitations
- Hopp–Woods is a window-interpreted, epitope-mapping scale derived from surface exposure tendencies, not a bulk physicochemical property and should not be interpreted as overall peptide hydrophilicity
- Proxy values represent model assumptions, not residue-specific measurements
- Unsupported residues block the feature in all modes
- Strict mode excludes partial-support residues; exploratory mode includes them with caveats
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | HW -0.5 | low | Used directly in sequence profile |
| R (Arg) | yes | direct canonical | HW 3.0 | low | Strongly hydrophilic |
| N (Asn) | yes | direct canonical | HW 0.2 | low | Slightly hydrophilic |
| D (Asp) | yes | direct canonical | HW 3.0 | low | Strongly hydrophilic |
| C (Cys) | yes | direct canonical | HW -1.0 | low | Weakly hydrophobic |
| Q (Gln) | yes | direct canonical | HW 0.2 | low | Slightly hydrophilic |
| E (Glu) | yes | direct canonical | HW 3.0 | low | Strongly hydrophilic |
| G (Gly) | yes | direct canonical | HW 0.0 | low | Neutral reference |
| H (His) | yes | direct canonical | HW -0.5 | low | Slightly hydrophobic on this scale |
| I (Ile) | yes | direct canonical | HW -1.8 | low | Hydrophobic reference |
| L (Leu) | yes | direct canonical | HW -1.8 | low | Hydrophobic reference |
| K (Lys) | yes | direct canonical | HW 3.0 | low | Strongly hydrophilic |
| M (Met) | yes | direct canonical | HW -1.3 | low | Moderately hydrophobic |
| F (Phe) | yes | direct canonical | HW -2.5 | low | Strongly hydrophobic aromatic |
| P (Pro) | yes | direct canonical | HW 0.0 | low | Neutral reference |
| S (Ser) | yes | direct canonical | HW 0.3 | low | Mildly hydrophilic |
| T (Thr) | yes | direct canonical | HW -0.4 | low | Slightly hydrophobic |
| W (Trp) | yes | direct canonical | HW -3.4 | low | Most hydrophobic aromatic |
| Y (Tyr) | yes | direct canonical | HW -2.3 | low | Hydrophobic aromatic |
| V (Val) | yes | direct canonical | HW -1.5 | low | Hydrophobic reference |
| Ahx | no | unsupported | none | high | No defensible Hopp-Woods value for this linker residue; blocks feature |
| AEEA | no | unsupported | none | high | No defensible Hopp-Woods value for this PEG-like linker residue; blocks feature |
| Aib | no | unsupported | — | high | No defensible Hopp–Woods value; blocks feature |
| AmPhe | partial | Phe-like proxy | HW -2.5 | medium | α-methyl effect not captured |
| beta-Ala | no | unsupported | — | high | β-backbone outside scale; blocks feature |
| Cit | partial | Gln-like proxy | HW 0.2 | medium | Ureido vs amide difference |
| Hyp | partial | polarity-adjusted Pro-like proxy | HW 0.3 | medium | Hydroxylated ring approximated |
| Nle | yes | curated value (Leu-like) | HW -1.8 | low–medium | Close aliphatic analog |
| Nva | yes | curated value | HW -1.5 | medium | Linear vs branched side chain |
| Orn | partial | Lys-like proxy | HW 3.0 | medium | Shorter side chain than Lys |
Computational Note
The Hopp–Woods profile assigns one value per residue position, but interpretation typically relies on short sliding windows rather than raw per-residue values. Canonical sequences use native constants; noncanonical sequences use curated hydropathy_hw values, provided the selected mode permits their inclusion.
Wimley–White Interfacial Free Energy
Peptalyzer™ computes the Wimley–White (WW) profile as a sum of whole-residue interfacial transfer free energies (POPC), not as a generic hydropathy average. Each residue contributes an experimentally derived free energy term describing transfer from water to a membrane interface. Canonical residues use the original Wimley–White constants directly. Noncanonical residues are included only when a curated hydropathy_ww value is available in the residue library. In the current implementation, Nle and Nva are treated as having curated values, while AmPhe, Cit, Hyp, and Orn are included through analog or proxy-based approximations (partial support). Aib and beta-Ala remain unsupported and block the feature if present. The model operates in strict mode (partial-support excluded) and exploratory mode (partial-support included with caution). Unlike KD or Hopp–Woods, the Wimley–White scale reflects membrane partitioning energetics, and should be interpreted as an additive free-energy model not a bulk hydrophobicity scale and should not be interpreted as solubility or aggregation tendency.
Canonical Handling
- Direct use of Wimley–White whole-residue free energies for all 20 amino acids
- Values are additive across the sequence. In practice, the profile is interpreted over sliding windows to capture local membrane-interacting segments
Noncanonical Handling
- Uses residue-library hydropathy_ww values when available
- Nle, Nva: curated values (treated as supported)
- AmPhe, Cit, Hyp, Orn: proxy/model values (partial support)
- Ahx, AEEA, Aib, and beta-Ala: unsupported (feature blocked)
Limitations
- The model describes interface partitioning, not bulk hydrophobicity
- Proxy values represent engineering approximations, not direct measurements
- Unsupported residues block the feature in all modes
- Strict mode excludes partial-support residues; exploratory mode includes them with explicit caveats
- Terminal contributions are included only when termini are free
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | WW 0.17 | low | Used directly in additive free-energy model |
| R (Arg) | yes | direct canonical | WW 0.81 | low | Highly unfavorable interface transfer |
| N (Asn) | yes | direct canonical | WW 0.42 | low | Polar residue contribution |
| D (Asp) | yes | direct canonical | WW 1.23 | low | Strongly unfavorable transfer |
| C (Cys) | yes | direct canonical | WW -0.24 | low | Slightly favorable interface partitioning |
| Q (Gln) | yes | direct canonical | WW 0.58 | low | Polar amide contribution |
| E (Glu) | yes | direct canonical | WW 2.02 | low | Highly unfavorable interface transfer |
| G (Gly) | yes | direct canonical | WW 0.01 | low | Near-neutral reference |
| H (His) | yes | direct canonical | WW 0.96 | low | Polar/ionizable contribution |
| I (Ile) | yes | direct canonical | WW -0.31 | low | Hydrophobic interface preference |
| L (Leu) | yes | direct canonical | WW -0.56 | low | Strong hydrophobic contribution |
| K (Lys) | yes | direct canonical | WW 0.99 | low | Unfavorable charged residue |
| M (Met) | yes | direct canonical | WW -0.23 | low | Moderately hydrophobic |
| F (Phe) | yes | direct canonical | WW -1.13 | low | Strong aromatic interface preference |
| P (Pro) | yes | direct canonical | WW 0.45 | low | Backbone-constrained but measured |
| S (Ser) | yes | direct canonical | WW 0.13 | low | Weakly polar contribution |
| T (Thr) | yes | direct canonical | WW 0.14 | low | Weakly polar contribution |
| W (Trp) | yes | direct canonical | WW -1.85 | low | Strongest aromatic interface preference |
| Y (Tyr) | yes | direct canonical | WW -0.94 | low | Aromatic with polar character |
| V (Val) | yes | direct canonical | WW 0.07 | low | Near-neutral small aliphatic |
| Ahx | no | unsupported | — | high | No defensible whole-residue Wimley–White value for this linker; blocks feature |
| AEEA | no | unsupported | — | high | No defensible whole-residue Wimley–White value for this PEG-like linker; blocks feature |
| Aib | no | unsupported | — | high | No defensible whole-residue WW value; blocks feature |
| AmPhe | partial | Phe-like proxy | WW -1.13 | medium | α-methyl effect not captured |
| beta-Ala | no | unsupported | — | high | β-backbone outside model; blocks feature |
| Cit | partial | Gln-like proxy | WW 0.6 | medium | Ureido vs amide difference |
| Hyp | partial | interpolated Pro/Ser proxy | WW 0.3 | medium | Hydroxylated ring approximated |
| Nle | yes | curated Met-like proxy | WW -0.23 | medium | Linear side chain approximation |
| Nva | yes | curated aliphatic proxy | WW 0.05 | medium | Linear vs branched side chain |
| Orn | partial | Lys-like proxy | WW 0.95 | medium | Shorter charged side chain |
Computational Note
The total Wimley–White score is computed as the sum of per-residue free energies, with optional terminal contributions (+1.15 for a free N-terminus, +1.2 for a free C-terminus). Profile visualizations (9- and 19-residue windows) are derived from these same values as windowed sums, without altering the underlying scale.
Solubility & Polarity Matrix
Peptalyzer™ places each peptide on a Solubility & Polarity Matrix using two coordinates: Htot (Kyte–Doolittle residue sum with optional terminal hydropathy constants when curated values are available) and fc (charge fraction computed from charged residues plus terminal charge groups, normalized by sequence length). The sequence is then classified as Polar, Intermediate, or Nonpolar using fixed thresholds. This feature does not use GRAVY or Henderson–Hasselbalch net-charge calculations for positioning; pI is displayed for context only and does not influence matrix coordinates.
Canonical residues contribute to Htot via Kyte–Doolittle values, and fc includes charged residue classes together with terminal charge-group contributions. For noncanonical residues, Htot uses curated hydropathy_kd metadata, and fc uses residue-library ionization_role/charge-role mappings; in the current set, Orn contributes as a basic analog while most other supported noncanonical residues are neutral on the fc axis. When active terminal hydropathy constants are missing, Peptalyzer™ uses a mixed fallback: fc remains termini-aware, while Htot stays residue-based. The matrix is computed only when required hydropathy and charge metadata are available; otherwise the feature is blocked.
Canonical Handling
- Htot is the Kyte–Doolittle residue sum across the sequence, plus curated terminal hydropathy constants when active termini are supported
- fc is a composition-based charge fraction that includes charged residues and terminal charge groups, normalized by sequence length
Noncanonical Handling
- Htot uses residue-library hydropathy_kd values for noncanonical residues, with optional curated terminal hydropathy constants applied when available
- fc uses residue-library ionization_role/charge_role mappings for noncanonical residues and includes terminal charge groups from active termini
- In the current noncanonical set, Orn is the only residue contributing as a basic noncanonical side chain on the fc axis
- Other currently supported noncanonical residues are treated as neutral on the fc side-chain axis
Limitations
- This is a composition-based classification for matrix placement, not a pH-dependent Henderson–Hasselbalch electrostatic model
- There is no strict/exploratory display mode for matrix placement itself; support/caveat messaging is metadata-driven
- Noncanonical sequences can be flagged as partial_support when hydropathy/charge mappings rely on analog or proxy assumptions rather than direct residue-specific measurements
- Htot is an absolute sum (not length-normalized), so values scale with peptide size
- The UI uses caveat-only warning markers and support messaging to indicate approximation/coverage limits
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | KD 1.8; fc neutral | low | Contributes to H only |
| R (Arg) | yes | direct canonical | KD -4.5; fc basic | low | Contributes to H and fc |
| N (Asn) | yes | direct canonical | KD -3.5; fc neutral | low | Contributes to H only |
| D (Asp) | yes | direct canonical | KD -3.5; fc acidic | low | Contributes to H and fc |
| C (Cys) | yes | direct canonical (excluded from fc) | KD 2.5; fc not counted | low | Counts only in H |
| Q (Gln) | yes | direct canonical | KD -3.5; fc neutral | low | Contributes to H only |
| E (Glu) | yes | direct canonical | KD -3.5; fc acidic | low | Contributes to H and fc |
| G (Gly) | yes | direct canonical | KD -0.4; fc neutral | low | Contributes to H only |
| H (His) | yes | direct canonical | KD -3.2; fc basic | low | Counted in fc |
| I (Ile) | yes | direct canonical | KD 4.5; fc neutral | low | Strong positive H contribution |
| L (Leu) | yes | direct canonical | KD 3.8; fc neutral | low | Strong positive H contribution |
| K (Lys) | yes | direct canonical | KD -3.9; fc basic | low | Contributes to H and fc |
| M (Met) | yes | direct canonical | KD 1.9; fc neutral | low | Contributes to H only |
| F (Phe) | yes | direct canonical | KD 2.8; fc neutral | low | Contributes to H only |
| P (Pro) | yes | direct canonical | KD -1.6; fc neutral | low | Contributes to H only |
| S (Ser) | yes | direct canonical | KD -0.8; fc neutral | low | Contributes to H only |
| T (Thr) | yes | direct canonical | KD -0.7; fc neutral | low | Contributes to H only |
| W (Trp) | yes | direct canonical | KD -0.9; fc neutral | low | Contributes to H only |
| Y (Tyr) | yes | direct canonical (excluded from fc) | KD -1.3; fc not counted | low | Counts only in H |
| V (Val) | yes | direct canonical | KD 4.2; fc neutral | low | Strong positive H contribution |
| Ahx | partial | Nle/Leu-like KD linker proxy | KD 3.8; neutral | high | ω-amino-acid linker; contributes to H only (no fc side-chain charge) |
| AEEA | partial | Gln-like KD PEG-linker proxy | KD -3.5; neutral | high | PEG-like linker; contributes to H only (no fc side-chain charge) |
| Aib | partial | Ala-like KD proxy | KD 1.5; neutral | high | Backbone constraint not captured |
| AmPhe | partial | Phe-like KD proxy | KD 2.8; neutral | medium | α-methyl effect not captured |
| beta-Ala | partial | neutral KD proxy | KD 0.0; neutral | high | β-backbone outside calibration |
| Cit | partial | Gln-like KD proxy | KD -3.5; neutral | medium | Ureido vs amide difference |
| Hyp | partial | Pro-like KD proxy | KD -2.5; neutral | medium | Hydroxylation approximated |
| Nle | partial | Leu-like KD proxy | KD 3.8; neutral | low–medium | Close aliphatic analog |
| Nva | partial | Val/Leu-like KD proxy | KD 4.0; neutral | medium | Linear vs branched side chain |
| Orn | partial | Lys-like KD + charge mapping | KD -3.8; basic | medium | Only residue affecting both axes |
Computational Note
Classification is based on fixed thresholds:
- Polar: H < 0 and fc ≥ 0.20
- Nonpolar: H > 20 and fc ≤ 0.05
- Intermediate: otherwise
pI is displayed in hover text for context, but does not affect the plotted coordinates.
Chou–Fasman Secondary Structure
Peptalyzer™ estimates secondary structure using the Chou–Fasman model, assigning each peptide a composition-based profile of helix, sheet, and coil propensities derived from residue-level parameters (Pα, Pβ, and coil/turn). These values are summed across the sequence to produce overall structure percentages and are also used to generate a beta-sheet hotspot trace based on 6-residue sliding windows. This feature relies entirely on canonical residue propensities and does not incorporate structural prediction algorithms or machine-learning models.
Canonical residues use the built-in Chou–Fasman parameters directly. For noncanonical residues, the model applies canonical analog mapping through the residue library (secondary_structure_cf). Residues with a defensible alpha-amino-acid analog are included, while those without a valid mapping are excluded. In the current implementation, noncanonical residues are either treated as supported through close analogs, included as approximations under partial_support, or blocked when no chemically meaningful mapping exists. The feature is computed only when all residues can be resolved under the selected mode (strict or exploratory).
Canonical Handling
- All residues use direct Chou–Fasman triplet values (Pα, Pβ, coil/turn)
- Sequence-level helix, sheet, and coil content is computed by summation
- Beta-sheet hotspots are derived from 6-residue sliding windows
Noncanonical Handling
- Residues are mapped to canonical analogs via secondary_structure_cf.residue
- Supported residues use close alpha-amino-acid analogs and are included in all modes
- Partial-support residues are excluded in strict mode and included only in exploratory mode
- Residues without a defensible analog block the feature
Limitations
- The model is based on canonical statistical propensities, not structural prediction
- No residue-specific Chou–Fasman parameters exist for noncanonical amino acids
- All noncanonical handling relies on analog mapping rather than experimental calibration
- Backbone-modified residues and strongly conformationally biased residues fall outside the model
- Beta-sheet hotspot thresholds are empirical and tuned for synthesis-oriented interpretation
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical residue | Pα 1.42; Pβ 0.83; coil 0.66 | low | Direct Chou-Fasman triplet |
| R (Arg) | yes | direct canonical residue | Pα 0.98; Pβ 0.93; coil 0.95 | low | Direct Chou-Fasman triplet |
| N (Asn) | yes | direct canonical residue | Pα 0.67; Pβ 0.89; coil 1.56 | low | Direct Chou-Fasman triplet |
| D (Asp) | yes | direct canonical residue | Pα 1.01; Pβ 0.54; coil 1.46 | low | Direct Chou-Fasman triplet |
| C (Cys) | yes | direct canonical residue | Pα 0.70; Pβ 1.19; coil 1.19 | low | Direct Chou-Fasman triplet |
| Q (Gln) | yes | direct canonical residue | Pα 1.11; Pβ 1.10; coil 0.98 | low | Direct Chou-Fasman triplet |
| E (Glu) | yes | direct canonical residue | Pα 1.51; Pβ 0.37; coil 0.74 | low | Direct Chou-Fasman triplet |
| G (Gly) | yes | direct canonical residue | Pα 0.57; Pβ 0.75; coil 1.56 | low | Direct Chou-Fasman triplet |
| H (His) | yes | direct canonical residue | Pα 1.00; Pβ 0.87; coil 0.95 | low | Direct Chou-Fasman triplet |
| I (Ile) | yes | direct canonical residue | Pα 1.08; Pβ 1.60; coil 0.47 | low | Direct Chou-Fasman triplet |
| L (Leu) | yes | direct canonical residue | Pα 1.21; Pβ 1.30; coil 0.59 | low | Direct Chou-Fasman triplet |
| K (Lys) | yes | direct canonical residue | Pα 1.16; Pβ 0.74; coil 1.01 | low | Direct Chou-Fasman triplet |
| M (Met) | yes | direct canonical residue | Pα 1.45; Pβ 1.05; coil 0.60 | low | Direct Chou-Fasman triplet |
| F (Phe) | yes | direct canonical residue | Pα 1.13; Pβ 1.38; coil 0.60 | low | Direct Chou-Fasman triplet |
| P (Pro) | yes | direct canonical residue | Pα 0.59; Pβ 0.55; coil 1.52 | low | Direct Chou-Fasman triplet |
| S (Ser) | yes | direct canonical residue | Pα 0.77; Pβ 0.75; coil 1.43 | low | Direct Chou-Fasman triplet |
| T (Thr) | yes | direct canonical residue | Pα 0.83; Pβ 1.19; coil 0.96 | low | Direct Chou-Fasman triplet |
| W (Trp) | yes | direct canonical residue | Pα 1.08; Pβ 1.37; coil 0.96 | low | Direct Chou-Fasman triplet |
| Y (Tyr) | yes | direct canonical residue | Pα 0.69; Pβ 1.47; coil 1.14 | low | Direct Chou-Fasman triplet |
| V (Val) | yes | direct canonical residue | Pα 1.06; Pβ 1.70; coil 0.50 | low | Direct Chou-Fasman triplet |
| Ahx | no | no defensible analog; blocked | no reliable data | high | Unsupported in both modes; ! shows blocking reason (omega-amino-acid linker outside alpha-backbone Chou–Fasman domain) |
| AEEA | no | no defensible analog; blocked | no reliable data | high | Unsupported in both modes; ! shows blocking reason (PEG-like linker outside alpha-backbone Chou–Fasman domain) |
| Aib | no | no defensible analog; blocked | no reliable data | high | Unsupported in both modes; ! shows blocking reason |
| AmPhe | partial | mapped to Phe | F → Pα 1.13; Pβ 1.38; coil 0.60 | high | Blocked in strict mode; exploratory only; ! shows approximation basis |
| beta-Ala | no | no defensible analog; blocked | no reliable data | high | Unsupported in both modes; ! shows blocking reason |
| Cit | partial | mapped to Gln | Q → Pα 1.11; Pβ 1.10; coil 0.98 | medium | Blocked in strict mode; exploratory only |
| Hyp | partial | mapped to Pro | P → Pα 0.59; Pβ 0.55; coil 1.52 | medium-high | Blocked in strict mode; exploratory only |
| Nle | yes | mapped to Leu | L → Pα 1.21; Pβ 1.30; coil 0.59 | low-medium | Supported in both modes; ℹ can list support basis |
| Nva | yes | mapped to Val | V → Pα 1.06; Pβ 1.70; coil 0.50 | medium | Supported in both modes; ℹ can list support basis |
| Orn | partial | mapped to Lys | K → Pα 1.16; Pβ 0.74; coil 1.01 | medium | Blocked in strict mode; exploratory only |
Computational Note
Secondary structure content is computed by summing residue-level Chou–Fasman propensities across the sequence. Beta-sheet hotspot detection is based on 6-residue sliding windows; risk bands are assigned from the rounded peak sheet value (max_sheet_val), with thresholds:
- Low: <30% sheet
- Moderate: 30–40% sheet
- High: >40% sheet
Topology Maps (Amphipathic Helical Map & β-Strand Zig-Zag Map)
Peptalyzer™ generates two topology projections: the Amphipathic Helical Map and the β-Strand Zig-Zag Map. These visualizations rely on canonical residue properties and are enabled through support gating combined with canonical analog mapping. The β-strand map requires Chou–Fasman support, while the helical map requires both Chou–Fasman and Eisenberg support for every residue. These features do not perform structural prediction; they project residue properties onto idealized geometries. Both maps are available only for sequences between 7 and 40 residues.
Canonical residues use direct Chou–Fasman and Eisenberg values. For noncanonical residues, canonical analog mapping is applied when defined in the residue library. Residues with accepted alpha-amino-acid analogs are included, while residues without a defensible mapping block the feature. In the current implementation, some noncanonical residues are treated as supported through close analogs, others are included only under partial_support in exploratory mode, and some are excluded entirely. The UI reflects this with ℹ for supported mappings and ! for exploratory use or blocking. The Support Summary follows the same logic.
Canonical Handling
- All residues use direct canonical Chou–Fasman propensities
- The Amphipathic Helical Map additionally uses canonical Eisenberg values
- No analog mapping is applied
Noncanonical Handling
- Residues are mapped to canonical analogs for both Chou–Fasman and Eisenberg when available
- Supported residues are included in all modes
- Partial-support residues are excluded in strict mode and included only in exploratory mode
- Residues without a defensible analog block the feature
Limitations
- These are topology projections, not structural predictions or experimental measurements
- The Amphipathic Helical Map requires both Chou–Fasman and Eisenberg support per residue
- The β-Strand Zig-Zag Map requires Chou–Fasman support only
- Noncanonical handling relies entirely on analog mapping
- Backbone-modified residues fall outside the model and are excluded
- Length constraints apply: <7 and >40 residues are not supported
- ℹ indicates supported noncanonical mapping; ! indicates approximation or blocking
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| R (Arg) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| N (Asn) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| D (Asp) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| C (Cys) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| Q (Gln) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| E (Glu) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| G (Gly) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| H (His) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| I (Ile) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| L (Leu) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| K (Lys) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| M (Met) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| F (Phe) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| P (Pro) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| S (Ser) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| T (Thr) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| W (Trp) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| Y (Tyr) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| V (Val) | yes | direct canonical residue | canonical CF + Eisenberg | low | Direct in both topology views |
| Ahx | no | no defensible analog | none | high | Blocked; ! shown (unsupported for both Chou–Fasman and Eisenberg in topology maps) |
| AEEA | no | no defensible analog | none | high | Blocked; ! shown (unsupported for both Chou–Fasman and Eisenberg in topology maps) |
| Aib | no | no defensible analog | none | high | Blocked; ! shown |
| AmPhe | partial | Phe analog | F analog (CF + Eisenberg) | high | Exploratory only; strict blocks |
| beta-Ala | no | no defensible analog | none | high | Blocked; ! shown |
| Cit | partial | Gln analog | Q analog (CF + Eisenberg) | medium | Exploratory only; strict blocks |
| Hyp | partial | Pro analog | P analog (CF + Eisenberg) | medium-high | Exploratory only; strict blocks |
| Nle | yes | Leu analog | L analog (CF + Eisenberg) | low-medium | Supported; ℹ may appear |
| Nva | yes | Val analog | V analog (CF + Eisenberg) | medium | Supported; ℹ may appear |
| Orn | partial | Lys analog | K analog (CF + Eisenberg) | medium | Exploratory only; strict blocks |
Computational Note
The β-Strand Zig-Zag Map is gated by Chou–Fasman support only. The Amphipathic Helical Map requires both Chou–Fasman and Eisenberg support. Both maps are displayed only for sequences of 7–40 residues.
Eisenberg Hydrophobic Moment
Peptalyzer™ computes an alpha-helical hydrophobic moment (μH) using the Eisenberg normalized consensus hydrophobicity scale. The calculation is performed over 11-residue sliding windows assuming a 100° rotation per residue, consistent with ideal alpha-helical geometry. This feature measures amphipathicity, not average hydropathy. The reported profile reflects local hydrophobic moment values, and the peak value corresponds to the maximum μH across the sequence.
Canonical residues use direct Eisenberg constants. For noncanonical residues, the model applies canonical analog mapping through the residue library. Residues with accepted alpha-amino-acid analogs are included, while residues without a defensible mapping block the feature. In the current implementation, some noncanonical residues are treated as supported through close analogs, others are included only under partial_support in exploratory mode, and some are excluded entirely. The UI reflects this with ℹ for supported mappings and ! for exploratory use or blocking. The Support Summary follows the same logic.
Canonical Handling
- All residues use direct Eisenberg normalized consensus hydrophobicity values
- These values are used to compute the hydrophobic moment within each window
- The reported peak corresponds to the maximum μH across the sequence
Noncanonical Handling
- Residues are mapped to canonical analogs via feature_analogs.hydropathy_eisenberg.residue
- Supported residues are included in all modes
- Partial-support residues are excluded in strict mode and included only in exploratory mode
- Residues without a defensible analog block the feature
Limitations
- The model assumes alpha-helical geometry and does not apply to other conformations
- It measures amphipathicity, not average hydrophobicity or solubility
- No residue-specific Eisenberg constants exist for noncanonical amino acids
- All noncanonical handling relies on analog mapping
- Backbone-modified residues fall outside the model and are excluded
- ℹ indicates supported noncanonical mapping; ! indicates approximation or blocking
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical residue | Eisenberg 0.62 | low | Direct value in μH calculation |
| R (Arg) | yes | direct canonical residue | Eisenberg -2.53 | low | Direct value in μH calculation |
| N (Asn) | yes | direct canonical residue | Eisenberg -0.78 | low | Direct value in μH calculation |
| D (Asp) | yes | direct canonical residue | Eisenberg -0.90 | low | Direct value in μH calculation |
| C (Cys) | yes | direct canonical residue | Eisenberg 0.29 | low | Direct value in μH calculation |
| Q (Gln) | yes | direct canonical residue | Eisenberg -0.85 | low | Direct value in μH calculation |
| E (Glu) | yes | direct canonical residue | Eisenberg -0.74 | low | Direct value in μH calculation |
| G (Gly) | yes | direct canonical residue | Eisenberg 0.48 | low | Direct value in μH calculation |
| H (His) | yes | direct canonical residue | Eisenberg -0.40 | low | Direct value in μH calculation |
| I (Ile) | yes | direct canonical residue | Eisenberg 1.38 | low | Strong hydrophobic contributor |
| L (Leu) | yes | direct canonical residue | Eisenberg 1.06 | low | Strong hydrophobic contributor |
| K (Lys) | yes | direct canonical residue | Eisenberg -1.50 | low | Strong polar/basic contributor |
| M (Met) | yes | direct canonical residue | Eisenberg 0.64 | low | Direct value in μH calculation |
| F (Phe) | yes | direct canonical residue | Eisenberg 1.19 | low | Hydrophobic aromatic contributor |
| P (Pro) | yes | direct canonical residue | Eisenberg 0.12 | low | Direct value in μH calculation |
| S (Ser) | yes | direct canonical residue | Eisenberg -0.18 | low | Direct value in μH calculation |
| T (Thr) | yes | direct canonical residue | Eisenberg -0.05 | low | Direct value in μH calculation |
| W (Trp) | yes | direct canonical residue | Eisenberg 0.81 | low | Hydrophobic aromatic contributor |
| Y (Tyr) | yes | direct canonical residue | Eisenberg 0.26 | low | Direct value in μH calculation |
| V (Val) | yes | direct canonical residue | Eisenberg 1.08 | low | Strong hydrophobic contributor |
| Ahx | no | no defensible analog | none | high | Blocked; ! shown (unsupported for Eisenberg hydrophobic-moment mapping) |
| AEEA | no | no defensible analog | none | high | Blocked; ! shown (unsupported for Eisenberg hydrophobic-moment mapping) |
| Aib | no | no defensible analog | none | high | Blocked; ! shown |
| AmPhe | partial | Phe analog | F analog → 1.19 | high | Exploratory only; strict blocks |
| beta-Ala | no | no defensible analog | none | high | Blocked; ! shown |
| Cit | partial | Gln analog | Q analog → -0.85 | medium | Exploratory only; strict blocks |
| Hyp | partial | Pro analog | P analog → 0.12 | medium-high | Exploratory only; strict blocks |
| Nle | yes | Leu analog | L analog → 1.06 | low-medium | Supported; ℹ may appear |
| Nva | yes | Val analog | V analog → 1.08 | medium | Supported; ℹ may appear |
| Orn | partial | Lys analog | K analog → -1.50 | medium | Exploratory only; strict blocks |
Computational Note
For each sequence position, Peptalyzer computes μH from the vector sum of Eisenberg values within an 11-residue sliding window using 100° angular increments, normalized by window length. Edge windows are clipped rather than padded, and the reported peak hydrophobic moment corresponds to the maximum μH across the sequence.
Aromaticity Index
Peptalyzer™ computes the Aromaticity Index directly from residue-library aromatic flags, not from hardcoded canonical residue lists. The base value is defined as aromatic_count / n, where residues with structural_flags.is_aromatic = true are counted. Canonical F, W, and Y are therefore included, and within the current noncanonical set, AmPhe is also included directly. All other supported noncanonical residues are treated as non-aromatic.
A companion output, Including Histidine (Neutral pH), is provided as a reporting variant. This value adds canonical histidine counts to the numerator, giving (aromatic_count + n_H) / n. This is not a speciation or titration model, but a fixed reporting convention. No analog mapping is used for noncanonical residues. In the Support Summary, this feature is fully supported, as it relies on explicit residue metadata rather than approximations.
Canonical Handling
- Base aromaticity counts residues flagged aromatic in canonical metadata (F, W, Y)
- H is not included in the base index
- A separate output includes histidine by adding canonical H counts directly
Noncanonical Handling
- Aromaticity is determined directly from structural_flags.is_aromatic
- No canonical analog mapping is applied
- In the current set, only AmPhe is treated as aromatic
- All other noncanonical residues are treated as non-aromatic
Limitations
- The histidine-inclusive value is a reporting convention, not a Henderson–Hasselbalch or tautomer model
- Only canonical histidine is included; no “histidine-like” inference is applied to noncanonical residues
- Aromaticity depends entirely on residue-library metadata; incorrect flags propagate directly
- No dedicated warning state is expected; the feature remains fully supported
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | is_aromatic = false | low | Not counted |
| R (Arg) | yes | direct canonical | is_aromatic = false | low | Not counted |
| N (Asn) | yes | direct canonical | is_aromatic = false | low | Not counted |
| D (Asp) | yes | direct canonical | is_aromatic = false | low | Not counted |
| C (Cys) | yes | direct canonical | is_aromatic = false | low | Not counted |
| Q (Gln) | yes | direct canonical | is_aromatic = false | low | Not counted |
| E (Glu) | yes | direct canonical | is_aromatic = false | low | Not counted |
| G (Gly) | yes | direct canonical | is_aromatic = false | low | Not counted |
| H (His) | yes | direct canonical | base: 0; variant: +1 | medium | Only counted in histidine-inclusive variant |
| I (Ile) | yes | direct canonical | is_aromatic = false | low | Not counted |
| L (Leu) | yes | direct canonical | is_aromatic = false | low | Not counted |
| K (Lys) | yes | direct canonical | is_aromatic = false | low | Not counted |
| M (Met) | yes | direct canonical | is_aromatic = false | low | Not counted |
| F (Phe) | yes | direct canonical | is_aromatic = true | low | Counted |
| P (Pro) | yes | direct canonical | is_aromatic = false | low | Not counted |
| S (Ser) | yes | direct canonical | is_aromatic = false | low | Not counted |
| T (Thr) | yes | direct canonical | is_aromatic = false | low | Not counted |
| W (Trp) | yes | direct canonical | is_aromatic = true | low | Counted |
| Y (Tyr) | yes | direct canonical | is_aromatic = true | low | Counted |
| V (Val) | yes | direct canonical | is_aromatic = false | low | Not counted |
| Ahx | yes | direct library | false | low | Non-aromatic linker (is_aromatic = false) |
| AEEA | yes | direct library | false | low | Non-aromatic PEG-like linker (is_aromatic = false) |
| Aib | yes | direct library | false | low | Non-aromatic |
| AmPhe | yes | direct library | true | low | Counted |
| beta-Ala | yes | direct library | false | low | Non-aromatic |
| Cit | yes | direct library | false | low | Non-aromatic |
| Hyp | yes | direct library | false | low | Non-aromatic |
| Nle | yes | direct library | false | low | Non-aromatic |
| Nva | yes | direct library | false | low | Non-aromatic |
| Orn | yes | direct library | false | low | Non-aromatic |
Computational Note
- Aromaticity Index: aromatic_count / n using residue-library aromatic flags
- Including Histidine (Neutral pH): (aromatic_count + n_H) / n
- Only residues with one_letter = H are included in the histidine-inclusive variant; no “histidine-like” inference is applied to noncanonical residues
Extinction Coefficients (ε280 and ε205)
Peptalyzer™ computes peptide extinction coefficients at 280 nm (ε280) and 205 nm (ε205) using two distinct models.
ε280 is a side-chain and disulfide-driven sum. Only Trp and Tyr contribute directly (W = 5500, Y = 1490), and each explicit user-defined disulfide bond adds 125. All other canonical residues contribute 0. Noncanonical handling uses direct residue-library values (epsilon280_sidechain) with no analog mapping. In the current residue set, all supported noncanonical amino acids define 0.0, so ε280 remains fully supported.
ε205 is backbone-dominant, computed as 2780 × (n−1) plus side-chain contributions. Canonical side-chain contributors are W, F, Y, H, M, and R. Noncanonical residues again use direct residue-library values (epsilon205_sidechain). In the current implementation, AmPhe and Cit contribute positive modeled terms, while all other noncanonical residues use 0.0. In the current noncanonical residue set, ε205 is computable and treated as partial_support when noncanonical residues are present.
The UI reflects this behavior through support status and sequence-specific notes: noncanonical ε205 handling is shown as partial_support with explanatory text about positive versus negligible side-chain contributions.
Canonical Handling
- ε280: 5500·W + 1490·Y + 125·(disulfide bonds)
- ε205: 2780·(number of peptide bonds) + Σ(side-chain ε205)
- Canonical side-chain ε205 contributors: W, F, Y, H, M, R
- All other canonical residues contribute 0 at both wavelengths
Noncanonical Handling
- Both ε280 and ε205 use residue-library fields directly (epsilon280_sidechain, epsilon205_sidechain)
- No canonical analog mapping is applied
- All noncanonical residues are fully supported at ε280
- ε205 remains partial_support due to modeled or assumed side-chain contributions
Limitations
- ε205 noncanonical side-chain values are model assumptions unless explicitly literature-grounded
- beta-Ala uses the standard peptide-bond backbone constant; no β-backbone correction is applied ε280 ignores potential aromatic absorbance in noncanonical residues if not explicitly defined
- Support is communicated via status plus sequence-specific notes that distinguish negligible versus positive modeled noncanonical side-chain contributions
- Support Summary flags ε205 as partial_support whenever noncanonical explanatory handling is active
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| R (Arg) | yes | direct canonical | ε280 0; ε205 1350 | low | ε205 side-chain contributor |
| N (Asn) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| D (Asp) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| C (Cys) | yes | direct canonical | ε280 0 (+125/disulfide); ε205 0 | low | Disulfide handled explicitly |
| Q (Gln) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| E (Glu) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| G (Gly) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| H (His) | yes | direct canonical | ε280 0; ε205 5200 | low | ε205 side-chain contributor |
| I (Ile) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| L (Leu) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| K (Lys) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| M (Met) | yes | direct canonical | ε280 0; ε205 1830 | low | ε205 side-chain contributor |
| F (Phe) | yes | direct canonical | ε280 0; ε205 8600 | low | Aromatic ε205 contributor |
| P (Pro) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| S (Ser) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| T (Thr) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| W (Trp) | yes | direct canonical | ε280 5500; ε205 20400 | low | Dominant contributor |
| Y (Tyr) | yes | direct canonical | ε280 1490; ε205 6080 | low | Contributor at both wavelengths |
| V (Val) | yes | direct canonical | ε280 0; ε205 0 | low | Backbone-only at ε205 |
| Ahx | partial | direct library | 0.0 / 0.0 | medium-high | Linker residue; side-chain ε terms treated as negligible in current model |
| AEEA | partial | direct library | 0.0 / 0.0 | medium-high | PEG-like linker; side-chain ε terms treated as negligible in current model |
| Aib | partial | direct library | 0.0 / 0.0 | medium | Negligible side-chain assumption |
| AmPhe | partial | direct library | 0.0 / 8600.0 | medium | Aromatic ε205 proxy; ! |
| beta-Ala | partial | direct library | 0.0 / 0.0 | medium-high | No β-backbone correction |
| Cit | partial | direct library | 0.0 / 400.0 | medium | Amide proxy; ! |
| Hyp | partial | direct library | 0.0 / 0.0 | medium | Negligible assumption |
| Nle | partial | direct library | 0.0 / 0.0 | low-medium | Aliphatic negligible |
| Nva | partial | direct library | 0.0 / 0.0 | low-medium | Aliphatic negligible |
| Orn | partial | direct library | 0.0 / 0.0 | medium | Amine not parameterized |
Computational Note
- ε280: sum of side-chain contributions plus 125 × (explicit disulfide bonds)
- ε205: 2780 × (n−1) using all residues (canonical and noncanonical), plus side-chain terms
- Residues with 0 side-chain values still contribute through the backbone term
Aliphatic Index
Peptalyzer™ computes the Aliphatic Index using the original Ikai formulation: 100 × (X_A + 2.9·X_V + 3.9·(X_I + X_L)), where each X is the mole fraction in the full sequence.
Canonical A, V, I, and L contribute weighted terms, while all other residues are included in sequence length but contribute 0. Noncanonical residues are handled through residue metadata, not inferred mapping. Nle→Leu and Nva→Val are treated as supported analogs. Aib→Ala is included as a partial-support approximation. AmPhe, beta-Ala, Cit, Hyp, and Orn are assigned explicit zero contribution and only affect normalization through sequence length.
In the Support Summary, status is shown as Supported, Partial, or Unsupported, with sequence-specific reason text and support-basis details when available.
Canonical Handling
- A, V, I, L contribute with Ikai weights 1.0, 2.9, 3.9, 3.9
- All other canonical residues contribute 0
- All residues are included in sequence length
Noncanonical Handling
- Nle → L and Nva → V are treated as supported analogs
- Aib → A is included as a partial-support approximation
- AmPhe, beta-Ala, Cit, Hyp, Orn use explicit zero-contribution metadata
- Residues without a defined aliphatic_index analog are unsupported and block the metric (currently Ahx and AEEA)
- No canonical analog mapping is inferred beyond residue metadata
Limitations
- The model remains strictly AVIL-based composition
- Zero contribution does not mean absence; residues still affect normalization
- beta-Ala is outside the alpha-amino-acid framework and is highly approximate
- Noncanonical contributions are either analog-based, explicitly zero, or unsupported when no analog is defined
- Output status is shown as Supported / Partial / Unsupported with sequence-specific reason text
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | weight 1.0 | low | Contributes to numerator |
| R (Arg) | yes | direct canonical | 0 | low | Denominator only |
| N (Asn) | yes | direct canonical | 0 | low | Denominator only |
| D (Asp) | yes | direct canonical | 0 | low | Denominator only |
| C (Cys) | yes | direct canonical | 0 | low | Denominator only |
| Q (Gln) | yes | direct canonical | 0 | low | Denominator only |
| E (Glu) | yes | direct canonical | 0 | low | Denominator only |
| G (Gly) | yes | direct canonical | 0 | low | Denominator only |
| H (His) | yes | direct canonical | 0 | low | Denominator only |
| I (Ile) | yes | direct canonical | weight 3.9 | low | Contributes to numerator |
| L (Leu) | yes | direct canonical | weight 3.9 | low | Contributes to numerator |
| K (Lys) | yes | direct canonical | 0 | low | Denominator only |
| M (Met) | yes | direct canonical | 0 | low | Denominator only |
| F (Phe) | yes | direct canonical | 0 | low | Denominator only |
| P (Pro) | yes | direct canonical | 0 | low | Denominator only |
| S (Ser) | yes | direct canonical | 0 | low | Denominator only |
| T (Thr) | yes | direct canonical | 0 | low | Denominator only |
| W (Trp) | yes | direct canonical | 0 | low | Denominator only |
| Y (Tyr) | yes | direct canonical | 0 | low | Denominator only |
| V (Val) | yes | direct canonical | weight 2.9 | low | Contributes to numerator |
| Ahx | no | unsupported (no aliphatic_index analog) | — | high | Omega-amino-acid linker; Aliphatic Index is unavailable for sequences containing Ahx in current model |
| AEEA | no | unsupported (no aliphatic_index analog) | — | high | PEG-like linker; Aliphatic Index is unavailable for sequences containing AEEA in current model |
| Aib | partial | Ala analog | weight 1.0 | medium | Forced Ala-like contribution |
| AmPhe | partial | zero-contribution | 0 | low-medium | Denominator only |
| beta-Ala | partial | zero-contribution | 0 | high | Outside alpha-AA model |
| Cit | partial | zero-contribution | 0 | low-medium | Denominator only |
| Hyp | partial | zero-contribution | 0 | medium | Denominator only |
| Nle | yes | Leu analog | weight 3.9 | low-medium | Supported analog |
| Nva | yes | Val analog | weight 2.9 | medium | Supported analog |
| Orn | partial | zero-contribution | 0 | medium | Denominator only |
Computational Note
- Numerator includes only weighted A/V/I/L contributions
- Denominator is total sequence length, including all residues
- Zero-contribution residues still affect the final value through normalization
Boman Index
Peptalyzer™ computes the Boman Index from residue-level binding free-energy constants using the standard formulation:
Boman = -Σ(values) / n
Canonical residues use direct Radzicka–Wolfenden constants without modification. Noncanonical residues are included only through explicit canonical analog mapping defined in the residue library. In the current implementation, Nle→Leu and Nva→Val are treated as supported, while AmPhe→Phe, Cit→Gln, Hyp→Pro, and Orn→Lys are handled as partial_support. Aib and beta-Ala are unsupported and block the feature.
In strict mode, any partial_support or unsupported residue disables the Boman calculation. In exploratory mode, partial_support residues are included using their mapped analog values. The UI reflects this behavior: ℹ indicates supported analog-based handling, while ! marks exploratory approximations or blocking cases. The Support Summary follows the same logic.
Canonical Handling
- Direct Radzicka–Wolfenden residue constants are used
- The index is computed as -Σ(values)/n
- No approximation or transformation is applied
Noncanonical Handling
- Inclusion requires feature_analogs.boman.residue mapping
- Nle → L and Nva → V are supported
- AmPhe → F, Cit → Q, Hyp → P, Orn → K are partial_support
- Aib and beta-Ala are unsupported and block the feature
- No residue-specific noncanonical constants are implemented
Limitations
- The model relies entirely on canonical residue constants
- No experimental Boman parameters exist for noncanonical residues partial_support residues are excluded in strict mode
- Unsupported residues block the calculation entirely
- ℹ indicates supported handling; ! indicates approximation or blocking
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | 1.81 | low | Used directly |
| R (Arg) | yes | direct canonical | -14.92 | low | Strong polar contributor |
| N (Asn) | yes | direct canonical | -6.64 | low | Direct value |
| D (Asp) | yes | direct canonical | -8.72 | low | Direct value |
| C (Cys) | yes | direct canonical | 1.28 | low | Direct value |
| Q (Gln) | yes | direct canonical | -5.54 | low | Direct value |
| E (Glu) | yes | direct canonical | -6.81 | low | Direct value |
| G (Gly) | yes | direct canonical | 0.94 | low | Direct value |
| H (His) | yes | direct canonical | -4.66 | low | Direct value |
| I (Ile) | yes | direct canonical | 4.92 | low | Hydrophobic contributor |
| L (Leu) | yes | direct canonical | 4.92 | low | Hydrophobic contributor |
| K (Lys) | yes | direct canonical | -5.55 | low | Direct value |
| M (Met) | yes | direct canonical | 2.35 | low | Direct value |
| F (Phe) | yes | direct canonical | 2.98 | low | Aromatic contributor |
| P (Pro) | yes | direct canonical | -0.94 | low | Direct value |
| S (Ser) | yes | direct canonical | -3.40 | low | Direct value |
| T (Thr) | yes | direct canonical | -2.57 | low | Direct value |
| W (Trp) | yes | direct canonical | 2.33 | low | Aromatic contributor |
| Y (Tyr) | yes | direct canonical | -0.14 | low | Direct value |
| V (Val) | yes | direct canonical | 4.04 | low | Hydrophobic contributor |
| Aib | no | unsupported | none | high | Blocks feature |
| AmPhe | partial | Phe analog | 2.98 | high | Exploratory only |
| beta-Ala | no | unsupported | none | high | Blocks feature |
| Cit | partial | Gln analog | -5.54 | medium | Exploratory only |
| Hyp | partial | Pro analog | -0.94 | medium-high | Exploratory only |
| Nle | yes | Leu analog | 4.92 | low-medium | Supported |
| Nva | yes | Val analog | 4.04 | medium | Supported |
| Orn | partial | Lys analog | -5.55 | medium | Exploratory only |
Computational Note
- Residue constants are summed across the sequence
- The mean is computed and multiplied by -1
- Supported noncanonical residues use canonical analog values
- Unsupported residues block the calculation entirely
Peptide Size Metrics
Peptalyzer™ computes peptide size outputs using two complementary model families: global size descriptors derived from molecular mass and chain length, and composition descriptors derived from residue-volume tables. This section covers the Estimated Molecular Volume (Peptide Sequence Identity card), Average Residue Volume (Molecular Characteristics card), and the full Molecular Size card — Equivalent Sphere Radius, Flexible-Chain Radius, and Expansion Ratio.
Canonical Handling
- Global size metrics are computed directly from sequence-derived mass and length
- Equivalent Sphere Radius and Flexible-Chain Radius are available for all valid canonical sequences
- Expansion Ratio classes are reported as: small model gap below 1.5, moderate model gap from 1.5 to below 2.5, and large model gap at 2.5 and above
- Average Residue Volume is computed from canonical residue-volume constants (ų)
Noncanonical Handling
- Global size metrics remain mass/length-derived and do not require analog mapping
- Terminal increments are applied independently of canonical or noncanonical residue identity
- For Average Residue Volume, Peptalyzer™ uses curated noncanonical residue volumes when defined, analog fallback only when a curated value is absent, and unsupported status only if neither is available
- In the current noncanonical set, curated residue volumes are defined for all supported tokens — Aib, AmPhe, beta-Ala, Cit, Hyp, Nle, Nva, Orn, Ahx, and AEEA — so Average Residue Volume is fully supported across this set
Limitations
- These are deterministic model descriptors, not structural predictions
- The compact model assumes protein-like density; the flexible model assumes polymer-scaling behavior
- Expansion Ratio is length-dependent and should not be treated as a folding-state classifier
- If an active terminal modification lacks a calibrated size increment, that terminus is treated as +0.0 ų in global size metrics and a caution note is shown
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | 88.0 ų | low | Included in average residue volume |
| R (Arg) | yes | direct canonical | 173.0 ų | low | Included in average residue volume |
| N (Asn) | yes | direct canonical | 114.0 ų | low | Included in average residue volume |
| D (Asp) | yes | direct canonical | 111.0 ų | low | Included in average residue volume |
| C (Cys) | yes | direct canonical | 108.0 ų | low | Included in average residue volume |
| Q (Gln) | yes | direct canonical | 143.0 ų | low | Included in average residue volume |
| E (Glu) | yes | direct canonical | 138.0 ų | low | Included in average residue volume |
| G (Gly) | yes | direct canonical | 60.0 ų | low | Backbone reference residue |
| H (His) | yes | direct canonical | 153.0 ų | low | Included in average residue volume |
| I (Ile) | yes | direct canonical | 166.0 ų | low | Included in average residue volume |
| L (Leu) | yes | direct canonical | 166.0 ų | low | Included in average residue volume |
| K (Lys) | yes | direct canonical | 168.0 ų | low | Included in average residue volume |
| M (Met) | yes | direct canonical | 162.0 ų | low | Included in average residue volume |
| F (Phe) | yes | direct canonical | 189.0 ų | low | Included in average residue volume |
| P (Pro) | yes | direct canonical | 112.0 ų | low | Included in average residue volume |
| S (Ser) | yes | direct canonical | 89.0 ų | low | Included in average residue volume |
| T (Thr) | yes | direct canonical | 116.0 ų | low | Included in average residue volume |
| W (Trp) | yes | direct canonical | 227.0 ų | low | Included in average residue volume |
| Y (Tyr) | yes | direct canonical | 193.0 ų | low | Included in average residue volume |
| V (Val) | yes | direct canonical | 140.0 ų | low | Included in average residue volume |
| Aib | yes | direct curated noncanonical | 115.2 ų | medium | Curated noncanonical volume |
| AmPhe | yes | direct curated noncanonical | 216.5 ų | medium | Curated noncanonical volume |
| beta-Ala | yes | direct curated noncanonical | 85.4 ų | medium-high | Curated noncanonical volume |
| Cit | yes | direct curated noncanonical | 160.3 ų | medium | Curated noncanonical volume |
| Hyp | yes | direct curated noncanonical | 116.0 ų | medium | Curated noncanonical volume |
| Nle | yes | direct curated noncanonical | 169.4 ų | low-medium | Curated noncanonical volume |
| Nva | yes | direct curated noncanonical | 142.7 ų | medium | Curated noncanonical volume |
| Orn | yes | direct curated noncanonical | 142.0 ų | medium | Curated noncanonical volume |
| Ahx | yes | direct curated noncanonical | 161.0 ų | medium | Curated noncanonical volume |
| AEEA | yes | direct curated noncanonical | 155.0 ų | medium | Curated noncanonical volume |
Computational Note
- Estimated Molecular Volume = 1.21 × average molecular weight + terminal volume increment (when calibrated)
- Equivalent Sphere Radius is computed from estimated volume as a sphere-equivalent compact-state radius (nm)
- Flexible-Chain Radius uses a length-only polymer scaling relation: 0.21 × N^0.57 (N = sequence residue count)
- Expansion Ratio = R_flexible / R_compact, classified as Small (<1.5), Moderate (1.5 to <2.5), Large (>=2.5)
- Average Residue Volume is computed as the mean residue volume across sequence tokens using canonical constants or curated noncanonical values (analog fallback only when curated values are absent)
Z-Scale Descriptors (Sandberg Framework)
Peptalyzer™ computes Z-scale descriptors as strict residue-level vectors (z1–z5) using direct per-residue constants from the internal Z-scale table. Canonical residues use direct values. Noncanonical residues are included only if they also have direct Z-scale vectors in that same table. There is no projection, interpolation, or virtual approximation fallback at runtime.
For supported sequences, Peptalyzer™ returns: a position-wise matrix of one z1–z5 vector per residue, whole-sequence per-axis means and population standard deviations, a deterministic interpretation block, and the profile plot built from the same matrix. If any residue lacks a direct Z-scale vector, Z-scale output is marked unsupported for that sequence.
Canonical Handling
Direct Sandberg-style z1–z5 vectors are used for all 20 canonical amino acids without approximation. Output is fully supported when all sequence tokens are covered by direct vectors.
Noncanonical Handling
- Noncanonical residues with direct Z-scale vectors in the current table are: Aib, beta-Ala, Cit, Hyp, Nle, Nva, and Orn
- Residues without a direct vector — currently Ahx, AEEA, and AmPhe — block the entire Z-scale feature for that sequence
- There is no partial fallback mode
Limitations
- Runtime is strict direct-coverage only; no virtual characterization or projection is applied in-app
- Z-scale descriptors are multivariate sequence descriptors, not direct physical observables
- Active terminal modifications are excluded from Z-scale descriptors and all derived summaries — means, standard deviations, ranges, and conclusion patterns
- Any missing, malformed, or non-finite direct vector value causes Z-scale output to be marked unsupported
Residue Handling Summary
| Residue | Used in Calculation | Approximation Strategy | Reference Value | Expected Deviation | Notes |
|---|---|---|---|---|---|
| A (Ala) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| R (Arg) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| N (Asn) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| D (Asp) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| C (Cys) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| Q (Gln) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| E (Glu) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| G (Gly) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| H (His) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| I (Ile) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| L (Leu) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| K (Lys) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| M (Met) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| F (Phe) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| P (Pro) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| S (Ser) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| T (Thr) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| W (Trp) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| Y (Tyr) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| V (Val) | yes | direct canonical | Z1–Z5 direct | low | Fully supported |
| Aib | yes | direct noncanonical | Z1–Z5 direct | medium | Supported via explicit table entry |
| beta-Ala | yes | direct noncanonical | Z1–Z5 direct | medium | Supported via explicit table entry |
| Cit | yes | direct noncanonical | Z1–Z5 direct | medium | Supported via explicit table entry |
| Hyp | yes | direct noncanonical | Z1–Z5 direct | medium | Supported via explicit table entry |
| Nle | yes | direct noncanonical | Z1–Z5 direct | low–medium | Supported via explicit table entry |
| Nva | yes | direct noncanonical | Z1–Z5 direct | medium | Supported via explicit table entry |
| Orn | yes | direct noncanonical | Z1–Z5 direct | medium | Supported via explicit table entry |
| Ahx | no | unsupported | none | high | Missing direct Z-scale vector; blocks feature |
| AEEA | no | unsupported | none | high | Missing direct Z-scale vector; blocks feature |
| AmPhe | no | unsupported | none | high | Missing direct Z-scale vector; blocks feature |
Computational Note
For a sequence of length n, each residue contributes a 5D vector (z1, z2, z3, z4, z5). The per-axis mean and population standard deviation are:
\[\bar{z}k = \frac{1}{n} \sum{i=1}^{n} z_{k,i}
\] \[
\sigma_k = \sqrt{\frac{1}{n} \sum_{i=1}^{n} \left(z_{k,i} – \bar{z}_k\right)^2}
\]
Axis range is reported as the observed min–max across residue values for each axis. Computation proceeds only when every sequence token has a valid direct 5-value Z-scale vector.
