Processing Raw Data and Fitting It Honestly
What to subtract, in what order, and how to tell whether a fit means anything.
The processing pipeline, in order#
Data processing in SPR is a short, fixed sequence. Doing the steps out of order, or skipping one, introduces artefacts that are then very difficult to diagnose because they no longer look like their cause.
- x-alignment. Shift each curve so that t = 0 is the moment of injection. Injection timing varies by a second or two between cycles, and unaligned curves smear the early association phase where the kinetic information is densest (Myszka, 1999).
- y-zeroing. Subtract the mean of a short baseline window immediately before injection, so every curve starts at zero. Use a window, not a single point, or you import that point’s noise into the whole curve.
- Reference subtraction. Subtract the reference flow cell trace from the active one. This removes bulk shift, drift, and non-specific binding common to both.
- Blank subtraction. Subtract the averaged buffer-only injection from every analyte curve. This removes systematic injection artefacts that survive step 3. Steps 3 and 4 together are double referencing (Myszka, 1999).
- Excluded-volume / solvent correction, where relevant — mandatory for DMSO-containing samples (small molecules and fragments).
- Region selection. Exclude the few seconds around injection start and stop where valve transients live. Exclude nothing else.
Fitting: local, global, and what the difference means#
Early biosensor analysis linearised the data and fitted straight lines. Morton, Myszka and Chaiken showed why that is a poor idea — linearisation distorts the error structure and hides model failure — and introduced numerical integration, which is why arbitrary interaction models became fittable at all (Morton et al., 1995). O’Shannessy and colleagues made the parallel case for non-linear least squares over transformation (O’Shannessy et al., 1993).
Global fitting
Fit every curve in a concentration series simultaneously, with ka, kd and Rmax shared across all of them and only the concentration differing. This is a far stronger test than fitting each curve alone, because one parameter set must explain the whole family (Morton et al., 1995; Myszka, 2000). A model that is wrong will usually fail visibly under a global fit while fitting each curve individually rather well.
Reading residuals
The residual plot — data minus fit, against time — is the most informative output of any fitting exercise, and it is the one most often omitted from publications (Rich & Myszka, 2008).
| Residual pattern | Usually means |
|---|---|
| Random scatter at instrument-noise amplitude | The model is consistent with the data. Not that it is uniquely correct |
| Systematic S-shape | Wrong model — work through the differential diagnosis in going beyond a 1:1 model |
| Deviation only in early association | Un-modelled bulk shift, or an x-alignment error |
| Deviation only in late dissociation | A slow second component: avidity, heterogeneity, or rebinding |
| Amplitude that grows with concentration | Transport limitation, or an analyte concentration error |
| A step at the injection boundary | Incomplete referencing |
How much should you trust a fitted number?#
Fitting software returns standard errors. Those errors describe the precision of the fit given the model, and they are almost always far smaller than the true uncertainty, because they take no account of the possibility that the model is wrong, the concentration is wrong, or the surface was decaying.
- Parameter correlation. ka and Rmax are strongly correlated when the concentration range does not approach saturation: the fit can trade a smaller Rmax against a larger ka almost freely. If your software reports a correlation matrix, look at it. Correlations above about 0.95 mean the two parameters are not independently determined.
- Concentration accuracy propagates directly. ka is derived from kobs = kaC + kd, so a 20% error in C is a 20% error in ka, with no fitting statistic reflecting it. This is the strongest argument for CFCA-based active concentration (kinetics and mass transport).
- Temperature is a variable, not a constant. Running the same interaction across a temperature series and applying van’t Hoff and Eyring analysis separates the enthalpic and entropic contributions to binding (Roos et al., 1998) — and, more prosaically, reveals when a “constant” you reported was really a property of the day’s room temperature.
- Independent repeats beat replicate injections. Repeated injections on the same surface measure reproducibility of injection. Fresh surfaces on different days measure reproducibility of the experiment, which is the number a reader needs.
The model zoo, and when each is defensible#
| Model | Parameters | Assumes | Justified when |
|---|---|---|---|
| 1:1 Langmuir | ka, kd, Rmax | One analyte, one identical site class, no transport | Default. Always start here (O’Shannessy et al., 1993) |
| 1:1 with mass transport | + kt | Transport and chemistry in series | Density series proves transport is present and you cannot lower density further (Myszka et al., 1998) |
| Bivalent analyte | ka1, kd1, ka2, kd2, Rmax | Analyte has two identical arms binding sequentially | Analyte is genuinely bivalent and you report the density dependence (Nieba et al., 1996) |
| Heterogeneous ligand | Two full parameter sets | Exactly two discrete site classes | Rarely. The physical claim of exactly two classes is strong — prefer distribution analysis (Svitel et al., 2003) |
| Two-state / conformational change | ka1, kd1, ka2, kd2, Rmax | Binding followed by an isomerisation | Only with independent evidence. See the warning in going beyond a 1:1 model |
| Steady-state affinity | KD, Rmax | Equilibrium reached at every concentration | Kinetics too fast to resolve; verify each curve actually plateaued (Myszka, 2000) |
| Distribution analysis | A regularised distribution | A continuum of site properties | Heterogeneity is the object of study, or discrete models keep failing (Svitel et al., 2003) |
The ordering here is deliberate. Move down the table only when an experiment — not a fit statistic — has ruled out the row above.
Sources cited on this page
Listed alphabetically. Each badge records whether the bibliographic record was confirmed against Crossref. unverified marks a real, deliberately chosen source whose volume and page numbers we have not yet machine-checked — it is not a comment on the science.
- Katsamba et al., 2006P. S. Katsamba, I. Navratilova, M. Calderon-Cacia, et al. (2006). Kinetic analysis of a high-affinity antibody/antigen interaction performed by multiple Biacore users. Analytical Biochemistry 352, 208–221. unverified
- Morton et al., 1995T. A. Morton, D. G. Myszka, I. M. Chaiken (1995). Interpreting complex binding kinetics from optical biosensors: a comparison of analysis by linearization, the integrated rate equation, and numerical integration. Analytical Biochemistry 227, 176–185. doi:10.1006/abio.1995.1268 verifiedIntroduced numerical integration to biosensor analysis, which is why arbitrary interaction models became fittable.
- Myszka et al., 1998D. G. Myszka, X. He, M. Dembo, T. A. Morton, B. Goldstein (1998). Extending the range of rate constants available from BIACORE: interpreting mass transport-influenced binding data. Biophysical Journal 75, 583–594. doi:10.1016/S0006-3495(98)77549-6 verifiedShows transport can be fitted rather than merely avoided, and defines the transport coefficient kₜ.
- Myszka, 1999D. G. Myszka (1999). Improving biosensor analysis. Journal of Molecular Recognition 12, 279–284. unverifiedOrigin of double referencing and blank-injection subtraction as standard practice.
- Myszka, 2000D. G. Myszka (2000). Kinetic, equilibrium, and thermodynamic analysis of macromolecular interactions with BIACORE. Methods in Enzymology 323, 325–340. doi:10.1016/S0076-6879(00)23372-7 unverified
- Nieba et al., 1996L. Nieba, A. Krebber, A. Plückthun (1996). Competition BIAcore for measuring true affinities: large differences from values determined from binding kinetics. Analytical Biochemistry 234, 155–165. doi:10.1006/abio.1996.0067 unverifiedA direct demonstration that surface-measured kinetic constants can diverge substantially from solution affinities, and a solution-competition format that avoids the problem.
- O’Shannessy et al., 1993D. J. O’Shannessy, M. Brigham-Burke, K. K. Soneson, P. Hensley, I. Brooks (1993). Determination of rate and equilibrium binding constants for macromolecular interactions using surface plasmon resonance: use of nonlinear least squares analysis methods. Analytical Biochemistry 212, 457–468. doi:10.1006/abio.1993.1355 verifiedThe case for fitting the sensorgram directly rather than linearising it.
- Papalia et al., 2006G. A. Papalia, S. Leavitt, M. A. Bynum, et al. (2006). Comparative analysis of 10 small molecules binding to carbonic anhydrase II by different investigators using Biacore technology. Analytical Biochemistry 359, 94–105. unverifiedA benchmark study in which many labs measured the same interactions; the spread between labs is the honest measure of SPR reproducibility.
- Rich & Myszka, 2008R. L. Rich, D. G. Myszka (2008). Survey of the year 2007 commercial optical biosensor literature. Journal of Molecular Recognition 21, 355–400. doi:10.1002/jmr.928 unverifiedreviewOne instalment of the long-running literature survey. Cited only for claims about the state of the published literature, which is what it measures.
- Roos et al., 1998H. Roos, R. Karlsson, H. Nilshans, A. Persson (1998). Thermodynamic analysis of protein interactions with biosensor technology. Journal of Molecular Recognition 11, 204–210. unverifiedVan’t Hoff and Eyring analysis of SPR data across temperature.
- Svitel et al., 2003J. Svitel, A. Balbo, R. A. Mariuzza, N. R. Gonzales, P. Schuck (2003). Combined affinity and rate constant distributions of ligand populations from experimental surface binding kinetics and equilibria. Biophysical Journal 84, 4062–4077. doi:10.1016/S0006-3495(03)75132-7 verifiedThe distribution analysis that replaces "the surface has one kind of site" with a measured spectrum of sites.