Skip to content

LC-MS/MS Annotation

LC-MS/MS annotation supports unknown compound review, fragmentation matching, and related-family grouping.

  • Upload mzML, mzXML, Waters RAW, or Bruker .d data.
  • Review precursor, fragment, and retention-time evidence.
  • Compare proposed structures against fragmentation support.
  • Group related compounds into LC-MS families.
  • Export reviewed evidence into Regentry.

Link to Core Concepts when explaining confidence scores and audit trails.


The MS / structure side of SpectraCheck’s pretrained layer fuses three independent signals into one calibrated candidate ranking; the deterministic structure verifier remains the arbiter of pass/fail. The release timeline gives chronological context.

MS models — CSI:FingerID, METLIN RT & DP4-AI fusion

Section titled “MS models — CSI:FingerID, METLIN RT & DP4-AI fusion”
  • CSI:FingerID (MS/MS → structure) — predict_msms_candidates() wraps SIRIUS / CSI:FingerID through its documented interface (env-configured REST service or CLI binary; injectable backend), returning ranked candidate structures + fingerprints. SIRIUS is never reimplemented or bundled, and on a host with no configured backend it degrades gracefully (available=False). (v0.15.0, 2026-06-07)
  • METLIN retention-time corroboration — predict_retention_times() + rt_corroboration() apply a Gaussian down-weight on the RT residual, so an RT-inconsistent candidate is demoted, never hard-filtered. (v0.15.0, 2026-06-07)
  • DP4-AI posterior + calibrated fusion — dp4_candidate_posterior() reuses the validated in-house DP4 scoring (Smith & Goodman 2010) for a calibrated NMR-candidate posterior, and fuse_candidates() combines NMR (DP4) + MS/MS (CSI) + RT into one ranking summing to 1.0 (RT as a multiplicative down-weight; signal weights renormalise when a signal is missing). Decision-support only — arbitrate() hands the top candidate to the structure verifier. (v0.15.0, 2026-06-07)

A candidate ranking now reports what it actually explains

Section titled “A candidate ranking now reports what it actually explains”

rms_error_ppm is computed over the peaks that paired within the matching window, so peaks the prediction missed badly are excluded from the error figure — which stops the figure responding to error. Measured on twelve ¹H shifts for one candidate:

true RMSE 0.140 -> reported 0.118 matched 11/12
true RMSE 0.540 -> reported 0.203 matched 8/12
true RMSE 2.418 -> reported 0.154 matched 6/12

A seventeen-fold degradation in the real fit moves the reported number from 0.118 to 0.154. The only thing that moved was the matched count, and the row emitted a matched-peak count with no denominator, so 6 was indistinguishable from 6 of 6.

The likelihood is not blind to this — unmatched peaks take a soft penalty, so the ranking is defensible and still identifies the correct candidate. It is a reporting defect, so the fix adds observed_peak_count, matched_fraction, low_coverage, error_basis, probability_is_calibrated and probability_basis, and does not touch the arithmetic. The minimum coverage constant is taken from where the measurement decouples, not from a round number.

No constants were substituted, and the direction of the calibration claim was corrected. An earlier note said an understated σ “saturates the posterior toward 1.0”. Measured, the top-candidate probability is 0.996 at an injected σ of 0.05 but 0.73 at the 0.42 the predictor actually achieves. Fitted against the predictor’s own held-out errors, the ¹H scale is marginally tighter than published — and the heavy-tail parameter is ≈ 1.23 against a published 14.18. That parameter is the load-bearing one: scoring a heavy-tailed predictor with a thin-tailed model drives the correct candidate toward zero, so the dangerous direction here is a confident false rejection, not a false confirmation. (v0.68.9, 2026-08-08)

Mass-spectrum evidence in the offline desktop installation

Section titled “Mass-spectrum evidence in the offline desktop installation”

ms_molecule_match is one of the verifier’s four tests, and it abstained on every check the desktop ever made because nothing on that machine could supply its evidence. The offline service now reads a processed centroid peak table through the platform’s own MS-peak parser — CSV, TSV or whitespace rows, comments and headers skipped, a percentage column tolerated, and a refusal that names the line number when a row carries only one number. Processed centroid peaks only: mzML and vendor formats go through the LC-MS import bridge, which is a server surface, and the refusal says so rather than returning an empty table.

Measured on a reference acquisition with the true structure: 2 of 4 tests → 3 of 4, and the verdict moves from inconclusive 0.562 to consistent 0.925.

The asymmetry is rendered, because it is not guessable from the result. This test weights itself by the matched fraction of the predicted pattern, so a candidate matching nothing scores zero significance and moves the posterior by nothing. Measured against a 114 Da molecular ion, ethanol, ethylene glycol and aspirin (46, 62 and 180 Da) each moved by exactly +0.000. A reader seeing one candidate lifted and three unchanged would reasonably conclude the three were ruled out; they were not, and the interface says to read an unchanged candidate as unsupported by the MS, never as ruled out by it. A malformed optional MS peak list is ignored rather than refused, because a structure check without MS is the normal case and must not fail on an optional field. (v0.75.1, 2026-09-07)

VersionDateHeadline
v0.75.12026-09-07Offline MS peak table reaches the verifier’s ms_molecule_match
v0.68.92026-08-08DP4 ranking reports coverage; the calibration claim is corrected
v0.15.02026-06-07MS models: CSI:FingerID, METLIN RT & DP4-AI candidate fusion