Skip to content

NMR Interpretation Guide

Use this guide when reviewing NMR evidence and deciding whether a SpectraCheck assignment is ready for report export.

  1. Upload the raw FID archive and confirm instrument metadata.
  2. Review 1H regions, integration, multiplicity, and solvent references.
  3. Review 13C assignments, solvent peaks, and carbonyl regions.
  4. Use COSY, HSQC, and HMBC evidence to validate connectivity.
  5. Resolve contradiction flags before approving an interpretation.

Every numbered step needs a product screenshot before publication.

Use one accessible example, such as caffeine or ibuprofen, and run it through the current production workflow. Capture every state a reviewer sees:

  • Upload accepted with raw source file and parameter file visible.
  • Processed spectrum with regions labeled.
  • Peak table with shift, multiplicity, integration, assignment, and confidence.
  • Evidence card for at least one assigned peak.
  • Contradiction or warning state, even if the example uses a seeded issue.
  • Final accepted interpretation ready for export.
ExperimentWhat it helps validate
COSYWhich protons are coupled to nearby protons.
HSQCWhich proton is attached to which carbon.
HMBCLonger-range proton-carbon connections used to support structure fragments.

Keep the explanation short enough for an analytical chemist who understands NMR but has not used MolTrace before.

A contradiction is not a failure; it is a review queue. Show what evidence disagrees, which assignment is affected, and what action the scientist can take:

  • Accept with rationale.
  • Reassign the peak.
  • Mark as impurity, solvent, reference, or unknown.
  • Request additional evidence before export.

The NMR scientist owns this page and must verify the example, screenshots, and language before release.


The NMR interpretation backend has shipped a substantial set of analysis capabilities. This section catalogs what is in production today; the release timeline at the end gives chronological context.

Global Spectral Deconvolution (GSD) — opt-in analysis backend

Section titled “Global Spectral Deconvolution (GSD) — opt-in analysis backend”

The opt-in POST /spectrum/analyze/gsd endpoint runs industry-standard Global Spectral Deconvolution on a processed spectrum and returns peaks auto-classified as compound | solvent | impurity | artifact | 13C_satellite. It ships behind a per-request experimental: true flag while the soak loop runs, and graduates per-tenant or platform-wide on a measured verdict — see the GSD experimental rollout section in the deployment guide.

  • Detection algorithm — single-pass detection via scipy.signal.find_peaks; per-peak fitting via lmfit Lorentzian / pseudo-Voigt; level-aware overlap resolution at levels 4–5; classification using the Fulmer / Gottlieb residual-solvent tables. (v0.4.0, 2026-05-27)
  • Algorithm semantics + envelope unification — cluster_into_environments groups adjacent same-category peaks within a nucleus-aware J-coupling window into one chemical-environment entry. Legacy raw-FID surfaces (/nmr/raw-fid/preview and /nmr/raw-fid/process) gain environments / environment_count / environment_counts so the FE renders both detectors with one component. A vectorized _pseudo_voigt_sum plus analytical jacobian gives an 8.5× speedup on dense ¹³C (60000006_13c fixture: 5.5 min → 39 s), bit-exact-equivalent. (v0.5.0, 2026-05-27)
  • Strict promotion gate cleared — on the NMRShiftDB2 corpus the sidecar cleared its strict production promotion gate (95 % solvent auto-detect plus median compound-environment-count delta ≤ 2). The HMDB-style validation framework forward-models a noisy Lorentzian spectrum from a published peak list and gates against environment-count and multiplet-line deltas on a 20-fixture mini-corpus. Default ¹H clustering window widened 20 Hz → 30 Hz to accommodate strong-coupling AB systems and constrained-ring geminal H-H couplings up to 25–30 Hz. (v0.6.0, 2026-05-28)
  • Per-peak QC metrics for legacy raw-FID — LegacyEnrichedPeak.fit_redchi / fit_rmse / fwhm_ppm / signal_to_noise / baseline_noise_sigma — the same regulatory-tier QC quintuple already published by the GSD endpoint via Peak.metadata. Both /nmr/raw-fid/preview and /nmr/raw-fid/process populate the quintuple before returning. (v0.6.1, 2026-05-28)
  • Real HMDB validation corpus — a 100-fixture real-instrument HMDB corpus (60 × ¹H + 40 × ¹³C; Bruker 59 / Varian 41; solvent mix Water/D₂O 85, CD₃OD 6, CDCl₃ 5, DMSO-d₆ 4). Result: 95/100 parse cleanly; 53/57 = 93 % solvent auto-detect on the subset with a known solvent reference. The literal Prompt 3 spec is satisfied across NMRShiftDB2 (19 fixtures; 100 % solvent), HMDB synthetic (20), and HMDB real-instrument (100; 95 % parseable, 93 % solvent). (v0.6.2, 2026-05-28)
  • Source-of-truth solvent/impurity expert system — moltrace.spectroscopy.classify.solvent_impurity sorts every peak into compound | solvent | residual_solvent | impurity | 13C_satellite | artifact using the Fulmer (2010) and Gottlieb (1997) residual-solvent + trace-impurity reference tables (14 deuterated solvents; ~30 common organic impurities across the seven Fulmer columns). Classification is a transparent additive evidence scheme — high for a solvent-table match or out-of-range shift, medium for a ¹³C-satellite pair at ±½ J(C–H) (125 Hz sp³ / 160 Hz sp²) or a line-width anomaly, low for sub-noise intensity — with an intensity-prominence gate so a dominant analyte resonance is never captured by a colliding impurity window. It is integration-ready with the GSD auto_classify categoriser, adding the explicit residual_solvent label that separates leftover process solvents from the bulk deuterated-solvent line. (v0.9.0, 2026-06-05)
  • IST baseline + optional JTF-Net — moltrace.spectroscopy.nus.reconstruct reconstructs a full spectrum from non-uniformly sampled data. The always-available iterative soft-thresholding (IST-S) baseline (Stern–Donoho–Hoch 2007; Hyberts 2012) is weights-free, numpy-only, and deterministic — the robust default for small-molecule 2-D spectra. An optional, lazily-loaded JTF-Net backend (Luo et al., Nat. Commun. 16, 2342, 2025) follows the same local-first device strategy (CUDA → MPS → CPU) and weights-cached-out-of-git policy as the NMRNet wrapper, and never fabricates a reconstruction — when torch / weights are absent it falls back to IST with a warning. The reference-free REQUIRER ratio (∈ [0, 1]; 1 = best) scores reconstruction quality against the measured NUS data, with no fully-sampled reference needed. (v0.10.0, 2026-06-06)

JTF-Net’s released weights were trained on protein multidimensional spectra and are treated as out-of-domain for small-molecule 2-D work, so reconstruct_jtfnet defaults to the IST baseline until it is re-validated or fine-tuned on small-molecule data.

Multiplet analysis and J-coupling refinement

Section titled “Multiplet analysis and J-coupling refinement”

The multiplet capability groups GSD-resolved peaks into multiplets, recognises multiplicity (s / d / t / q / p / sext / sept / dd / dt / td / ddd / m), and recovers the underlying J couplings.

  • Multiplet detection plus synthetic overlay — POST /spectrum/analyze/multiplets takes a GSD peak list and returns recognised multiplets with recovered J couplings. The forward modeller generate_synthetic_multiplet is publicly exposed so the FE can overlay predicted-vs-observed peaks (light red) on the spectrum view. Algorithm: spatial cluster at 30 Hz → first-order Pascal-triangle match → dd analytical inversion / dt-td-ddd J-set enumeration with scipy.optimize.least_squares refinement → “m” fallback for unstructured clusters. Validation: 8 quinine multiplets resolved with J within 0.3 Hz of literature; a known hidden 11.4 Hz coupling benchmark recovered where standard peak picking misses it. (v0.7.0, 2026-05-28)
  • Multiplet J-coupling → unified confidence layer — the recovered J-couplings feed the unified candidate-confidence engine as the 40th evidence layer (multiplet_jcoupling). POST /candidates/compare/jcoupling returns per-candidate labels (strong | partial | weak | poor_j_agreement plus j_coupling_contradiction) so the FE can render a J-agreement badge per candidate. A contradiction (observed J above a threshold the candidate topology cannot produce) caps the score at 0.25. Purely additive: existing callers unchanged when no multiplet input is supplied. (v0.7.1, 2026-05-28)
  • Opt-in Karplus 3J refinement — Layer 40’s topological J-predictor gains an opt-in, conformer-averaged Karplus refinement for sp³ vicinal (³J) couplings (RDKit ETKDGv3 plus MMFF). When enabled (use_karplus=True), the flat 7.0 Hz aliphatic_vicinal placeholder is replaced by a geometry-aware estimate. Default-off and byte-for-byte identical when the flag is omitted. (v0.7.2, 2026-05-28)
  • Karplus validation corpus — an 8-molecule curated literature validation corpus (karplus_jcoupling_corpus_v1.json) and a pytest accuracy gate: mean absolute error 0.44 Hz (median 0.26, max 1.41), with clean separation between conformationally locked diaxial systems (mean 9.5 Hz, all ≥ 8.49 Hz) and mobile/averaged systems (mean 6.9 Hz, all ≤ 7.14 Hz) with no overlap. (v0.7.3, 2026-05-28)
  • Opt-in Haasnoot–Altona generalized Karplus plus honest negative result — a second selectable relation (karplus_method=haasnoot_altona). Per individual conformer it is more literature-faithful (recovers trans-decalin diaxial at 11.64 Hz, above the generic 10.26 Hz ceiling), but the corpus study — shipped as a regression gate — shows HLA does not improve averaged discrimination under the unweighted conformer model, openly documented. (v0.7.4, 2026-05-30)
  • Boltzmann conformer-population weighting (sugar blind-spot fix) — opt-in karplus_conformer_weighting field (uniform | boltzmann, default uniform) weights each conformer by its MMFF-energy Boltzmann population at 298.15 K. Measured corpus effect: β-D-galactose recovers from 8.49 → ~10.1 Hz onto its literature value; locked-vs-mobile separation widens (generic: +1.35 → +2.28 Hz). Once conformers are population-weighted, the generic relation discriminates better than HLA — the sugar gap was a weighting problem, not an equation one. (v0.7.5, 2026-05-30)
  • Karplus corpus scaled to 18 molecules — a new 18-molecule v2 corpus (9 locked diaxial plus 9 mobile/averaged, including five new pyranosides) graded across the {generic, haasnoot_altona} × {uniform, boltzmann} grid shows generic/boltzmann is the only one of the four that cleanly separates locked from mobile at scale. Within-tolerance 1.00, mean abs error 0.57 Hz, locked-vs-mobile separation +1.84 Hz. (v0.7.6, 2026-05-31)
  • NMRNet wrapper plus HOSE-code fallback — predict_shifts(smiles, nuclei) returns predicted ¹H / ¹³C shifts (ppm) with per-atom uncertainty. Two backends: the NMRNet SE(3)-equivariant model (Xu et al., Nat. Comput. Sci. 5, 292, 2025) as an optional, lazily-loaded backend (in-process or remote GPU microservice), and a HOSE-code / NMRShiftDB2 topological fallback (spheres 6 → 1) as the default. NMRNet never fabricates a prediction — it activates only when configured. Exposed via POST /spectrum/predict/shifts. (v0.7.8, 2026-06-01)

    Correction (v0.64.0, 2026-08-07). For most of the period above, the “HOSE-code / NMRShiftDB2 fallback” described in this bullet was not running against NMRShiftDB2. It had silently degraded to a bundled 16-molecule curated seed table: on four drug-like molecules 22.6–44.4 % of atoms resolved to a bare element prior and the median ¹³C uncertainty was 35.0 ppm. Nothing was hidden — every fallback appended a per-atom warning — but no caller aggregated them, so a prediction that was mostly element averages was indistinguishable from a resolved one. The published NMRNet test-set MAEs were also presented as “headline accuracy” for a function that runs the fallback in production; those figures are now scoped to the nmrnet path. See Reference table coverage below for what the fallback does today and what a deployment must stage for it to hold.

  • NMRNet wrapper rework: local-first device strategy — reworked from microservice-first to local-first (Apple-Silicon dev): device resolution CUDA → MPS → CPU (CPU baseline, MPS best-effort with a clean CPU fallback), lazy torch, per-atom uncertainty from the conformer ensemble (std across n_conformers; null at n=1), Zenodo/HF-mirror weights acquisition (cached, SHA-256). HOSE fallback now requires ≥ 3 references per matched sphere. The QM9-NMR gate targets the paper’s QM9NMR MAE (0.020 / 0.262 ppm). NMRNet is never vendored. (v0.7.9, 2026-06-01)

  • Multi-test ASV scorer — verify_structure(spectrum, proposed_smiles, prior_confidence=0.5, tests=None, options=None) scores how well a proposed structure explains an experimental 1-D NMR spectrum and combines several independent tests into one auditable posterior confidence. Four tests ship: PredictionBoundsTest, AssignmentsTest, HSQC2DRangesTest, MSMoleculeMatchTest, each returning a TestResult (score, significance, quality = score · tanh(significance/3), diagnostic). Bayesian log-odds combination (logit(p_post) = logit(prior) + Σ quality_i · ln10); verdict thresholds 0.80 (consistent) / 0.20 (inconsistent). Tests with no data abstain rather than fabricate evidence; a per-test error degrades to an abstain. Grounded in published ASV / CASE literature (Golotvin & Williams; Elyashberg et al.); no vendor scoring scheme is reproduced. (v0.8.0, 2026-06-03)

Spectrum retrieval — vector plus set similarity

Section titled “Spectrum retrieval — vector plus set similarity”
  • FAISS HNSW similarity layer — moltrace.spectroscopy.similarity provides a Gaussian-smoothed 256-D spectral encoding [v_1H(128); v_13C(128)] with FAISS HNSW L2 retrieval, plus a Kuhn-Munkres set-similarity score (scipy.optimize.linear_sum_assignment; unmatched peaks allowed → robust to insertion/deletion). Performance: top-100 from 45 k in ≈ 2 ms (target was < 1 s). Implements the NMR-Solver methodology (Jin et al., arXiv:2509.00640, 2025) from the published equations. (v0.8.1, 2026-06-03)
  • POST /spectrum/retrieve endpoint — the similarity layer becomes a typed API. The endpoint matches a query spectrum (¹H/¹³C shift lists or a SMILES) against the server-configured FAISS index (MOLTRACE_SIMILARITY_INDEX) and returns the top-k nearest reference spectra by L2 distance. Graceful index_available=false when unset; one spectrum.retrieve audit event per call. (v0.8.2, 2026-06-03)
  • Precedent-grounded structure proposal — moltrace.spectroscopy.ai.rag wraps Anthropic Claude in a retrieval layer over the similarity index: it retrieves the nearest known spectra, asks the model for candidate structures grounded in that precedent, and applies a cite-or-drop hallucination guard (a candidate that neither cites a real retrieved analogue nor structurally matches one is dropped before verification). The deterministic verifier — never the LLM — decides pass/fail; the model’s self-confidence is advisory and is never used as the verifier prior. The full prompt + raw completion + retrieved ids are captured for the audit trail. The LLM, index, resolver, verifier, and recorder are all injectable, so the pipeline runs on a CPU-only host with no network. Exposed via POST /spectrum/reason — see Backend / API Contract. (v0.16.0, 2026-06-07)

Reference table coverage, and what happens without it

Section titled “Reference table coverage, and what happens without it”

The shift predictor’s accuracy is a property of the deployment, not only of the code, because the HOSE reference table is a NMRShiftDB2 derivative that is gitignored (CC BY-SA) and staged at deploy time.

  • The table, and the measured improvement it bought. Built from the full NMRShiftDB2 NMReDATA export (64 723 records → 49 618 molecules / 495 215 assignments, 14 s), prior-fallback share falls 22.6–44.4 % → 0.0 % across the test panel, pooled median ¹³C σ 35.00 → 1.88 ppm and ¹H σ 0.52 → 0.26 ppm — an 18.6× improvement in ¹³C uncertainty with no model, no GPU and no new dependency. The table is built from assignments, not peak lists: records emit a molblock rather than SMILES, because rebuilding through SMILES → AddHs re-orders hydrogens and would attach ¹H shifts to the wrong protons — a healthy-looking table that is wrong everywhere and undetectable downstream. Per-shift provenance ships with it (AtomShift.source = nmrnet | hose | element_prior, plus kb_source / kb_records / prior_fallback_fraction / median_uncertainty_ppm). Gate on the fraction; do not parse warning strings. (v0.64.0, 2026-08-07)
  • Made shippable. Loading the table re-parsed 49 618 molblocks with RDKit on every process start — 51 s and 193 MB, incompatible with a scale-to-zero service, and the real reason a GPU sidecar had looked necessary. Because lookup only ever returns mean, population stdev and count, a per-bucket Welford (n, mean, M2) accumulator is lossless for the entire public contract: 193 MB / 47.3 s → 14 MB / 1.1 s, with no RDKit at load. Science is identical to 1e-9 ppm, pinned by a test comparing every lookup and every predicted shift against the table it was built from. (v0.64.1, 2026-08-07)
  • A deploy can no longer silently ship the seed predictor. Because the table is gitignored, a clean CI checkout built images whose predictor fell back to the 16-molecule seed, while a manual deploy from a checkout that happened to hold the file baked it in — so which deploy ran last decided production accuracy. Four independent guards now make that loud: the deploy workflow stages the table from Cloud Storage and verifies a tracked sha256; the Dockerfile refuses to build without it (REQUIRE_HOSE_KB=0 is the deliberate development opt-out); /health and /admin/deployment publish a hose_kb block (configured / path_present / loaded / source / reference_count) that degrades health when set-and-missing or unset-in-production; and startup validation records the seed fallback as a startup issue. Unset in development stays a legitimate, quiet configuration. See Deployment & Hosting. (v0.69.5, 2026-08-13)

Measured accuracy, on molecules the predictor has never seen

Section titled “Measured accuracy, on molecules the predictor has never seen”

Held-out measurement replaced quoted third-party benchmarks. spectroscopy/eval/shift_accuracy.py splits NMRShiftDB2 by molecule (SHA-256 of the structure, so a published figure is reproducible), builds the table from the training split alone, and scores the disjoint remainder — leakage is the design problem, because scoring a predictor on molecules already in its table collapses the error to ~0 and the number is worthless.

44 668 train / 4 950 test molecules; 445 702 training reference atoms:

ncoverageMAEmedianp90p95max
¹³C37 00599.6 %3.44 ppm1.548.2212.07146.1
¹H12 508100 %0.332 ppm0.1510.821.2214.2

A matched environment beats the element prior by 13.7× on ¹³C (3.28 vs 44.98 ppm) — the premise of the method, measured rather than assumed.

Correction to the v0.64.x figures. Those entries reported a median predicted σ of 1.88 ppm and called the predictor “sharper than the error model that consumes it”. σ is the claimed uncertainty, not the measured error. Measured ¹³C MAE is 3.44 ppm — above DP4’s 2.306 ppm scale; the median error, 1.54 ppm, is below it. The distribution is strongly right-skewed, so only quoting both is honest.

Calibration is the finding that matters. ¹³C σ is well calibrated at 2–5 ppm, conservative above that, and optimistic by ~3× in the tightest bin — exactly where DP4 weights most heavily and where the verifier scored highest. (v0.66.0, 2026-08-07)

Prediction intervals with a coverage guarantee (conformal calibration)

Section titled “Prediction intervals with a coverage guarantee (conformal calibration)”

Platt and temperature scaling calibrate a classifier’s probabilities; neither gives a regression predictor a coverage guarantee. spectroscopy/eval/conformal.py fits Mondrian split-conformal bands (nucleus × reported-σ decile) whose guarantee is distribution-free and finite-sample, assuming only exchangeability between calibration and evaluation — which the molecule-level split provides.

Measured on held-out NMRShiftDB2 (39 628 train / 5 040 calibration / 4 950 evaluation molecules, 393 760 reference atoms):

targetnucleusncoveragemean half-width
90 %¹³C36 84490.03 %6.95 ppm
90 %¹H12 50890.61 %0.714 ppm
95 %¹³C36 84494.72 %9.18 ppm
95 %¹H12 50894.93 %0.951 ppm

At 90 % both nuclei meet target; at 95 % there is a 0.28 pp shortfall on ¹³C, reported rather than rounded away.

The finding: σ is differentially mis-scaled, not merely mis-scaled. The ratio of conformal half-width to mean reported σ runs 8.66× in the tightest ¹³C band down to 1.77× in the widest — a 4.90× monotone spread, where a correctly-scaled σ would give a constant. A constant mis-scaling could be repaired by multiplying σ by one number; this cannot. And it is worst precisely where the arbiter leaned hardest. (v0.68.0, 2026-08-08)

What the verifier weighs — and what it is measured to get wrong

Section titled “What the verifier weighs — and what it is measured to get wrong”

The ASV scorer above is unchanged in shape: four tests, Bayesian log-odds combination, thresholds 0.80 / 0.20. What changed is how much each piece of evidence is allowed to be worth, and — for the first time — a measurement of how often it confirms a wrong structure.

  • Significance now comes from the conformal interval, not from the claimed σ. _significance_from_half_width replaces _significance_from_sigma at its single call site in PredictionBoundsTest, with the anchor read off the live calibration rather than restated as a constant. The mapping compresses — most-to-least significant falls from 6.757× to 4.382× — so tight atoms lose weight they had not earned and wide atoms get some back. Verdicts on clear cases do not flip: the interval changes how much evidence is worth, not which way it points. With no calibration supplied every match falls back to the σ basis, and details.significance_basis records how many atoms used each basis plus the calibration fingerprint, so a run scored on the weaker basis is visible in the audit record. (v0.68.1, 2026-08-08)
  • A match five lines could have explained is no longer worth five-sixths of a match. On held-out data, 26.5 % of in-window ¹³C resonances and 32.5 % of ¹H resonances have a rival line strictly closer to the prediction than their own, and narrowing the window cuts exposure 33 % while moving misassignment only 2.3 pp — the ambiguity is intrinsic at this predictor’s accuracy, so it belongs in the scoring model. _ambiguity_weight is a normalised likelihood under the same Gaussian the merit function already uses: one candidate scores exactly 1.0 (an unambiguous match is untouched), k equidistant candidates exactly 1/k, and it is order-independent. It attenuates significance, never score — five candidate lines do not mean the structure is wrong, they mean the observation says little either way. A 40 % floor is applied as an affine rescale and is recorded in the code as a policy choice, not a measured constant. Net effect: the shift test had been over-claiming its evidence by roughly a third on ¹³C and two-fifths on ¹H, and a ¹H-only verification that previously read consistent on a perfect match now reads inconclusive — the intended direction, and a real change in what the platform tells a user. (v0.68.7, 2026-08-08)
  • AssignmentsTest stops pricing every atom on the same ruler. It used a flat tolerance twice — as the candidate radius and as the width of the merit Gaussian — so an atom predicted to ±1.5 ppm scored 0.88 for a 2 ppm miss outside its interval while an atom predicted to ±22 ppm scored 0.32 for a 6 ppm hit inside its own. Both now scale by the resonance’s own conformal interval. Retention rises from 95.06 % → 99.07 % (¹³C) and 90.96 % → 99.25 % (¹H); each lost pairing had been penalised twice, since the resonance’s integral was also counted as unexplained impurity — a correct structure marked down for the predictor’s uncertainty. With no calibration supplied the behaviour is byte-identical to before. (v0.68.5 / v0.68.6, 2026-08-08)
  • How often a wrong structure wins, measured. spectroscopy/eval/decoys.py generates the mistakes a chemist actually makes (heteroatom misread, homologue, misplaced methyl, regioisomer, inverted stereocentre) — a randomly drawn molecule is rejected trivially and measures nothing. Over 1 657 held-out molecules / 5 639 decoy pairs, ¹³C shift lists through DP4: false-confirmation rate 39.5 % (1 214 / 3 073 scored pairs). Read it precisely — this is not “MolTrace confirms wrong structures 39 % of the time”. It is the ¹³C-shift-list layer alone through DP4: no 2-D, no MS, no multiplicity, and not the multi-test verify_structure arbiter. The load-bearing row is the regioisomer at 37.7 % — the one decoy class that shares the truth’s formula and so survives every upstream filter. Stereochemistry is a demonstrated hard boundary: 279 of 282 stereo decoys are literally indistinguishable (identical connectivity → identical HOSE code → identical prediction). The originally published 38.1 % credited an exact 0.5/0.5 DP4 tie to the truth; a tie is DP4 failing to separate, not the truth winning, so the reported figure is now 39.5 % and decoy_strict_win_rate publishes 38.1 % beside it. (v0.68.2, 2026-08-08; corrected 2026-08-27)
  • The arbiter is markedly better than that layer — on a pilot, not a publishable rate. Scoring the same decoys through verify_structure on 13 real Bruker ¹³C fixtures gives 11 scored pairs, 10 truth wins, 1 decoy win — 9.1 % false confirmation, median margin +0.485 log-odds, with the verifier calling truth and decoy “consistent” in 2 of 11. n is 11 and 9 of the 13 fixtures leak into the training split, so this is a pilot and may not be cited as a rate. Margins are log-odds rather than posterior confidence, because the shared prior logit cancels and the same evidence therefore reports the same margin at any prior. Synthesising spectra instead of using real ones was tried and refuted: r = −0.106 with 35 % of pairs ranked in the opposite direction. (v0.69.4, 2026-08-09)
  • 2-D evidence separates what shift lists cannot. Over 2 840 held-out regioisomer pairs, predicted HMBC correlations separate 2 812 / 2 840 = 99.0 %. The reason is not that “2-D is more informative”: a regioisomer already has a different predicted ¹³C list, but a shift is a continuous value requiring accurate prediction and held-out ¹³C MAE (3.44 ppm) exceeds DP4’s own scale (2.306 ppm), while an HMBC cross-peak is a near-binary observation whose position follows from topology alone. Stated as a bound, not an accuracy: this measures separation in predicted-correlation space, since only 218 of 64 723 NMRShiftDB2 records carry HMBC blocks (0.34 %) — far too few for a held-out score. Real spectra have missing weak 3-bond correlations, noise and overlap. HMBC topology is blind to stereochemistry exactly as HOSE codes are. (v0.68.8, 2026-08-08)
  • A candidate generator, held separate from the arbiter. spectroscopy/case/enumeration.py returns every constitutional isomer consistent with an HRMS molecular formula and, optionally, per-carbon hydrogen counts from an HSQC — exhaustive rather than sampled, so “the true structure is not in this list” is a real statement. Validated against a known answer: C₄H₁₀O yields exactly the seven textbook isomers. HSQC carbon–hydrogen counts prune hard (C₃H₆O 9 → 1, C₄H₁₀O 7 → 1, C₅H₁₂O 14 → 1, C₄H₈O₂ 122 → 15). Cost is exponential and measured (5 heavy atoms 0.05 s, 6 → 0.52 s, 7 → 22.2 s); past its bound it refuses with a named cause and returns no candidates, because a truncated list presented as complete invites the conclusion that the true structure is absent when it was simply never reached. The budget is a node count, not a timeout, so an identical call gives an identical result on any machine. It is a generator, never an arbiter; stereochemistry is deliberately not enumerated. (v0.69.1, 2026-08-08)

Detection, quantitation, and what a peak table is allowed to claim

Section titled “Detection, quantitation, and what a peak table is allowed to claim”

A series of measured corrections to peak detection and to the labels drawn from it. Several re-baselined published numbers; each is stated with what moved.

  • The ¹³C detection threshold sat at 1.4 σ. _detection_noise measured the baseline on the peak-free lower half of the sorted spectrum and then applied 1.4826 — the MAD-to-σ constant for a symmetric distribution — to half of one, returning 0.589–0.624× the truth on synthetic noise of known σ. The height gate is noise × 3.5, so the real threshold was 1.4× MAD, below any conventional limit of detection: four ¹³C acquisitions saturated the 220-line ceiling and were reported as up to 188 distinct signals. Reflecting the peak-free half about the median rebuilds a symmetric distribution of the same width, so the ordinary constant applies and no new one is invented (0.998–1.080× the truth across peak densities 0–10 %). Scored against a corpus that plants lines of known height in noise of known width: recall on separated ≥10 σ lines 35/35 unchanged, false positives per ¹³C acquisition 207 → 19, saturating acquisitions 4/23 → 0/23. ¹H is untouched. The A/B envelope that should have caught this was itself encoding the bug — it pinned |live − captured| ≤ tolerance, treating a fall and a rise as one event — and now judges a drop by whether any captured peak above the quantitation floor is no longer found. (v0.72.0, 2026-08-27)
  • Detection and quantitation are different claims. On a real ¹³C acquisition, 47 of 55 signals stood between 1.8 and 6.3× the baseline noise, and every one of the 27 sitting above 220 ppm — outside the range carbon-13 shifts occupy at all — was among them. Six sevenths of the table was the detection floor and nothing on screen said so. Every signal now carries snr and quantifiable, and the table splits into “signals you can measure” and “detected, but not strong enough to measure”. On that acquisition the split is 8 against 47, and the 8 carry 98 % of the signal; all 27 implausible shifts fall in the detected-only half with no rule mentioning ppm. Signal-to-noise is computed from the observed apex, never a fitted amplitude. (v0.73.1, 2026-08-27)
  • Say what the analysis cannot separate. Two lines closer than the detector’s minimum separation come back as one signal — a hard limit of the method, not a confidence caveat. Every result carries resolution_hz and every signal width_hz. Width is the only observable that shows a merge happened (a merged pair fits 3.3–4.5× the true width against 1.0–1.3× for a single line), but it is deliberately not flagged automatically: across 588 fitted lines the width distribution has a long tail (p90 4.96×, p99 222×) and 14 % exceed 3×, so an automatic “this may be two lines” would cry wolf on one signal in seven. The reported figure is the detector’s minimum separation; the true limit is coarser, at about four linewidths. (v0.72.1, 2026-08-27)
  • Model selection, not more fitting, is what separates merged lines. spectroscopy.peaks.deconvolve asks whether two Lorentzians explain a window better than one by more than noise allows, requiring the 99.9th percentile of χ² with 3 dof. Over 40 seeds at 150.9 MHz: a single line kept as one 40/40; a pair 1.0 linewidths apart found 39/40; pairs at 1.5–4.0 linewidths 38–39/40; median error in recovered centres 0.00 Hz. So pairs are recovered from one linewidth apart where the detector needs four, with no false splitting. It became answerable only once the noise estimate was unbiased (v0.72.0) — a test against a noise level itself wrong by 40 % decides nothing. Wired into the offline desktop path only, where it adds a column and takes nothing away; gsd_peak_pick is shared by SpectraCheck, qNMR and the verifier and a change there is a change to every peak list this platform has produced. A signal below the limit of quantitation is never split. (v0.73.0, 2026-08-27)
  • Two components on top of each other are one line. Three strong carbons were reported as fitting into more than one line, their extra components sitting 0.2 to 0.8 Hz from the main line against linewidths of 1 to 3 Hz. Nothing is resolvable at that separation: a real line is slightly asymmetric, and least squares models that asymmetry with a second component almost on top of the first — which the earlier Voigt guard could not see, because a Voigt is symmetric. Components closer than half a linewidth are now rejected, and the broadest component sets that scale, not the narrowest, or an artefact defines the floor meant to catch it. Signals fitting as more than one line fall from 10 of 173 to 3 across six acquisitions, while recovery of genuine pairs stays at 11–12 of 12 at every separation from 1.0 to 4.0 linewidths. (v0.74.2, 2026-08-29)
  • A peak now carries the width of its own line, not of the group it sits in. width_ppm is the extent of the whole cluster, wings included — on real acquisitions 3 to 16 times the width of the line inside it — so any threshold written in linewidths and fed that number is wrong by that factor. Peaks now also carry line_width_ppm and line_width_hz (FWHM), published index-aligned with the peak list; Hz is None when the spectrometer frequency was not supplied rather than fabricated from a default field. It is a measurement, honest only where the trace is sampled (1.06× true at 16 points per FWHM, 3.50× at 2), and no threshold is drawn on it: on real isolated lines the natural spread within one spectrum already reaches a median 1.11× for ¹³C and 1.87× for ¹H, so no fixed multiple separates a merge from a genuinely broad resonance. (v0.74.3, 2026-08-29)
  • “Broad” is a comparison, and it had nothing to compare against. Multiplicity chose s vs br s by testing the cluster extent against a fixed 0.12 ppm. That extent is non-monotonic in line separation — it climbs, collapses over roughly 0.65–0.85 FWHM as the flank walk terminates at the saddle between two lines, then recovers — and a fixed ppm floor is nucleus-blind, so 70 % of ¹³C single components (76 of 109) were reported as broad, including the CDCl₃ solvent lines, about the sharpest in the sample. A signal is now broad when its FWHM reaches 2.5× the median line width of its own spectrum, a multiple taken from where the measured distribution separates; ¹³C is insensitive to any choice from 2.0 to 5.0, and that insensitivity is the evidence a boundary belongs there. Measuring against the spectrum’s own lines also makes the test self-normalising, so the same spectrum sampled more finely no longer changes its own label. Re-baselined visibly: across 296 corpus signals 104 labels move, only between s and br s — no d, t or m changed and no peak list changed length. This does not make br s a merge detector. (v0.74.4, 2026-08-29)
  • A fitted line may not be wider than the data that constrained it. The first change in this series to the shared fitter, so it reaches every product that reads a peak list. _fit_single_with_model bounded σ at the width of the peak’s own fit window, and since both fitted models define fwhm = 2σ that permitted a line exactly twice as wide as the window it was fitted on — while amplitude is the integral to infinity, so a fit pinned there reports mostly area that was never measured. Measured across the corpus before the change: 35 of 428 fitted lines were wider than their own window, ratios running to exactly 2.000. The ceiling is now fwhm ≤ span — not a tuned number, but the point at which a fit stops claiming to have measured something wider than it looked at; 393 of 428 lines were already inside it. Re-baselined visibly: peak counts unchanged on 22 of 22 acquisitions, nine acquisitions change their widest line with the runaways coming down (43.0 → 21.5 Hz, 33.8 → 19.8), and all 137 tests across the fifteen fitter suites still pass. fit_window_ppm is now recorded on every fitted peak, because the width of the data a fit actually saw cannot be recovered afterwards — with an area that integrates to infinity, this field is what separates measured from extrapolated. What this did not fix, stated plainly: summed fitted areas still over-recover — 1.831× the true trace integral on the reference acquisition, down from 2.150×. Capping a component’s width limits how much of its neighbour’s envelope it can absorb; it does not stop it absorbing. (v0.75.4, 2026-09-07)

Multiplicity and J couplings — what a label is allowed to claim

Section titled “Multiplicity and J couplings — what a label is allowed to claim”
  • The plausible-J window was not a chemical bound. A first-order label was reported when every adjacent-line spacing fell inside a window whose upper edge was 60.0 Hz — not a statement about ¹H–¹H coupling, since no proton–proton scalar coupling comes close. When the detector clustered two unrelated signals, the pair was handed back as a doublet carrying a J that cannot exist: it shipped on two of nineteen golden fixtures at J = 45.6 Hz and J = 43.5 Hz, in compounds containing neither fluorine nor phosphorus. The new edge was measured, not guessed: dumping all 773 adjacent-line spacings from 126 multiplets splits into two regimes with nothing between them — real couplings top out at 18.06 Hz, and the next values are those two spurious pairs. The edge is now 30.0 Hz, agreeing with the independently tuned ¹H window that answers the same physical question. Any edge in [19, 43] Hz classifies all 177 golden peaks identically, which is the point rather than a caveat: a threshold on a fitted quantity belongs where the density is zero. A second, latent defect surfaced while narrowing it — the dd hypothesis tested J1 + J2 and J1 − J2 against a coupling ceiling, which rejects real dd patterns the moment the ceiling means what it says. (v0.69.11, 2026-08-22)
  • A multiplet must have the intensities of one, not just the spacings. Labels were chosen from inter-line spacings alone, so any two lines whose gap fell inside the J window became a doublet however unequal they were. In the quantifiable half of one real ¹³C table, four signals read d with line ratios of 186:1, 101:1, 237:1 and 62:1 — a strong carbon beside a weak line a quarter of a ppm away, with the “coupling” being the gap. A doublet is one nucleus split by one neighbour: its two lines are equal. All four now read m. The check accepts three intensity families because all three occur — binomial (spin-½ neighbours), trinomial (deuterium is spin-1), and uniform (n different couplings give all lines equal, so a dd is 1:1:1:1). The same spectrum’s DMSO-d₆ septet measures 1.0 : 3.1 : 6.2 : 7.3 : 6.2 : 3.1 : 1.0 and survives unchanged; omitting the uniform family had reclassified quinine’s H10 vinyl ddd as m. Conservative by construction — an unmatched pattern falls through to m, and nothing gains a more confident label than before. (v0.74.0 / v0.74.1, 2026-08-29)
  • A withdrawn label must take its coupling with it. Gating only the first-order label left two routes to a coupling open, and four signals kept 9.5–28.9 Hz couplings after their doublet labels had been withdrawn — the label corrected, the number beside it left standing, which is the same claim in a different cell. A spacing between two independent carbons is a shift difference; those four now report no couplings at all. (v0.74.1, 2026-08-29)

read_processed_spectrum() reads Bruker pdata/N/1r and JCAMP-DX. It is deliberately not just a second loader: processed data arrives already apodized, phased, baseline-corrected and referenced by someone else, and a quantitation claim over a spectrum an unknown operator processed sits in a different evidentiary class from one over a FID MolTrace processed itself. So the reader records the difference — domain='frequency' / processed_by='vendor', a processing_provenance block (window function, line broadening, both phase corrections, baseline mode) and a referencing block stating the basis on which the ppm axis was established, read from the file and never assumed. When it cannot be established the axis is point index and established=False says so, rather than a plausible-looking scale. Refusals name their cause: a JCAMP file with a Hz axis but no carrier frequency is refused rather than returned as Hz labelled ppm, which would put every ¹³C peak at ~190× its true shift with nothing to signal it. Verified against ground truth rather than structure — every structural assertion would still pass with a subtly wrong axis. nmrML is deliberately not built: no fixture exists and nmrglue has no reader, so it would be an unverifiable parser written from a spec. (v0.66.0, 2026-08-07)

Offline analysis in the desktop installation

Section titled “Offline analysis in the desktop installation”

The MolTrace desktop shell runs a packaged copy of the science service as a local child process and analyses an acquisition with no network call at all. The engines are the platform’s own — verify_structure, the GSD fitter, classify_peaks, the integration methods — so the guarantees above apply unchanged. What follows is what is specific to running them on a chemist’s machine.

  • It reads a spectrum off the machine it runs on. The service previously took the ppm axis and intensities as arrays — 3.6 MB of JSON per spectrum to analyse a file already sitting on the same machine as the caller. fid.open takes a path instead: 132 bytes on the wire, 200 acquisitions in 1.9 s. Reading is bounded by the caller’s own authority — the service runs as that user and can open nothing they could not already open; what the transport credential buys is that nothing else on the machine can ask. It returns multiplets, not peaks, which is correctness rather than presentation: on one acquisition 30 fitted lines resolve to 8 multiplets, and a raw line list shown to a chemist misstates how many signals the spectrum contains. Every result carries its own limits emitted by the engine rather than added by an interface, so a caller that receives bare numbers cannot render them bare. (v0.70.0, 2026-08-24)
  • It opened 7 of 23 real acquisitions. open_spectrum called the processed-spectrum reader and nothing else, so an acquisition carrying only raw time-domain data was refused — with the reader’s own developer-facing sentence shown to a chemist. Measured across every acquisition in the repository, 7 carry a processed spectrum and 16 carry only the FID, and every 400–600 MHz acquisition was among the refused. It now falls through and opens 23 of 23. Which reader produced the numbers is reported (processing = instrument or moltrace), because a spectrum computed locally uses this application’s phasing and baseline settings rather than the spectrometer’s, and a difference from the chemist’s own printout would otherwise look like a defect. A saturated detector now says so — four instrument-processed ¹³C acquisitions came back at exactly the 220-line ceiling and were reported as 68 to 188 distinct signals; saturated marks the count as a floor rather than a finding. (v0.70.1, 2026-08-27)
  • The spectrum is drawn, not just tabulated. A peak table with no trace beside it cannot be checked, and on the acquisitions where the detector saturates the table is exactly what should not be trusted. The reduction is the part that had to be right: a min/max envelope, not every Nth point — on a 524 288-point acquisition a stride left the tallest peak at 19.9 % of its real height, while keeping each bucket’s minimum and maximum reproduced it at 100 %. Both edges are kept because negative excursions are how a chemist sees bad phasing. Highest ppm first, since an NMR spectrum is read right to left. The window follows the signal and states the trim when a meaningful part is off screen — a trimmed axis that does not admit it is a claim that nothing lies outside. (v0.71.0, 2026-08-27)
  • A signal’s share is the trace under it, not a sum of fitted areas. A sharp CH₃ reported 1.4 H where the molecule has 3, and the premise behind the investigation was wrong: the line’s own fitted area is right to 2.1 %. What was wrong is what it was divided by. The desktop fits one independent peak per detected apex with nothing requiring the fits to partition the spectrum, so the denominator was inflated 2.15× — one peak carried only 29.7 % of its reported area inside its own window. relative_area is now the integral of the baseline-subtracted trace under each signal, over windows built from each line’s own centre and fitted half-width; the windows tile, so nothing is counted twice and the shares sum to 1.000000. Against C₄H₈O as 3/2/1/2: 1.394 / 2.202 / 1.682 / 2.722 → 2.951 / 2.054 / 1.008 / 1.987. Three hypotheses died first, each by measurement rather than argument — under-sampling, a width-dependent area bias in the fitter, and the σ floor; the load-bearing bound was the upper one, fixed separately in v0.75.4. (v0.74.12, 2026-09-07)
  • Windows bounded by neighbouring centres, because edge clamping can invert. Integration windows do not merely overlap — a sharp line’s window can sit entirely inside a broad neighbour’s, and clamping edge-to-edge then places the boundary past the contained window’s own upper edge and inverts it to negative width, integrating to exactly 0.0. Two signals across the corpus were silently given a share of zero, one of them a residual-solvent line the proton-count readout subtracts from its denominator. Both checks running at the time were true and blind: “overlaps: 0” holds because a zero-width window overlaps nothing, and “shares sum to 1.000000” holds because the sum is normalised by itself; the invariant that catches it is that every listed signal holds a positive share. Each window is now bounded by the midpoints to its neighbouring centres, which cannot invert because a multiplet’s centre always lies strictly between them — disjoint by construction rather than by repair. Zero-share signals go from 2 to 0 of 264. The hand-rolled trapezoid was also replaced by the platform’s own integrate_sum — measured equivalent (0.054 H vs 0.055 H worst error), so the duplicate goes but nothing improved. (v0.74.13, 2026-09-07)
  • The structure feeds back into the measurement. Areas were reported as a share of the listed signals because a proton count needs a denominator only a structure supplies. Given a structure, structure.inventory makes each share a proton count, scaled to the NON-LABILE hydrogens — OH, NH and SH exchange with the solvent, and normalising a spectrum that never showed them against a count that includes them puts every other signal low by the missing fraction (for glycerol, 37.5 %). The total is not evidence and the panel says so in its own alert: the scale is chosen so the measured signals add up to the structure’s hydrogen count, so the bottom row agrees for any structure with that many non-exchanging hydrogens. What carries information is the per-signal residual. The rows are regions of the shift axis, not assignments, and the readout discloses what it left out of the denominator — on one reference acquisition 48.7 % of the listed area was classified as solvent or impurity, with the sentence “if any of these is your compound, every count below is wrong.” Two categorisers meet here and only one answers the question: classify_peaks decides whether a signal is the compound at all, while the inventory buckets by chemical class via categorize_peak; handing the first one’s words to the second matched none of its category sets and returned 0.0 for every observed row while the total stayed right — a table that looks computed and says nothing. This is not a fifth test; the verifier remains the sole arbiter and it never moves a verdict. (v0.74.10, 2026-08-31)
  • A contaminant whose shape contradicts its label is moved back to the compound. classify_peak matches one line by position and cannot see a multiplet. Water in CDCl₃ is a singlet at 1.56 ppm, and a nine-line coupled multiplet centred on 1.583 is not water however well its tallest line matches — that call was taking 27 % of one acquisition’s area out of every proton count derived from it. A contaminant call is now overturned when the signal shows more resolved lines than that specific contaminant can produce. Three ways the first version was wrong, each found by measuring: it reclassified ¹³C (a deuterated solvent’s own carbon is split by coupling to deuterium — CDCl₃ is a 1:1:1 triplet — and the table’s patterns are proton patterns, so it promoted solvent carbons into the analyte; now ¹H only); it counted lines that were not apart (a six-line “multiplet” spanning 36.3 Hz with a 43.0 Hz linewidth has no resolved coupling at all); and one suspicion about over-picking was simply wrong, and measuring it kept a correct promotion that reasoning would have thrown away. Four signals move across the whole 264-multiplet corpus. Only ever toward the compound; nothing is demoted — being wrong here hands a chemist a signal to explain, which they can see and judge, while the opposite silently deletes evidence. Confined to the offline module: the shared classifier still decides. (v0.74.11, 2026-09-03)
  • All four of the verifier’s tests now run offline. Two of the four abstained on every check the desktop ever made, because nothing on the machine could supply their evidence — and they said “none was supplied”, which reads as a missing file a chemist could go and provide when what was missing was the ability to accept one. ms.open reads a processed centroid peak table through the platform’s own parse_ms1_peak_text, so ms_molecule_match runs: on a reference acquisition, 2 of 4 tests → 3 of 4, and the verdict moves from inconclusive 0.562 to consistent 0.925. The asymmetry is rendered, because it is not guessable from the result — this test’s weight is proportional to the matched fraction, so a candidate whose predicted pattern matches nothing scores zero significance and moves the posterior by exactly +0.000; a reader seeing one candidate lifted and three unchanged would reasonably conclude the three were ruled out, and they were not. (v0.75.1, 2026-09-07)
  • The fourth test, and the only one that can argue back. nmr2d.open reads a processed cross-peak table and hsqc_2d_ranges runs. It is qualitatively unlike the mass-spectrum test: it weights itself by how many correlations the structure predicts rather than by how many the data matched, so a structure whose predicted rectangles are empty scores negative at full weight. Measured against ethanol’s own HSQC — ethanol inconclusive 0.415 → consistent 0.830, aspirin 0.227 → inconsistent 0.036, benzene 0.500 → inconsistent 0.119. The page states that contrast directly beneath the mass-spectrum caveat that says the opposite. The experiment filter is part of the correctness: COSY correlates proton to proton and HMBC spans more than one bond, so either would mark down a correct structure for evidence that was never about it — only HSQC and HMQC reach the verifier, and a mixed table reports what was set aside and why. Axis order is asserted rather than assumed, because reversing it makes every correlation fall outside every rectangle and marks the true structure down — a silent failure that looks like a confident refutation. (v0.75.2, 2026-09-07)
  • What the offline screens were saying that was not so. An audit of three newly landed offline features returned eighteen defects sharing one shape: the numbers were computed correctly and then described wrongly. The reference-library hit rate on screen was measured on a different task — a clean leave-one-out record against record, while the function queries with the detected multiplet centres of a real acquisition. Re-measured at the operating point the function actually runs at: 3 of 15 first and 4 of 15 inside the top five, against a displayed figure roughly double that. The old measurement was not sloppy — it was rigorous about the wrong operating point, which is the harder error to see, because a clean number still reads as evidence. Also corrected: a ranking that reported a stable order while the winner changed (3 of 19 corpus cases trade first and second without moving the sorted margin); a maximal 100 % margin printed where only one candidate matched anything at all, which is the absence of a contest rather than a robust ordering; a refusal that printed the chemist’s own filesystem path; and four caveats that described the previous product, including a page that warned its confidence must never pick a winner and then sorted the list by it. (v0.74.8, 2026-08-30)
  • Two of the four checks could never run here, and the screen blamed the data. A test carrying zero weight looked like a test that ran: the assignments test switches itself off once an acquisition’s unexplained integral reaches 25 %, which is true for 9 of 11 acquisitions in this corpus with a stated structure — so a verdict a chemist reads as resting on two tests rested on one. The screen now gives the figure and the threshold. And the first version of that sentence named a cause it had not measured: it said solvent peaks account for the unexplained integral, and across the seven zeroed cases solvent-labelled area accounts for the whole of it in none of them, with one acquisition zeroed at 30 % unexplained and no non-compound area at all. A cause is a measurement, not an inference from the example you happened to open. (v0.74.9, 2026-08-31)
  • A truncated Bruker parameter file could hang the reader that takes uploads. A guard added to the desktop reader had been applied to only one of the two readers; nmrcheck.fid, the one behind the upload routes, still called the vendor parser unguarded at two places — both inside try/except Exception, which is exactly what made them look safe, because a try cannot catch a loop. Measured through the archive entry point, an archive whose acqus was truncated to 90 % never returned, while 40/50/60/70/100 % completed in about six seconds. Severity stated exactly rather than alarmingly: those handlers run in the threadpool, so unlike the desktop service this does not stop the process — it permanently consumes one threadpool worker per bad upload, against a bounded pool. The check is imported, not reimplemented, and runs before the call and outside the surrounding except, which would otherwise have turned a named refusal into a silent fallback. The guard test constructs its truncation rather than sampling one, because cutting at a fraction only hangs when the cut lands inside an array parameter’s values — 90 % hung on one fixture and was fine on another, which is how this defect was once called refuted on the strength of a single offset that happened to work. Verified to hang on 5 of 5 Bruker acquisitions in the repository before the guard. (v0.74.7, 2026-08-30)

A chronological summary; see each subsection above for substantive detail.

VersionDateHeadline
v0.75.42026-09-07A fitted line may not be wider than the data that constrained it
v0.75.22026-09-07HSQC cross-peak table reaches the verifier’s fourth test offline
v0.75.12026-09-07Mass-spectrum peak table reaches ms_molecule_match offline
v0.74.132026-09-07Integration windows bounded by neighbouring centres (they could invert)
v0.74.122026-09-07A signal’s share is the trace under it, not a sum of fitted areas
v0.74.112026-09-03A contaminant whose shape contradicts its label returns to the compound
v0.74.102026-08-31Proton counts from a supplied structure, scaled to non-labile hydrogens
v0.74.92026-08-31A zero-weight test named honestly, and a cause that was measured
v0.74.82026-08-30Eighteen offline-readout defects: right numbers, wrong descriptions
v0.74.72026-08-30Truncated-parameter-file guard on the reader that takes uploads
v0.74.42026-08-29”Broad” measured against the spectrum’s own median line width
v0.74.32026-08-29Per-line FWHM published beside the cluster extent
v0.74.22026-08-29Components closer than half a linewidth are one line
v0.74.12026-08-29A withdrawn multiplicity label takes its coupling with it
v0.74.02026-08-29A multiplet must have the intensities of one, not just the spacings
v0.73.12026-08-27Detection and quantitation split into two tables (snr, quantifiable)
v0.73.02026-08-27Model-selection deconvolution separates lines the detector merges
v0.72.12026-08-27resolution_hz / width_hz — saying what cannot be separated
v0.72.02026-08-27The ¹³C detection threshold sat at 1.4 σ (207 → 19 false positives)
v0.71.02026-08-27The spectrum is drawn — min/max envelope, not every Nth point
v0.70.12026-08-27Offline analysis opens 23 of 23 acquisitions, not 7
v0.70.02026-08-24Offline service reads a spectrum by path and returns multiplets
v0.69.112026-08-22The plausible-J window becomes a chemical bound (60 → 30 Hz)
v0.69.52026-08-13A deploy cannot silently ship the 16-molecule seed predictor
v0.69.42026-08-09The arbiter’s own false-confirmation pilot (9.1 %, n = 11)
v0.69.12026-08-08Exhaustive constitutional-isomer enumerator (generator, never arbiter)
v0.68.82026-08-08HMBC separates 99.0 % of the regioisomers ¹³C shifts get wrong
v0.68.72026-08-08Ambiguity-weighted evidence in PredictionBoundsTest
v0.68.62026-08-08AssignmentsTest prices each atom on its own conformal interval
v0.68.52026-08-08A retracted match-tolerance defect, and the real one beside it
v0.68.22026-08-08Decoy generator + measured false-confirmation rate (39.5 %)
v0.68.12026-08-08The verifier weighs evidence by the conformal interval
v0.68.02026-08-08Mondrian split-conformal prediction intervals (90.03 % / 90.61 %)
v0.66.02026-08-07Processed-spectrum ingest + first held-out accuracy measurement
v0.64.12026-08-07Reference table made deployable (193 MB / 47 s → 14 MB / 1.1 s)
v0.64.02026-08-07Shift predictor switched on; prior-fallback share 22.6–44.4 % → 0 %
v0.16.02026-06-07Retrieval-augmented reasoning over the spectral index
v0.10.02026-06-06NUS reconstruction (IST baseline + JTF-Net)
v0.9.02026-06-05Solvent/impurity expert system
v0.8.22026-06-03POST /spectrum/retrieve endpoint (similarity retrieval contract)
v0.8.12026-06-03FAISS HNSW spectrum retrieval (vector + set similarity)
v0.8.02026-06-03Multi-test ASV verification scorer
v0.7.92026-06-01NMRNet wrapper reworked (local-first, conformer-ensemble uncertainty)
v0.7.82026-06-01NMRNet chemical-shift prediction wrapper + HOSE-code fallback
v0.7.62026-05-31Karplus validation corpus scaled to 18 molecules
v0.7.52026-05-30Boltzmann conformer-population weighting (sugar blind-spot fix)
v0.7.42026-05-30Opt-in Haasnoot–Altona Karplus + honest negative result
v0.7.32026-05-28Karplus vicinal-³J validation corpus + accuracy gate
v0.7.22026-05-28Opt-in Karplus 3J refinement for Layer 40 vicinal couplings
v0.7.12026-05-28Multiplet J-coupling → unified-confidence evidence layer
v0.7.02026-05-28Multiplet analysis with GSD-enhanced J-coupling
v0.6.22026-05-28100-fixture real-instrument HMDB corpus
v0.6.12026-05-28Per-peak QC metrics + legacy parity
v0.6.02026-05-28Validation framework + strict promotion gate cleared
v0.5.02026-05-27Algorithm semantics + envelope unification
v0.4.02026-05-27Prompt 3 GSD backend launch