AI Model Lifecycle
SpectraCheck’s AI predictions are produced by versioned, governed models — never a single hard-coded predictor. This page catalogs the model-lifecycle backend that makes every AI-assisted result reproducible, auditable, and continuously improving: a model registry and inference router, a licence-aware datasets pipeline, an evaluation harness with a dominance gate, domain fine-tuning, closed-loop feedback, active learning, and the MLOps monitoring + deployment gate.
Across every stage the deterministic structure verifier remains the sole arbiter of correctness — models sharpen ranking and routing, they never override the science. The release timeline at the end gives chronological context.
Backend capabilities
Section titled “Backend capabilities”Model registry & inference router
Section titled “Model registry & inference router”- Append-only model registry —
moltrace.spectroscopy.ai.registryversions every model artifact (role, semantic version, SHA-256, training-data lineage = dataset-snapshot hash + row count, lifecycle status) with deterministic content addressing. Entries are immutable, duplicate ids are rejected, and lifecycle changes append as events; promotion toproductionauto-retires the incumbent for the same (role, nucleus). One pluggable store backs PostgreSQL in production and SQLite in tests. (v0.12.0, 2026-06-07) - 5-layer inference router —
InferenceRouter.predict_shifts_routedresolves each atom LoRA → NMRNet → HOSE fallback (the fine-tuned adapter is used only when a production adapter exists for the nucleus and the conformer-ensemble uncertainty is within its validated confidence band). Every prediction carries a complete, deterministicmodel_versionsmap that feeds the audit trail verbatim — one prediction, one immutable provenance record — so a result is reproducible bit-for-bit from the registry + lineage, and a reviewer sees which artifact produced each number and why one layer was chosen. (v0.12.0, 2026-06-07)
Public datasets pipeline
Section titled “Public datasets pipeline”- Licence-aware ingestion + version pinning —
moltrace.spectroscopy.data.datasets_pipelineturns the canonical public datasets (NMRShiftDB2, HMDB, BMRB, MassBank EU, GNPS, METLIN, QM9-NMR, 2DNMRGym, AIST SDBS) into a deduplicated, version-pinned corpus. Each source pins its upstream version + licence and is content-hashed — a changed upstream hash raisesUpstreamChangedErrorrather than being silently accepted — and non-redistributable sources (SDBS, METLIN) are never written into a redistributable corpus. Normalization is RDKit-standardised SMILES + InChIKey with dedup by(InChIKey, spectral-hash); invalid records are quarantined with reasons, not dropped. (v0.13.0, 2026-06-07) - Leakage-free frozen holdout —
freeze_splitsproduces deterministic, seeded train/val/test splits grouped by InChIKey skeleton so a molecule never straddles splits. The test split is the sacred evaluation holdout — experimental-only, checksummed, and returned as a hash-exclusion set;assert_training_excludes_holdoutis the guard every fine-tune must call. Computed (QM9) records are train-only and are dropped when they share a molecule with the eval set. (v0.13.0, 2026-06-07)
Evaluation harness & dominance gate
Section titled “Evaluation harness & dominance gate”-
The ten-metric dominance gate —
moltrace.spectroscopy.eval.harnesspromotes a model version only when its full metric vector dominates the incumbent on a frozen, checksum-locked 100-spectrum gold set (60 NMRShiftDB2 + 20 HMDB + 20 in-house). The ten metrics: top-1 / top-3 structure accuracy, ¹H / ¹³C shift MAE, ECE, false-confirmation rate, retrieval recall@k, error-vs-uncertainty AUROC, perturbation robustness, reviewer-agreement rate, and end-to-end latency p50 / p95. No model ships on a single improved number, and the safety-critical metrics (false-confirmation rate, calibration) may never regress. That rule held only where the metric was measured:evaluatereturned 0.0 forfalse_confirmation_ratewhenever the gold set held no wrong structures — the best possible value on a zero-tolerance safety metric, earned by measuring nothing. It is nowNone, withn_wrong_structurescarrying the denominator, and absence is fail-closed in all three directions (candidate missing, incumbent missing, neither reporting). See A safety metric with no denominator. (v0.14.0, 2026-06-07)Correction (v0.69.6, 2026-08-14). The 100-spectrum gold set described above did not exist.
gate_for_cihad never been called outside its own tests, because there was no gold set for it to score against — so the dominance machinery documented here was, in practice, gating nothing. A real frozen set now exists and is much smaller than the one described: 21 records, built byscripts/build_nmr_gold_set.pyfrom the repository’s own real NMRShiftDB2 Bruker fixtures (expected peak lists joined to molblocks from the bundled NMReDATA source), checksummed attests/fixtures/nmr_gold_set/gold_set_v1.json, both nuclei. See The gold set that exists below. -
CI promotion gate —
gate_for_ciwraps evaluate + dominance against the production incumbent with exit codes0promotable /1not promotable /2gold-set checksum drift;GoldSet.assert_integrity()aborts the run if the holdout size or checksum drifts, so it can never be silently contaminated. (v0.14.0, 2026-06-07)
Domain fine-tuning — LoRA
Section titled “Domain fine-tuning — LoRA”- LoRA domain adapters with K-fold validation —
moltrace.spectroscopy.ai.finetunetrains a LoRA adapter (low rank r=8–16; train the adapter, freeze the base) on accumulated reviewer-validated in-house spectra, validates it with K-fold cross-validation (GAMP 5 Appendix D11), and registers it with full lineage. Each run freezes an immutable, content-addressed training snapshot that subtracts the holdout at freeze time — a leaked record raisesHoldoutLeakageError. Therun_idis a path-independent content address of the full manifest (hyperparameters, per-fold + aggregate MAE, snapshot hash, base id, code git sha, adapter SHA-256); GPU-hours and cost are logged. (v0.17.0, 2026-06-07) - Gated registration, never auto-promoted —
register_if_eligibleevaluates the dominance gate on the frozen gold set and registers the adapter ascandidatealways, promotes toshadowonly if it does not regress the incumbent, and never auto-promotes toproduction(human sign-off required). Registration refuses if the snapshot’s gold checksum disagrees with the live gold set — no adapter is registered without complete lineage. The trainer is injectable, so the whole pipeline runs deterministically on a CPU-only host; adapter weights are cached out of git, with only the SHA-256 + lineage persisted. (v0.17.0, 2026-06-07)
Hyperparameter optimization, calibration & contradiction detection
Section titled “Hyperparameter optimization, calibration & contradiction detection”- Bayesian hyper-parameter optimization —
optimize_hyperparametersruns Optuna (seeded TPE) over the five LoRA knobs, budgeted to ~10 trials (a budget, not a sweep). Each trial runs a full K-fold CV scored by a mean-CV-MAE + calibration objective, every trial is logged, and the resultingHPOStudyis reproducible (its id excludes wall-clock time). The best config feedsfinetune_lora, so a trained adapter is traceable back to the search that chose its hyper-parameters. (v0.18.0, 2026-06-07) - Confidence calibration as a promotion gate — temperature / Platt calibration heads, with ECE enforced as a first-class gate: a candidate whose gold-set ECE exceeds
max_eceis not promotable even if it dominates on accuracy. A model that is more accurate but lies about its confidence does not ship. (v0.18.0, 2026-06-07) - Contradiction detection — a deterministic detector plus a trained, calibrated classifier flag internal inconsistencies: no single consistent structure (from verifier verdicts), cross-modal NMR-vs-MS disagreement, and intra-spectral impossibilities (integration vs proton count, multiplicity vs coupling neighbours via the n+1 rule, a shift outside its plausible window). It complements, never replaces, the deterministic verifier — surfacing to the reviewer and feeding the active-learning queue. (v0.18.0, 2026-06-07)
Leak-proof cross-validation
Section titled “Leak-proof cross-validation”- GroupKFold across every CV loop — the fold partitioner is group-aware: whole molecules (the InChIKey connectivity skeleton) or an explicit sample / batch key are assigned to a single fold, so a physical sample’s multiple scans never straddle the train and eval sides of a fold. This closes the cross-batch leakage that makes naive per-spectrum K-fold report optimistic, untrustworthy metrics, and is threaded through fine-tuning, HPO, and the contradiction trainer. With no grouping signal it reproduces the historical per-record split byte-for-byte, and every run manifest records its
cv_strategy(group_kfold/kfold, group key, group count) for auditable lineage. (v0.18.1, 2026-06-07)
Closed-loop feedback & RLHF
Section titled “Closed-loop feedback & RLHF”- In-app feedback capture —
moltrace.spectroscopy.feedbackrates every AI output (predicted shift, proposed structure, peak label, purity call) with thumbs up/down + an optional free-text correction + a structured reason, persisted as an immutable, content-addressed event carrying the exactmodel_versionsthat produced it. Corrections fan out to the labeled-example store (training data), bare overrides to the active-learning queue, and usage / override analytics roll up where the model is weakest. (v0.19.0, 2026-06-07) - RLHF reward model — advisory, never authoritative — corrections + accept/reject signals become Bradley-Terry preference pairs and train a deterministic reward model that advisorily re-ranks the reasoner’s candidates and prioritizes the annotation queue. Verifier supremacy is enforced structurally: a verifier-accepted candidate always ranks above a rejected one — reward only orders within a verdict class. (v0.19.0, 2026-06-07)
- A/B champion vs challenger, no auto-deploy — deterministic sticky hash routing (shadow = never served / canary = controlled served fraction), dominance-gated promotion with reviewer-acceptance + override guards, human sign-off required before the registry is ever mutated, and instant rollback modeled as a routing-layer canary-kill that never touches the append-only registry. (v0.19.0, 2026-06-07)
- Structured “why was it wrong?” taxonomy — a closed 7-value reason vocabulary (
wrong_shift,wrong_multiplicity,wrong_structure,missed_impurity,wrong_integration,calibration_off,other) shared by the in-app control, the API contract, and the feedback engine, so override analytics segment by exactly where the model is weak. Surfaced through the AI-inference API — see Backend / API Contract. (v0.19.1, 2026-06-07)
Active-learning loop
Section titled “Active-learning loop”- The flywheel —
moltrace.spectroscopy.ai.active_learningturns every reviewer override into labeled training data and actively chooses the most informative spectra for scarce expert attention.capture_overriderecords full provenance in one append-only, idempotent record (raw-FID hash, processed spectrum + content hash, the producingmodel_versions, the AI output, the human correction, reviewer id + timestamp);disagreement_scoreblends vote-split on the top-1 structure, predicted-shift variance, and confidence spread across model variants;build_annotation_queueranks candidates by disagreement, greedily de-dups near-identical spectra, and slices to a labeling budget. (v0.20.0, 2026-06-08) - Self-triggering retrains + loop yield —
retraining_triggerfires on a monthly schedule or a volume of new labels and wires the fine-tune chain (snapshot → LoRA → gated registration);loop_yield_metricsreports labeled examples / month, the override-rate trend over consecutive windows (a falling trend is direct, auditable evidence the model is improving), and accuracy lift per retrain — emitted to the audit trail for the operations dashboard. (v0.20.0, 2026-06-08)
MLOps: monitoring, drift & the deployment gate
Section titled “MLOps: monitoring, drift & the deployment gate”- Continuous drift monitoring + lineage dashboard —
moltrace.spectroscopy.opsruns four monitors and returns a worst-ofok | warn | breachreport: input drift (population-stability index of nucleus / field / solvent + molecular-weight PSI vs the training snapshot — a large PSI flags new chemistry the model never saw), confidence drift, override-rate drift (reusing the loop-yield trend), and latency (p50 / p95 vs SLO), paging an injectable alerter on each breach. The registry-backed lineage dashboard reads the registry as the source of truth — per production model: its version, training-snapshot hash, gold metric vector, promotion record + supersession, and current live drift status. (v0.21.0, 2026-06-08) - Fail-closed deployment gate —
run_deployment_gateallows a deploy only if all four checks pass: dominance (no safety regression), audit-chain integrity (provenance intact), test-suite-green, and data-leakage (the training snapshot is bound to the gold checksum and its record hashes are disjoint from the holdout). Every input defaults to the failing state, so an under-specified call is blocked — it fails closed. It is wired into CI as adeployment-gatejob thedeployjob depends on, so a regression in the release-control logic fails CI before it can ever let a bad model reach production. (v0.21.0, 2026-06-08)
The engine seam — where a number becomes provenanced
Section titled “The engine seam — where a number becomes provenanced”For most of the period catalogued above, MolTrace had two AI/ML layers and they did not touch each other. One was the science described on this page — inference router, append-only registry, dominance-gated harness, LoRA, active learning, the verifier-constrained reward model. The other was 76 REST routes of AI governance — model cards, training runs, calibration assessments, drift alerts, canary deployments — in which every number arrived from the caller. POST /ai/predictions read confidence_score out of the request body and, when it was absent, recorded a hard-coded 0.82; downstream — the calibration record, the review queue, the drift alert — that number was indistinguishable from a measured one. The governance layer was not wrong. It was a correct, audit-grade system of record built before the engine it was meant to record.
- One import boundary.
nmrcheck/ai_engine_adapter.pyis the single seam. The stores keep zero engine imports — they receive results, never engines — and a test asserts it, because a store that reaches past the adapter is a store whose numbers have no guaranteed provenance. Three rules it enforces: fail loud, degrade recorded (an engine that cannot run raises and the route answers503; it never substitutes a caller-supplied number for a computed one); provenance is mandatory (every result carriesmodel_versionsas artifact id → SHA-256, and a result with an empty map is refused inside the adapter rather than stored with a gap); and lazy import, so the ~800-route app still builds without the ML extras. Requests for shift prediction and candidate ranking may no longer supplyconfidence_score,uncertainty_json,ood_statusormodel_versions— the engine derives them, and a submitted value would be recorded as if a model had produced it. (v0.67.0, 2026-08-08) - Confidence is now on the arbiter’s own scale. A prediction’s confidence is derived from the verifier’s uncertainty→significance mapping rather than inventing a second notion of certainty. Two consequences are deliberate: a prediction at exactly the reference σ scores 0.870, not 1.0; and the 35 ppm median ¹³C σ that reached production before the reference table landed scores 0.143 and reports
out_of_domain. That is the defect this closes — the uncertainty was always honest and always reported, it was simply never aggregated, so a per-atom warning nobody read let an abstention stand in for a prediction indefinitely. The0.82default is removed: a service with no engine wired records no confidence rather than a plausible-looking one, and says so in its warnings. (v0.67.0, 2026-08-08) - Approving a model now changes what the router serves. The router resolves what to serve from the append-only registry, and nothing in the product ever wrote those tables — so approving a deployment candidate flipped a status and changed nothing about which artifact actually answered a prediction. The governance surface and the serving path were two systems that both called their subject “the deployed model”. They are now linked, not merged: merging would either destroy the registry’s append-only guarantee (lifecycle changes are appended, so a promotion is reconstructable and a retirement cannot be edited away) or require rewriting a table 35 routes read. Approval is the single writer of production status and carries an optional
registry_promotionblock — role, nucleus, semantic version, dataset-snapshot hash, row count — because those are facts only the promoter holds and are not derivable from the artifact row; omit it and the artifact is approved without changing what serves traffic, which is a different decision and now says so.GET /ml/model-artifactsreports the registry’s live status, so what a reviewer sees as deployed is what the router would resolve. Promotion is refused, with the approval left standing, when the artifact carries no content hash or when a semantic version is reused against different bytes; the approval commits before the promotion is attempted, so a failed promotion leaves the router serving the incumbent. Migration0044adds the link, nullable and not backfilled, because NULL is the truthful state of every artifact approved before this existed. (v0.67.1, 2026-08-08)
The gold set that exists, and what it is not
Section titled “The gold set that exists, and what it is not”scripts/build_nmr_gold_set.py joins the repository’s own real NMRShiftDB2 Bruker fixtures into a frozen, checksummed 21-record set across both nuclei, and moltrace-eval-gate scores the production path on it — each record’s real FID through read_fid → verify_structure → predict_shifts, never a reimplementation, honouring the measured rule that synthesised spectra do not transfer.
The first minted vector ships as benchmarks/nmr/incumbent_metric_vector.json and pins, among others, reviewer_agreement_rate = 0.571 — the verifier confirms 12 of 21 correct structures on real spectra. That number is now a floor a commit cannot silently lower.
It is named a regression sentinel, not an accuracy claim, and the CLI says so on every run:
- Several gold molecules sit in the shipped reference table’s training data, so absolute MAE here partially measures memorisation. Held-out numbers remain
eval/shift_accuracy.py’s job. - With no decoy records yet,
false_confirmation_rateis unmeasured (n_wrong_structures = 0) — the CLI warns on that gap every run rather than letting it read as a perfect score. - The default
--mode regressionpasses an unchanged model; dominance’s strict-improvement requirement stays for--mode promotion. A candidate that drops a measured safety metric still hard-fails, and cross-reference-table comparisons are refused rather than reported as regressions.
Wired into the deploy workflow after the reference-table staging step, informational until the incumbent is re-minted on CI hardware and reviewed. (v0.69.6, 2026-08-14)
A safety metric with no denominator is not a perfect score
Section titled “A safety metric with no denominator is not a perfect score”false_confirmation_rate is safety-critical with zero tolerance. It returned 0.0 whenever the gold set held no wrong structures — the best possible value, earned by measuring nothing — and that value then flowed into every consumer as though it had been observed.
- Nullable alone would have been worse than the bug.
dominatesskipped any metric absent from either side, so a candidate that simply stopped reporting the metric promoted over an incumbent that measured it, and the metric did not appear in the deltas at all — converting a wrong record into no record. Absence is therefore fail-closed for safety-critical metrics in all three directions, emitted as a delta withmeasured=Falseandregressed=True, so every existing consumer ofregressedblocks without an edit. Ordinary metrics keep skipping, and that split is pinned by a test so the two rules are not merged later. - Three consumers moved in the same commit, or the guard would have been half-applied. A/B testing and ops monitoring both rendered an unmeasured metric as a “safety-critical regression”, which sends an operator to the model when the fix is the gold set. And the fine-tuning path required every metric key while the emitter omits unmeasured ones — so it returned
Nonefor every real snapshot, and the caller readsNoneas “no incumbent” and promotes without comparing anything. That fail-open predated the change and widened with each nullable metric added; a singleNULLABLE_METRICSlist is now the source of truth for which are optional. - Open, deliberately. The API-side
dominance_verdictstill passes when neither side reports a safety-critical metric, so the platform’s two promotion gates diverge in that one direction. And a gold set with zero wrong structures now blocks every promotion — the intended forcing function — but no reason string yet tells the operator to fix the gold set rather than the model. (v0.69.4, 2026-08-09)
Conformal coverage as a promotion metric
Section titled “Conformal coverage as a promotion metric”Conformal calibration measured coverage; nothing gated on it. The gold metric vector now carries conformal_coverage_deficit (largest shortfall below the stated target) and conformal_interval_width (mean half-width), both lower-is-better. Coverage is expressed as a shortfall precisely so “more is better” can never apply to it — over-coverage is not a win, it is paid for in width.
Adding a metric changes what can ship, so two failure modes closed in the same change. An unmeasured metric must not read as a perfect one, so both fields are None when unmeasured and non-finite values are dropped. And a third safety-critical metric would refuse every promotion during rollout, because the gate refuses when one evaluation reports a safety-critical metric and the other does not — so coverage ships with a zero tolerance but deliberately stays out of the safety-critical set until evaluations report it across the board, with the docstring recording the condition for promoting it. _is_number now rejects NaN: it passed isinstance while every comparison against it was False, so a NaN metric looked present to the asymmetry check while registering neither a regression nor an improvement — a gate reporting itself as applied while guarding nothing. The width tolerance is 0.16 ppm = 3 % of the measured 5.371 ppm count-weighted mean half-width, the same fraction of scale the existing MAE tolerance uses, rather than a new convention. (v0.69.3, 2026-08-08)
The gate names the measure it actually blocked on
Section titled “The gate names the measure it actually blocked on”The corpus conveyor reuses the existing fail-closed dominance gate rather than growing a second one — but that gate wrote its refusals in prose fixed to the reaction model, so a blocked corpus candidate rendered “Measure treated as the blocking one: Citation support recall” directly above “Safety-flag recall is missing or out of range; failing closed” — two names for one measure, one of them belonging to a different model.
Nothing was computed wrongly: the correct value was compared, the fail-closed behaviour held, and the verdict was right in every case. What was wrong was the sentence a reviewer reads while deciding whether a refusal was a near miss or a missing number. It could not be fixed downstream — rewriting the string client-side is exactly the summarising the frontend contract forbids, and it would put the interface in the business of authoring why a candidate was refused. The gate now takes a blocking_metric_label, defaulting so the existing reaction call sites say exactly what they said before, and the conveyor passes its candidate’s own recorded measure. The decision itself is unchanged — same rule, same tolerance handling, same fail-closed behaviour; only the noun differs, and it is a display label, never a rename on the wire. (v0.69.2, 2026-08-08)
Corpus governance — the conveyor from an extracted record to a canary
Section titled “Corpus governance — the conveyor from an extracted record to a canary”The knowledge corpus feeds extraction models, so the promotion machinery on this page applies to it. Three changes finish that governance work; the corpus and search surface itself is documented in Regentry → Regulatory knowledge corpus.
- Promoting a dataset version takes two different people. Promotion is the point where curated records start training something, and one actor could previously do it by setting a status field. Approval now comes from the authenticated principal — never a value the caller supplies, the same rule e-signature applies to a signer identity — and a unique constraint on (dataset version, approver) makes the rule unbypassable rather than conventional: one human with two sessions still has one user id. A machine credential is refused outright, because it is not a person and two calls carrying it are the same principal. Patching status straight to approved now fails; if a single edit can promote, the rule was decoration. Scoped to promotion deliberately — two-person on every extracted record would not survive a 10 000-record corpus.
- The conveyor reaches a canary.
knowledge_deployment_candidatesbinds an approved dataset version to a trained artifact and its eval result — the object there was previously nothing to promote from. Each step is gated on the one before: a candidate needs a two-person approved version, a canary needs a passed gate, and promotion needs a canary even when the gate passed. A canary with no gate behind it would be a deployment mechanism wearing a governance label. The gate is not a second dominance rule — it calls the platform’s existing fail-closed gate, where the blocking dimension must not regress at all and a missing or non-finite measure blocks rather than being skipped; what differs by model is which measure is blocking, so it is recorded rather than assumed. This is a different conveyor from the reaction benchmark gate and from/ml/deployment-candidates; neither covered the knowledge corpus. - A record can show which passage it came from. Every extracted record carries
locators— citation label, page, section, paragraph, and the quoted passage — resolved from the record’s citations rather than copied onto it, so “show me where this came from” is one mechanism in the product instead of two that drift, and one query serves a whole list. Migrations0046and0047. (v0.68.10, 2026-08-08)
Release timeline
Section titled “Release timeline”A chronological summary; see each subsection above for substantive detail.
| Version | Date | Headline |
|---|---|---|
| v0.69.6 | 2026-08-14 | The accuracy gate finally has a gold set to gate on (21 records) |
| v0.69.5 | 2026-08-13 | Deploy-time staging + four guards on the reference table |
| v0.69.4 | 2026-08-09 | A safety metric with no denominator is not a perfect score |
| v0.69.3 | 2026-08-08 | Conformal coverage becomes a promotion metric |
| v0.69.2 | 2026-08-08 | The promotion gate names the measure it actually blocked on |
| v0.68.10 | 2026-08-08 | Two-person dataset promotion, gated conveyor to a canary |
| v0.68.4 | 2026-08-08 | Sources are superseded, never edited (revision binding) |
| v0.68.3 | 2026-08-08 | Search stops returning facts a reviewer rejected |
| v0.67.1 | 2026-08-08 | Approving a model now changes what the router serves |
| v0.67.0 | 2026-08-08 | The AI/ML engine layer is wired to the product; the 0.82 default removed |
| v0.21.0 | 2026-06-08 | MLOps: monitoring, drift detection & the fail-closed deployment gate |
| v0.20.0 | 2026-06-08 | Closed active-learning loop |
| v0.19.1 | 2026-06-07 | Structured feedback reason taxonomy → AI-inference API |
| v0.19.0 | 2026-06-07 | Closed-loop feedback: capture, RLHF reward model & A/B rollout |
| v0.18.1 | 2026-06-07 | Leak-proof GroupKFold cross-validation |
| v0.18.0 | 2026-06-07 | Bayesian HPO, calibration head & contradiction detection |
| v0.17.0 | 2026-06-07 | LoRA domain fine-tuning pipeline |
| v0.14.0 | 2026-06-07 | Evaluation harness: the ten metrics + dominance gate |
| v0.13.0 | 2026-06-07 | Public-datasets pipeline: ingestion, versioning & frozen splits |
| v0.12.0 | 2026-06-07 | Model registry + 5-layer inference router |