Changelog
This page tracks recent and forthcoming platform updates. Previously consolidated entries have been distributed into the feature pages where their substance lives:
- NMR analysis, multiplet detection, J-coupling refinement, chemical-shift prediction, structure verification, solvent/impurity classification, NUS reconstruction, spectrum retrieval, and retrieval-augmented reasoning → SpectraCheck → NMR Interpretation.
- Quantitative region integration (Sum / Edited Sum / Peaks) and qNMR purity (internal-standard & PULCON) → SpectraCheck → qNMR Quantification.
- MS models — CSI:FingerID, METLIN retention-time corroboration, and DP4-AI candidate fusion → SpectraCheck → LC-MS/MS Annotation.
- The AI model lifecycle — model registry + inference router, datasets pipeline, evaluation harness, LoRA fine-tuning, closed-loop feedback, active learning, and MLOps monitoring + deployment gate → SpectraCheck → AI Model Lifecycle.
- Audit trail + GxP controls supporting 21 CFR Part 11 → Core Concepts → Audit trails.
- ICH + FDA impurity engines — Q3A/B thresholds, Q3C(R8) residual solvents, Q3D(R2) elemental impurities, M7(R2) mutagenic impurities, and the FDA CPCA nitrosamine classifier — plus the deterministic-first Phase 0 foundation → Regentry → ICH Guidelines and Overview.
- The dossier impurity-assessment workflow — the unified
impurities/assessendpoint, the engines wired behind the dossier assessment endpoints (with product dose + route on the dossier), and the nitrosamine cumulative-risk rollup → Regentry → Impurity assessment. - Per-user dossier access control + data isolation — owner-scoped reads/writes (migration 0015), the cross-module bridge gates, and privileged surveillance → Regentry → Access control.
- AI-decision governance for EU GMP Annex 22 (draft) — the tamper-evident per-dossier decision log, the HITL gate, and auto-recording from CPCA / M7 / Q3D → Regentry → AI-decision governance.
- Readiness-report rehydration + content-hash provenance → Regentry → Report Generation.
- Enterprise SSO (OIDC federation, JIT, enforce-SSO), SCIM 2.0 provisioning + soft auto-deprovisioning, MFA & passkeys (TOTP + WebAuthn/FIDO2) with step-up re-authentication, policy-as-code authorization (centralized PDP + deny-by-default baseline), and rotating-refresh session hardening with reuse detection → Security Policy → Access controls and → Session management.
- Argon2id credential hashing (rehash-on-login), KMS envelope encryption for sensitive fields (BYOK seam), the secrets-provider seam + CI secret-scanning gate, and TLS / HSTS + security response headers → Security Policy → Encryption, → Secrets management, and Deployment & Hosting → Transport security.
- Tamper-evident audit-chain hardening (hash-chained
audit_events, HMAC-signed checkpoints, signed high-water tail-truncation guard, admin verify / anchor + reconciliation alerts) → Security Policy → Audit trail integrity. - 21 CFR Part 11 e-signature hardening (server-attributed identity §11.100, content-bound
signature_digest§11.70, durable JSON / HTML manifestation §11.50, step-up re-auth §11.200) → Compliance → Electronic signatures. - ALCOA+ hardening for controlled records (queryable
reason_for_change, soft-delete reversibility, strict raw-vault immutability, immutable-by-design audit tables) → Compliance → ALCOA+. - GAMP 5 / CSA validation lifecycle — regenerable validation package per system release (traceability + IQ / OQ / PQ-from-CI + change-control + signature manifestations), CI evidence-ingestion seam, validated-state change-control gate → Compliance → Computer system validation.
- The opt-in GSD experimental-backend rollout policy (per-call telemetry → aggregate rollup → flip-readiness verdict → per-tenant graduation) → Deployment & Hosting → GSD experimental backend rollout.
- Secure-SDLC CI gates (SAST / SCA / IaC scanning with CRITICAL-blocks and triage SLAs), the signed supply chain (CycloneDX SBOMs, keyless SLSA provenance, verify-at-deploy), and zero-trust CI hardening (SHA-pinned actions, least-privilege workflow permissions, IaC posture-drift gate) → Security Policy → Secure development lifecycle, → Supply-chain integrity, and → Zero-trust CI hardening and IaC posture.
- API abuse protection — the token-bucket rate limiter (
429+Retry-After+X-RateLimit-*), the request-body size cap (413), and the WAF edge runbook → Security Policy → API abuse protection and Backend / API Contract → Rate limiting. - Coordinated vulnerability disclosure and penetration testing —
/.well-known/security.txt, the VDP with safe harbor, the pen-test program, threat model, and findings register → Security Policy → Vulnerability disclosure and penetration testing and Backend / API Contract → Public unauthenticated endpoints. - SIEM sink seam and security detections (impossible travel, privilege escalation, cross-tenant access, audit-chain break) plus the admin alert / detection-run endpoints → Security Policy → Security monitoring and detections and Backend / API Contract → Admin / operations endpoints.
- The incident-response program — severity model, roles, containment levers, runbooks, the breach-notification deadline engine, and the endpoints IR depends on → Security Policy → Incident-response program and Backend / API Contract → Endpoints that incident response depends on.
- The SOC 2 / ISO 27001 control-evidence register (a self-assessment of control coverage — neither a SOC 2 report nor an ISO/IEC 27001 certificate is held), the inherited-vs-operational control boundary, and the Trust Center with its sub-processor register → Compliance → Framework control-evidence register, → What the register does not claim, and → Trust Center.
- The 2026-07 infrastructure migration off Render to Google Cloud — the backend on Cloud Run (
moltrace-backend, projectmoltrace-prod,us-central1, scale-to-zero), Cloud SQL for PostgreSQL 16 on a private IP over Direct VPC egress, Cloud Storage / Secret Manager / Cloud KMS / Artifact Registry / Cloud Build, keyless CI/CD deploys via Workload Identity Federation, and the Vercel frontend behind the same-origin/api/backendproxy → Deployment & Hosting → Platform hosting (Google Cloud), → CI/CD keyless deploy, and Security Policy → Infrastructure and hosting. - The Cloud Storage raw-FID vault backend that keeps the write-once ALCOA+ vault working on serverless (create-only
if_generation_match=0writes, SHA-256 verification, bucket retention + versioning as the WORM mechanism) → Deployment & Hosting → Raw-FID vault on serverless and Security Policy → Write-once raw-evidence vault on serverless. - Backup and disaster-recovery resilience — what is backed up, the RTO / RPO objectives, the restore-integrity verifier, and the restore-drill cadence → Deployment & Hosting → Backup & disaster recovery, → Recovery objectives, and → Restore-integrity verification.
- Repho Phase C heavy-ML reaction engines as default-off governed guests — the six engines and their site-installed extras, the surfaces that need no heavy dependency (yield predictions, route scores, forward checks), the deliberately unwired generative paths, the absent SDL execution surface, and the capability honesty readout → Reaction Optimization → Phase C engines and the capability readout, → Enabling the optional Phase C engines, and → Reading Phase C predictions, route scores, and forward checks.
- GDPR data-subject requests and the right to erasure — the library-only DSAR / erasure planner (Art. 15 discovery, Art. 17 per-store plan, Art. 12(3) deadline), the classified personal-data map with its four dispositions, the pseudonymisation-is-never-erasure invariant, and the plainly stated limit that identity cannot be erased from the immutable audit ledger (crypto-shredding is a documented seam, not a capability) → Privacy Policy → Data-subject requests and the right to erasure and Compliance → Privacy: data-subject requests and residency.
- Data residency — single-region hosting with no tenant region pinning (no EU-pinned deployment available), and the honest per-item status of what pinning would require → Privacy Policy → Data residency.
- The security-documentation accuracy sweep from the retired Render deployment to Google Cloud — the corrected shared-responsibility posture, and the two postures the migration genuinely improved (private-IP Cloud SQL with no public interface, keyless Workload Identity Federation deploy authority in place of stored deploy-hook secrets) → Security Policy → Documentation accuracy sweep after the Google Cloud migration.
- The three open Medium findings the migration opened in the security findings register — the per-instance rate limiter on multi-instance Cloud Run, the missing container-image vulnerability scan, and the database RPO gap against the documented ≤ 5 min objective → Security Policy → Open findings from the infrastructure migration, → Rate limiting across multiple Cloud Run instances, Deployment & Hosting → Container-image vulnerability scanning gap, and → Where the stated RPO is not met today.
- The chemical-shift predictor’s reference table as deployed state — the 495 215-assignment NMRShiftDB2 index that replaced a silently-degraded 16-molecule seed table, the Welford re-encoding that made it shippable (193 MB / 47 s → 14 MB / 1.1 s), and the four guards that stop a deploy shipping the seed predictor → SpectraCheck → Reference table coverage and Deployment & Hosting → Reference table staging.
- MolTrace’s own held-out accuracy measurement (¹³C MAE 3.44 ppm / ¹H 0.332 ppm), the correction that withdrew the earlier “sharper than the error model” claim, and processed-spectrum ingest with recorded processing provenance → SpectraCheck → Measured accuracy and → Processed-spectrum ingest.
- Distribution-free conformal prediction intervals (90.03 % / 90.61 % measured coverage), the finding that the predictor’s σ is differentially mis-scaled, and the switch of the verifier’s evidence weighting onto the interval → SpectraCheck → Prediction intervals and → What the verifier weighs.
- The first measurement of how often a wrong structure is confirmed — 39.5 % for the ¹³C-shift-list layer through DP4, 9.1 % as a pilot for the multi-test arbiter, stereochemistry recorded as a hard capability boundary — plus ambiguity-weighted evidence, per-atom conformal rulers, the HMBC separation bound (99.0 %), and the exhaustive constitutional-isomer generator → SpectraCheck → What the verifier weighs.
- Peak-detection and quantitation honesty — the ¹³C detection threshold that sat at 1.4 σ (207 → 19 false positives per acquisition), the split of a peak table into measurable and detected only, published resolution and per-line widths, model-selection deconvolution, and the fitter bound that stops a line being reported wider than the data that constrained it → SpectraCheck → Detection, quantitation, and what a peak table is allowed to claim.
- Multiplicity and J-coupling bounds — the plausible-J window narrowed from 60 Hz to a measured 30 Hz chemical bound, the intensity test that stops a 237:1 line pair being reported as a doublet, and the rule that a withdrawn label takes its coupling with it → SpectraCheck → Multiplicity and J couplings.
- Offline analysis in the desktop installation — reading a spectrum by path, drawing the trace as a min/max envelope, a signal’s share as the trace integral over tiling windows, proton counts scaled to non-labile hydrogens, contaminant reclassification by shape, and all four verifier tests running with no network call → SpectraCheck → Offline analysis and LC-MS/MS → Mass-spectrum evidence offline.
- Relative integrals disclosed as ratios rather than proton counts, and a DP4 ranking that reports the coverage its error figure was computed over → qNMR → Relative integrals are not proton counts and LC-MS/MS → A candidate ranking reports what it explains.
- The AI/ML engine seam — the single import boundary that made a recorded confidence a computed one, the removal of the hard-coded
0.82default, and the link that makes approving a model change what the router serves → AI Model Lifecycle → The engine seam. - The gold set that now exists (21 records, a regression sentinel and not an accuracy claim), the fail-closed treatment of an unmeasured safety metric, conformal coverage as a promotion metric, and the refusal that names its own blocking measure → AI Model Lifecycle → The gold set that exists, → A safety metric with no denominator, and Reaction Optimization → A refused candidate names its own measure.
- Regulatory knowledge-corpus provenance — search that excludes reviewer-rejected records, sources that are superseded rather than edited, and the two-person dataset promotion conveyor gated through to a canary → Regentry → Regulatory knowledge corpus and AI Model Lifecycle → Corpus governance.
- The rule that a measured level with no applicable limit is undetermined, never a pass — the ICH Q3C route gate on the dossier path, the Option 1 / Option 2 limit labelling, the assumed-Q3D-route disclosure, and the qualification that now travels into the CTD bundle with the rows it qualifies → Regentry → ICH Guidelines.
- Ordered rule-set versions published beside the content they address, and the version-currency catalogue that lets an installation tell whether its science matches the workspace’s → Regentry → Rule-set versions and Backend / API Contract → Version currency.
- Signed offline entitlement statements, their two-level key handling, and the guarantee that no commercial term can make a regulated record unreadable → Security Policy → Offline entitlement statements and Compliance → A licence term may never make a regulated record unreadable.
- Denial bodies minimised on every client path rather than only in the browser (
PUBLIC_CODES9 → 14, the 401/403 pair hoisted into the OpenAPI contract), and WebAuthn verification widened to several exact origins under one RP-ID → Security Policy → Denial bodies, → One RP-ID, several exact origins, and Backend / API Contract → The 401 / 403 body. - Build provenance and serving behaviour — the code revision that reached every regulated result as
"unknown", the installer pinned by digest, Cloud Build retries, and the FID pipeline moved off the event loop with a persisted report cache → Deployment & Hosting → Reference table staging and build provenance and → Serving behaviour under load. - Owner-scoped compound-registry reads (with shared as a setting), colleague review of FID runs, and the usage-event instrumentation behind the ROI figures → Backend / API Contract → Access-scope changes and → Usage events behind the ROI figures.
- API contract surface (new endpoints, audit events, admin actions, cross-cutting request gates) → Backend / API Contract, including the Phase C reaction optimization endpoints.
- Documentation, brand, and navigation history → Brand Identity & Navigation → Shipped.
The canonical, version-pinned source of truth for backend releases remains moltrace_backend/CHANGELOG.md.
Recent updates
Section titled “Recent updates”Everything through v0.78.0 (2026-08-22) has been distributed into the feature pages listed above. That includes the whole v0.64.0 – v0.78.0 range: the shift predictor and its deployed reference table, MolTrace’s first held-out accuracy and conformal-coverage measurements, the false-confirmation work and the evidence-weighting changes it drove, the peak-detection and quantitation corrections, the offline desktop analysis path, the AI/ML engine seam and the gold set the promotion gate now scores against, the regulatory “undetermined, never a pass” series, ordered rule-set versions, signed offline entitlement statements, and the total sanitization of 401 / 403 denial bodies. Nothing is pending redistribution. New releases land here first, then move into the relevant feature page when stable.
Two notes for anyone reconciling these pages against the upstream changelog:
- Several entries here are corrections to what this documentation previously said, not only additions. The largest are the 100-spectrum evaluation gold set described on the AI Model Lifecycle page — which did not exist, so the dominance gate it describes was in practice gating nothing until a real 21-record set was built — and the “HOSE-code / NMRShiftDB2 fallback” on the NMR Interpretation page, which for most of that period was running against a bundled 16-molecule seed table. The word calibrated has also been removed from the general definition of a confidence score in Core Concepts, because it does not apply to all three of the numbers a user sees.
- One upstream reference is stale and is not corrected here. The v0.77.0 entry cites migration
0050for the device identity key; the migration on disk is0051_device_identity_key, and0050is an unrelated change. These pages publish0051.