Core Concepts
This page is the shared vocabulary for the platform. Link back here whenever another page mentions confidence scores, ICH thresholds, audit trails, or Bayesian optimization.
Confidence scores
Section titled “Confidence scores”A confidence score is MolTrace’s estimate that a structure assignment, impurity flag, or interpretation is supported by the available evidence. A score of 87% is not certainty.
Not every percentage on a MolTrace screen is the same quantity, and the word “calibrated” does not apply to all of them. Three are worth telling apart:
- A verifier verdict confidence is a posterior from an explicit Bayesian log-odds combination over the four independent structure tests, starting from a stated prior. It is auditable end to end — each test’s score, significance and diagnostic are recorded — and the thresholds that turn it into
consistent/inconclusive/inconsistentare fixed at 0.80 and 0.20. - A DP4 candidate share is a relative ranking across the candidates supplied, not a calibrated probability, and the API says so in the result itself (
probability_is_calibrated,probability_basis). Ranking three candidates gives shares that sum to 1.0 whether or not any of them is right, and a candidate that matched no observed peak scores exactly zero — so a 100 % share can mean the only candidate that matched anything, rather than a confident win. See LC-MS/MS Annotation → What a ranking actually explains. - A prediction confidence is derived from the same uncertainty→significance mapping the verifier uses, rather than from a second notion of certainty invented for the purpose. A consequence worth knowing: a prediction at exactly the reference uncertainty scores 0.870, not 1.0.
Because these are different scales, a single numeric cut-off applied across them does not mean the same thing on each. Read a confidence beside the basis reported with it.
Use confidence as a triage signal, not an approval decision. Users must review low-confidence results and any contradiction flag before approving an export. (clarified v0.67.0 / v0.68.9, 2026-08-08)
ICH thresholds
Section titled “ICH thresholds”ICH thresholds are regulatory limits that determine when impurities, solvents, elemental impurities, or mutagenic risks must be reported, identified, qualified, or controlled. MolTrace uses the selected guideline, jurisdiction, dose context, and project metadata to calculate the thresholds shown in Regentry.
For ICH Q3A-style impurity work, the threshold category typically answers a different question:
| Threshold type | Example threshold context | What the user must do |
|---|---|---|
| Reporting | 0.03% | Include the impurity in the submission report. |
| Identification | 0.05% | Provide a structural identification or justified rationale. |
| Qualification | 0.05% or 1 mg/day | Provide toxicological safety data or documented qualification. |
Final threshold values depend on the applicable guideline, drug substance or product context, maximum daily dose, jurisdiction, and customer SOPs. Do not publish a numeric table until a regulatory owner has verified it against the official source document.
Audit trails
Section titled “Audit trails”An audit trail is the chronological record of what happened to a project, file, analysis, interpretation, report, or approval. It should show who acted, what changed, when it happened, the software version, and what evidence supported the decision.
MolTrace treats raw file upload, processing, review decisions, contradiction resolution, export generation, and approval as audit-relevant events. Audit events are permanent records and should not be editable or deleted by ordinary users. Export access should live in Settings -> Audit Log -> Export.
The backend supporting these controls ships in moltrace.spectroscopy.audit: a tamper-evident, cryptographically chained audit trail (a SHA-256 hash chain catches insertion, deletion, and reordering; a keyed HMAC catches any content tampering), electronic-signature primitives designed per 21 CFR Part 11 §11.50 / §11.70, append-only log sinks (in-memory, durable JSON-Lines, or a PostgreSQL / AWS QLDB backend), periodic chain verification, a configurable retention floor (default 7 years), and capture of AI model-weight checksums so any AI-assisted result stays reproducible and traceable. These controls support a customer’s 21 CFR Part 11 workflows — MolTrace does not claim the product is itself compliant; full computerized-system validation remains the customer’s responsibility. (v0.11.0, 2026-06-06)
Bayesian optimizer
Section titled “Bayesian optimizer”The Bayesian optimizer recommends the next experiment by balancing what is already known with what still needs to be learned. Instead of testing every condition in a grid, it builds a probabilistic model of the reaction landscape and suggests conditions likely to improve yield, selectivity, impurity profile, or another objective.
Human-in-the-loop review
Section titled “Human-in-the-loop review”MolTrace assists expert scientists; it does not replace them. Any AI-supported interpretation that affects a report, regulatory decision, or submitted evidence requires human review and sign-off before export.
Human review is required for contradiction resolution, impurity confirmation, and export sign-off. The platform should preserve enough evidence to support alignment with AI governance expectations, including model versioning, explainability, review status, and performance monitoring.
Raw FID archive
Section titled “Raw FID archive”A raw FID archive is the original instrument evidence bundle preserved before processing. The FID is the raw time-domain signal collected by an NMR instrument; processing applies Fourier transform and other steps to produce the frequency-domain spectrum users inspect.
MolTrace stores the original FID read-only and performs processing on a derived copy. Supported source evidence includes Bruker fid plus acqus, Agilent fid, and JCAMP-DX .jdx. Keeping the immutable archive allows reprocessing, inspection, and audit review without losing the source evidence.