Interpreting Optimization Results
Use this page after an optimization run has enough experimental evidence to recommend next conditions.
What to review
Section titled “What to review”- Recommended condition and expected objective gain
- Uncertainty around the recommendation
- Constraint satisfaction, including impurity thresholds
- Convergence indicators
- Human notes explaining accepted, modified, or rejected recommendations
Reading Phase C predictions, route scores, and forward checks
Section titled “Reading Phase C predictions, route scores, and forward checks”Phase C adds three reviewable records to a reaction project. All three are advisory decision support, and each carries its own limitations on the record rather than in a footnote. (v0.63.0, 2026-07-23)
A yield-prediction run
Section titled “A yield-prediction run”A run is a surrogate fit on this project’s own completed experiments — no cross-project or vendor corpus is involved — evaluated over the candidate conditions you submitted.
- The backend is named on the run and on every prediction. Read it before you read the numbers: it tells you whether a heavy model or a lightweight fallback produced them.
- How many experiments trained it is recorded, as is whether you asked for verified outcomes only. A run fit on unverified outcomes is not a verified-evidence run, and should never be presented as one.
- Each prediction carries a mean and a spread, plus per-prediction warnings. Warnings are where degraded conditions are disclosed — for example, a condition value that appears nowhere in the training experiments, which the featurisation encodes as zero and flags rather than silently absorbing.
- Capability provenance is stored verbatim with the run, so a reviewer months later can reconstruct which engine decision produced the number.
A route score
Section titled “A route score”The route you supply is scored by the frozen safety and green-chemistry engines — the same engines the rest of the module uses, not a separate scoring heuristic.
- Reagents are screened alongside the reaction steps.
- An atom economy that cannot be weighed is refused, not estimated.
- An unrecognised or unknown risk severity ranks worse than critical, so a route can never score milder than what was actually screened.
- A Mermaid rendering of the route tree is persisted with the score, so the reviewed artifact and the drawn tree cannot drift apart.
- Every record is flagged as requiring human review. A score ranks options for review; it is never a safety determination or a synthesis instruction.
A forward check
Section titled “A forward check”A forward check takes a prediction that already exists — from a model, a colleague, or an external tool — and cross-checks it against the frozen engines before anyone acts on it. Flagged chemistry is surfaced rather than hidden, unrecognised risk again ranks worse than critical, and the record is flagged as requiring human review.
The capability readout
Section titled “The capability readout”When a Phase C surface is empty or missing, check the readout rather than guessing. Per capability it reports enabled (flag), available (dependency probe), active (usable for a real decision), the missing modules, and a named reason. Two readings are intentional: the yield GNN reads inactive without a supplied benchmark-gate artifact even when its flag is on and its dependency is installed, and the retrosynthesis and forward_prediction capabilities read unavailable rather than degrading, because the generative half has no lightweight fallback. That does not affect the route-score and forward-check records above, which are produced by endpoints that need no extra at all.