What happens after the decision: a learning surface over executed work.
Automation ends where measurement usually stops. The Observatory is the component of CORTEX that treats the aftermath of every automated act — the extraction that was accepted, the arbitration that was overridden, the case that was reopened, the document that had already been submitted — as the primary scientific material of the programme. It converts an append-only ledger of decision records into supervised, unsupervised, sequential and causal learning problems, and returns the resulting estimates to the operating surface as early warnings, precedent retrieval, duplicate containment and drift-governed retraining. Every figure below is produced by the same harness that produces the evaluation results in §7; none of it depends on an operational secret.
Headline observatory indicators
Four indicators summarise whether post-decision learning is compounding. Trends span the last ten evaluation windows under the canonical configuration.
The decision ledger — what is recorded, and why it is recorded that way
Learning from operations fails for a structural reason far more often than for a modelling reason: the record of what happened is incomplete, mutable, or labelled at the wrong moment. The ledger is designed against those three failure modes before any estimator is fitted.
Every terminal act emits an immutable record ρ = ⟨case, φ, decision, evidence-closure, agent-set, confidence-vector, actor, t, latency⟩. The record is the atomic unit of learning: nothing that was not recorded can be learned, and nothing recorded can be silently rewritten. Records are content-addressed, so a model trained on window w can be re-derived exactly from the same hashes months later.
A decision is not supervised at emission time but at settlement time, when the administrative outcome becomes observable (accepted, reworked, contested, cancelled, impeded). The Observatory therefore maintains a delayed-label buffer and only promotes a record into the training corpus once its label horizon has closed, which removes the optimistic bias that arises when partially settled cases are treated as successes.
For every automated proposal that a human overrode, both branches are retained: what the agent network proposed and what the operator decided. This pairing is the only source of counterfactual signal available without running experiments on live cases, and it is what makes disagreement learnable rather than merely auditable.
Each ranked-attention item carries whether it was opened, how long it was inspected, and whether it changed the outcome. Interventions that never change the outcome are noise in the operator's day, and the ranking model is retrained explicitly to suppress them.
Learning problems induced by the ledger
The ledger admits six distinct learning formulations. They are listed separately because they carry different evaluation protocols, different failure modes and different governance obligations — collapsing them into a single 'AI layer' is precisely the abstraction that makes such systems unauditable.
Regularised linear regression and gradient-boosted trees estimate the remaining hours to terminal state for every open case. The linear model is retained as the reportable estimator because its standardised coefficients are an auditable statement of causation-adjacent structure; the boosted model is retained as a performance ceiling to quantify how much accuracy interpretability costs (empirically 0.04 R²).
A calibrated classifier predicts P(rework | case state) at every phase boundary. Calibration is enforced with isotonic regression and verified against a reliability diagram: an uncalibrated 0.8 that behaves like a 0.55 is operationally worse than no prediction at all, because it silently reallocates human attention to the wrong cases.
Documents and cases are embedded and compared under cosine similarity with a locality-sensitive index. Three regimes are separated: byte-identical re-uploads, semantically identical documents differing only in scan geometry, and legitimately re-issued documents for the same asset. Only the first two are suppressible; the third must be preserved because it carries the legal history of the case.
Given an open problematic case, a k-nearest-neighbour query over the settled corpus returns the most similar historical cases together with the action sequence that resolved them and the resulting cycle time. The Observatory therefore answers 'what worked last time' with evidence rather than with a generated recommendation.
Case traces are modelled as sequences over the phase alphabet; low-likelihood traces under the fitted model are surfaced as process anomalies. Anomalies are not errors by definition — many are legitimate rare variants — so they are routed to DISCOVERY, where a human decides whether the variant is promoted into the process model or treated as a defect.
Where randomisation is impossible, the effect of an intervention (earlier document request, alternative source, pre-emptive escalation) is estimated with propensity-score matching over the settled corpus, and reported with its overlap diagnostics. Effects whose matched support is thin are published as insufficiently identified rather than as findings.
Predictive surface — estimating the case before it fails
Figures O1–O4 constitute the predictive surface. All estimates are produced under a temporally blocked protocol (train on window w, evaluate on w+1); random splitting is rejected because it leaks future process states into training and inflated R² by roughly 0.2 during harness calibration.
Failure taxonomy and resolution structure
Cases that did not succeed are the most information-dense objects in the corpus. Each is attributed to one of six failure classes from the evidence closure rather than from triage notes, and the empirical distribution of resolutions per class is what determines whether a class is automatable, escalatable, or structurally human.
F₅ (external source timeout) resolves under plain retry in 72% of occurrences and is therefore a scheduling problem, not an intelligence problem. F₄ (restriction conflict) escalates in two thirds of occurrences and must not be automated further: the correct engineering response is to reduce the latency of escalation, not its frequency. F₆ (stale catalogue) is dominated by auto-heal, which is the empirical signature of a working DISCOVERY loop. Classifying failures by their resolution structure — rather than by their symptom — is what turns a defect list into an engineering programme.
Duplicates, near-duplicates and legitimate re-issues
Duplicate work is the most under-measured waste class in document-heavy operations, because the second submission is usually indistinguishable from the first at the interface level. The Observatory separates three regimes and treats only the first two as suppressible.
- R₁ identical — byte-equal artefact re-submitted. Suppressed with operator confirmation; the confirmation supervises τ.
- R₂ near-identical — same content under different scan geometry, compression or crop. Detected in embedding space, never by hash.
- R₃ re-issue — a genuinely new document for the same asset. Must be preserved: it carries the legal history of the case, and suppressing it is a correctness failure, not an efficiency gain.
The suppression threshold τ is not chosen to maximise F1. It is chosen against an explicitly asymmetric cost: a false suppression destroys evidence and is bounded at 0.1% by policy, whereas a missed duplicate merely costs a redundant execution. τ is therefore set at the highest operating point satisfying the false-suppression bound, and the resulting recall (0.83) is reported as a consequence of that constraint rather than as a tuned result.
Preventive loop — turning post-mortems into pre-flight checks
The scientific value of the Observatory is not the retrospective dashboard; it is the closed loop in which a settled case changes how the next similar case is executed. Six mechanisms implement that loop, each with its own measurable lead indicator.
A case is flagged when its predicted residual time crosses the SLA horizon, not when the SLA is already breached. Lead time between flag and breach is itself a metric — a warning that fires two hours before the deadline is technically correct and operationally worthless.
Each failed case is attributed to a dominant cause class through the failure taxonomy F₁–F₆ using the evidence closure, not through free-text triage notes. Attribution stability across annotators is reported as Cohen's κ.
When a new case matches the signature of a previously failed cohort, the framework surfaces the failure pattern before execution and proposes the corrective action that closed the cohort. This converts post-mortem knowledge into a pre-flight check.
Duplicate submission is measured as a first-class waste category. Suppression is proposed, never executed silently: the operator confirms, and the confirmation is itself logged as supervision for the similarity threshold τ.
Every rework loop is priced in machine cost and human minutes, so process changes can be ranked by expected saving instead of by intuition. The ranking is recomputed each window and its realised versus predicted saving is tracked as a forecast-quality metric.
An action sequence that resolves a cohort above a pre-registered success threshold and sample size is promoted to a playbook, versioned, and thereafter proposed automatically. Demotion is symmetric: a playbook whose realised success falls below threshold is retired and its retirement recorded.
Drift governance and retraining discipline
A model that was correct in window w is an unvalidated hypothesis in window w+1. Drift is monitored on the feature distribution, on the label distribution and on calibration independently, because each degrades for different reasons and demands a different response.
Every deployed estimator ships a card stating training window, feature list, excluded features, known failure modes, calibration status and the population on which it must not be used.
Retraining is triggered by drift (PSI ≥ 0.25), by a calibration breach, or by the scheduled window boundary — never ad hoc. Each retrain is a versioned artefact with a promotion decision recorded against a champion–challenger comparison.
A challenger model runs in shadow for a full window, scoring live cases without influencing them, and is promoted only if it dominates the champion on both the accuracy metric and the attention-precision metric.
Ranked attention is audited for systematic starvation: no case cohort may remain unranked beyond a bounded number of windows, regardless of its predicted value, otherwise the ranking becomes a self-fulfilling backlog.
No model output is terminal. Overrides are unrestricted, always logged with the acting identity, and treated as high-value labels rather than as violations of the automation.
The Observatory stores the features required by the published models and nothing else; free-text and identifying content are excluded at ingestion rather than filtered at query time.
Reference implementation
The Observatory contract is deliberately small: ingest immutable records, refuse to train on unsettled labels, separate duplicates from re-issues, retrieve precedent, and let a governance verdict — not a schedule or an intuition — decide when a model is replaced.
class Observatory:
"""Post-decision learning surface over the CORTEX ledger."""
LABEL_HORIZON_DAYS = 14
DUP_TAU = 0.94 # cosine threshold for near-duplicate
PSI_RETRAIN = 0.25 # drift budget before mandatory retrain
def ingest(self, record: DecisionRecord) -> None:
"""Records are immutable; supervision arrives later."""
self.ledger.append(record.freeze())
self.pending_labels.schedule(record.case_id,
horizon=self.LABEL_HORIZON_DAYS)
def settled_corpus(self, window: Window) -> Frame:
"""Only cases whose outcome horizon has closed may train a model."""
return self.ledger.filter(window=window, settled=True)
def duplicates(self, case: Case) -> list[Match]:
"""Three regimes: identical, near-identical, legitimately re-issued."""
cand = self.index.query(case.embedding, k=25)
return [m for m in cand
if m.score >= self.DUP_TAU and not m.is_reissue]
def precedent(self, case: Case, k: int = 8) -> list[Playbook]:
"""Prescriptive retrieval: what resolved the most similar cases."""
neigh = self.index.query(case.embedding, k=k, settled_only=True)
return rank_by_realised_gain(p.playbook for p in neigh)
def governance_check(self, window: Window) -> Verdict:
psi = population_stability(self.reference, self.live(window))
calib = expected_calibration_error(self.scores(window))
if psi >= self.PSI_RETRAIN or calib > 0.05:
return Verdict.RETRAIN(reason=("drift" if psi >= self.PSI_RETRAIN
else "calibration"))
return Verdict.HOLD- O-I₁ — no record is mutated after emission; corrections are new records referencing the original.
- O-I₂ — no estimator trains on a case whose label horizon has not closed.
- O-I₃ — no suppression executes without a logged human confirmation.
- O-I₄ — every published figure is reproducible from (seed, window, manifest).
- O-I₅ — every model in production has a current card and a live drift verdict.
Three questions remain unresolved and are stated as such: whether delayed labelling can be shortened with surrogate outcomes without introducing optimistic bias; whether precedent retrieval degrades operator judgement over time by anchoring; and whether the attention-ranking model, once it changes which cases are worked first, invalidates the very distribution it was fitted on. The third is a genuine feedback pathology and is currently mitigated only by a bounded exploration quota.
