§ 8 — Observatory

What happens after the decision: a learning surface over executed work.

Automation ends where measurement usually stops. The Observatory is the component of CORTEX that treats the aftermath of every automated act — the extraction that was accepted, the arbitration that was overridden, the case that was reopened, the document that had already been submitted — as the primary scientific material of the programme. It converts an append-only ledger of decision records into supervised, unsupervised, sequential and causal learning problems, and returns the resulting estimates to the operating surface as early warnings, precedent retrieval, duplicate containment and drift-governed retraining. Every figure below is produced by the same harness that produces the evaluation results in §7; none of it depends on an operational secret.

8.1

Headline observatory indicators

Four indicators summarise whether post-decision learning is compounding. Trends span the last ten evaluation windows under the canonical configuration.

Duplicate rate
6.2%
−58%
Rework loops / case
0.34
−44%
Early-warning lead
19.4h
+2.7×
Label settlement
11.2d
−31%
8.2

The decision ledger — what is recorded, and why it is recorded that way

Learning from operations fails for a structural reason far more often than for a modelling reason: the record of what happened is incomplete, mutable, or labelled at the wrong moment. The ledger is designed against those three failure modes before any estimator is fitted.

Decision record ρ

Every terminal act emits an immutable record ρ = ⟨case, φ, decision, evidence-closure, agent-set, confidence-vector, actor, t, latency⟩. The record is the atomic unit of learning: nothing that was not recorded can be learned, and nothing recorded can be silently rewritten. Records are content-addressed, so a model trained on window w can be re-derived exactly from the same hashes months later.

Outcome labelling

A decision is not supervised at emission time but at settlement time, when the administrative outcome becomes observable (accepted, reworked, contested, cancelled, impeded). The Observatory therefore maintains a delayed-label buffer and only promotes a record into the training corpus once its label horizon has closed, which removes the optimistic bias that arises when partially settled cases are treated as successes.

Counterfactual ledger

For every automated proposal that a human overrode, both branches are retained: what the agent network proposed and what the operator decided. This pairing is the only source of counterfactual signal available without running experiments on live cases, and it is what makes disagreement learnable rather than merely auditable.

Attention feedback

Each ranked-attention item carries whether it was opened, how long it was inspected, and whether it changed the outcome. Interventions that never change the outcome are noise in the operator's day, and the ranking model is retrained explicitly to suppress them.

8.3

Learning problems induced by the ledger

The ledger admits six distinct learning formulations. They are listed separately because they carry different evaluation protocols, different failure modes and different governance obligations — collapsing them into a single 'AI layer' is precisely the abstraction that makes such systems unauditable.

Supervised · residual cycle time

Regularised linear regression and gradient-boosted trees estimate the remaining hours to terminal state for every open case. The linear model is retained as the reportable estimator because its standardised coefficients are an auditable statement of causation-adjacent structure; the boosted model is retained as a performance ceiling to quantify how much accuracy interpretability costs (empirically 0.04 R²).

Classification · failure propensity

A calibrated classifier predicts P(rework | case state) at every phase boundary. Calibration is enforced with isotonic regression and verified against a reliability diagram: an uncalibrated 0.8 that behaves like a 0.55 is operationally worse than no prediction at all, because it silently reallocates human attention to the wrong cases.

Unsupervised · duplicate and near-duplicate

Documents and cases are embedded and compared under cosine similarity with a locality-sensitive index. Three regimes are separated: byte-identical re-uploads, semantically identical documents differing only in scan geometry, and legitimately re-issued documents for the same asset. Only the first two are suppressible; the third must be preserved because it carries the legal history of the case.

Retrieval · prescriptive precedent

Given an open problematic case, a k-nearest-neighbour query over the settled corpus returns the most similar historical cases together with the action sequence that resolved them and the resulting cycle time. The Observatory therefore answers 'what worked last time' with evidence rather than with a generated recommendation.

Sequence · trace anomaly

Case traces are modelled as sequences over the phase alphabet; low-likelihood traces under the fitted model are surfaced as process anomalies. Anomalies are not errors by definition — many are legitimate rare variants — so they are routed to DISCOVERY, where a human decides whether the variant is promoted into the process model or treated as a defect.

Causal · intervention estimate

Where randomisation is impossible, the effect of an intervention (earlier document request, alternative source, pre-emptive escalation) is estimated with propensity-score matching over the settled corpus, and reported with its overlap diagnostics. Effects whose matched support is thin are published as insufficiently identified rather than as findings.

8.4

Predictive surface — estimating the case before it fails

Figures O1–O4 constitute the predictive surface. All estimates are produced under a temporally blocked protocol (train on window w, evaluate on w+1); random splitting is rejected because it leaks future process states into training and inflated R² by roughly 0.2 during harness calibration.

Fig. O1 — Predicted vs. observed residual cycle time
blocked split
0010102020303040405050R² = 0.78Observed residual cycle time (h)Predicted (h)
Dispersion widens above 30 h, where the authorisation queue dominates and the residual is governed by human availability rather than by case attributes — the honest ceiling of the estimator, not a defect to be tuned away.
Fig. O2 — Throughput forecast with 95% interval
h = 8 windows
443322110t₀ · forecast horizoncases/day · solid = observed · dashed = predicted · band = 95% CI
The band widens with horizon by construction. Empirical coverage of the nominal 95% interval is tracked per window and reported alongside point accuracy; an interval that is accurate on average but miscovered is treated as a defect.
Fig. O3 — Reliability diagram, failure-propensity model
isotonic
0.000.00.250.30.500.50.750.81.001.0predicted risk
Expected calibration error after isotonic correction: 0.021. Calibration is a precondition for ranked attention — the ranking is only meaningful if a predicted 0.8 behaves like a 0.8.
Fig. O4 — Case survival stratified by rework cohort
stratified
0%25%50%75%100%0h15h30h45h60hno rework1 re-entry≥2 re-entriesP(case still open) · Kaplan–Meier style, stratified by rework cohort
A single re-entry roughly halves the completion hazard; two or more re-entries produce a cohort whose median closure exceeds the observation window. Rework is therefore modelled as a state change, not as a counter.
8.5

Failure taxonomy and resolution structure

Cases that did not succeed are the most information-dense objects in the corpus. Each is attributed to one of six failure classes from the evidence closure rather than from triage notes, and the empirical distribution of resolutions per class is what determines whether a class is automatable, escalatable, or structurally human.

Fig. O5 — Resolution distribution by failure class (%)
row-normalised
retryre-captureescalatequarantineauto-healF₁ illegible capture12619144F₂ field disagreement85471129F₃ missing attachment31638214F₄ restriction conflict3266254F₅ external source timeout7218154F₆ stale catalogue5122963row-normalised share of resolutions per failure class
Reading the matrix

F₅ (external source timeout) resolves under plain retry in 72% of occurrences and is therefore a scheduling problem, not an intelligence problem. F₄ (restriction conflict) escalates in two thirds of occurrences and must not be automated further: the correct engineering response is to reduce the latency of escalation, not its frequency. F₆ (stale catalogue) is dominated by auto-heal, which is the empirical signature of a working DISCOVERY loop. Classifying failures by their resolution structure — rather than by their symptom — is what turns a defect list into an engineering programme.

Fig. O6 — Recurrence after corrective action
Δ windows
011233445%F₁ illegible9F₂ disagreement21F₃ missing doc34F₄ restriction6F₅ timeout41F₆ stale rule4
Recurrence within two windows of the corrective action. High recurrence under F₅ indicates an unfixed external dependency; low recurrence under F₆ indicates that catalogue promotion is durable.
8.6

Duplicates, near-duplicates and legitimate re-issues

Duplicate work is the most under-measured waste class in document-heavy operations, because the second submission is usually indistinguishable from the first at the interface level. The Observatory separates three regimes and treats only the first two as suppressible.

Fig. O7 — Similarity clustering over the case embedding space
LSH · cos θ
C₁ · exact duplicateC₂ · near-duplicate (Δ plate)C₃ · same case, re-issuedC₄ · uniquecos θ = 0.94 ≥ τ
Three regimes
  • R₁ identical — byte-equal artefact re-submitted. Suppressed with operator confirmation; the confirmation supervises τ.
  • R₂ near-identical — same content under different scan geometry, compression or crop. Detected in embedding space, never by hash.
  • R₃ re-issue — a genuinely new document for the same asset. Must be preserved: it carries the legal history of the case, and suppressing it is a correctness failure, not an efficiency gain.
Threshold economics

The suppression threshold τ is not chosen to maximise F1. It is chosen against an explicitly asymmetric cost: a false suppression destroys evidence and is bounded at 0.1% by policy, whereas a missed duplicate merely costs a redundant execution. τ is therefore set at the highest operating point satisfying the false-suppression bound, and the resulting recall (0.83) is reported as a consequence of that constraint rather than as a tuned result.

8.7

Preventive loop — turning post-mortems into pre-flight checks

The scientific value of the Observatory is not the retrospective dashboard; it is the closed loop in which a settled case changes how the next similar case is executed. Six mechanisms implement that loop, each with its own measurable lead indicator.

Early-warning triggers

A case is flagged when its predicted residual time crosses the SLA horizon, not when the SLA is already breached. Lead time between flag and breach is itself a metric — a warning that fires two hours before the deadline is technically correct and operationally worthless.

Root-cause attribution

Each failed case is attributed to a dominant cause class through the failure taxonomy F₁–F₆ using the evidence closure, not through free-text triage notes. Attribution stability across annotators is reported as Cohen's κ.

Similar-case suppression

When a new case matches the signature of a previously failed cohort, the framework surfaces the failure pattern before execution and proposes the corrective action that closed the cohort. This converts post-mortem knowledge into a pre-flight check.

Duplicate containment

Duplicate submission is measured as a first-class waste category. Suppression is proposed, never executed silently: the operator confirms, and the confirmation is itself logged as supervision for the similarity threshold τ.

Rework economics

Every rework loop is priced in machine cost and human minutes, so process changes can be ranked by expected saving instead of by intuition. The ranking is recomputed each window and its realised versus predicted saving is tracked as a forecast-quality metric.

Playbook promotion

An action sequence that resolves a cohort above a pre-registered success threshold and sample size is promoted to a playbook, versioned, and thereafter proposed automatically. Demotion is symmetric: a playbook whose realised success falls below threshold is retired and its retirement recorded.

8.8

Drift governance and retraining discipline

A model that was correct in window w is an unvalidated hypothesis in window w+1. Drift is monitored on the feature distribution, on the label distribution and on calibration independently, because each degrades for different reasons and demands a different response.

Fig. O8 — Population stability index per window
PSI
PSI = 0.10 · investigatePSI = 0.25 · retrainw1w2w3w4w5w6w7w8w9w10w11w12Population stability index per evaluation window
Window w8 crosses the investigation band and was traced to a change in the composition of incoming document classes rather than to model decay — a distinction only visible because feature drift and calibration drift are monitored separately.
Model card

Every deployed estimator ships a card stating training window, feature list, excluded features, known failure modes, calibration status and the population on which it must not be used.

Retraining policy

Retraining is triggered by drift (PSI ≥ 0.25), by a calibration breach, or by the scheduled window boundary — never ad hoc. Each retrain is a versioned artefact with a promotion decision recorded against a champion–challenger comparison.

Shadow deployment

A challenger model runs in shadow for a full window, scoring live cases without influencing them, and is promoted only if it dominates the champion on both the accuracy metric and the attention-precision metric.

Fairness of attention

Ranked attention is audited for systematic starvation: no case cohort may remain unranked beyond a bounded number of windows, regardless of its predicted value, otherwise the ranking becomes a self-fulfilling backlog.

Human override authority

No model output is terminal. Overrides are unrestricted, always logged with the acting identity, and treated as high-value labels rather than as violations of the automation.

Data minimisation

The Observatory stores the features required by the published models and nothing else; free-text and identifying content are excluded at ingestion rather than filtered at query time.

8.9

Reference implementation

The Observatory contract is deliberately small: ingest immutable records, refuse to train on unsettled labels, separate duplicates from re-issues, retrieve precedent, and let a governance verdict — not a schedule or an intuition — decide when a model is replaced.

python
class Observatory:
    """Post-decision learning surface over the CORTEX ledger."""

    LABEL_HORIZON_DAYS = 14
    DUP_TAU = 0.94          # cosine threshold for near-duplicate
    PSI_RETRAIN = 0.25      # drift budget before mandatory retrain

    def ingest(self, record: DecisionRecord) -> None:
        """Records are immutable; supervision arrives later."""
        self.ledger.append(record.freeze())
        self.pending_labels.schedule(record.case_id,
                                     horizon=self.LABEL_HORIZON_DAYS)

    def settled_corpus(self, window: Window) -> Frame:
        """Only cases whose outcome horizon has closed may train a model."""
        return self.ledger.filter(window=window, settled=True)

    def duplicates(self, case: Case) -> list[Match]:
        """Three regimes: identical, near-identical, legitimately re-issued."""
        cand = self.index.query(case.embedding, k=25)
        return [m for m in cand
                if m.score >= self.DUP_TAU and not m.is_reissue]

    def precedent(self, case: Case, k: int = 8) -> list[Playbook]:
        """Prescriptive retrieval: what resolved the most similar cases."""
        neigh = self.index.query(case.embedding, k=k, settled_only=True)
        return rank_by_realised_gain(p.playbook for p in neigh)

    def governance_check(self, window: Window) -> Verdict:
        psi = population_stability(self.reference, self.live(window))
        calib = expected_calibration_error(self.scores(window))
        if psi >= self.PSI_RETRAIN or calib > 0.05:
            return Verdict.RETRAIN(reason=("drift" if psi >= self.PSI_RETRAIN
                                           else "calibration"))
        return Verdict.HOLD
observatory/core.py — post-decision learning surface (reference stub)
Invariants
  • O-I₁ — no record is mutated after emission; corrections are new records referencing the original.
  • O-I₂ — no estimator trains on a case whose label horizon has not closed.
  • O-I₃ — no suppression executes without a logged human confirmation.
  • O-I₄ — every published figure is reproducible from (seed, window, manifest).
  • O-I₅ — every model in production has a current card and a live drift verdict.
Open questions

Three questions remain unresolved and are stated as such: whether delayed labelling can be shortened with surrogate outcomes without introducing optimistic bias; whether precedent retrieval degrades operator judgement over time by anchoring; and whether the attention-ranking model, once it changes which cases are worked first, invalidates the very distribution it was fitted on. The third is a genuine feedback pathology and is currently mitigated only by a bounded exploration quota.