§ 7 — Results

Experimental evaluation of the CORTEX Framework.

We evaluate the framework along seven orthogonal dimensions on a synthetic but representative enterprise corpus instrumented to reproduce the modality heterogeneity, source-disagreement spectrum and longitudinal drift observed in production deployments. Reported figures decompose extraction error from validation error, expose tail-latency contributors at stage granularity, and quantify the marginal contribution of each module through a pre-registered ablation protocol with bootstrap-derived 99% confidence intervals. Quantitative figures shown below are illustrative of the evaluation harness; statistically significant results against the monolithic baseline are deferred to the upcoming technical report (TR-2026-01).

7.1

Headline indicators

A compact summary of the four primary indicators tracked across the evaluation harness. Trend lines depict the most recent ten evaluation windows under the canonical configuration (OPTIC · NEXUS · AEGIS enabled).

Extraction F1
0.912
+0.18
Latency p95
184ms
−42%
Throughput
2.4k/h
+3.1×
Escalations
4.8%
−61%
7.2

Evaluation Dimensions

The seven evaluation dimensions define a multi-axis quality envelope. Each dimension is reported independently and aggregated into the quality profile in Fig. R3.

Accuracy

Token-, entity- and record-level F1 measured at every module interface, decomposing extraction error from validation error.

Performance

Compute footprint per pipeline stage, normalised by document complexity and reported as work-units per validated decision.

Latency

End-to-end p50 / p95 / p99 wall-clock, partitioned by stage to expose tail-latency contributors and queueing effects.

Scalability

Sustained throughput as a function of corpus size and concurrency, fit against an Amdahl-bounded reference curve.

Robustness

Performance degradation under synthetic noise injection, adversarial perturbation, and source-disagreement scenarios.

Auditability

Provenance completeness and justification coverage measured as the fraction of decisions admitting full causal replay.

Operational Efficiency

Human-hours saved per audit cycle and reduction in escalation rate relative to a monolithic baseline.

Reproducibility

Bit-identical re-execution rate under fixed (seed, corpus, manifest) triples; deviations are themselves treated as candidate determinism faults.

Cost Profile

Aggregate compute, storage and human-review cost per validated decision, partitioned by stage and reported against a monolithic reference baseline.

7.3

Preliminary figures

Figures R1–R5 illustrate the evaluation surface. Distributions are computed from a synthetic harness sized to match the operating envelope of the target deployment; statistical tests against the monolithic baseline are reported in the technical appendix.

Fig. R1 — End-to-end throughput (sliding window)
Preliminary
1007550250docs/min · 60s rolling window
Fig. R2 — Latency distribution
Preliminary
p50p95end-to-end latency (ms, log-bin)
Fig. R3 — Quality profile vs. baseline
Preliminary
AccuracyLatencyThroughputRobustnessAuditabilityCost
Fig. R4 — Median stage cost
Preliminary
018355370msacquire22extract64validate41restrict18emit9
Fig. R5 — Error reduction by module ablation
Preliminary
0285583110%Baseline (monolith)100− AEGIS78− NEXUS56− OPTIC28Full framework18
7.4

Process-level bottleneck attribution

Document-level benchmarks conflate machine latency with organisational waiting time. Table R6 decomposes the observed case cycle time over the phase alphabet Φ, separating the fraction attributable to automated execution from the fraction attributable to queueing and human authorisation. Residence times are attributed once per phase visit; before this correction was applied, summing the duration carried by every emitted event row inflated waiting time by the mean event multiplicity per visit (2.4–3.0×) and produced physically impossible estimates in the first evaluation window. Rework denotes the share of cases re-entering a phase already left.

Phase φ ∈ ΦMedian residencep95Machine shareReworkDominant bottleneck
φ₁ · Intake & acquisition0.8 h4.1 h94%3.1%B₁ optical latency
φ₂ · Extraction & typing0.3 h1.2 h98%1.4%B₁ optical latency
φ₃ · Cross-source arbitration2.6 h18.4 h41%11.7%B₂ source divergence
φ₄ · Restriction evaluation0.2 h0.9 h99%0.6%—
φ₅ · Human authorisation9.4 h62.0 h0%8.9%B₃ handoff
φ₆ · Terminal emission0.4 h2.2 h88%2.0%B₄ catalogue drift

Tab. R6 — Cycle-time decomposition over the phase alphabet. The dominant contributor to end-to-end latency is not generative decoding but the authorisation queue at φ₅: automation compresses the machine-bound phases by an order of magnitude and thereby shifts the critical path onto human attention, which is precisely the regime the operator console and the ranked-attention loop are designed to address.

7.5

Predictive layer — residual cycle time

The residual cycle time of open cases is estimated by an interpretable regularised linear model under a temporally blocked protocol: train on evaluation window w, test on w+1. Random splitting was explicitly rejected — it leaks future process states into training and inflated R² by roughly 0.2 during calibration. Reported coefficients are standardised and constitute the auditable statement of which process attributes drive delay.

Blocked evaluation (w → w+1)
  • MAE
    mean absolute error on residual cycle time
    6.9 h
  • RMSE
    penalises the long-queue tail at φ₅
    11.4 h
  • R²
    against a persistence baseline at 0.38
    0.71
  • Coverage
    empirical coverage of the 90% interval
    92.4%
Fig. R7 — Standardised coefficient magnitude
Preliminary
023456890βprior re-entries82phase age74NEXUS disagreement61pending attachments47human annotations33OPTIC mean confidence21
7.6

Operational surface — attention economics

Autonomy is evaluated not as the removal of the human but as the reallocation of human attention. The two dependent variables attached to the operator console measure whether the framework reduces the cognitive load per case without displacing the decision itself.

Interventions / 100 cases
18.6

−67% against the manual baseline; each intervention remains a signed, replayable ledger event.

Authorisation latency
3.2 h

Median time between proposal and human decision under ranked attention ordering.

Attention precision
0.81

Fraction of ranked-attention items that indeed required a human decision.

Silent-loss rate
0.00%

Cases leaving the pipeline without a terminal ledger entry; invariant I₃ forbids a non-zero value.

7.7

Data-science surface — forecast, survival and drift

The final block of the evaluation is prospective rather than retrospective: it asks whether the framework can state, in advance and with quantified uncertainty, what the operation is about to do. Figures R8–R11 are produced by the Observatory harness (§8) under the same blocked protocol used throughout §7.

Fig. R8 — Predicted vs. observed residual cycle time
Preliminary
0010102020303040405050R² = 0.78Observed residual cycle time (h)Predicted (h)
Fig. R9 — Throughput forecast with 95% interval
Preliminary
443322110t₀ · forecast horizoncases/day · solid = observed · dashed = predicted · band = 95% CI
Fig. R10 — Case survival by rework cohort
Preliminary
0%25%50%75%100%0h15h30h45h60hno rework1 re-entry≥2 re-entriesP(case still open) · Kaplan–Meier style, stratified by rework cohort
Fig. R11 — Feature drift per evaluation window
Preliminary
PSI = 0.10 · investigatePSI = 0.25 · retrainw1w2w3w4w5w6w7w8w9w10w11w12Population stability index per evaluation window

Forecast quality is reported jointly as point accuracy and interval coverage; an estimator that is accurate on average while systematically miscovering its interval is treated as defective, because ranked attention consumes the interval, not the point. The complete specification of the learning surface — decision ledger, delayed labelling, failure taxonomy, duplicate containment, precedent retrieval and retraining governance — is given in §8, the Observatory.

7.8

Interactive panel — windowed bottleneck exploration

Aggregate figures hide regime changes. The panel below exposes the same evaluation series under a selectable observation window and bottleneck focus, so the reader can inspect the trajectory attributable to a single bottleneck class rather than the pooled average. Hovering any vertex reveals the underlying data label; the dashed continuation is an ordinary-least-squares projection over the selected window and is illustrative, not a committed forecast.

windowfocus
1088154270cases/h · 30d · All phases · dashed = OLS projection

Fig. R12 — Filterable throughput surface. Selecting B₃ (handoff) reproduces the central finding of Tab. R6: the authorisation queue dominates the critical path and is the only class whose projected slope does not flatten under the current configuration.

7.9

Operational impact — the agent console in production

The deployment surface is not a batch job but a supervised console: each operator receives a bounded terminal bound to a single agent, a single phase subset, and an explicit grant set. The operator observes the agent executing on their own workload in real time, may suspend or escalate at any point, and cannot exceed the grants attached to their role. Autonomy is therefore cadenced rather than unconstrained — the machine advances the case, the human retains the authority, and every emitted line is a signed ledger event admitting full causal replay.

operations substrate · live case CS-4471dashed channels = typed inter-agent messages
T-01 · OPTIC
φ₁ · φ₂
$ optic.decode(page=6096) → tokens=42 conf=0.90
$ optic.crop(region='header') ok latency=136ms
$ WARN low-contrast field 'issuer' → re-render @300dpi
operator · intake deskread:scanemit:record
T-02 · NEXUS
φ₃
$ nexus.arbitrate(sources=3) agreement=0.85
$ CONFLICT field='total' src_a=11 src_b=323 → escalate
$ nexus.merge(strategy='weighted-quorum') ✓
operator · arbitrationread:sourcespropose:merge
T-03 · AEGIS
φ₄
$ aegis.evaluate(policy='RES-14') → ALLOW
$ DENY rule R-08: missing authorisation signature
$ aegis.seal(ledger_event=7377) hash=0x90f6
compliance officereval:policyblock:emit
T-04 · HELIOS
Φ (observe)
$ helios.trace(case=CS-5026) phases=6 rework=4
$ helios.psi(window=w2) = 0.029 stable
$ helios.bottleneck() → φ₅ authorisation queue
process analystread:ledger
Governance by grant

Each terminal carries an explicit capability set (read / propose / block). Actions outside the set are refused at the AEGIS boundary and recorded as denied attempts.

Cadenced autonomy

Agents execute continuously but surface every irreversible step for authorisation, keeping the human on the decision rather than on the transcription.

Technological contact

The operator watches their own agents work, which converts automation from an opaque substitution into an observable instrument and materially raises adoption.

Measured gain

Median hands-on time per case falls from 42 min to 6 min in the illustrative harness, with interventions concentrated on the 18.6% of cases that genuinely require judgement.