Experimental evaluation of the CORTEX Framework.
We evaluate the framework along seven orthogonal dimensions on a synthetic but representative enterprise corpus instrumented to reproduce the modality heterogeneity, source-disagreement spectrum and longitudinal drift observed in production deployments. Reported figures decompose extraction error from validation error, expose tail-latency contributors at stage granularity, and quantify the marginal contribution of each module through a pre-registered ablation protocol with bootstrap-derived 99% confidence intervals. Quantitative figures shown below are illustrative of the evaluation harness; statistically significant results against the monolithic baseline are deferred to the upcoming technical report (TR-2026-01).
Headline indicators
A compact summary of the four primary indicators tracked across the evaluation harness. Trend lines depict the most recent ten evaluation windows under the canonical configuration (OPTIC · NEXUS · AEGIS enabled).
Evaluation Dimensions
The seven evaluation dimensions define a multi-axis quality envelope. Each dimension is reported independently and aggregated into the quality profile in Fig. R3.
Token-, entity- and record-level F1 measured at every module interface, decomposing extraction error from validation error.
Compute footprint per pipeline stage, normalised by document complexity and reported as work-units per validated decision.
End-to-end p50 / p95 / p99 wall-clock, partitioned by stage to expose tail-latency contributors and queueing effects.
Sustained throughput as a function of corpus size and concurrency, fit against an Amdahl-bounded reference curve.
Performance degradation under synthetic noise injection, adversarial perturbation, and source-disagreement scenarios.
Provenance completeness and justification coverage measured as the fraction of decisions admitting full causal replay.
Human-hours saved per audit cycle and reduction in escalation rate relative to a monolithic baseline.
Bit-identical re-execution rate under fixed (seed, corpus, manifest) triples; deviations are themselves treated as candidate determinism faults.
Aggregate compute, storage and human-review cost per validated decision, partitioned by stage and reported against a monolithic reference baseline.
Preliminary figures
Figures R1–R5 illustrate the evaluation surface. Distributions are computed from a synthetic harness sized to match the operating envelope of the target deployment; statistical tests against the monolithic baseline are reported in the technical appendix.
Process-level bottleneck attribution
Document-level benchmarks conflate machine latency with organisational waiting time. Table R6 decomposes the observed case cycle time over the phase alphabet Φ, separating the fraction attributable to automated execution from the fraction attributable to queueing and human authorisation. Residence times are attributed once per phase visit; before this correction was applied, summing the duration carried by every emitted event row inflated waiting time by the mean event multiplicity per visit (2.4–3.0×) and produced physically impossible estimates in the first evaluation window. Rework denotes the share of cases re-entering a phase already left.
| Phase φ ∈ Φ | Median residence | p95 | Machine share | Rework | Dominant bottleneck |
|---|---|---|---|---|---|
| φ₁ · Intake & acquisition | 0.8 h | 4.1 h | 94% | 3.1% | B₁ optical latency |
| φ₂ · Extraction & typing | 0.3 h | 1.2 h | 98% | 1.4% | B₁ optical latency |
| φ₃ · Cross-source arbitration | 2.6 h | 18.4 h | 41% | 11.7% | B₂ source divergence |
| φ₄ · Restriction evaluation | 0.2 h | 0.9 h | 99% | 0.6% | — |
| φ₅ · Human authorisation | 9.4 h | 62.0 h | 0% | 8.9% | B₃ handoff |
| φ₆ · Terminal emission | 0.4 h | 2.2 h | 88% | 2.0% | B₄ catalogue drift |
Tab. R6 — Cycle-time decomposition over the phase alphabet. The dominant contributor to end-to-end latency is not generative decoding but the authorisation queue at φ₅: automation compresses the machine-bound phases by an order of magnitude and thereby shifts the critical path onto human attention, which is precisely the regime the operator console and the ranked-attention loop are designed to address.
Predictive layer — residual cycle time
The residual cycle time of open cases is estimated by an interpretable regularised linear model under a temporally blocked protocol: train on evaluation window w, test on w+1. Random splitting was explicitly rejected — it leaks future process states into training and inflated R² by roughly 0.2 during calibration. Reported coefficients are standardised and constitute the auditable statement of which process attributes drive delay.
- MAEmean absolute error on residual cycle time6.9 h
- RMSEpenalises the long-queue tail at φ₅11.4 h
- R²against a persistence baseline at 0.380.71
- Coverageempirical coverage of the 90% interval92.4%
Operational surface — attention economics
Autonomy is evaluated not as the removal of the human but as the reallocation of human attention. The two dependent variables attached to the operator console measure whether the framework reduces the cognitive load per case without displacing the decision itself.
−67% against the manual baseline; each intervention remains a signed, replayable ledger event.
Median time between proposal and human decision under ranked attention ordering.
Fraction of ranked-attention items that indeed required a human decision.
Cases leaving the pipeline without a terminal ledger entry; invariant I₃ forbids a non-zero value.
Data-science surface — forecast, survival and drift
The final block of the evaluation is prospective rather than retrospective: it asks whether the framework can state, in advance and with quantified uncertainty, what the operation is about to do. Figures R8–R11 are produced by the Observatory harness (§8) under the same blocked protocol used throughout §7.
Forecast quality is reported jointly as point accuracy and interval coverage; an estimator that is accurate on average while systematically miscovering its interval is treated as defective, because ranked attention consumes the interval, not the point. The complete specification of the learning surface — decision ledger, delayed labelling, failure taxonomy, duplicate containment, precedent retrieval and retraining governance — is given in §8, the Observatory.
Interactive panel — windowed bottleneck exploration
Aggregate figures hide regime changes. The panel below exposes the same evaluation series under a selectable observation window and bottleneck focus, so the reader can inspect the trajectory attributable to a single bottleneck class rather than the pooled average. Hovering any vertex reveals the underlying data label; the dashed continuation is an ordinary-least-squares projection over the selected window and is illustrative, not a committed forecast.
Fig. R12 — Filterable throughput surface. Selecting B₃ (handoff) reproduces the central finding of Tab. R6: the authorisation queue dominates the critical path and is the only class whose projected slope does not flatten under the current configuration.
Operational impact — the agent console in production
The deployment surface is not a batch job but a supervised console: each operator receives a bounded terminal bound to a single agent, a single phase subset, and an explicit grant set. The operator observes the agent executing on their own workload in real time, may suspend or escalate at any point, and cannot exceed the grants attached to their role. Autonomy is therefore cadenced rather than unconstrained — the machine advances the case, the human retains the authority, and every emitted line is a signed ledger event admitting full causal replay.
Each terminal carries an explicit capability set (read / propose / block). Actions outside the set are refused at the AEGIS boundary and recorded as denied attempts.
Agents execute continuously but surface every irreversible step for authorisation, keeping the human on the decision rather than on the transcription.
The operator watches their own agents work, which converts automation from an opaque substitution into an observable instrument and materially raises adoption.
Median hands-on time per case falls from 42 min to 6 min in the illustrative harness, with interventions concentrated on the 18.6% of cases that genuinely require judgement.
