§ 6.1
Research Question
Can a society of autonomous agents built on generative optical models — agents that perceive, negotiate, and self-regulate rather than execute a fixed script — simultaneously dominate monolithic vision–language baselines on efficiency under contention, factual error rate, and audit-grade traceability? The question is deliberately tri-objective: any single axis is trivially optimisable in isolation, and the scientific interest lies exclusively in whether a point exists that is non-dominated on all three at once.
§ 6.2
Central Thesis
T. Autonomous multi-agent orchestration over generative OCR attains a jointly non-dominated point on the efficiency × error × auditability frontier that no monolithic end-to-end extractor of comparable compute budget reaches. The thesis decomposes into three falsifiable claims. (T₁ · Efficiency) Agent-level bottleneck detection with dynamic rerouting raises sustained throughput under contention relative to a statically scheduled pipeline of identical parameter budget. (T₂ · Error) Consensus arbitration across independent evidence sources measurably reduces factual error and hallucinated field values relative to single-decoder generative extraction at matched latency. (T₃ · Auditability) Provenance closure is preserved under agent autonomy: every emitted decision admits complete causal replay despite non-deterministic inter-agent negotiation. The composite claim is that T₁, T₂ and T₃ hold jointly, not merely severally.
§ 6.3
Bottleneck Taxonomy
Four bottleneck classes are measured explicitly and treated as the dependent variables of the study. (B₁ · Optical latency) Cost of generative decoding per page, reported at p50 / p95 / p99 with stage-level attribution, since generative OCR shifts the dominant cost from I/O to inference. (B₂ · Source divergence) Rework and arbitration cost incurred when independent sources disagree, measured as excess evidence lookups per contested field. (B₃ · Agent handoff) Queueing delay and message overhead at inter-agent boundaries, modelled as an M/M/c network and bounded against the Amdahl-serial fraction of the topology. (B₄ · Schema and rule drift) Longitudinal degradation as document distributions and restriction catalogues evolve, quantified as accuracy decay per evaluation window against a frozen baseline.
§ 6.4
Agent Communication Protocol
Agents exchange typed messages over an asynchronous channel algebra rather than through direct invocation. Four message classes are defined: CLAIM (an agent asserts a field value with calibrated confidence), CHALLENGE (a peer contests a claim and supplies counter-evidence), CONGESTION (an agent reports saturation, triggering re-routing or back-pressure upstream), and DISCOVERY (an agent surfaces an unmodelled pattern or candidate restriction for catalogue promotion). Negotiation terminates under a proven convergence bound: confidence is monotonically non-increasing across CHALLENGE rounds, so the protocol cannot cycle. This message algebra is what makes autonomy evaluable rather than anecdotal — every negotiation is a recorded, replayable artefact.
§ 6.5
Hypothesis
We hypothesise that a multi-agent pipeline composed of generative optical perception (OPTIC), consensus arbitration (NEXUS) and symbolic restriction evaluation (AEGIS) attains a strictly superior point in the reliability–auditability frontier than monolithic baselines of comparable parameter budget, while preserving throughput within 15% of an equivalently sized end-to-end neural extractor. The hypothesis is operationalised through a pre-registered ablation protocol detailed in §7.
§ 6.6
Objectives
(O₁) Specify the framework formally as a typed agent network with per-agent contracts and a convergent message algebra. (O₂) Implement reference agents under those contracts on a generative optical backbone. (O₃) Evaluate accuracy, hallucination rate, latency, throughput and auditability against a synthetic but representative enterprise corpus. (O₄) Characterise the four bottleneck classes B₁–B₄ through fault and load injection at every agent boundary. (O₅) Release reproducible artefacts — corpora, configurations and evaluation harness — under an open scientific licence.
§ 6.7
Workflow
Document acquisition → schema-bound extraction → cross-source consistency analysis → symbolic restriction evaluation → decision-support emission with full provenance. Each arrow is a typed channel; each node a pure function over its declared input envelope. Side-effects (caching, source lookups, telemetry) are isolated in effect-aware wrappers that preserve the purity of the underlying transformation.
§ 6.8
Pipeline
Typed, unidirectional, with explicit boundaries between stages. Each stage is a total function over its declared input type, returning a typed result envelope whose disjoint sum encodes success, partial success and recoverable failure. The dataflow graph is intentionally acyclic; iteration is expressed through provenance-annotated reprocessing rather than back-edges, preserving the algebraic properties required for compositional reasoning.
§ 6.9
Validation Strategy
Three nested evaluation loops are exercised in sequence. (V₁) Per-stage unit evaluation under controlled distributions, isolating extraction error from validation error. (V₂) End-to-end pipeline evaluation on held-out corpora drawn from the operating envelope of the target deployment. (V₃) Module ablation, in which each agent is independently disabled to quantify its marginal contribution to the end-to-end metric of interest. Statistical significance is reported at α = 0.01 with bootstrap-derived confidence intervals.
§ 6.10
Error Handling
Partial failure is modelled as a first-class citizen of the type system rather than as exceptional control flow. Each stage returns a typed result envelope distinguishing success, partial success, and recoverable failure; uncoverable failures are routed to a quarantine bucket whose contents are surfaced through the same observability surface as the canonical decision stream, guaranteeing that no document silently disappears from the audit trail.
§ 6.11
Architecture Description
Three cooperating agents over a typed pipeline orchestrator. Provenance is propagated as a first-class artefact through every stage boundary, materialised as an immutable DAG whose nodes carry cryptographic hashes of the contributing evidence. The architecture is described through five complementary views (orchestrator graph, high-level pipeline, component decomposition, data-flow algebra, validation topology) constituting a 4+1 style description.
§ 6.12
Processing Logic
Stateless transformations are preferred wherever the semantics permit; stateful effects (cache lookups, source resolution, telemetry emission) are encapsulated in observable wrappers exposing explicit retry, time-out and back-pressure controls. This separation of pure logic from effectful boundary allows the entire framework to be re-evaluated deterministically against any persisted evidence snapshot.
§ 6.13
Cross-Source Validation
Validation is formulated as an agreement functional a(R, E) over the extracted record R and an indexed family of independent evidence sources E = {S₁, …, Sₙ} weighted by an empirically calibrated reliability prior. The functional returns both a scalar confidence and a provenance subgraph that records every contributing source, transformation, and arbitration decision, supporting closed-form replay against the original evidence snapshot.
§ 6.14
Restriction Analysis
Operational restrictions are encoded as a declarative constraint set C and evaluated symbolically against the validated record R; violations are surfaced with machine-readable rationale and natural-language justification. The constraint system is monotone in C — adding a restriction never converts a reject into an approve — which guarantees that constraint catalogues can be extended without invalidating prior decisions, a property required for stable longitudinal audit.
§ 6.15
Reproducibility Protocol
Every reported figure is paired with a deterministic seed, an immutable corpus snapshot identifier, a stage-level configuration manifest, and a content-addressed hash of the evaluation harness. Re-execution against the same triple (seed, corpus, manifest) is required to yield bit-identical artefacts; deviations are themselves treated as findings and are investigated as candidate determinism faults rather than discarded as noise. All artefacts are released under an open scientific licence in alignment with ACM and IEEE reproducibility badging criteria.
§ 6.16
Threats to Validity
Four threats are pre-registered. (Internal) Agent autonomy introduces non-determinism in negotiation order; we mitigate through seeded scheduling and report variance across ten repetitions rather than single runs. (Construct) Hallucination is not directly observable, so it is operationalised as a field value unsupported by any evidence source in the provenance closure — a conservative proxy that may under-count fluent but source-consistent errors. (External) The corpus is synthetic-but-representative; generalisation to jurisdictions with materially different document conventions is asserted only as a hypothesis for future work. (Statistical) With four bottleneck classes and three claims, multiple comparison inflates false-positive risk; all reported significance is Holm–Bonferroni corrected at α = 0.01.
§ 6.17
Process Instrumentation Model
CORTEX is evaluated not on isolated documents but on complete regulatory dossiers traversing a multi-phase administrative workflow. The unit of analysis is therefore the case: an asset-bound folder that accumulates documents, human interventions and state transitions until a terminal regularisation status is reached. Every case is instrumented as an append-only event stream e = (case_id, phase, event_type, actor_class, t) with four event families — PHASE_FIRST_ENTRY, PHASE_REENTRY, PHASE_EXIT and FIELD_MUTATION — complemented by annotation events emitted by human reviewers. Because a phase may be visited several times, phase residence time is attributed once per visit rather than once per emitted row; failing to enforce this visit-level deduplication inflates waiting time by a factor equal to the mean event multiplicity per visit (empirically 2.4–3.0) and was the dominant measurement artefact identified during harness calibration. All identifiers are pseudonymised at ingestion and no organisational, vendor or system name enters the research corpus.
§ 6.18
Process-Mining Layer
The event stream induces a directly-follows graph G_p = (Φ, ⟶) over the phase alphabet Φ, from which the study derives cycle time per case, residence time per phase, rework ratio (fraction of cases re-entering a phase already left), inter-phase queueing delay, and the empirical distribution of terminal outcomes. Conformance is measured as the fraction of observed traces admitted by the declarative process model; deviating traces are not discarded but promoted to DISCOVERY messages, which is how the agent network learns unmodelled process variants. This layer supplies the ground truth against which bottleneck class B₃ (agent and human handoff) is quantified: it separates machine latency from organisational waiting time, a distinction absent from document-level benchmarks.
§ 6.19
Predictive Cycle-Time Layer
A supervised regressor estimates the residual cycle time of an open case from features observable at prediction time: current phase, phase age, number of prior re-entries, count and class of attached documents, count of human annotations, confidence profile emitted by OPTIC, and NEXUS disagreement mass. The baseline estimator is a regularised linear model, deliberately chosen for interpretability: its coefficients constitute a directly auditable statement about which process attributes drive delay, which matters more for operational adoption than a marginal accuracy gain. Performance is reported as MAE, RMSE and R² under a temporally blocked split (train on window w, evaluate on w+1) rather than a random split, since random splitting leaks future process states into training and systematically overstates predictive skill on workflow data.
§ 6.20
Automation of the Documentary Lifecycle
The applied substrate of the study is the end-to-end regularisation lifecycle of asset documentation: issuance, retrieval of official records, transfer evidence, inspection reports and complementary attachments. Each document class is acquired through a typed adapter that abstracts over the underlying channel (authenticated portal, structured export, machine-readable feed or scanned artefact), so the research claims remain independent of any particular provider. Acquisition is executed in bounded lots with idempotent retry, per-lot quarantine, and a three-way operator decision on failure (retry the same lot, halt with partial reporting, or skip and route to an error ledger). The scientific interest of the lot abstraction is that it makes throughput and failure containment measurable at a granularity that matches how the work is actually organised.
§ 6.21
Operator Platform and Human-in-the-Loop
Autonomy is not framed as the removal of the human but as the reallocation of human attention. The framework exposes an operator console in which the decision remains human while the technical execution is delegated: the operator selects a scope, inspects the proposed transitions, and authorises execution, with live progress, pause/resume/finalise controls, and an exportable execution report. Every automated act is signed with the acting identity and materialised as an annotation event in the same ledger used for scientific measurement, so operator behaviour is a first-class observable rather than an untracked exogenous factor. Two dependent variables are attached to this surface: attention cost (interventions per hundred cases) and authorisation latency (time between proposal and human decision).
§ 6.22
Analytical Copilot
The observability ledger, the process-mining layer and the predictive layer are exposed to a retrieval-grounded conversational agent that answers managerial questions — where the queue is forming, which phase regressed this week, which document class dominates rework — by composing typed queries over the ledger rather than by free-form generation. Every answer is emitted with the query executed, the row count it aggregated, and the evaluation window, which subordinates the copilot to the same provenance-closure invariant (I₃) as the extraction pipeline. Ungrounded answers are refused by construction: if no query supports the assertion, the agent returns an explicit insufficiency verdict instead of a fluent guess.
§ 6.23
Data Governance and Anonymisation
The research corpus is derived from an operational domain but is released only in de-identified, synthesised form. Three controls are applied before any artefact leaves the boundary: identifier pseudonymisation with a per-release salt, suppression of every organisational, vendor and platform name, and distributional resampling of rare categorical values that could act as quasi-identifiers. Credentials, tokens and endpoints are excluded from the artefact by static scanning in the continuous-integration pipeline. The published corpus preserves the statistical structure required to reproduce every reported figure while carrying no operational secret, satisfying the double-blind requirement of review venues.
§ 6.24
Post-Decision Learning Loop
Measurement does not stop at emission. Every terminal act writes an immutable decision record carrying the evidence closure, the participating agent set, the confidence vector, the acting identity and the realised latency; supervision is attached later, when the administrative outcome settles. This delayed-label discipline is methodologically essential — training on unsettled cases treats provisional success as final and produces optimistic estimates that collapse in deployment. The settled ledger induces six distinct learning formulations (residual cycle-time regression, calibrated failure-propensity classification, duplicate and near-duplicate detection, prescriptive precedent retrieval, sequence-level trace anomaly detection, and matched causal estimation of interventions), each with its own evaluation protocol and governance obligation. The complete specification of this surface is given in §8, the Observatory.
§ 6.25
Drift Governance of Learned Components
Any learned component is treated as a hypothesis with an expiry date. Three drift channels are monitored independently because they degrade for different reasons and demand different responses: feature drift via the population stability index, label drift via the settled-outcome distribution, and calibration drift via expected calibration error on the reliability diagram. Retraining is triggered by a governance verdict (PSI ≥ 0.25, ECE > 0.05, or window boundary) rather than by discretion, and a challenger model must dominate the champion on both predictive accuracy and attention precision across a full shadow window before promotion. Every promotion, demotion and override is a ledger event, which keeps the learning system inside the same provenance-closure invariant (I₃) as the extraction pipeline.
§ 6.26
Feedback Pathologies
A system that reorders human work also perturbs the distribution it learns from. Three pathologies are pre-registered as threats specific to the Observatory. (P₁ · Ranking feedback) Once ranked attention determines which cases are worked first, observed cycle times are no longer samples from the original process; a bounded exploration quota keeps a randomised residual of unranked cases in the corpus to preserve identifiability. (P₂ · Anchoring) Precedent retrieval may degrade independent operator judgement over time; the effect is monitored as the divergence between operator decisions with and without a displayed precedent. (P₃ · Suppression bias) Duplicate suppression removes future evidence about duplicates; every suppression therefore remains a logged, reversible record rather than a deletion.