Research Programme

The research programme behind the CORTEX Framework.

A multi-year investigation into generative optical document understanding, autonomous agent negotiation, and restriction-aware decision support in production-grade enterprise environments. The programme is organised around a single tri-objective question — whether autonomous agents over generative OCR can be simultaneously efficient, factually reliable and fully auditable — and is structured to admit reproducible empirical evaluation, open release of artefacts, and incremental extension by the broader research community. Throughout this report the framework is treated as an object of scientific study rather than as a product: every design decision is justified against an explicit thesis and every claim is paired with a falsification protocol.

§ 1.1

Problem Statement

Scientific question

Can a society of autonomous agents built on generative optical models — agents that perceive, negotiate, self-regulate and escalate novelty rather than execute a fixed script — simultaneously dominate monolithic vision–language baselines on efficiency under contention, factual error rate, and audit-grade traceability? The question is deliberately tri-objective: each axis is trivially optimisable in isolation, and the scientific interest lies exclusively in whether a configuration exists that is non-dominated on all three at once.

Operational motivation

Generative OCR removed the extraction ceiling and replaced it with two new liabilities: inference cost now dominates the critical path, and fluent models hallucinate field values that no evidence supports. Enterprise workflows cannot absorb either. Reliable document intelligence in this regime requires perception, consensus and restriction reasoning to be separated into autonomous agents that can measure and reroute around their own contention while keeping every emitted decision causally replayable — a structural property absent from monolithic production systems.

§ 1.2

Bottlenecks under study

Four bottleneck classes constitute the dependent variables of the programme. Each is instrumented at the agent boundary, injected under controlled load, and reported with stage-level attribution rather than as an aggregate figure.

B₁ · Optical latency

Generative decoding shifts the dominant cost from I/O to inference. Measured as cost per page at p50 / p95 / p99 with per-agent attribution, against a statically scheduled baseline of identical parameter budget.

B₂ · Source divergence

When independent evidence sources disagree, arbitration consumes lookups and wall-clock. Measured as excess evidence retrievals and CHALLENGE rounds per contested field, correlated against final decision correctness.

B₃ · Agent handoff

Autonomy costs coordination. Queueing delay and message overhead at inter-agent boundaries are modelled as an M/M/c network and bounded against the Amdahl-serial fraction of the topology.

B₄ · Schema and rule drift

Document distributions and restriction catalogues evolve. Measured as accuracy decay per evaluation window against a frozen baseline, and as the latency with which DISCOVERY messages promote a new pattern into the catalogue.

§ 1.3

Scope

In scope

Formal specification and reference implementation of the autonomous agent network; generative optical perception kernels; consensus arbitration across heterogeneous evidence; symbolic restriction reasoning; inter-agent message algebra with bottleneck detection; ablation and reproducibility infrastructure.

Out of scope

Domain-specific business policy authoring; vendor and integrator selection; jurisdictional regulatory interpretation; end-user product UI; deployment-specific cost optimisation. These concerns are intentionally externalised to keep the scientific claims falsifiable independently of any particular operational context.

Evaluation

Extraction accuracy decomposed at token, entity and record granularity; hallucination rate as the fraction of emitted fields unsupported by the provenance closure; consensus consistency under heterogeneous source reliability; end-to-end latency at p50 / p95 / p99 with agent-level decomposition; sustained throughput against an Amdahl-bounded reference; error reduction by agent ablation; auditability as the fraction of decisions admitting full causal replay.

§ 1.4

Applied substrate — the documentary regularisation lifecycle

The programme is grounded in a concrete class of enterprise processes: the regularisation of asset documentation, in which a case accumulates official records, transfer evidence, inspection reports and complementary attachments while traversing a multi-phase administrative workflow until a terminal status is reached. The domain is chosen because it exhibits, in a single setting, every property the thesis requires — heterogeneous document modalities, independent evidence sources that routinely disagree, restriction catalogues that evolve, and an authorisation step that keeps a human in the loop by design. All identifiers, organisational names and platform references are removed at ingestion; the released corpus is de-identified and synthesised, preserving statistical structure without carrying operational content.

Case as unit of analysis

The unit of measurement is not the page but the case: an asset-bound dossier whose lifecycle spans days to weeks and dozens of events. Document-level accuracy is necessary but insufficient — a pipeline can be locally accurate and globally slow, and only case-level instrumentation exposes that gap.

Evidence heterogeneity

Document classes arrive through structurally different channels — interactive authenticated portals, bulk structured exports, machine-readable feeds and scanned artefacts of variable quality. Each is admitted through a typed adapter so that no scientific claim is contingent on a particular provider or channel.

Human authorisation as a phase

The decision remains human. Authorisation is modelled explicitly as a phase of the workflow with its own residence-time distribution, which is what allows the study to state precisely how much of the end-to-end latency automation can and cannot remove.

Macro-level management view

Beyond per-case execution, the operational surface must answer aggregate questions: where the queue is forming, which document class dominates rework, which phase regressed this evaluation window. These are treated as first-class research outputs, not as reporting by-products.

Platform for the operator

The framework is instantiated as a console through which an operator executes routine work at scale — scope selection, proposal review, bounded-lot execution with pause, resume and finalise control, and an exportable execution report — with every act signed into the same ledger used for scientific measurement.

Grounded analytical agent

The ledger, the process-mining layer and the predictive layer are exposed to a retrieval-grounded conversational agent that answers bottleneck questions by composing typed queries, returning the executed query and aggregated row count with every assertion, and refusing to answer where no query supports the claim.