Three autonomous agents. One negotiated decision.
The CORTEX Framework is decomposed into three autonomous agents, each governed by an explicit type contract, each independently evaluable through the unit-evaluation harness described in §6, and each able to contest the outputs of its peers. The decomposition is intentionally symmetric — every agent exposes an interface boundary, a computational kernel, a provenance fabric and an emission boundary — which permits the same orchestration, observability and ablation infrastructure to be reused without modification. Coordination is not scripted: agents exchange CLAIM, CHALLENGE, CONGESTION and DISCOVERY messages, renegotiating contested fields and rerouting around saturated peers until the protocol converges. Together they form the canonical configuration of the framework; in isolation each constitutes a falsifiable scientific contribution in its own right.
OPTIC — Generative Optical Perception
OPTIC is the generative optical perception agent of CORTEX. It admits heterogeneous source documents — scanned PDFs, photographed forms, semi-structured layouts, machine-readable XML and structurally noisy email bodies — through a single typed acquisition interface and reads them with a vision–language decoder rather than a classical character recogniser, emitting schema-bound structured records directly from pixels. Perception is decomposed into three composable subsystems: a layout analyser that recovers the physical and logical structure of the page, a generative typed-entity decoder that projects that layout onto a domain-specific schema, and a confidence calibrator that attaches a per-field reliability score derived from agreement across redundant decoding passes at different temperatures. Because generative decoding can produce fluent but unsupported values, OPTIC never asserts a field unilaterally: every extraction is emitted as a CLAIM carrying calibrated confidence and a pixel-level evidence span, leaving adjudication to NEXUS.
Formalises document intelligence as a typed, side-effect-free transformation from unstructured corpora to schema-bound representations admitting compositional reasoning, ablation and audit-grade replay.
token-, entity- and record-level agreement on held-out corpora
expected calibration error of per-field confidence
per page, batched, generative decoder at fixed budget
values with no evidence span in the provenance closure
# OPTIC — Document Intelligence Kernel
from typing import Mapping
from cortex.types import (
SourceRef, DocumentEnvelope, ExtractedRecord,
LayoutTree, Typed, ProvenanceTree,
)
def acquire(source: SourceRef) -> DocumentEnvelope:
"""Typed, source-agnostic acquisition boundary."""
raw = source.fetch()
meta = extract_metadata(raw)
return DocumentEnvelope(
payload = raw,
mime = meta.mime,
provenance = meta.provenance,
)
# Composable, pure transformations
pipeline = compose(
parse_pdf,
segment_layout, # → LayoutTree
project_to_schema, # → Mapping[str, Typed]
calibrate_confidence, # redundant-decoder agreement
)
record: ExtractedRecord = pipeline(envelope)
# record.fields :: Mapping[str, Typed]
# record.confidence :: Mapping[str, float ∈ [0,1]]
# record.layout :: LayoutTree
# record.provenance :: ProvenanceTreeNEXUS — Consensus Arbitration
NEXUS introduces validation as a first-class pipeline stage rather than a post-hoc quality check. Given an extracted record R and an indexed family of independent evidence sources E = {S₁, …, Sₙ}, it computes an agreement functional a(R, E) weighted by an empirically calibrated reliability prior over the sources and returns a confidence-weighted verdict together with a closed provenance subgraph. The module is designed for fan-in/fan-out topologies in which redundant sources attenuate the variance of any individual extractor; disagreement is resolved through a principled arbitration procedure rather than through ad-hoc heuristics, eliminating the silent failure modes characteristic of first-source-wins strategies.
Provides formal consistency guarantees across heterogeneous evidence sources and converts validation from an implicit assumption of the pipeline into an explicit, evaluable kernel.
contested fields resolved to the settled ground truth
evidence retrievals per contested field vs. uncontested
mean CHALLENGE rounds before confidence monotonicity halts
verdicts admitting complete causal replay (invariant I₃)
# NEXUS — Cross-Source Validation Kernel
from nexus.types import (
ExtractedRecord, Source, Verdict, Evidence,
)
from nexus.priors import RELIABILITY_PRIOR
THRESHOLD: float = 0.78 # empirically calibrated
def validate(
record: ExtractedRecord,
sources: list[Source],
) -> Verdict:
"""Agreement functional a(R, E) under heterogeneous reliability."""
evidence: list[Evidence] = [
s.lookup(record.key) for s in sources
]
# Weighted agreement under reliability prior
weights = [RELIABILITY_PRIOR[s.id] for s in sources]
score = weighted_agreement(record, evidence, weights)
return Verdict(
consistent = score >= THRESHOLD,
confidence = score,
evidence = evidence,
arbitration = arbitrate_disagreement(record, evidence),
provenance = build_dag(record, evidence, weights),
)
verdict = validate(record, sources=[s1, s2, s3])
# verdict.consistent :: bool
# verdict.confidence :: float ∈ [0,1]
# verdict.provenance :: DAG[Source, Transform]AEGIS — Symbolic Guardrail
AEGIS encodes operational restrictions as a declarative constraint set C and evaluates it symbolically against the validated record R. The output is an auditable decision-support emission carrying the decision status, the set of violated constraints, and a machine-readable rationale paired with a natural-language justification suitable for downstream human review. The constraint system is constructively monotone in C — adding a restriction never converts a reject into an approve — which guarantees that constraint catalogues can be extended without invalidating prior decisions, a property required for stable longitudinal audit and regulatory compliance.
Enables transparent, restriction-aware decision support with full traceability and longitudinal stability under constraint catalogue evolution.
violations detected against the frozen restriction catalogue
approvals wrongly blocked under the monotone constraint set
decisions carrying machine- and human-readable justification
windows from DISCOVERY to promoted restriction
# AEGIS — Symbolic Restriction Evaluator
from aegis.types import (
ExtractedRecord, Verdict, Constraint,
Decision, Justification, Status,
)
from aegis.catalogue import RESTRICTION_SET
def evaluate(
record: ExtractedRecord,
verdict: Verdict,
rules: frozenset[Constraint] = RESTRICTION_SET,
) -> Decision:
"""Monotone symbolic evaluation of C against (R, verdict)."""
violations: set[Constraint] = {
c for c in rules if not c.holds(record, verdict)
}
status: Status = (
Status.APPROVE if not violations and verdict.consistent
else Status.REVIEW if verdict.confidence >= 0.6
else Status.REJECT
)
return Decision(
status = status,
violations = frozenset(violations),
rationale = Justification.from_violations(violations),
provenance = verdict.provenance.extend(rules),
)
decision = evaluate(record, verdict, RESTRICTION_SET)
# decision.status :: {approve, review, reject}
# decision.violations :: Set[Constraint]
# decision.rationale :: JustificationHELIOS — Process Observability
HELIOS is the observability agent of CORTEX and the instrument through which the framework becomes an object of empirical process science rather than a black box. It consumes the append-only event ledger produced by the operational substrate — phase entries, re-entries, exits, field mutations, human annotations and agent messages — and reconstructs, for each case, a canonical trace whose residence times are attributed once per visit rather than once per emitted row. That distinction is not cosmetic: a phase visited n times emits up to 3n event rows carrying the same duration, and naive summation inflates waiting time by the mean event multiplicity, an artefact measured at 2.4–3.0× during harness calibration and responsible for physically impossible cycle-time estimates in the first evaluation window. From the deduplicated trace HELIOS induces a directly-follows graph over the phase alphabet Φ and emits cycle time, residence time per phase, rework ratio, inter-phase queueing delay and conformance against the declarative process model. Traces the model does not admit are not discarded but escalated as DISCOVERY messages, which is the mechanism by which the agent network learns unmodelled process variants instead of silently mis-measuring them.
Separates machine latency from organisational waiting time, giving bottleneck class B₃ an operational definition that document-level benchmarks cannot express, and supplies the measurement substrate on which every reported process figure rests.
cases with an unbroken phase sequence in the ledger
inflation removed by attributing residence per visit
observed traces admitted by the declarative process model
flagged traces confirmed as defects rather than rare variants
# HELIOS — Process Observability Kernel
from helios.types import (
Event, Trace, PhaseVisit, ProcessGraph, Metrics,
)
FAMILIES = ("PHASE_FIRST_ENTRY", "PHASE_REENTRY",
"PHASE_EXIT", "FIELD_MUTATION", "ANNOTATION")
def canonicalise(events: list[Event]) -> Trace:
"""Normalise, order and repair a raw event stream."""
ordered = sorted(events, key=lambda e: (e.case_id, e.t))
return Trace(
case_id = ordered[0].case_id,
events = [normalise_text(e) for e in ordered],
)
def visits(trace: Trace) -> list[PhaseVisit]:
"""Duration attributed ONCE per visit, not per event row.
A visit is keyed by (phase, visit_ordinal); the three event
families emitted for the same visit carry an identical
duration and must not be summed independently.
"""
seen: dict[tuple[str, int], PhaseVisit] = {}
for e in trace.events:
if e.family not in ("PHASE_FIRST_ENTRY",
"PHASE_REENTRY", "PHASE_EXIT"):
continue
key = (e.phase, e.visit_ordinal)
seen.setdefault(key, PhaseVisit(
phase = e.phase,
ordinal = e.visit_ordinal,
duration = e.phase_duration_s, # counted once
))
return list(seen.values())
def measure(traces: list[Trace]) -> tuple[ProcessGraph, Metrics]:
graph = directly_follows(traces) # G_p = (Φ, ⟶)
return graph, Metrics(
cycle_time = percentiles(cycle_times(traces)),
residence = residence_by_phase(traces),
rework_ratio = rework(traces), # re-entries / visits
queue_delay = handoff_delay(graph),
conformance = conformance(traces, DECLARATIVE_MODEL),
deviations = as_discovery(non_conforming(traces)),
)ORACLE — Predictive Cycle-Time Estimation
ORACLE closes the loop between measurement and operation. It estimates the residual cycle time of an open case from features observable strictly at prediction time — current phase, phase age, prior re-entries, count and class of attached evidence, number of human annotations, the confidence profile emitted by OPTIC and the disagreement mass reported by NEXUS — and returns both a point estimate and an interval. The default estimator is a regularised linear model, chosen deliberately over higher-capacity alternatives: in an audited operational setting the coefficient vector is itself a scientific artefact, a directly falsifiable statement about which process attributes drive delay, and interpretability dominates a marginal accuracy gain that no reviewer or operator could act upon. Evaluation uses a temporally blocked split — train on window w, evaluate on w+1 — because random splitting leaks future process states into training and systematically overstates predictive skill on workflow data, a failure mode we observed and corrected during calibration. Predictions are never actuated autonomously: they are surfaced to the operator console as ranked attention under the same provenance-closure invariant that governs extraction.
Turns retrospective process mining into a forward-looking, auditable decision aid, and provides the empirical instrument through which schema and rule drift (B₄) is detected as a measurable degradation rather than inferred anecdotally.
residual cycle time under a temporally blocked split
against a persistence baseline at 0.38
empirical coverage of the nominal 90% interval
median time between flag and SLA breach
# ORACLE — Residual Cycle-Time Estimator
import numpy as np
from oracle.types import CaseView, Prediction, WindowReport
FEATURES = (
"phase_id", "phase_age_h", "reentries", "attachments",
"annotations", "optic_mean_conf", "nexus_disagreement",
)
def featurise(case: CaseView) -> np.ndarray:
"""Strictly prediction-time observables — no future leakage."""
return np.array([case.get(f) for f in FEATURES], dtype=float)
def evaluate_blocked(windows: list[list[CaseView]]) -> list[WindowReport]:
"""Train on window w, evaluate on w+1. Never a random split:
random splitting leaks future process states and inflates R².
"""
reports = []
for w in range(len(windows) - 1):
X_tr = np.stack([featurise(c) for c in windows[w]])
y_tr = np.array([c.residual_h for c in windows[w]])
model = ridge_fit(X_tr, y_tr, alpha=1.0)
X_te = np.stack([featurise(c) for c in windows[w + 1]])
y_te = np.array([c.residual_h for c in windows[w + 1]])
y_hat = model.predict(X_te)
reports.append(WindowReport(
window = w + 1,
mae = mae(y_te, y_hat),
rmse = rmse(y_te, y_hat),
r2 = r2(y_te, y_hat),
coefficients = dict(zip(FEATURES, model.coef_)),
))
return reports
def predict(case: CaseView, model) -> Prediction:
mu = float(model.predict(featurise(case)[None, :])[0])
return Prediction(
residual_hours = mu,
interval = residual_interval(model, mu),
drivers = top_coefficients(model, k=3),
emitted_to = "operator_console", # never actuated alone
)