Technology

Platform architecture, governance controls, and enterprise implementation reference
Mode: LIVE · Governed Autonomy · Deterministic Evidence · Read-Only Intelligence Fail-Closed · Replayable · Policy-Bound · Tenant-Isolated
Technology Pack - Provenance
Refreshed: 2026-07-20 13:43:31 UTC
SOURCE
Platform architecture + bp Sphere runtime
Scope: v3 unified technology reference
POLICY
Platform governance policy
EVIDENCE
Architecture specifications + runtime telemetry
RECORDS
153 agents, LIVE mode
ACTION
Review technology reference documentation
CapabilitiesOntology + Enterprise ContextEnterprise Memory + Knowledge ManagementValue Attribution Engine
DOC-00 · v3 unified

Executive Summary

What bp Sphere guarantees, and how the live implementation aligns to enterprise governance

bp Sphere is a deterministic, policy-bound Decision Operating System operating alongside BP systems of record.

Operating mode: LIVE · Domains visible: 10 · Entities tracked: 0.

┌──────────────────────────────────────────────────────────────┐
│ USER / EXPERIENCE LAYER                                      │
│ 50+ UI Surfaces · CFO · FP&A · Tower Leaders · Evidence UX   │
├──────────────────────────────────────────────────────────────┤
│ CONVERSATIONAL GATEWAY                                       │
│ Runtime-configured LLM · Explain · Drill · Simulate         │
├──────────────────────────────────────────────────────────────┤
│ bp Sphere API GATEWAY (FastAPI — 300+ Routers)                    │
│ Middleware: Correlation → TenantSep → Auth → Governance      │
├──────────────────────────────────────────────────────────────┤
│ AGENT FABRIC + AGENT MESH                                    │
│ 153 Agents · 8 Towers · 13 Categories                       │
│ Agent Runtime · Supervisor · HITL Queue · Tick Engine        │
├──────────────────────────────────────────────────────────────┤
│ EVENT BUS                                                    │
│ NATS JetStream · CDC Triggers · SSE Streaming                │
├──────────────────────────────────────────────────────────────┤
│ GOVERNANCE CORE                                              │
│ YAML Policies · RBAC/PBAC · AL1-AL4 Autonomy Engine         │
├──────────────────────────────────────────────────────────────┤
│ ONTOLOGY & SEMANTIC LAYER                                    │
│ 4-Layer Object Model · Trigger Engine · Cross-Domain Bridges │
├──────────────────────────────────────────────────────────────┤
│ DATA PLANE + KNOWLEDGE GRAPH                                 │
│ PostgreSQL 5 Domains · Neo4j Knowledge Graph · Evidence Vault│
├──────────────────────────────────────────────────────────────┤
│ INTEGRATION LAYER                                            │
│ 114 SOR Adapters · 10 Fabrics · Circuit Breakers              │
├──────────────────────────────────────────────────────────────┤
│ KUBERNETES CLUSTER                                           │
│ Deployments · StatefulSets · HPA · Health Probes · Ingress   │
├──────────────────────────────────────────────────────────────┤
│ INFRASTRUCTURE                                               │
│ VPC · TLS Termination · Persistent Volumes · Secrets Mgmt    │
└──────────────────────────────────────────────────────────────┘
  • Fail-closed execution — ExceptionHandlerMiddleware catches all unhandled errors, returns structured JSON, increments iris_unhandled_errors_total Prometheus counter
  • Deterministic replay — SHA-256 sealed evidence packs with inputs_hash, outputs_hash, evidence_hash, mandatory backfill in LIVE mode
  • Policy-bound autonomy — 5 YAML policy files (global, p2p, o2c, treasury, model_risk) with AL1-AL4 gating, materiality thresholds ($500K+ requires dual approval)
  • Tenant-isolated data plane — TenantSeparationMiddleware rejects cross-tenant requests (HTTP 403), BP plc-owned config boundary in config/tenants/bp/
  • Drill-level provenance from KPI → execution → evidence artifacts across 30+ model modules and 200+ SQLModel entities
  • Ontology-driven canonical object model — 4-layer design (core, finance, decision, evidence) with trigger engine and cross-domain bridges
  • Knowledge graph reasoning — Neo4j-backed GraphRAG with inference engine for contextual agent decisions
bp Sphere isbp Sphere is not
Decision Operating System sidecar for enterprise SORsERP replacement
Governed orchestration layer for mission and domain agentsUnbounded autonomous executor
Evidence-first and replayable by designBlack-box analytics dashboard
Policy-driven, role-aware intelligence surfaceChatbot wrapper over business data
CapabilityWhat is ImplementedWhere it appears
Mission-control finance UX CFO, FBT, and mission pages use KPI-first operations surfaces with L1-L4 drilldown drawers /ui/cfo-command-center, mission pages, tower views
Enterprise Context Context contracts assemble transaction, historical, policy, risk, and organizational context before a governed decision is made /ui/enterprise-intelligence-fabric, /ui/crest-context-studio
Enterprise Memory P2P decisions expose institutional knowledge, similar-case recall, and knowledge evolution in decision moments Decision Intelligence drawer and P2P scenario drilldowns
Knowledge Management Knowledge registries and learning fabric convert resolved decisions into reusable knowledge and policy improvement proposals /ui/learning-fabric, domain knowledge registries
Decision replay + evidence integrity Replay timeline, hash lineage, and evidence verification are available from KPI and execution drilldowns /ui/decision-replay-studio, evidence tabs, replay APIs
Value Attribution Protected value, explicit impact, and attribution confidence are tied back to decisions and evidence rather than headline KPI claims /ui/value-attribution, decision ledger and mission attribution surfaces
Learning Fabric + CREST Learn-stage surfaces for pattern capture, policy proposals, and context reconstruction are live /ui/learning-fabric, /ui/crest-context-studio
Skill Fabric Skill registry, lifecycle, and replay views support modular, policy-bound capability execution /ui/skill-fabric, /ui/skills-registry, /ui/skill-replay

Implementation stance: finance docs now describe only deployed capabilities wired to live routes and drilldown contracts.

Agents
153
153 implemented
API Endpoints
300+
across 60+ routers
UI Surfaces
50+
HTML pages
Finance Towers
8
R2R, P2P, I2C, FPA, Treasury, Governance, Platform, Upstream
SOR Services
114
across 5 domain databases
Fabrics
10
ai_mesh, workflow, guardian, trust, ...
ComponentCountDetails
Middleware Layers6Correlation, TenantSeparation, CORS, Auth, Governance, ExceptionHandler
SQLModel Entities120+Across 24 model modules (_agents, _r2r, _p2p, _i2c, _governance, _treasury, etc.)
YAML Config Files50+Platform + Tenant + Domain + Mission configs
Services40+Supervisor, evidence, policy, tick engine, SOR adapters, HITL, etc.
Startup Steps22From DB init through evidence backfill loop
Prometheus MetricsActiveCustom counters, histograms, and SLO targets (P95 < 500ms, P99 < 2s)
Fabrics10ai_mesh, control_plane, data, federated_gateway, guardian, integration, payment_gateway, saga, trust, workflow
Ontology YAML Files8+kernel + finance domain + tenant packs
Knowledge GraphActiveNeo4j-backed GraphRAG + inference engine
Top to Bottom Visual Stack
bp Sphere Platform Architecture Layering
Experience -> Intelligence -> Decision -> Execution -> Governance -> Context -> Integration -> Infrastructure
8. Experience & Command Center
CFO command center, domain workspaces, and conversational interfaces for explainable, drillable, role-adaptive decision operations.
▼ ▼ ▼
7. Intelligence Services
Enterprise simulation, forecasting, risk prediction, optimization, scenario modeling, and what-if analysis operating on governed decision context.
▼ ▼ ▼
6. Decision Fabric
Reasoning outputs converted into governed decisions with decision ledger records and evidence vault traceability for replay and audit.
▼ ▼ ▼
5. Agentic Execution Layer
Agent runtime, agent mesh, and mission engine orchestrating enterprise work (P2P, O2C, R2R, Treasury, FP&A) with deterministic event triggers.
▼ ▼ ▼
4. Governance Core
Policy engine, supervisory control, autonomy boundaries (AL1-AL4), RBAC/PBAC, authority queues, and segregation-of-duties controls.
▼ ▼ ▼
3. Enterprise Context Layer
Ontology + knowledge graph + canonical models (invoice, PO, payment, asset) to provide semantic consistency and cross-domain reasoning context.
▼ ▼ ▼
2. Integration & Data Plane
SOR connectivity and data/event fabric: ERP, Ariba, Maximo, Workday, Salesforce, bank APIs via adapters, API gateway, schema harmonization, and NATS streaming.
▼ ▼ ▼
1. Infrastructure Layer
Kubernetes/k3s, compute (CPU/GPU), Postgres, object storage, Redis, secrets, and observability (Prometheus, Grafana, Loki, Jaeger).

Latest platform capability update

CapabilityStatusRuntime proofSurface / endpoint
Experience Orchestration LayerSurface availableShared decision objects render across desktop, mobile, and chat surfaces./ui/mission/p2p, /ui/mobile-cockpit, /ui/enterprise-chat
Governed Chat InterfaceSurface availableChat resolves role, mission, case, evidence, policy, action contract, replay, and feedback./ui/enterprise-chat
Workforce Intelligence Integration FabricSurface availableAssistant registry, connector packages, certification, compliance mappings, identity resolution, failover drills, and workforce coordination APIs./ui/workforce-intelligence-fabric, /ui/assistant-operations-center
Token Economy Control PlaneSurface availableLLM calls are routed, cached, denied, downgraded, or escalated before spend occurs; token value is attributed to decisions./ui/ai-cost-usage, /api/iaf/decision-intelligence/token-economy/decision
Guardrail Policy EngineSurface availablePolicy evaluation fails closed in production posture; non-fail-closed failures route to escalation instead of silent execution./ui/policy-registry, /ui/control-plane
Decision Replay + LLM DeterminismSurface availableExecution seed, prompt/context hash, evidence hash, policy version, and replay records are persisted and validated./ui/decision-journey, /ui/single-decision-audit, /ui/decision-replay-studio
Externally Verifiable EvidenceSurface availableEvidence bundles are signed with asymmetric verification and can be checked through the verify endpoint./ui/evidence, /api/iaf/decision-intelligence/evidence/verify
Runtime Adoption AuditEndpoint availableAgents/workflows are audited for identity, policy, evidence, replay, value, token, and learning seam adoption./api/iaf/decision-intelligence/runtime-adoption-audit
Skill-Binding Learning PromotionEndpoint availableSkill learning overlays require evidence, confidence, and runtime effect before promotion into runtime metadata./api/iaf/decision-intelligence/skill-learning-overlays/{id}/validate
Autonomy Demotion EnforcementEndpoint availableTelemetry demotion guidance writes supervisory state, event logs, and owner review queue items./api/iaf/decision-intelligence/autonomy/demotions/enforce
Shared Case-State PropagationEndpoint availableOne case resolution updates analyst state, supervisor state, value records, replay reference, and CFO rollup readiness./api/iaf/decision-intelligence/case-state/{case_id}/rollup
Observability ReadinessEndpoint available, env-boundProduction telemetry posture exposes OTEL endpoint status, Prometheus/failure-mode surfaces, and SLO targets./api/iaf/decision-intelligence/observability/readiness

Bangalore runtime proof additions

CapabilityCurrent implementation
Continuous Finance Command CenterContinuous Close is positioned as a finance digital twin: actuals + forecast accruals + forecast provisions + open risks + controls + evidence = current financial position.
Architect-run CLI proofbpsphere now, bpsphere close-now, bpsphere financial-position, bpsphere digital-twin, bpsphere close-runtime, bpsphere runtime, and bpsphere provenance prove runtime status without the browser.
Provenance badges live API query · seeded BP-shaped runtime baseline · estimate/model output with disclosed basis · deterministic replay/policy path.
Self-disclosing demo dataHeadline workshop magnitudes are no longer presented as unexplained static numbers. CLI and UI proof paths disclose data basis, query timestamp, metric lineage, calculation method, replay ID, and source posture.
Production SOR boundaryProduction SOR access remains fail-closed until BP supplies credentials and feeds. This prevents claiming SAP/S4/CFIN/BlackLine production connectivity without validated access.

External production bindings

  • SAP/SOR phase 2-3: RFC/OData reads, CDC streams, governed write-back, and master-data sync require customer credentials and agreed write boundaries.
  • Live Entra tenant/Graph sync and vendor assistant connector credentials are external deployment bindings.
  • Production OTEL/SLO alerting requires environment configuration and failure-mode drills in dev/prod.
  • Agent de-stubbing and route-level adoption breadth continue as capability migration work, not platform architecture gaps.

Previously shipped implementation baseline

AreaCurrent implementation
Runtime modesIRIS_RUNTIME_MODE drives dev, demo, shadow, production behavior through runtime profiles.
Finance AI model controlModel selection is centralized on OPENAI_MODEL; current active setting resolves to gpt-5.4-mini.
External truth layerBP databook disclosures are ingested into DB-backed reconciliation tables and shown in FP&A via Reported vs bp Sphere surfaces.
FP&A mission stackReconciled variance, simulation presets, LLM narrative, evidence sections, and decision-oriented V2 mission panels are deployed.
Architectural Capability View

bp Sphere Technology Pack

Architecture, governance, context, trust, and learning capabilities across the live bp Sphere system.

Trust Posture

Replay determinism and evidence coverage from the live runtime.

0.0% replay determinism · 0.0% evidence coverage

Context Sophistication

Semantic and memory-backed context available to decisions.

0 graph-backed decisions · 0 memory-assisted decisions

Workflow Sophistication

Mission and workflow orchestration running through the transaction spine.

0 workflow instances · 0 objects under spine

Operating Scope

Shared platform deployed through tenant and domain overlays.

10 canonical domain packs · 114 SOR adapters

ERCS Gate

Measured release posture from replay and evidence health.

BLOCKED · 0.0/100

Enterprise Architecture (5-Layer)
Experience Layer · CFO command center, mission workspaces, conversational UI
Intelligence Layer · Forecasting, simulation, anomaly detection, AI reasoning
Decision Fabric · Decision ledger, evidence vault, policy enforcement
Agentic Execution Layer · Agents, missions, triggers, orchestration
Data & Infrastructure Layer · Integration fabric, ontology, data services, Kubernetes
Latest Runtime Capability Update
CapabilityStatusRuntime proofSurface / endpoint
Experience Orchestration LayerSurface availableShared decision objects render across desktop, mobile, and chat surfaces./ui/mission/p2p, /ui/mobile-cockpit, /ui/enterprise-chat
Governed Chat InterfaceSurface availableChat resolves role, mission, case, evidence, policy, action contract, replay, and feedback./ui/enterprise-chat
Workforce Intelligence Integration FabricSurface availableAssistant registry, connector packages, certification, compliance mappings, identity resolution, failover drills, and workforce coordination APIs./ui/workforce-intelligence-fabric, /ui/assistant-operations-center
Token Economy Control PlaneSurface availableLLM calls are routed, cached, denied, downgraded, or escalated before spend occurs; token value is attributed to decisions./ui/ai-cost-usage, /api/iaf/decision-intelligence/token-economy/decision
Guardrail Policy EngineSurface availablePolicy evaluation fails closed in production posture; non-fail-closed failures route to escalation instead of silent execution./ui/policy-registry, /ui/control-plane
Decision Replay + LLM DeterminismSurface availableExecution seed, prompt/context hash, evidence hash, policy version, and replay records are persisted and validated./ui/decision-journey, /ui/single-decision-audit, /ui/decision-replay-studio
Externally Verifiable EvidenceSurface availableEvidence bundles are signed with asymmetric verification and can be checked through the verify endpoint./ui/evidence, /api/iaf/decision-intelligence/evidence/verify
Runtime Adoption AuditEndpoint availableAgents/workflows are audited for identity, policy, evidence, replay, value, token, and learning seam adoption./api/iaf/decision-intelligence/runtime-adoption-audit
Skill-Binding Learning PromotionEndpoint availableSkill learning overlays require evidence, confidence, and runtime effect before promotion into runtime metadata./api/iaf/decision-intelligence/skill-learning-overlays/{id}/validate
Autonomy Demotion EnforcementEndpoint availableTelemetry demotion guidance writes supervisory state, event logs, and owner review queue items./api/iaf/decision-intelligence/autonomy/demotions/enforce
Shared Case-State PropagationEndpoint availableOne case resolution updates analyst state, supervisor state, value records, replay reference, and CFO rollup readiness./api/iaf/decision-intelligence/case-state/{case_id}/rollup
Observability ReadinessEndpoint available, env-boundProduction telemetry posture exposes OTEL endpoint status, Prometheus/failure-mode surfaces, and SLO targets./api/iaf/decision-intelligence/observability/readiness

Production bindings still external

  • SAP/SOR phase 2-3: RFC/OData reads, CDC streams, governed write-back, and master-data sync require customer credentials and agreed write boundaries.
  • Live Entra tenant/Graph sync and vendor assistant connector credentials are external deployment bindings.
  • Production OTEL/SLO alerting requires environment configuration and failure-mode drills in dev/prod.
  • Agent de-stubbing and route-level adoption breadth continue as capability migration work, not platform architecture gaps.
Decision Lifecycle
Event
Mission
Execution
Policy
Decision
Evidence
Ledger
Learning Feedback
Architectural Capability Map

Shared Platform Layer

Reusable runtime capabilities: event ingestion, context graph, agent orchestration, policy gates, evidence vault, replay, learning, observability, and action routing. These are tenant-neutral and should not carry bp-specific process assumptions.

Boundary: platform owns runtime primitives, control contracts, schemas, telemetry, and reusable services.

bp Tenant Overlay

bp-specific configuration: source-system endpoints, mission packs, policies, approval matrices, glossary terms, evidence mappings, role models, SOR access rules, and validation datasets.

Boundary: bp overlay owns business semantics, policies, data access, evidence sources, and operational ownership.

Knowledge Management + Enterprise Memory

Decision memory, knowledge registries, GraphRAG, and learning fabric convert institutional knowledge into reusable operating logic.

Live surfaces: Decision Memory · Learning Fabric

Value Attribution

Value is linked to decisions through recorded impact, protected value, attribution confidence, and double-counting controls.

Live surfaces: Value Attribution · Decision Ledger

Trust, Evidence, and Replay

Every governed decision can carry evidence hashes, lineage, replay support, and policy/version provenance.

Live surfaces: Evidence Vault · Decision Replay Studio

Skill Fabric + Agentic Runtime

Skills, missions, triggers, and bounded autonomy are composed into reusable enterprise execution flows.

Live surfaces: Skill Fabric · Runtime Overview

Policy + Authority Control

Human authority, thresholds, queues, inline policy context, and supervisory control remain part of the runtime contract. Policy now sits as a hard gate between the lifecycle and autonomy gates in the supervisory chain.

Live surfaces: Policy Library · Decision Records · Policy Studio

Enterprise Runtime Certification Suite (ERCS) — Measured Gate

ERCS Certification Suite

Twelve certification categories — ontology execution, policy governance, evidence integrity, agentic fabric, replay determinism, supervisor governance, external signal intelligence, simulation runtime, value attribution, AI adoption, institutional learning, infrastructure resilience — each backed by runtime paths and tests. Production releases for bp pass through an explicit ERCS gate before publish; the UI must show measured score and failed controls rather than imply universal Platinum status.

Current measured gate: BLOCKED · score 0.0/100 · replay 0.0% · evidence 0.0%

Example. A finance pack passes 11 of 12 categories but fails the simulation-fidelity test (forecast drift > tolerance). The ERCS release gate blocks promotion, the failing category posts to the certification inbox, and operators see the exact assertion that failed — not a generic "deploy blocked". Illustrative — categories and thresholds are configurable per release.

Live surfaces: ERCS Scorecard · ERCS Proof Suite · Certification Inbox

Strategic Intelligence Layer

External signal fabric ingests commodity (Brent / WTI / gas), FX, logistics, weather, geopolitical, regulatory and supplier-risk feeds. A correlation engine walks ontology relationships to surface cross-domain impact, drives forecast drift detection, and generates governance-grounded strategic recommendations with confidence scoring.

Example. An external signal (e.g. Brent crude) is ingested as bpsphere.strategic.signal_ingested. The correlation engine resolves cross-domain impact through ontology relationships (treasury, capex, downstream, procurement, forecast drift) and emits bpsphere.strategic.impact_detected followed by bpsphere.strategic.recommendation_generated. Affected missions are surfaced on the executive copilot with citations to decision_id, policy_version, and evidence_hash. Illustrative — signal source, agent count, and timing vary by tenant and live data state.

Live surfaces: Strategic Intelligence · Executive Copilot

Simulation & Digital Twin Runtime

Enterprise digital twin engine with scenario branching, counterfactual replay over the existing replay tape, policy simulation against historical decisions, and stress-test fixtures (CFO liquidity, supply-chain Gulf delay, payment-hold threshold). Simulation runs anchor back to original decision IDs so value attribution survives replay.

Example. "What if Brent −12%, DSO +7d, FX −4% next quarter?" The scenario engine projects cash flow, working capital, covenant exposure, treasury pressure and margin compression along three branches. A separate counterfactual replays a prior period’s payment-hold decisions under an alternate threshold and quantifies the risk delta — without touching production state. Illustrative — signal mix, branch count, and threshold values are configurable.

Live surfaces: Simulation Runtime · Replay Theater

Governed Institutional Learning

Pgvector-backed memory graph with similar-case retrieval. Learning proposals are bounded edits (prioritization weights, escalation thresholds, confidence calibration only — never rules or ontology) and pass through a replay-reproducibility gate plus Trust-fabric attestation before any apply. Review, apply and rollback permissions are split four-eye.

Example. A supervisor consistently overrides low-confidence invoice holds (all released as safe). The system proposes raising the auto-release confidence threshold — a bounded prioritization edit. The replay-reproducibility test confirms the original decisions still verify under the proposed threshold. Reviewer (learning.evolution_review) approves; a different role (learning.apply) applies. Policy rules unchanged, ontology unchanged, the original signed bundle still verifies. Illustrative — thresholds and counts vary by feedback volume.

Live surfaces: Learning Proposals · Institutional Memory

KPI Lineage + Live Data Credibility

Every dashboard KPI traces UI → SOR query → GL posting → vendor / entity reference → freshness SLA, with partial-coverage disclosure when an upstream is unavailable. Closes the "show me the evidence" question at the CFO axis.

Example. A DSO drill-down opens its live lineage: iaf_customer_invoices rows aggregated by aging bucket, evidence-coverage percentage shown, vendor and customer references resolved against iaf_customers and iaf_vendors. When an upstream SOR is unreachable, the lineage row marks data_source_type=UNAVAILABLE and the KPI value carries a partial-coverage disclosure. The same path is implemented for DPO and CCC. Illustrative — specific values vary by tenant and live data state.

Live surfaces: Decision Trace · Value Attribution · Adoption Intelligence

Event Fabric Hardening & Auto-Rollback

NATS JetStream durable subjects across bpsphere.signal.*, bpsphere.strategic.*, bpsphere.simulation.*, bpsphere.learning.*. EventDispatcher with idempotency ledger and dead-letter queue. AutoRollbackWorker subscribes to bpsphere.regression.detected and reverses a published pack when z-score breaches threshold. Guardian and Trust fabrics subscribe to supervisory events on bpsphere.>.

Example. A bad pack publish drives p99 latency on a finance mission past its rolling-window threshold. Observability detects the z-score breach, emits bpsphere.regression.detected, AutoRollbackWorker rewinds to the prior signed pack and emits bpsphere.governance.rollback_executed with the rationale plus before/after metrics on the audit log. Duplicate deliveries are absorbed by the idempotency ledger. Illustrative — latency and z-score thresholds are configured per mission; rollback timing depends on stream lag.

Live surfaces: Runtime Observability · Mission Pack Manager · Operator Workflow

Cross-cutting Platform Substrate

Saga Workflow Orchestration

Forward + backward recovery for cross-SOR transactions. Steps declare compensations; on failure the engine plays compensations in reverse with bounded retries. Step dependencies, retry policy, and persistent state are all first-class. Used internally by mission orchestration; available to tenant pack authors via the SDK.

Example. A four-step payment workflow reserves a budget, posts a journal entry, notifies the vendor, and updates the AP ledger. If the vendor notification fails, the saga compensates by reversing the journal entry and releasing the budget. The original transaction never partially commits. Illustrative — step shape and compensation policy are configurable per workflow.

Live surfaces: shared_capabilities/workflows/saga.py · Composition Studio

Transaction Spine + Idempotency

Every cross-SOR action passes through an idempotency registry keyed by transaction id + request hash. Replays, duplicate webhooks, and retried API calls deduplicate cleanly. Paired with the EventDispatcher’s ledger-backed idempotency, the platform absorbs at-least-once delivery without double-actioning the upstream SOR.

Example. A network partition retries the same NATS-delivered "post invoice" message three times. The transaction spine recognises the identical idempotency key on the second and third deliveries and replies with the original action result — SAP receives exactly one journal posting. Illustrative — key shape and retention window are configurable per SOR.

Live surfaces: shared_capabilities/transaction_spine/idempotency.py

LLM Gateway with Cost-Aware Routing

Five routing strategies (primary-fallback, cost-optimised, latency-optimised, round-robin, quality-threshold) across nine providers (Anthropic, OpenAI, Groq, Cohere, Ollama, NVIDIA, Mistral, Grok, vLLM). Per-request max_cost budgets, prompt versioning with DRAFT/ACTIVE/ARCHIVED lifecycle and built-in A/B testing.

Example. A cost-optimised routing strategy sends a routine vendor-classification prompt to the cheapest available model that meets the configured quality threshold; an SLA-critical CFO summarisation prompt is force-routed to a flagship model under the latency-optimised policy. Prompt A/B variants are picked deterministically per decision_id so replay is reproducible. Illustrative — routing policy and quality thresholds are configurable per tenant.

Live surfaces: shared_capabilities/ai/llm/gateway.py · shared_capabilities/ai/prompts/manager.py

Observability Stack — OTEL + SLOs + Per-Tenant Rate Limit + Residency

OpenTelemetry auto-instrumentation for FastAPI / SQLAlchemy / psycopg2 / HTTPx is on by default. Canonical platform SLOs (availability 99.9%, p95 latency, error rate) are recorded against every /api/platform/* request and surfaced at /api/platform/slo. A sliding-window rate limiter shapes per-tenant / per-API-key traffic. Data-residency policy is enforced as HTTP 451 on cross-border violations when a tenant declares a region allow-list.

Example. A tenant configured with allowed_regions: [EU, UK] in tenant_metadata.yaml receives a 451 from /api/platform/* when the ingress geo-IP resolves the request to US; the violation is audit-logged with the tenant, framework rule, and source region. Operator can run the same policy in log_only mode first to stage rollout. Illustrative — region allow/deny lists and enforcement mode are configurable per tenant.

Live surfaces: /api/platform/slo · shared_capabilities/observability/ · shared_capabilities/security/http.py · shared_capabilities/trust/compliance/residency.py

bp Databricks Streaming Lineage

Dedicated databricks_sor.event_ingest.DatabricksEventIngestor subscribes to bpsphere.decision.logged, bpsphere.agent.*.evidence.committed, bpsphere.policy.attestation.created, bpsphere.gate.decision and writes through to the databricks_event_lineage Delta table — closing the streaming Postgres → Lakehouse gap for bp analytics.

Example. A duplicate-payment agent decides to hold an invoice. The decision streams into databricks_event_lineage carrying the original decision_id, the gate verdict, the evidence_hash, and the attached value attribution — queryable from the bp analytics workbench alongside historical decisions in the same Delta table. Illustrative — latency and retention horizon depend on stream load and Lakehouse retention policy.

Live surfaces: Tenant Overlay Manager · Observability

Supervisory Control Surface

Mission- and tenant-scope kill switches with 4-hour gate-escalation SLA and 1-hour kill-switch review SLA. Operator playbooks (e.g. pack promotion checklist) run synchronously with bounded steps and survive mid-chain agent failure with context inheritance and evidence sealing intact.

Example. During an SOR brownout, an operator pauses a mission with a one-click action via the supervisory control surface. No new agents activate, in-flight escalations stay visible in the supervisor queue, replay history stays sealed, and existing escalation SLAs continue ticking. When the SOR recovers the operator resumes the mission and the queue drains naturally — no decisions silently rerun. Illustrative — SLA values are configurable in the supervisory policy pack.

Live surfaces: Operator Workflow · Certification Inbox · Composition Studio

Supporting Architecture Views

Runtime Overview · Decision Memory · Enterprise Intelligence Fabric · Technology → Runtime Mapping · Value Attribution · CREST

Platform-tier surfaces: Platform Home · ERCS Scorecard · ERCS Proof Suite · Strategic Intelligence · Simulation · Executive Copilot · Learning Proposals · Institutional Memory · Composition Studio · Observability · Replay Theater · Certification · Tenant Overlay · Policy Studio · Ontology Studio · Pack Manager · Operator Workflow · Value Attribution · Decision Trace · Adoption Intelligence

Ontology + Enterprise Context

Semantic spine, business objects, runtime context bundles, and cross-domain references that normalize meaning before execution.

Why it matters: Makes decisions context-aware and tenant-consistent instead of prompt-only or system-local.

CREST Context Reconstruction

Context reconstruction layer connecting enterprise signals, institutional memory, policy context, and decision-ready state.

Why it matters: Lets agents and humans work from reconstructed enterprise context, not fragmented application screens.

Enterprise Memory + Knowledge Management

Decision memory, knowledge registries, GraphRAG, and learning fabric that preserve institutional knowledge beyond individuals.

Why it matters: Turns prior decisions and domain knowledge into reusable operating intelligence.

Value Attribution Engine

Decision-linked financial impact, protected value, effort avoided, and attribution confidence with double-counting controls.

Why it matters: Connects automation and governed judgment to cash, cost, risk, and working-capital outcomes.

Trust Fabric + Replay

Evidence vault, deterministic replay, lineage, integrity hashes, and replay-readiness checks across runtime decisions.

Why it matters: Provides proof that decisions are reproducible, governable, and audit-ready.

Skill Fabric + Agentic Runtime

Skill composition, mission orchestration, event-driven triggers, and bounded autonomy across agents and roles.

Why it matters: Moves from isolated assistants to governed enterprise execution.

Policy + Authority Control

Role-aware authority thresholds, supervisory control, HITL queues, inline policy context, and fail-closed decision gating.

Why it matters: Keeps humans in control at the right boundary while still scaling automation.

Event Fabric + SOR Sidecar

NATS/event routing, SOR adapters, workflow instances, and mission triggers running beside enterprise systems of record.

Why it matters: Lets bp Sphere operate as a decision sidecar rather than an ERP replacement.

bp Tenant — Live State

Every count and identifier below is bp tenant specific. Platform capabilities described elsewhere (148-pod SOR fabric, 9-provider LLM gateway, EIF, FIL, Context Studio) are available to bp; this page shows what bp tenant currently exercises.

Agent definitions
557
tenants/bp/program/agents
Tenant extension packs
2
inherit finance-context-pack
EIF case packs
6
JE / Contract / Credit / Maint / Spend / R2R
SORs highlighted
11
+ dynamic via LiveSORClient
bp Sphere pods
8
prod / dev / worker / showcase / smoke
Production DB
440 GB
iris_agentic_finance_db on finance-pg
Decision revisions
~22M
iaf_agent_execution_revisions (signed)
Certification
Runtime-gated
release gate uses measured ERCS evidence
DeploymentRoleReplicas
bp Sphere productionsor-iris-agentic-finance1/1
bp Sphere production workersor-iris-agentic-finance-worker1/1
Apex sub-tenant surfacesor-iris-agentic-finance-apex1/1
bp Sphere devsor-iris-agentic-finance-dev2/2
bp Sphere dev workersor-iris-agentic-finance-dev-worker1/1
bp Sphere showcasesor-iris-agentic-finance-showcase1/1
bp Sphere showcase workersor-iris-agentic-finance-showcase-worker1/1
bp Sphere compact smokesor-iris-agentic-finance-compact-smoke1/1
SurfaceVolumePurpose
iaf_agent_execution_revisions~22 M rowsSigned decision audit (replay-ready)
iaf_policy_evaluation_logs~79 M rowsPolicy gate firings per request
iaf_agent_monitor_events~92 M rowsLive agent telemetry stream
iaf_ontology_runtime_links~60 M rowsObject ↔ instance binding history
iaf_agent_executions~20 M rowsAgent run inventory
iaf_finance_events~30 M rowsFinance event stream
Public endpoint

bp tenant currently served at https://prod.sphere-iris.com via Cloudflare worker (quick tunnel → in-cluster ingress → bp Sphere shell → bp Sphere dev service). Other tenants on the same fabric have their own worker URLs and are not represented in this tech pack.

Bangalore runtime proof path

The bp tenant now exposes an architect-run command-line path: bp-now, bp-close-now, bp-financial-position, bp-digital-twin, bp-close-runtime, bp-runtime, bp-provenance. Every headline runtime number must disclose provenance badges: ● live query · ◐ seeded baseline · ◇ estimate/model · ◆ deterministic replay/policy. This is intentionally self-disclosing: the runtime query is live, workshop baselines are BP-shaped seeded data, and production SOR connectivity remains fail-closed until BP credentials and feeds are supplied.

Continuous Finance Command Center

Continuous Close is now documented as a finance digital twin, not a close dashboard. The hero proof asks: If BP closed the books right now, what would the financial statements look like? The runtime assembles actuals, forecast accruals, forecast provisions, reconciliations, controls, evidence, confidence, open exceptions, and replay into a current financial position.

LLM provider posture

9 providers are wired at the platform level (Anthropic, OpenAI, Gemini, Cohere, Grok, Groq, NVIDIA, Ollama, vLLM). bp tenant runs OpenAI via environment-configured model in production today; the gateway will route to a fallback provider on request, gated by eval_harness response-quality scoring before display.

Frequently Asked Questions

Q: Why does this differ from the platform-wide tech pack?
A: This pack scopes every count to bp. Multi-tenant claims (other tenants, marketplace breadth, shared infrastructure scale) are described as platform capabilities the bp tenant inherits, not as bp's own counts.
Q: What is the difference between bp tenant and the 148-pod SOR fabric?
A: The 148 SOR pods are a shared platform service representing the bp source-system landscape (SAP, Ariba, ServiceNow, Treasury, etc.) as production-equivalent APIs. bp tenant currently integrates with 11 of those SORs explicitly through the bp Sphere runtime code path, plus dynamic resolution via LiveSORClient.
Q: Is the showcase deployment counted in bp production?
A: No. Showcase and compact-smoke are isolated environments used for demos and smoke testing. They share the bp code base and ontology but run on smaller datasets.

bp Tenant — SOR Integration Estate

Registered SOR endpoints
114
from service_ports.json
Highlighted BP services
11
critical workshop integration examples
Activation model
DNS / adapter target swap
same LiveSORClient path
                bp tenant bp Sphere runtime
                        │
                        │ LiveSORClient (HTTP / async)
                        │ same code path used at tenant activation
                        ↓
        ┌────────────┬───┴──────┬──────────┬──────────┐
        ↓            ↓          ↓          ↓          ↓
   FINANCE      PROCUREMENT   RISK      MARKET     DATA
   ───────      ───────────  ──────    ───────    ─────
   sor-treasury sor-ariba   sor-cyber-risk        sor-databricks
                sor-coupa   sor-ai-gov-trust
                            sor-audit-controls-fabric

                                       FEEDS
                                       ─────
                                       sor-commodity-price-feed
                                       sor-fx-rate-feed
                                       sor-macro-indicator-feed

                  OPERATIONS
                  ──────────
                  sor-servicenow
SOR (in-cluster service)bp integration purpose
sor-ai-gov-trustAI governance signals
sor-aribaProcurement: suppliers, requisitions, RFQs
sor-audit-controls-fabricSOX / control attestations
sor-commodity-price-feedBrent, gas, LNG market data
sor-coupaBSM / sourcing
sor-cyber-riskSupplier cyber posture
sor-databricksAnalytics + ML data plane
sor-fx-rate-feedFX rate market data
sor-macro-indicator-feedMacro economic indicators
sor-servicenowWork orders, incidents, change management
sor-treasuryCash, FX exposure, sanctions
Why these are real wires, not external mocks

Each SOR runs as its own production-shape FastAPI service in the same cluster (its own port, its own database, its own validation logic). bp Sphere calls them over real HTTP through LiveSORClient — the same code that will call bp's actual SAP, Ariba, ServiceNow, and Treasury endpoints at tenant activation. Only the DNS target inside each adapter changes.

🔍 Illustrative operating example — bp Sphere → sor-ariba call pattern (live)
# iris_runtime/src/iris_sor/services/sor_client.py (simplified)

async with LiveSORClient(target="sor-ariba") as client:
    suppliers = await client.get(
        path="/api/v1/suppliers",
        params={"realm_id": "bp", "active": True},
    )
    if not suppliers:
        raise SorDataError("ariba supplier set empty for bp realm")

# At tenant activation, the only change:
#   sor-ariba.platform-services.svc.cluster.local
# → bp.ariba.cloud (or whatever bp's actual Ariba endpoint is)
# Code path, evidence pack assembly, FIL sealing are unchanged.
Full registered SOR estate

The service registry currently exposes 114 registered SOR endpoints covering finance, procurement, treasury, trading, tax, HR, risk, planning, market data, reporting, and operational systems. The table above is the critical BP workshop subset because those services directly support the current mission demos: governance, procurement, controls, market signals, cyber risk, Databricks, FX, macro, ServiceNow, and treasury. Additional endpoints such as SAP S/4HANA, SAP GRC, Workday, BlackLine, Bloomberg, Endur, Murex, Kyriba, Salesforce, ServiceNow GRC, Avalara, Vertex, OneStream, Snowflake, and Power BI are already present in the registry and can be bound into bp workflows through configuration.

Frequently Asked Questions

Q: Why does the page show 11 services and also over 100 registered SOR endpoints?
A: The 11 services are the highlighted in-cluster SORs used to explain the current bp workshop flows. The broader estate comes from service_ports.json, which currently registers 114 SOR endpoints. That registry includes the wider enterprise service landscape and is what the runtime uses for adapter discovery and endpoint resolution.
Q: How does bp tenant decide which SOR to call for which decision?
A: The EIF case pack (e.g. credit_intelligence) declares which SORs are required for evidence assembly. The agent runtime resolves the SOR list, makes the calls in parallel, and gates the pack on evidence completeness before passing to the policy layer.
Q: What happens if an SOR is down during a bp decision?
A: SorDataError is raised, the partial pack is sealed by FIL with a degraded-mode marker, and the decision either falls back to a cohort precedent or escalates to HITL. Recovery is observable in the FIL revision audit.

bp Tenant Pack Inheritance

        Platform shared pack (registry)
        ────────────────────────────────
           finance-context-pack
                  ▲
                  │ parent_pack_id
                  │
        ┌─────────┴────────────────────┐
        ↓                              ↓
   bp-upstream-ap-pack            bp-trading-credit-pack
   (tenant_extension)             (tenant_extension)
   beta · 87% coverage            beta · 84% coverage

   Adds:                          Adds:
     Objects:                       Objects:
       Contract                       Counterparty
       Milestone                      Exposure
       AFE                            MarketSignal
       Waiver
                                    Decisions:
     Decisions:                       dynamic_credit_limit_review
       contract_milestone_              proactive_credit_renewal_
         invoice_validation               intervention
       maintenance_consumption_
         service_verification         Policies:
                                        BP.CREDIT.LIMIT_THRESHOLD
     Policies:                          BP.CREDIT.EARLY_WARNING
       BP.P2P.CONTRACT_VARIANCE
       BP.P2P.MILESTONE_BILLING
FieldValue
pack_idbp-upstream-ap-pack
pack_typetenant_extension
parent_pack_idfinance-context-pack (shared)
maturitybeta
certification statecurated
coverage pct0.87
owner teambp-finance-transformation
business stewardbp-p2p-governance
technical stewardbp-context-stewardship
referencestenant_extensions/bp/upstream_ap/context_pack.yaml
FieldValue
pack_idbp-trading-credit-pack
pack_typetenant_extension
parent_pack_idfinance-context-pack (shared)
maturitybeta
certification statecurated
coverage pct0.84
owner teambp-credit-risk
business stewardbp-credit-governance
technical stewardbp-context-stewardship
referencestenant_extensions/bp/trading_credit/context_pack.yaml
🔍 Illustrative operating example — Pack inheritance YAML (live registry)
# platform/registries/context_marketplace.yaml

packs:
  - pack_id: bp-upstream-ap-pack
    pack_type: tenant_extension
    parent_pack_id: finance-context-pack
    maturity: beta
    coverage_pct: 0.87
    includes:
      objects: [Contract, Milestone, AFE, Waiver]
      decisions:
        - contract_milestone_invoice_validation
        - maintenance_consumption_service_verification
      policies:
        - BP.P2P.CONTRACT_VARIANCE
        - BP.P2P.MILESTONE_BILLING
    references:
      - tenant_extensions/bp/upstream_ap/context_pack.yaml

  - pack_id: bp-trading-credit-pack
    pack_type: tenant_extension
    parent_pack_id: finance-context-pack
    maturity: beta
    coverage_pct: 0.84
    includes:
      objects: [Counterparty, Exposure, MarketSignal]
      decisions:
        - dynamic_credit_limit_review
        - proactive_credit_renewal_intervention
      policies:
        - BP.CREDIT.LIMIT_THRESHOLD
        - BP.CREDIT.EARLY_WARNING
    references:
      - tenant_extensions/bp/trading_credit/context_pack.yaml
Tenant Certification posture

bp tenant uses a measured 13-domain Tenant Certification rubric across context, evidence, policy, replay, agent runtime, security, and operations. The publish-blocker gate prevents merging code that drops any domain below threshold, and the certification surface should show the current score and failing controls rather than a static perfect-status claim.

Frequently Asked Questions

Q: Why two bp packs instead of one?
A: The two packs cover materially different decision surfaces: bp-upstream-ap-pack covers P2P / upstream AP (Contract, Milestone, AFE, Waiver objects), while bp-trading-credit-pack covers trading credit exposure (Counterparty, Exposure, MarketSignal). Splitting them allows independent stewardship and lifecycle management.
Q: Can bp extend further?
A: Yes. New tenant_extension packs can be added under tenant_extensions/bp/{name} with a parent_pack_id referencing the shared platform pack. The new pack inherits all parent objects, decisions, and policies, and adds bp-specific extensions on top.
Q: How is pack drift detected?
A: Context Studio's /context/ops/scorecard endpoint computes coverage and confidence per pack on every refresh; drift below target triggers a workqueue item for the bp stewards listed above.

bp Enterprise Alignment & Readiness

bp Sphere operates as the governed finance decision, control, evidence, and accountability runtime inside bp guardrails. bp Sphere is presented as the decision and control runtime inside tenant guardrails, not as a competing data or AI platform.

Alignment score
100.0/100
PASS
bp constructs
9
Yalla / Nexus / LaunchPad / UDP / OneData / identity
Data products
5
certified product lineage
Semantic objects
10
KDO + semantic authority
Positioned asNot positioned as
Decision RuntimeData Runtime consumerGovernance RuntimeIdentity-aware RuntimeTransformation RuntimeA replacement for UDP, OneData, Foundry, Databricks, SAP, Nexus, Yalla, Purview, or bp identity platforms.A standalone AI platform that creates competing data, governance, or integration layers.
bp constructDecision runtimeBoundaryRequired from bpWorkshop proof
Yalla golden pathDeployment and platform operations runtimebp-specific deployment path is tenant configuration; core runtime remains platform-neutral.Approved hosting pattern, namespaces, service accounts, pipeline standards, SRE/SLO expectations.Runtime topology and operator surfaces show deployable services, health, recovery, and release gates.
Nexus registryAgent Registry RuntimeGeneric agent inventory fields are platform-level; Nexus IDs/status are bp tenant overlay metadata.Nexus registration schema, naming rules, lifecycle states, duplicate-check process.Agent cards expose owner, risk class, MCP/A2A status, LaunchPad status, red-team status, TDR readiness, and replay coverage.
AI LaunchPadAI Governance RuntimeLaunchPad gates are bp-specific policy states attached to generic governance lifecycle.Risk triage, approval gates, AI inventory fields, data sensitivity taxonomy, model/use-case acceptance criteria.AI Governance Registry maps every use case and agent to owner, risk class, approval gate, evaluation posture, and audit trail.
UDP + OneDataEnterprise Data RuntimeData runtime models data products and certified datasets generically; UDP/OneData labels are bp tenant overlay.Data domains, product catalog, owner/steward/custodian mappings, certification levels, marketplace metadata.Certified Data Catalog and Data Product Marketplace show source -> bronze -> silver -> gold -> decision lineage.
FIM + Finance Data OfficeFIM-Native Finance Context LayerFIM remains BP's curated finance backbone; bp Sphere consumes FIM products for decisions, controls, evidence, and governed action.FIM product access, FDO owner/steward metadata, ADH/HWG lineage, quality dimensions, refresh cadence, row-level security rules.FIM Context Layer shows source ERP -> FDF/ADH/HWG -> FIM product -> quality checks -> bp Sphere decision -> evidence pack -> governed write-back.
Foundry Key Domain ObjectsKDO + Semantic Authority ManagerObject manager is generic; bp KDO names and semantic-authority ownership are tenant-scoped.Canonical KDO list, object owners, source-system authority, lineage rules, quality thresholds.Object lineage shows SAP/Ariba/Databricks/Foundry -> KDO -> agent -> decision -> action -> replay.
Purview monitoring + red-team testingEvaluation and Compliance RuntimeEvaluation metrics are platform-level; Purview sinks and bp red-team statuses are tenant metadata.Monitoring sink, red-team criteria, prompt/tool-injection policies, evaluation thresholds.Evaluation Runtime tracks quality, hallucination, tool-call accuracy, policy compliance, cost, latency, and red-team readiness.
Identity governanceIdentity Governance RuntimeOIDC/OAuth/RBAC/ABAC concepts are platform-level; bp IdP, entitlement, and recertification specifics are tenant-scoped.IdP integration path, app registration, consent, entitlement rules, recertification cadence, privileged access controls.Identity page shows user identity, agent identity, entitlements, access certification, SoD, MFA, and drift handling.
Technical Design ReviewDesign Review and Readiness RuntimeTDR workflow is a tenant-specific design authority overlay over generic approval workflow.TDR template, architecture review gates, security review gates, data design authority workflow, exception process.TDR Readiness view maps architecture, APIs, data, controls, observability, recovery, evaluation, and approval state.
Data productDomainCertificationOwnershipQualityLineage
Journal Data Product
journal_product
FinanceGoldFinance Data Owner
R2R Data Steward / UDP Custodian
94SAP S/4HANA ACDOCA/BKPF -> UDP bronze/silver/gold -> KDO Journal -> Decision Runtime
Invoice Data Product
invoice_product
ProcurementGoldProcurement Data Owner
P2P Data Steward / UDP Custodian
91SAP ECC/S4 + Ariba -> UDP bronze/silver/gold -> KDO Invoice -> Decision Runtime
Counterparty Credit Product
counterparty_product
CustomerSilverCredit Risk Data Owner
O2C Data Steward / UDP Custodian
88Customer master + AR + market feeds -> UDP silver -> KDO Counterparty -> Credit Decision Runtime
Treasury Liquidity Product
treasury_product
TreasuryGoldTreasury Data Owner
Liquidity Data Steward / UDP Custodian
92Bank/cash/FX systems -> UDP gold -> KDO Cash Position -> Treasury Decision Runtime
Control Evidence Product
control_product
ControlsGoldSOX Control Owner
Control Steward / Audit Evidence Custodian
93Policy/control/evidence registries -> UDP gold -> KDO Control -> Decision Runtime
KDOSemantic authoritySourcesPoliciesDecisions
InvoiceProcurement Data Owner
P2P Data Steward / UDP Custodian
SAP ECCSAP S/4HANAAribaThree Way MatchDuplicate Payment ControlApproval AuthorityApprove InvoiceHold PaymentRequest EvidenceEscalate Supplier Risk
JournalFinance Data Owner
R2R Data Steward / UDP Custodian
SAP S/4HANACFINBlackLineMaterialityJournal ApprovalSegregation of DutiesPost JournalEscalate JournalRequest SupportAdjust Close Readiness
VendorProcurement Data Owner
Supplier Master Steward / SAP/Ariba Custodian
SAP Vendor MasterAriba SupplierCyber Risk FeedSupplier OnboardingBank VerificationSanctions ScreeningApprove SupplierBlock SupplierRequest Master Data Remediation
ContractLegal / Procurement Owner
Contract Steward / Contract Repository Custodian
Ariba ContractsDocument RepositoryLegal CLMContract CoverageMilestone BillingPrice VarianceValidate ContractRequest EvidenceEscalate Leakage
Cost CenterFinance Data Owner
Management Reporting Steward / ERP Custodian
SAP ControllingOrganization HierarchyCost Center OwnershipPosting AuthorityApprove CodingRoute ApprovalEscalate Miscode
AssetAsset Operations Owner
Asset Integrity Steward / Maximo / SAP Asset Custodian
MaximoSAP Asset AccountingServiceNowAsset CriticalityMaintenance ApprovalCapitalizationEscalate Asset RiskApprove Work OrderRequest Inspection Evidence
CounterpartyCredit Risk Data Owner
Credit Data Steward / UDP Custodian
Customer MasterTreasuryExternal RatingsMarket FeedsCredit LimitCollateral RequirementSanctions ScreeningIncrease LimitReduce LimitRequire CollateralBlock Exposure
ControlSOX Control Owner
Control Steward / Control Registry Custodian
Control RegistryPolicy RegistryAudit FindingsSOXSODAuthority MatrixBlock TransactionEscalate FailureApprove Override
PolicyPolicy Owner
Policy Steward / Policy Repository Custodian
Policy RegistrySharePointAuthority MatrixVersion ControlApproval WorkflowEvaluate PolicySimulate PolicyEscalate Conflict
DecisionDecision Owner
Runtime Governance Steward / Replay Custodian
Decision RuntimeEvidence RuntimePolicy RuntimeHuman AccountabilityReplay RequiredEvidence RequiredApproveRejectEscalateReplay
AgentOwnerRiskLaunchPadNexusMCP / A2ARed team / TDR
Duplicate Invoice Agent
agent-p2p-duplicate-invoice
P2P Control OwnerHighConditional approval for workshopNEXUS-PENDING-BP-001Compliant adapter contract
Registered handoff map
Prompt/tool injection scenarios defined
Ready for BP template mapping
Journal Agent
agent-r2r-journal-risk
ControllerHighConditional approval for workshopNEXUS-PENDING-BP-002Compliant adapter contract
Registered handoff map
Evidence-gated unsafe action tests defined
Ready for BP template mapping
Control Assurance Agent
agent-control-assurance
SOX OwnerHighConditional approval for workshopNEXUS-PENDING-BP-003Compliant adapter contract
Registered control escalation map
Authority bypass tests defined
Ready for BP template mapping
🔍 Illustrative operating example — Full-estate registry backfill
The production registry backfills every active bp agent with governance metadata. Explicit manifest agents retain BP-alignment provenance; the rest are marked derived_default until BP plc-owned LaunchPad, Nexus, TDR, and red-team approvals are supplied.
FBT controls
3,124
tenant context
SOX controls
1,005
control runtime
SOD population
5,868
identity and controls
Procurement invoices
2,750,000
$62bn reference
Authentication

bp IdP / Entra-compatible OIDCOAuth 2.0MFAReduced sign-on

Authorization

RBACABACData classificationAction permissionsAgent workload identity

Governance

Entitlement assignmentAccess recertificationImmediate removal on role changePrivileged access controlsAccess drift detectionSoD runtime

ContractSourceTargetFrequencyQuality SLAOwner
sap_to_udp_journalSAP S/4HANAUDP Journal Data ProductEvent + batch reconciliationGold certified before CFO consumptionFinance Data Owner
ariba_to_udp_invoiceAriba + SAPUDP Invoice Data ProductNear-real-time events + daily certificationDuplicate and lineage checks before agent actionProcurement Data Owner
identity_to_runtime_entitlementsbp Identity Governance Platformbp Sphere Identity RuntimeJIT + recertification feedFail-closed on missing entitlementIdentity Platform Owner
Review areaArtifactStatusEvidence
ArchitectureEnterprise Runtime TopologycompleteLayered runtime, adapter boundaries, event flow, and deployment caveat visible from UI
DataUDP / OneData product contractscompleteData products expose owner, steward, custodian, quality, SLA, lineage, and consumers
IdentityHuman, agent, service, tool identity modelconditionalWorkshop simulation present; production requires BP tenant app registration, consent, SCIM/JIT, and recertification feeds
SecurityPolicy, access, prompt/tool injection, and sensitive data controlsconditionalRuntime controls visible; production requires Purview sink, CyberArk/secret controls, and BP red-team acceptance
IntegrationAPI, event, batch, MCP contractscompleteRepresentative contracts show schema, frequency, SLA, owner, retries, and failure mode
EvaluationAgent evaluation gatescompleteEvaluation runtime defines accuracy, hallucination, policy compliance, tool success, latency, and cost gates
RecoveryFailure handling and replay-assisted recoveryconditionalTopology and replay model visible; production requires BP RTO/RPO and DR test cadence
ApprovalLaunchPad / Nexus / TDR mappingconditionalTenant overlay records statuses; production requires BP plc-owned approval IDs and lifecycle states
🔍 Illustrative operating example — TDR status
READY FOR BP TEMPLATE MAPPING · owner: Architecture Governance Owner · design authority: Data Design Authority + Enterprise Architecture Review Board
Evaluation metrics

AccuracyRelevanceTask successTool-call accuracyRedundant callsPlanning qualityReasoning coherenceHallucination ratePolicy complianceOverride rateEscalation rateCostLatency

Evaluation gates

No production promotion without evaluation baselineNo write-back without policy and evidence testsNo high-risk agent without red-team scenariosNo autonomous action outside approved scope

Transformation runtime

DeploymentsHypercareTrainingRole MappingData CleansingBusiness ReadinessParallel RunAdoptionGo-Live RiskTraining RiskData RiskControl RiskHeadcount ImpactCycle Time ImpactControl ImpactAutomation Impact

🔍 Illustrative operating example — Correct workshop wording
OneData and UDP govern trusted data. Nexus and LaunchPad govern AI assets. Identity governs access. bp Sphere turns those governed data products, policies, controls, events, and human accountabilities into explainable finance decisions, actions, evidence, replay, and transformation outcomes.
CapabilitiesEvent Fabric + SOR SidecarPolicy + Authority Control
DOC-01

bp Sphere — Canonical Enterprise Architecture Blueprint

Distributed intelligence control plane — architecture, topology, and governance model

bp Sphere (bp Sphere runtime) is a distributed, policy-bound intelligence runtime with a dedicated Enterprise Control Plane deployed alongside bp systems of record.

It is not an ERP replacement. It is not a data warehouse. It is not a chatbot wrapper.

It is a governed orchestration layer that:

  • Observes enterprise signals from systems of record
  • Reasons through registered, policy-bound agents
  • Produces reproducible, deterministic decisions
  • Seals immutable evidence for every execution
  • Operates agents, policies, runtime health, evidence, replay, learning, and risk through the Enterprise Control Plane

Every decision is deterministic. Every agent is policy-registered. Every execution is replayable.

🔍 Illustrative operating example — Invoice block decision (P2P)
Scenario:
  Supplier invoice: $12.4M
  Status: Blocked (price mismatch)

bp Sphere flow:
  Signal:
    - SAP CFIN emits invoice_blocked
  Context:
    - Supplier has 3 prior disputes
    - Contract variance threshold: 2.0%
    - Current variance: 3.8%
  Decision:
    - Materiality > $500K so dual approval required
    - Simulate release risk vs hold risk
  Recommendation:
    - Hold invoice
    - Escalate to P2P lead
  Evidence:
    - inputs_hash seals supplier, contract, and invoice snapshot
    - outputs_hash seals recommendation and simulation outputs
    - policy_reference: p2p_policy.yaml#variance_rule
  Outcome:
    - Decision logged and replayable
  ┌──────────────────────────────┐
  │    BP Systems of Record      │
  │  SAP · Treasury · CFIN       │
  │  Market Data · Ariba ·       │
  │  BlackLine                   │
  └──────────────┬───────────────┘
                 │  (Read-Only, CDC / REST / Batch)
                 ▼
  ┌──────────────────────────────┐
  │ bp Sphere Intelligence Plane │
  │ Agent Fabric + Governance    │
  └──────────────┬───────────────┘
                 │
                 ▼
  ┌──────────────────────────────┐
  │ Immutable Evidence Vault     │
  │ SHA-256 Sealed               │
  └──────────────┬───────────────┘
                 ▼
  ┌──────────────────────────────┐
  │    BP Decision Makers        │
  │  CFO · Treasurer · Controller│
  └──────────────────────────────┘
Strict Separation Principle

Observe → Reason → Recommend. Never execute in source systems.

🔍 Illustrative operating example — Signal to decision flow
Event source:
  SAP CFIN -> journal_posted

bp Sphere processing:
  1. CDC event enters the integration layer
  2. Event is routed to agent.run.request on NATS
  3. ReconciliationAgent and AnomalyDetectionAgent are triggered
  4. Agents detect GL vs subledger mismatch: $2.1M
  5. Decision is emitted: flag discrepancy and recommend investigation
  6. Evidence seals journal_id, reconciliation snapshot, and variance calculation
  ┌────────────────────────────────────────────────┐
  │ EXPERIENCE FABRIC                              │
  │ CFO Command Centre · 50+ UI Surfaces           │
  │ Decision UX · Evidence Explorer · Ontology UI  │
  ├────────────────────────────────────────────────┤
  │ CONVERSATIONAL GATEWAY                         │
  │ Runtime-configured LLM · Explain · Drill       │
  │ Simulate · Recommend                           │
  ├────────────────────────────────────────────────┤
  │ AGENT FABRIC + MESH                            │
  │ 153 Agents · 25 Categories · Agent Mesh       │
  │ Supervisor · HITL · Tick Engine · Role Router  │
  ├────────────────────────────────────────────────┤
  │ GOVERNANCE CORE                                │
  │ Policy Engine · RBAC/PBAC · Kill Switch        │
  │ Autonomy Gating AL1-AL4                        │
  ├────────────────────────────────────────────────┤
  │ ONTOLOGY & SEMANTIC LAYER                      │
  │ 4-Layer Object Model · Trigger Engine          │
  │ Cross-Domain Bridges · Tenant Packs            │
  ├────────────────────────────────────────────────┤
  │ EVIDENCE & CONTEXT LAYER                       │
  │ Decision Ledger · Canonical Entities           │
  │ SHA-256 Replay · Knowledge Graph (Neo4j)       │
  ├────────────────────────────────────────────────┤
  │ EVENT BACKBONE                                 │
  │ NATS JetStream · Async Subjects                │
  │ Agent Execution Events                         │
  ├────────────────────────────────────────────────┤
  │ FABRIC ORCHESTRATION                           │
  │ 10 Fabrics · Workflow · Guardian · Trust   │
  │ AI Mesh · Saga · Federated Gateway             │
  ├────────────────────────────────────────────────┤
  │ INTEGRATION LAYER                              │
  │ 114 SOR Adapters · CDC · REST · Batch          │
  └────────────────────────────────────────────────┘
🔍 Illustrative operating example — Agent execution inside the fabric
Agent:
  name: ReconciliationAgent
  domain: R2R

Trigger:
  kpi.threshold.crossed (GL mismatch)

Execution:
  - Fetch GL balance from canonical snapshot
  - Fetch subledger totals from canonical snapshot
  - Compute variance deterministically

Decision output:
  - mismatch_detected: true
  - variance_amount: $2.1M
  - severity: HIGH

Policy check:
  - materiality > $1M -> escalate

Result:
  - policy.violation event emitted
  - decision record created
Production Environment

Deployed on Kubernetes (Production Cluster)

LayerSpecification
ClusterMulti-node Kubernetes cluster (control plane + worker nodes)
NetworkDedicated VPC subnet · TLS ingress termination · Internal load balancer
StoragePersistent Volume Claims for PostgreSQL + Evidence Vault
SecretsKubernetes Secrets isolated per tenant (JWT, signing keys, DB credentials)
ScalingHorizontal Pod Autoscaling enabled · Rolling deployment strategy
Namespaceplatform (production) · platform-demo (workshop)
  ┌──────────────────────────────────────────────────┐
  │ Kubernetes Cluster                               │
  │                                                  │
  │  ┌──────────────────────────────────────────┐    │
  │  │ API Gateway Pod (iris-api)               │    │
  │  │ FastAPI · 300+ routes                    │    │
  │  │ Uvicorn workers=2 · uvloop               │    │
  │  └──────────────────┬───────────────────────┘    │
  │                     │                            │
  │  ┌──────────────────▼───────────────────┐        │
  │  │ Agent Runtime Pods                   │        │
  │  │ 153 Agents · Supervisor · HITL Queue  │        │
  │  └──────────────────┬───────────────────┘        │
  │                     │                            │
  │  ┌──────────────────▼───────────────────┐        │
  │  │ Evidence Service                     │        │
  │  │ Decision Ledger + SHA-256 Sealing    │        │
  │  └──────────────────┬───────────────────┘        │
  │                     │                            │
  │  ┌──────────────────▼──────────┐  ┌───────────┐ │
  │  │ PostgreSQL (5 Domains)      │  │ NATS      │ │
  │  │ 200+ SQLModel entities      │  │ JetStream │ │
  │  │ + Neo4j Knowledge Graph     │  │ Port 4222 │ │
  │  └──────────────┬──────────────┘  └───────────┘ │
  │                 │                                │
  │  ┌──────────────▼──────────────┐                 │
  │  │ Ontology & Semantic Layer   │                 │
  │  │ 4-Layer Object Model        │                 │
  │  └─────────────────────────────┘                 │
  └──────────────────────────────────────────────────┘
Primary Database
  • Canonical entities (200+ SQLModel)
  • Agent registry & execution metadata
  • Decision ledger (immutable)
  • Policy store & evaluation logs
  • KPI definitions & thresholds
  • Evidence vault artefacts
Replica Database
  • Read-only workloads
  • Simulation & scenario engine
  • Analytics isolation
  • Dashboard aggregations
  • Streaming replication from primary
Knowledge Graph
  • Neo4j graph store
  • Entity relationships & inference
  • GraphRAG retrieval pipeline
  • Ontology schema validation
Tenant Isolation
  • Schema-level scoping with realm_id on all core tables
  • Middleware enforcement — cross-tenant access returns HTTP 403
  • Config boundary: config/tenants/bp/
  • Separate database per environment (iris_agentic_finance vs _demo)

NATS JetStream provides the asynchronous event backbone for agent orchestration, governance escalation, and real-time streaming to UI surfaces.

SubjectPurpose
agent.run.requestTrigger agent execution
agent.run.completedExecution result with evidence reference
policy.violationGovernance breach escalation
evidence.sealedImmutable evidence pack committed
kpi.threshold.crossedKPI alert trigger
cfin.signal.receivedSAP CFIN change data capture
Asynchronous orchestration — fire-and-forget agent dispatch
Replayable streams — JetStream persistence for audit
Agent decoupling — no direct inter-agent calls
Governance hooks — escalation triggers on policy violation
🔍 Illustrative operating example — Event flow over NATS
Step 1:
  subject: cfin.signal.received
  payload: invoice_blocked

Step 2:
  subject: agent.run.request
  agents:
    - InvoiceRiskAgent
    - SupplierHistoryAgent

Step 3:
  subject: agent.run.completed
  result:
    - risk_score: HIGH

Step 4:
  subject: policy.violation
  escalation: required

Step 5:
  subject: evidence.sealed
  decision_id: DEC-1245
LevelNameBehaviourHuman Role
AL-1AdvisoryRecommends onlyHuman decides
AL-2SupervisedProposes actionHuman approves
AL-3Governed AutoActs within policyHuman monitors
AL-4Full AutoExecutes + reportsHuman audits

Current operating state: AL-1 / AL-2

AL-3 Promotion Requirements
  • Deterministic replay verification passed
  • Shadow parallel runs completed (v1 vs v2 comparison)
  • Policy coverage > 95%
  • CFO sign-off
🔍 Illustrative operating example — AL2 versus AL3
AL2 (current):
  - Agent proposes invoice release
  - Human approves before execution

AL3 (future):
  - Agent auto-releases invoices below $100K
  - Policy and evidence are sealed automatically
  - Human monitors the decision stream
WORKSHOP Mode
  • Snapshot-based data
  • Signal injection allowed
  • Time manipulation enabled
  • Evidence retained 90 days
  • Safe experimentation environment
LIVE Mode
  • Real enterprise signals only
  • Injection blocked (HTTP 403)
  • Fail-closed on all errors
  • Evidence retention: 7 years
  • SOX-grade traceability

Current mode: LIVE

Six-layer fail-closed middleware chain. Outermost layer executes first.

#MiddlewarePurpose
1ExceptionHandlerMiddlewareFail-closed outer layer — structured JSON errors, no stack trace leakage, Prometheus counter
2AuthMiddlewareJWT validation — 200 req/min rate limit, public paths: /health, /ready, /metrics
3GovernanceMiddlewareRoute-level permission enforcement via GovernanceGateChain (30+ mappings)
4CORSMiddlewareOrigin validation, credential passthrough
5TenantSeparationMiddlewareCross-tenant rejection (HTTP 403 on mismatch)
6CorrelationMiddlewareTrace propagation — X-Correlation-ID, X-Request-ID, X-Trace-ID
Guarantees
  • No stack trace leakage to clients
  • Cross-tenant request rejection
  • Structured JSON error responses
  • Prometheus counter increment on every failure
Prometheus Metrics
  • iris_unhandled_errors_total
  • agent_execution_duration_seconds
  • policy_violation_total
  • autonomy_level_distribution
SLO Targets
  • API latency P95 < 500ms
  • API latency P99 < 2s
  • Cross-tenant violations: 0
  • Evidence sealing compliance: 100%
Structured Logging

Every log entry carries: correlation_id · tenant_id · agent_id · decision_id · replay_trace_id

CREST is the strategic layer between enterprise systems and agent execution. It shifts bp Sphere from AI tooling to an enterprise context operating system.

  Typical stack
  Enterprise Systems -> Lake/RAG -> LLM Agents -> App
  (stateless AI, weak enterprise context)

  bp Sphere + CREST stack
  Enterprise Systems -> CREST Context Engine -> Context Graph -> Agent Fabric -> Decisions
CREST EngineRole in bp SphereWhy it matters
Context BuilderBuilds context bundles per mission/entity/objectiveAgents reason with enterprise situational context, not raw records
Ontology + PolicyProvides semantic meaning and policy constraintsDecisions remain interpretable and governed
Decision MemoryStores decision-context-policy-outcome lineageCreates replayable enterprise decision intelligence
Institutional KnowledgeCaptures tribal knowledge into reusable artifactsPreserves enterprise memory beyond individuals
Learning LoopTurns outcomes/overrides into improvement proposalsDrives compounding context quality over time
Context Flywheel

Decision -> Evidence -> Memory -> Better Context -> Better Decisions

Leadership narrative

AI models generate answers; enterprise context generates intelligence.

🔍 Illustrative operating example — CREST context bundle
Context bundle:
  entity: invoice INV-88921

Includes:
  - Supplier risk profile
  - Contract terms
  - Historical disputes
  - Payment behavior
  - Market price benchmarks

Used by:
  - InvoiceRiskAgent
  - PaymentOptimizationAgent
bp Sphere ISbp Sphere IS NOT
A policy-governed intelligence sidecarAn ERP replacement
A deterministic decision operating systemA transactional processor
A replayable audit-grade reasoning engineAn unbounded autonomous executor
An enterprise control plane with evidenceA black-box analytics tool

Frequently Asked Questions

Q: Does bp Sphere replace SAP?
A: No. It reads from SAP and related systems. It never writes back.
Q: What if an SOR becomes unreachable?
A: The platform operates on last verified snapshot and flags staleness. No speculative execution.
Q: What happens if bp Sphere fails?
A: Fail-closed. No decision emitted. No action taken. Kill switch halts all agents instantly.
Q: Can agents scale horizontally?
A: Yes. Agent runtime pods scale independently via Kubernetes HPA.
Architecture seamLatest implementation
Experience orchestrationDesktop, mobile, and chat render the same decision object, evidence, policy boundary, action contract, replay, and value context.
Assistant portabilityWIIF keeps enterprise intelligence inside the platform while Claude, Copilot, Gemini, OpenAI, internal assistants, and future tools remain channels.
Runtime proof/runtime-adoption-audit audits identity, policy, evidence, replay, value, token, and learning seam adoption.
Command-line proofbpsphere exposes the same runtime from command line: health, agents, policies, Continuous Close, digital twin, replay, evidence, value, runtime observability, and provenance.
Continuous Close architectureThe close story is now an event-driven Continuous Financial Position Engine, not a static dashboard: financial events, reconciliation, accrual/provision intelligence, controls, certification, and learning run as one chain.
  • Architecture now explicitly separates runtime mode from deployment environment.
  • Demo/showcase, dev, and prod are distinct deployments; public workers can be pointed independently to demo, dev, or prod surfaces.
  • Mission pages now pre-render initial tab state server-side, fixing cache/path drift on FP&A deep links.
CapabilitiesCREST Context Reconstruction
DOC-02

Conversational Command Layer (Guided Demo)

Governed narration · Deterministic click orchestration · Evidence-first credibility

The Conversational Command Layer (CCL) is not a chatbot wrapper and not a pre-recorded demo.

It is a governed orchestration and narration layer that:

  • Translates live bp Sphere runtime state into role-adaptive explanation
  • Drives a deterministic UI click-path (“guided mode”) across command surfaces
  • Binds every claim to policy, provenance, and evidence
  • Ensures executive storytelling remains audit-grade (replayable, hash-sealed)

CCL exists to solve a credibility problem: “Show me this is live, governed, and real.” CCL answers with inline proof.

Governed NarrationDeterministic Click OrchestrationEvidence-First Credibility
User (UI / API / Guided Mode)
        |
        v
CCL Router (role + intent + mode)
        |
        |-- Explain / Drill / Simulate / Recommend
        |
        |-- Click Orchestration Plan (deterministic)
        |
        |-- Evidence Binding Plan (required)
        |
        +-- Policy Safety Gate (required)
        |
        v
Domain Copilots + Scenario Engine
        |
        v
Response + Step Plan + Proof Bar
(mode . confidence . sources . policy_version . replay_id . evidence_hash)
Core invariant: No narration step is emitted without proof bindings.
User Query (UI / API)
        |
        v
+-----------------------------------------------------------+
| Enterprise Chat Command Layer (api/enterprise_chat.py)    |
|  UI   /ui/enterprise-chat                                 |
|  POST /api/iaf/enterprise-chat/ask                        |
|  Routes: Journal . Recon . Invoice . Cash . Treasury . FP&A|
|  Emits: capability . confidence . policy . evidence . path |
+-------------------------+---------------------------------+
                          |
                          v
+-----------------------------------------------------------+
| FP&A Copilot (api/fpa_copilot.py)                         |
|  POST /api/iaf/fpa-copilot/ask        (JSON)              |
|  POST /api/iaf/fpa-copilot/ask-stream (SSE)               |
|  Model: gpt-5.4-mini . runtime-configured               |
|  Modes: Explain . Drill . Simulate . Recommend            |
+-----------------------------------------------------------+
| CFO Intelligence (api/cfo_intelligence.py)                |
|  POST /api/iaf/cfo/chat                                   |
|  POST /api/iaf/cfo/scenario/interpret                     |
|  POST /api/iaf/cfo/exposure/compute                       |
+-----------------------------------------------------------+
| CFO Scenario Engine (api/cfo_scenario.py)                 |
|  GET /baseline  POST /simulate  POST /monte-carlo         |
|  GET /guardrails  GET /capital-allocation                  |
|  POST /evidence-pack  GET /recommendations                |
+-------------------------+-------------------------------+
                          |
                          v
Evidence Pack (SHA-256) + Governance Metadata + Replay Pointer

CCL sits above these runtime services and produces: a user response, a guided click plan, and a proof bar (always-on credibility controls).

Guided mode is implemented as a deterministic plan, not an animation.

Each guided run emits a ClickPlan:

step_id
target_surface (URL / route)
ui_element_id (stable selector)
preconditions (SOR freshness, heartbeat, policy)
expected_state (KPI value, toggle, evidence)
evidence_pointer (pack_id + hash)
fallback_step (if preconditions fail)

Determinism Rules for Guided Mode

  • Plan generated from stable templates + runtime context (no free-form wandering)
  • Identical inputs → identical click plan
  • Every step must be verifiable from live telemetry
  • If verification fails → step blocked with reason (never “hand-wave”)

This makes the demo replayable and challenge-proof.

Every response (and every guided step) includes the following metadata:

mode
(Explain/Drill/Simulate/Recommend)
agent_id + agent_version
confidence + confidence_method
sources (SOR provenance list)
assumption_version
policy_version
replay_id
evidence_pack_id + evidence_hash
llm_model + latency_ms
data_freshness (per-domain staleness flags)
UI rule: In guided mode, the proof bar cannot be hidden.

This is what neutralizes “this is a gimmick.”

Before CCL can emit a recommendation, scenario result, “next click” instruction, approval request, or narrative conclusion — it must pass:

1. Tenant boundary check
2. Role permission check
3. Domain scope check
4. Policy evaluation check
5. Evidence availability check
6. Freshness staleness check (or explicit stale disclaimer)
If any gate fails:
  • Response is produced in Explain mode
  • Clearly states what cannot be asserted and why
  • Includes governance metadata and next safe step

Fail-closed, not “best effort guessing.”

Guardrails (Enforced)
  • gearing ≤ 20%
  • VaR ≤ $0.25B
  • FCF ≥ $15B
  • dividend cover ≥ 1.3x
  • net debt / EBITDA ≤ 2.5x
  • liquidity ≥ 3 months
  • OpEx cut ≤ 15%
  • capex deferral ≤ 40%
Scenario Evidence Pack Requirements

CFO scenario packs are SHA-256 sealed and include:

  • model_version = 2.0.0
  • sensitivity_matrix = v3.2
  • policy_version = FP&A-GOV-v3.2
  • baseline snapshot reference
  • slider inputs + derived deltas
  • guardrail breach probabilities
  • replay pointer

Monte Carlo runs are governed and reproducible:

Seeded RNG (seed = 42)
Configurable runs (100 – 50K)
40-bin histogram
Breach probability tracking per guardrail
Results stored as evidence artifacts

No unseeded randomness is allowed in executive demo mode.

SSE event protocol:

event: meta — immediate: mode, policy_version, replay_id
event: token — streamed response chunks
event: done — stats: latency_ms, tokens, evidence_pack_id, hash

SSE ensures:

  • Perceived responsiveness
  • Explicit end-of-run proof
  • Structured capture for audit playback

Every LLM call is recorded in LLMAuditRecord.

Governance flag: ENABLE_LLM_AUDIT is non-overridable.

Captured fields:

prompt_hash
response_hash
model_id
policy_version
replay_id
user_role
latency_ms

This ensures forensic traceability of copilot behavior.

CCL adapts output by role, not by “tone only.”

RoleDefault ModeOutput PatternProof Depth
CFOExplain / Simulateheadline → risk → guardrails → valueFULL
Tower LeadDrill / Recommendexceptions → drivers → actionsHIGH
AnalystDrillevidence-first → worklistHIGH
OperatorRecommendchecklist + approvals + SLAMED/HIGH

Rule: Higher authority → higher proof visibility.

PhaseCapabilityExit Criteria
1Scripted guided mode + narration + replay metadata100% steps governance-ready
2Controlled Q&A with scenario triggeringpolicy-safe Q&A coverage > 95%
3Role-adaptive flow generationdeterminism retained + replay proven
Key requirement: Even Phase 3 cannot introduce non-deterministic click paths.
  • CCL is an explanation and orchestration layer, not an animation feature.
  • Guided mode drives live UI and validates run context at each step.
  • Any action requiring approval still routes through HITL.
  • Proof metadata must remain visible to defeat “demo theater” risk.

Frequently Asked Questions

Q: Is the guided experience pre-recorded?
A: No. Guided mode reads live runtime state and validates evidence/replay bindings per step.
Q: Can guided mode execute actions autonomously?
A: Only within policy bounds and configured autonomy level. Approval-required steps route through HITL.
Q: How do we prove this is not a gimmick?
A: We expose run_id, agent_version, policy_version, evidence_pack_id, evidence_hash, replay_id, and data freshness inline for every guided claim.

Instead of saying: "We have a copilot."

bp Sphere's Conversational Command Layer is a governed narration and deterministic click-orchestration system that turns live runtime state into role-adaptive explanation — with every claim bound to policy, evidence hashes, replay IDs, and source provenance visible inline.

CredibleAuditableLiveGovernedNot Demo Theater
  • Conversational Finance AI uses a single runtime model setting via OPENAI_MODEL.
  • Preset simulations use deterministic finance math for numbers and LLM-generated executive narrative for explanation.
  • Current production-safe default has been tuned back to gpt-5.4-mini for reliability and latency over gpt-5-mini.
CapabilitiesOntology + Enterprise ContextCREST Context Reconstruction
DOC-03

Ontology & Semantic Intelligence Layer

4-layer object model · Trigger engine · Cross-domain bridges · Tenant packs

The bp Sphere ontology provides a 4-layer semantic object model that defines how business entities are structured, related, triggered, and audited across the platform.

Unlike static data schemas, the ontology is active — it drives agent bindings, trigger rules, and cross-domain bridges.

4-Layer DesignTrigger EngineCross-Domain BridgesTenant Packs

The ontology is organized in 4 layers. Lower layers provide universal abstractions; upper layers bind to agents and produce auditable evidence.

  ┌────────────────────────────────────────────────┐
  │ EVIDENCE LAYER                                 │
  │ EvidencePack · Decision · AuditTrail           │
  ├────────────────────────────────────────────────┤
  │ DECISION LAYER                                 │
  │ AgentExecution · PolicyEvaluation              │
  │ Recommendation · Confidence · ReasoningChain   │
  ├────────────────────────────────────────────────┤
  │ DOMAIN LAYERS (extensible per domain)          │
  │ Finance: Invoice · Supplier · PurchaseOrder    │
  │          Payment · BankAccount · GLPosting     │
  │ HSE:     Permit · Incident · RiskRegister     │
  │ HR:      Employee · TimeEntry · Payroll       │
  ├────────────────────────────────────────────────┤
  │ CORE / KERNEL LAYER                            │
  │ Party · OrganizationUnit · Agreement           │
  │ Account · EvidencePack                         │
  └────────────────────────────────────────────────┘
Bottom-Up Inheritance

Domain objects inherit from core types — e.g. Supplier extends Party

Top-Down Binding

Evidence layer binds to domain objects — every Invoice decision produces an EvidencePack

Core / Kernel Objects

ObjectSubtypesKey Attributes
PartySupplier, Customer, Employee, Contractorparty_id, name, type, status
OrganizationUnitCompanyCode, BusinessUnit, CostCenter, Plantorg_id, name, hierarchy
AgreementContract, FrameworkAgreement, ServiceAgreementagreement_id, parties, terms
AccountGLAccount, BankAccountaccount_id, type, currency
EvidencePackpack_id, execution_id, inputs_hash

Finance Domain Objects

ObjectKey AttributesBound Agents
Invoiceinvoice_id, supplier_id, amount, currency, statusInvoiceValidationAgent, DuplicateDetectionAgent
Suppliersupplier_id, name, country, risk_scoreSupplierRiskAgent, SupplierPerformanceAgent
PurchaseOrderpo_id, supplier_id, amount, statusPOComplianceAgent, POApprovalAgent
Paymentpayment_id, invoice_id, amountPaymentExecutionAgent, PaymentFraudDetectionAgent
BankAccountaccount_id, bank_name, currency, balanceLiquidityMonitoringAgent, CashForecastAgent
GLPostingjournal_id, account, amount, posting_dateJournalValidationAgent, ReconciliationAgent

The trigger engine binds ontology events to agent execution. When an object lifecycle event occurs, matching triggers fire and route to the bound agents.

  Object Event (e.g. Invoice.created)
              │
              ▼
  Trigger Resolution API
  POST /api/iaf/ontology/resolve-triggers
              │
              ▼
  Matched Triggers
      ├── InvoiceValidationAgent (0.95)
      ├── DuplicateDetectionAgent (0.90)
      └── InvoiceApprovalAgent (conditional)
              │
              ▼
  Agent Execution (policy-gated)
RouteMethodPurpose
/api/iaf/ontology/summaryGETOntology statistics
/api/iaf/ontology/objectsGETAll ontology objects
/api/iaf/ontology/objects/{name}GETObject detail
/api/iaf/ontology/relationshipsGETCross-domain relationships
/api/iaf/ontology/actionsGETAvailable actions
/api/iaf/ontology/triggersGETTrigger rules
/api/iaf/ontology/lifecyclesGETObject lifecycles
/api/iaf/ontology/resolve-triggersPOSTResolve triggers for event
TabPurposeFeatures
Object RegistryBrowse objectsFilterable with layer, domain, attributes
Relationship GraphCross-domain relationshipsVisual entity graph
Trigger EngineTest trigger rulesLive resolution with demo payloads
Action BindingsAgent-to-object bindingsWhich agents triggered by which objects
Object LifecycleState machinesValid state transitions per type
PackScopeExtensions
FinanceR2R, P2P, TreasuryInvoice workflows, payment rules, close calendar
ProcurementP2P, Supplier MgmtPO validation, supplier risk triggers
HRWorkforce, PayrollEmployee lifecycle, timesheet triggers

Tenant Isolation: Packs loaded from tenants/bp/config/ cannot modify platform-level objects.

  P2P Tower (Invoice)              R2R Tower (GLPosting)
       │                                 │
       └── ontology bridge ──────────────┘
                    │
  ReconciliationAgent traverses:
    Invoice → Payment → BankAccount → GLPosting
Invoice → Payment

P2P to Treasury

Supplier → Contract

P2P to Governance

Employee → CostCenter

HR to FPA

CREST extends semantic modeling into context reconstruction so agents retrieve decision-ready bundles instead of raw records.

  Event / Mission Objective
           │
           ▼
  Context Builder (CREST)
    ├── Ontology entities
    ├── Policy constraints
    ├── Enterprise memory (similar decisions)
    └── Active signals
           │
           ▼
  Context Package -> Agent runtime -> Evidence + Learn
Context ElementPurposeSource
Canonical entitiesNormalize business meaning before executionOntology kernel + domain packs
Policy contextConstrain autonomy and decisionsPolicy engine + AL gates
Enterprise memoryRecall prior similar outcomes and SME-validated patternsDecision history + learning records
Signal contextInject real-time state into reasoningMission/event streams
  • The ontology is now visible in live agent surfaces through entity_type and ontology_refs on recommendations and traces.
  • Lease, payment, obligation, and finance objects are tied into drawer-level drilldowns rather than narrative-only cards.
  • Cross-object drill paths now include transaction → trace → evidence navigation.
CapabilitiesEnterprise Memory + Knowledge Management
DOC-04

Knowledge Graph & AI Reasoning

Neo4j graph store · GraphRAG · Inference · Chain-of-thought

bp Sphere maintains a Neo4j-backed knowledge graph for contextual reasoning, combining entity relationships, semantic search, and multi-hop inference.

Neo4j StoreGraphRAGInferenceChain-of-Thought
ComponentModulePurpose
KnowledgeGraphai.knowledge.coreEntity/triple store with search and BFS path-finding
GraphRetrieverai.knowledge.retrieverSemantic + graph hybrid retrieval
GraphRAGai.knowledge.retrieverCombined graph + vector RAG
OntologySchemaai.knowledge.ontology.schemaEntity/triple validation with constraints
CoTReasonerai.reasoning.chain_of_thoughtExplainable reasoning with self-consistency
Entity
  • id, name, entity_type
  • properties (key-value metadata)
  • embedding (vector representation)
Triple
  • subjectpredicateobject
  • properties (edge metadata)
  • confidence (relationship score)
CapabilityImplementationUse Case
Multi-hopBFS path findingIndirect entity relationships
Zero-shot CoTreason(ZERO_SHOT_COT)Step-by-step prompting
Self-Consistencyself_consistency()Majority voting
Reflectionreason_with_reflection()Critique + refinement
Validationcheck_logical_flow()Reasoning chain integrity
ConstraintDescriptionExample
CARDINALITYMin/max relationshipsInvoice must have 1 Supplier
DOMAINValid source typessupplies_to from Supplier only
RANGEValid target typessupplies_to to Organization only
REQUIREDMandatory propertiesEntity must have name
SYMMETRICBidirectionalpartners_with implies reverse
TRANSITIVEInferred chainsreports_to through hierarchy
LayerHow it worksFinance relevance
Decision MemoryStores context, policy, recommendation, outcome, and evidence links per decisionSupports traceable P2P/CFO drilldown and replay
Enterprise MemoryIndexes repeated patterns and validated institutional rulesEnables similar-case guidance in mismatch and exception scenarios
Learning FabricConverts evidence-backed outcomes into policy/skill improvement proposalsReduces repeat exceptions and improves autonomy safety
  Decision Execution
       │
       ▼
  Evidence Pack + Hashes
       │
       ▼
  Memory Update (similar-case index)
       │
       ▼
  Learning Proposal (policy/skill/routing)
       │
       ▼
  Governed Approval -> Future runs improve
  • Knowledge graph patterns remain part of the architecture narrative, but current finance credibility is driven primarily by reconciled models, policy metadata, and evidence linkage.
  • Where graph reasoning is not the primary live path, the pack now needs to treat it as architecture capability rather than overclaiming default execution for every flow.

Evidence Intelligence Fabric & Context Studio

                Evidence Intelligence Fabric (EIF)
                ─────────────────────────────────
                        Shared platform service
                                │
        ┌──────────┬────────────┼────────────┬──────────┐
        ↓          ↓            ↓            ↓          ↓
    Discovery  Resolution    Lineage       Pack      Viewer
    Service    Service       Service       Service   Service

                Cross-system     Per-record
                evidence         lineage:
                discovery        source / time
                                 confidence /
                                 match_posture /
                                 governing_status

                                │
                                ↓
                bp Tenant EIF Case Packs (6):
                ─────────────────────────────────
                journal_entry           → EIF-PACK-JE-*
                contract_validation     → EIF-PACK-CONTRACT-*
                credit_intelligence     → per-decision
                maintenance_verification → EIF-PACK-MAINT-4402
                spend_intelligence      → EIF-PACK-SPEND-0712
                close_intelligence      → EIF-PACK-CLOSE-0630
EIF case aliasSurfaceEvidence pack ID
journal_entryJournal Entry StudioEIF-PACK-JE-*
contract_validationContract & Milestone ValidationEIF-PACK-CONTRACT-*
credit_intelligenceDynamic Credit Intelligence(per-decision pack)
maintenance_verificationMaintenance Consumption VerificationEIF-PACK-MAINT-4402
spend_intelligenceSpend IntelligenceEIF-PACK-SPEND-0712
close_intelligenceR2R Anomaly / Close IntelligenceEIF-PACK-CLOSE-0630
                Context Studio  /  context-graph service
                ─────────────────────────────────────────
                Stack:  FastAPI + SQLAlchemy + Neo4j + NATS
                State:  live in cluster (port 8080)

                24 REST routes:

                Registry / Marketplace
                  GET  /context/registry
                  GET  /context/registry/revisions
                  POST /context/registry/object
                  POST /context/registry/relationship
                  GET  /context/marketplace

                Coverage / Confidence / Discovery
                  GET  /context/coverage          (6-dim rubric)
                  GET  /context/confidence        (per-attribute)
                  GET  /context/discovery/engine
                  GET  /context/discovery/findings

                Graph traversal
                  GET  /context/graph/path
                       (e.g. Supplier → Invoice → Payment)

                ContextOps
                  GET  /context/ops/scorecard
                  GET  /context/ops/workqueue
                  GET  /context/ops/drift
                  GET  /context/ops/contract

                Strategy
                  GET  /context/strategy/digital-twin-scope

                Mutations (stewardship workflow)
                  POST /context/discovery/findings/{id}/promote
                  POST /context/discovery/findings/{id}/assign-steward
                  POST /context/discovery/findings/{id}/reject
                  POST /context/ops/workqueue/{id}/promote
                  POST /context/ops/workqueue/{id}/assign-steward
                  POST /context/ops/workqueue/{id}/reject
🔍 Illustrative operating example — Live coverage scorecard (finance domain)
GET /context/coverage?domain=finance

→ HTTP 200
{
  "coverage_contract": {
    "id": "iris-context-coverage-v1",
    "dimensions": [
      "object_coverage", "relationship_coverage",
      "event_coverage", "policy_coverage",
      "decision_pattern_coverage", "evidence_lineage_coverage"
    ],
    "weights": { "object_coverage": 0.20, ... },
    "quality_gates": {
      "minimum_confidence_pct": 0.75,
      "minimum_freshness_sla_hours": {
        "critical": 24, "high": 72, "standard": 168
      }
    }
  },
  "scorecard": [{
    "domain": "finance",
    "coverage_pct": 0.7219,
    "target_pct": 0.90,
    "trust_score": 0.725,
    "governance": {
      "current_state": "curated",
      "next_state": "certified"
    }
  }]
}
🔍 Illustrative operating example — Live graph traversal example
GET /context/graph/path?source_object=Supplier&target_object=Payment&domain=finance

→ HTTP 200
{
  "domain": "finance",
  "source_object": "Supplier",
  "target_object": "Payment",
  "path": ["Supplier", "Invoice", "Payment"],
  "edges": [
    { "from": "Supplier", "to": "Invoice",
      "reason": "Supplier context explains invoice demand, dispute, and fraud posture." },
    { "from": "Invoice", "to": "Payment",
      "reason": "Invoice approval or hold changes payment execution timing." }
  ],
  "path_confidence": 0.9
}
Companion: Functional Integrity Layer (FIL)

Every EIF pack generated for a bp tenant decision is sealed by FIL: signed evidence hash, lineage row, and replay key. bp tenant currently holds ~22M signed revisions in iaf_agent_execution_revisions — every one of which can be replayed deterministically.

Frequently Asked Questions

Q: Where does EIF run?
A: The EIF code lives in iris_runtime/src/iris_sor/services/evidence_intelligence_fabric.py and is invoked by every use case launch page. Context Studio runs as the separate context-graph deployment in platform-services (port 8080), backed by Neo4j and NATS.
Q: Is Context Studio bp-specific?
A: No. Context Studio is shared platform infrastructure. bp tenant's data flows through it; the registry, marketplace, coverage scorecard, and lineage revisions reflect bp pack contents alongside the shared finance-context-pack.
Q: How is evidence proven authentic?
A: Each EIF pack record carries six fields: source, record_id, version, retrieved_at, confidence, match_posture, excerpt, governing_status. The pack is then signed by FIL and stored immutably; the same pack can be regenerated and verified bit-for-bit on replay.
CapabilitiesSkill Fabric + Agentic RuntimeEvent Fabric + SOR Sidecar
DOC-05

Agent Runtime & Orchestration Fabric

Deterministic execution · Policy gating · Evidence-coupled mission delivery

bp Sphere operates a single canonical execution contract across all 153 agents.

  • No mission-specific shortcuts
  • No bypasses around policy
  • No execution without ledger registration
  • No LLM invocation outside governance envelope

Registry version: 1.2.0 · Runtime agents: 153 registered, 153 implemented · LLM provider: shared_capabilities.llm.OpenAIClient · Operating mode: LIVE

  Signal / Schedule / SOR Event
              │
              ▼
  Mission Router (API Layer)
              │
              ▼
  Governance Pre-Check
              │
              ▼
  Agent Runtime Queue (NATS Subject)
              │
              ▼
  Agent Worker Pod (Horizontally Scalable)
              │
              ▼
  Decision Envelope Sealing
              │
              ▼
  Evidence Vault + Decision Ledger
Event-driven via NATS JetStream
Stateless execution containers
Idempotent via decision_id
Deterministic output enforcement

Step 1 — Trigger

SOR event (CDC) KPI breach Scheduled tick Manual invocation Conversational query

Every trigger receives: correlation_id · tenant_id · mission_id · trace_id

Step 2 — Mission Router

  • Validates mission scope
  • Resolves domain
  • Loads agent registry entry
  • Verifies active deployment version

No dynamic agent injection allowed.

Step 3 — Policy Pre-Check

Role authorisation
Sensitivity classification
Materiality thresholds
Autonomy level
Authority domain scope
SOD rules

If denied → BLOCKED. If escalation required → AUTHORITY QUEUE. No execution occurs before this gate.

Step 4 — Agent Execution Envelope

  • Context snapshot frozen
  • Inputs canonicalised — hash pre-computed
  • LLM call (if required) wrapped in policy-bound interface
  • Output normalised to schema
  • Determinism check applied

No raw LLM output is ever returned directly.

Step 5 — Decision Resolution

PROPOSED EXECUTED ESCALATED BLOCKED

Each outcome sealed with: inputs_hash · outputs_hash · evidence_hash · policy_id · autonomy_level · sor_provenance[]

FieldPurpose
decision_idImmutable global reference
agent_versionRelease traceability
inputs_hashCanonical input hash
outputs_hashDeterministic output verification
evidence_hashEvidence integrity
replay_pointerSnapshot reference
policy_versionGovernance context at execution time
Replay Validation Process
  1. Retrieve input snapshot
  2. Re-run agent under identical policy version
  3. Compare outputs_hash
  4. Flag divergence if mismatch

LIVE mode requires 100% replay compliance.

All 153 agents are declarative and registered. No agent may execute without registry presence, valid deployment activation, and health status = READY.

GroupFieldsPurpose
Identityid, name, versionUnique identification and versioning
Classificationtower, category, decision_type, criticalityFinance tower + functional category + risk
Governanceautonomy_level, authority_scope, escalation_conditions, hitl_thresholds, policiesPolicy binding + escalation rules
Operationssla_seconds, timeout_seconds, kill_switch, is_criticalPerformance SLA + safety controls
Runtimecapabilities, dependencies, kpi_idsDependency graph + KPI associations
Ownershipowner_business, owner_technicalBusiness and technical accountability
58
P2P Tower
21
R2R Tower
17
FP&A Tower
13
I2C Tower
11
Governance
8
Treasury
4
Upstream
3
Platform

13 Functional Categories: ORCHESTRATION · EXECUTION · VALIDATION · ANALYSIS · MONITORING · EXCEPTION · ADVISORY · ENRICHMENT · DETECTION · ROUTING · AUDIT · GOVERNANCE · OPTIMIZATION

Agent Workers
  • Kubernetes deployments
  • Scalable via HPA
  • Resource-limited per pod
  • Independent crash recovery
Event Subjects
  • At-least-once delivery
  • Idempotency via decision_id
  • Back-pressure management
  • Priority queue for critical agents

Critical agents (is_critical = true) receive priority queue routing.

If any of the following occurs during execution:

Policy denial
Contract violation
Context staleness
LLM exception
Schema mismatch
Hash integrity failure

Result: Execution = BLOCKED · Decision logged · Evidence sealed · No action emitted. Fail-closed always.

LLM output is advisory input to a deterministic post-processor. No free-form action allowed.

ControlEnforcement
RoutingAll calls via shared_capabilities.llm.OpenAIClient
Token limitsEnforced per agent configuration
Prompt injectionInput sanitisation + context isolation
Network accessNo external network calls from LLM context
Output validationSchema validation mandatory before acceptance
Audit trailPrompt + response logged in evidence vault
EndpointMethodPurpose
/api/iaf/registry/agentsGETList all agents with metadata
/api/iaf/registry/runtime/{key}GETRuntime state for agent
/api/iaf/registry/deploymentsPOSTCreate deployment
/api/iaf/registry/deployments/{id}/validatePOSTValidate deployment
/api/iaf/registry/deployments/{id}/activatePOSTActivate (requires deploy:agents)

Activation requires: deploy:agents permission · shadow validation pass · drift check pass

  1. Single execution contract across all towers
  2. No mission-specific override paths
  3. Policy evaluated before execution
  4. Deterministic replay enforced
  5. Evidence sealed before mission UI render
  6. Kill switch globally halts execution
  7. Autonomy level strictly enforced
  8. No SOR write-back capability
Metrics
  • agent_execution_duration_seconds
  • agent_error_total
  • policy_denial_total
  • autonomy_level_distribution
  • replay_divergence_total
Heartbeat Model
  • Agents emit runtime heartbeats
  • Mission strips show live agents only
  • Stale agents auto-flagged
  • SSE streaming to monitor UI

“bp Sphere operates a deterministic, policy-bound orchestration fabric where every agent executes under a canonical contract, is replay-verifiable, horizontally scalable, and evidence-sealed before surfacing to leadership.”

The agent mesh provides topology-aware orchestration across the runtime fleet. Rather than static dispatch, the mesh evaluates agent capabilities, roles, load, and cost to determine optimal task routing — ensuring every request reaches the most appropriate agent profile while respecting governance constraints and operational SLAs.

  Task Request
       │
       ▼
  Role Router (6 strategies)
       │
       ├── Capability Match
       ├── Role-Based
       ├── Round Robin
       ├── Least Loaded
       ├── Priority
       └── Cost Optimized
       │
       ▼
  Agent Profile Selection
       │
       ▼
  Mesh Execution (topology-aware)
       │
       ▼
  Result + Execution Trace

AgentCapability Types (14)

#CapabilityDescription
1CODE_GENERATIONProduce source code from specification or prompt
2CODE_REVIEWAnalyse code for defects, style, and security
3DATA_ANALYSISStatistical and exploratory data examination
4DOCUMENT_EXTRACTIONStructured extraction from unstructured documents
5SEARCH_RETRIEVALIndexed lookup across knowledge bases and SOR data
6REASONINGMulti-step logical inference and deduction
7PLANNINGTask decomposition and sequencing
8SUMMARIZATIONCondensation of verbose input to key points
9TRANSLATIONCross-language or cross-format transformation
10MATHNumerical computation and formula evaluation
11SAFETY_CHECKContent safety and policy compliance validation
12VERIFICATIONOutput correctness and integrity verification
13CONVERSATIONMulti-turn dialogue management
14TOOL_USEExternal tool invocation and result integration

TaskType Categories (8)

#Task TypeRouted Capabilities
1CODINGCODE_GENERATION, CODE_REVIEW, TOOL_USE
2ANALYSISDATA_ANALYSIS, MATH, REASONING
3EXTRACTIONDOCUMENT_EXTRACTION, SUMMARIZATION
4SEARCHSEARCH_RETRIEVAL, REASONING
5REASONINGREASONING, PLANNING, VERIFICATION
6CREATIVECONVERSATION, TRANSLATION, SUMMARIZATION
7SAFETYSAFETY_CHECK, VERIFICATION
8GENERALCONVERSATION, TOOL_USE, REASONING

AgentProfile Schema

FieldTypeDescription
idstrUnique profile identifier
namestrHuman-readable profile name
capabilitieslist[AgentCapability]Supported capability set
roleslist[str]Assigned role bindings for routing
priorityintDispatch priority (lower = higher priority)
cost_per_requestfloatEstimated cost per invocation (USD)
avg_latency_msfloatAverage response latency in milliseconds
max_concurrentintMaximum concurrent task slots
healthyboolCurrent health status from heartbeat

Pre-Configured Profiles

Code Agent

CODE_GENERATION
CODE_REVIEW
TOOL_USE

Data Agent

DATA_ANALYSIS
MATH
DOCUMENT_EXTRACTION

Safety Agent

SAFETY_CHECK
VERIFICATION
REASONING

Search Agent

SEARCH_RETRIEVAL
SUMMARIZATION
REASONING

Runtime controlLatest status
Policy enforcementGuardrailPolicyEngine blocks under fail-closed posture and escalates non-fail-closed failures without silent bypass.
Token economyLLM calls are evaluated before spend: deterministic skip, cache, deny, downgrade, escalate, or proceed.
Autonomy demotionTelemetry guidance writes supervisory state, event log, and owner review queue.
Runtime observability CLIbpsphere close-runtime and bpsphere runtime expose runtime metrics with provenance metadata for unassisted workshop validation.
Protected-dashboard fallbackWhen protected dashboard APIs require auth, bpsphere runtime can use the public runtime proof payload and marks the source/fallback in the returned payload.
  • Runtime behavior has been migrated off scattered env branches onto profile-backed settings.
  • Startup/runtime verification now exposes a snapshot contract through /api/runtime-profile plus startup logging.
  • Core agentic surfaces now show decision/action/evidence sections rather than only KPI tiles and chat.
CapabilitiesSkill Fabric + Agentic RuntimeEvent Fabric + SOR Sidecar
DOC-06

Agent Mesh & Automation Platform

Topology-aware orchestration · Capability routing · Task queue · HITL · Workflow engine

The bp Sphere Agent Mesh provides topology-aware orchestration for 153 agents, combining capability-based routing with a full automation platform for task queuing, HITL escalation, and multi-step workflow execution.

14 Capabilities 6 Routing Strategies 8 Task Types 8 Automation Domains
  Task Request → Role Router
      ├── CAPABILITY_MATCH (best fit)
      ├── ROLE_BASED (role priority)
      ├── ROUND_ROBIN (fair distribution)
      ├── LEAST_LOADED (load balancing)
      ├── PRIORITY (highest priority)
      └── COST_OPTIMIZED (lowest cost)
              │
      RoutingDecision → Mesh Execution
      → output + tokens + execution_order

14 Agent Capabilities

CODE_GENERATION
CODE_REVIEW
DATA_ANALYSIS
DOC_EXTRACTION
SEARCH_RETRIEVAL
REASONING
PLANNING
SUMMARIZATION
TRANSLATION
MATH
SAFETY_CHECK
VERIFICATION
CONVERSATION
TOOL_USE

8 Task Types

CODING
ANALYSIS
EXTRACTION
SEARCH
REASONING
CREATIVE
SAFETY
GENERAL
  Issue/Event → AutomationFacade
      ├── Domain Client (NOC/Finance/HR/...)
      ├── Learning Client (recommendations)
      └── Task Queue (priority-based)
              │
      Worker → Auto-resolve or Escalate
      Feedback → Learning Client
DomainScopeCapabilities
NOCNetwork opsAlert triage, auto-remediation
FINANCEFinancial opsInvoice processing, reconciliation
HRHuman resourcesOnboarding, timesheet validation
RETAILRetail opsInventory alerts, pricing anomalies
SUPPLY_CHAINSupply chainLogistics tracking, demand planning
LEGALLegal opsContract review, risk flagging
HSEHealth/SafetyIncident triage, permit validation
HEALTHCAREHealthcarePatient flow, resource scheduling
StepPurposeBehavior
ACTIONExecute taskRuns action handler
CHECKValidatePre/post conditions
DECISIONBranchContext-based routing
WAITPauseTimer or event resume
HUMANHITLWaits for human decision
PARALLELConcurrentMultiple steps
NOTIFYAlertStakeholder notification
ROLLBACKUndoCompensating transactions
ai_mesh
control_plane
data_fabric
federated_gw
guardian
integration
payment_gw
saga
trust
workflow

Skill Fabric introduces reusable capability modules between agents and tools. Agents orchestrate; skills provide domain expertise; tools execute SOR actions.

  Mission/Event
      │
      ▼
  Agent Orchestrator
      │
      ▼
  Skill Selector (context + role + policy)
      │
      ├── Primary skill
      ├── Supporting skills
      └── Guardrail/evidence skills
      ▼
  Tool calls + evidence assembly
Surface/APIPurpose
/ui/skill-fabricSkill runtime and composition visibility
/ui/skills-registrySkill catalog and metadata lookup
/ui/skill-replayExecution replay for skill runs
/api/iaf/skill-fabric/registry/{skill_id}Resolve skill metadata and contract
/api/iaf/skill-fabric/lifecycle/{skill_id}Skill lifecycle and governance state
  • Scenario presets are first-class named inputs instead of freeform demo prompts.
  • Trigger realism has been expanded with mission-specific scenario/question catalogs and live recommendation/trace/evidence flows.
  • In bp Sphere domain apps, multi-stack drawers now surface real transaction objects instead of static trace placeholders.
CapabilitiesPolicy + Authority Control
DOC-07

Human-in-the-Loop & Authority Queue Fabric

Governed delegation · SLA-enforced escalation · Evidence-bound approvals

Human-in-the-Loop (HITL) is not a fallback mechanism. It is a formal authority orchestration layer.

  • Risk-bearing decisions are authorised by the correct role
  • Materiality thresholds are enforced deterministically
  • SLAs are measurable and enforceable
  • Escalations are automatic and replayable
  • All approval actions are sealed as evidence

No execution may cross materiality boundaries without policy-defined authority.

  Agent Decision Candidate
              │
              ▼
  Policy + Materiality Engine
              │
              ├── AL-permitted ──────────▶ Execute + Seal Evidence
              │
              └── Approval Required
                      │
                      ▼
            Authority Queue (Event-backed)
                      │
              Role-Routed Assignment
                      │
                      ▼
            Approve / Modify / Reject / Timeout
                      │
                      ▼
        Decision Ledger + Evidence Sealing
Stateful queue with persistent state
Event-driven via NATS subject
SLA-monitored with auto-escalation
Deterministically resumable
StateDescriptionTransition
CREATEDAgent produced approval-required decisionPolicy gate requires HITL
PENDING_APPROVALRouted to authorised roleAwait action
APPROVEDAuthority approvesResume execution
MODIFIEDAuthority edits proposalRe-evaluate under policy
REJECTEDAuthority deniesClose + seal
ESCALATEDSLA breach or risk triggerRoute to higher authority
EXPIREDNo action within SLAForced escalation or closure
Idempotent
Logged
Evidence-sealed
Replay-verifiable
Decision TypeMaterialityDefault ApproverEscalationSLA
Working capital action< $100kControllerTreasury Lead24h
Working capital action$100k–$1MTreasury LeadCFO12h
Policy overrideAnyCFOCFO Delegate + Internal Audit8h
Cross-entity closeAnyR2R Tower LeadCFO6h
High-risk complianceAnyCompliance LeadCFO + Legal4h

Authority rules defined in config/tenants/bp/approval_chains.yaml. No runtime override permitted outside config.

Each approval request includes: created_timestamp · sla_deadline · risk_classification · authority_level

If current_time > sla_deadline then:

  1. Escalate to next authority level
  2. Emit alert event (NATS + UI notification)
  3. Log escalation evidence
  4. Notify via UI + notification service

Escalation is deterministic and replayable.

Distinct-user rule:

Submitter cannot approve own decision.

AspectImplementation
EnforcementSODRule model + POST /api/iaf/policies/check-sod
Violation resultAutomatic rejection + governance violation logged
EvidenceViolation sealed in evidence vault
OverrideNo override allowed without dual-approval rule

HITL state persisted via hitl_store.py:

Redis

Primary store

In-memory LRU

Fallback (non-prod)

PostgreSQL

Ledger snapshot

On Restart:
  • Pending approvals reloaded
  • SLA recalculated
  • Escalation timer re-armed
  • No approval request lost
Approved
  • Resume execution at envelope
  • Seal execution evidence
Modified
  • Re-run policy evaluation
  • Recompute hashes
  • New decision_id version
Rejected
  • Close decision
  • Seal rejection evidence
  • Emit governance record

All flows sealed in GovernanceOutboxRecord.

SurfaceRoutePurpose
Confirmation Queue/ui/confirmation-queueCentralised approval dashboard
My Work/ui/my-workRole-specific pending actions
Work Item Detail/ui/work-item/{item_id}Decision summary, evidence, controls

Displays: decision summary · materiality · risk classification · agent identity · SLA countdown · evidence snapshot · Modify / Approve / Reject controls

No UI action bypasses governance middleware.

ControlEnforcement
approve:decisions permissionGovernanceMiddleware
SOD enforcementPolicy engine
Kill switchSupervisor
Escalation auto-routingSLA engine
Audit immutabilityEvidence vault
Replay consistencyDecision envelope hash
KPITargetPurpose
Approval latency p95< 6hVelocity discipline
Expired approvals< 1%Escalation reliability
Override rate< 5%Policy calibration
Replay-ready HITL records100%Audit defensibility
Escalation correctness100%Authority routing integrity

Monitored via Prometheus + Governance dashboards.

If any of the following occurs:

Approver unavailable
Identity provider unreachable
Redis failure
Policy mismatch

Result: Queue remains pending · Execution blocked · No action emitted · Escalation triggered when SLA met. Fail-closed always.

  1. No risk-bearing action bypasses authority routing
  2. Every approval is evidence-sealed
  3. Every rejection is logged
  4. Escalations are deterministic
  5. No self-approval permitted
  6. SLA breaches auto-escalate
  7. HITL history replayable for 7 years

“bp Sphere operates a formal authority orchestration fabric where material decisions are SLA-governed, separation-of-duties enforced, escalation deterministic, and every approval is evidence-sealed for replay and audit.”

Frequently Asked Questions

Q: Can business users bypass the authority queue?
A: No. Governance middleware blocks any direct execution route. Approval routing is mandatory for material decisions.
Q: What if an approver is unavailable?
A: SLA breach triggers deterministic escalation to next configured authority level with full evidence continuity.
Q: Is approval history auditable?
A: Yes. All approval events (requested, approved, modified, rejected, expired, escalated) are sealed in the evidence vault and replay-verifiable.
Q: Does HITL reduce automation value?
A: No. HITL enables controlled promotion from AL-1/2 to AL-3 safely by proving governance discipline.
  • Demo and prod authority behavior are now distinguished primarily by runtime profile and role forcing, not by separate code paths.
  • Showcase remains the right public demo shell; prod should keep the broader shell and not inherit demo role forcing.
CapabilitiesPolicy + Authority Control
DOC-08

Enterprise Control Plane Architecture

Operating control centers · Deterministic enforcement · Role separation · Controlled autonomy promotion

The Enterprise Control Plane is the operating and governance layer above the operational data plane. Agents, signals, workflows, evidence retrieval, recommendations, and human approvals perform work; the Control Plane governs, pauses, recovers, audits, and tunes that work.

CentralizedDeterministicInterruptibleReplay-Verifiable
    Agent Execution Request
            │
            ▼
    1. Registry Validation
            │
            ▼
    2. Autonomy Gate Assessment
            │
            ▼
    3. Policy Evaluation Engine
            │
            ├── Permitted ──────────────▶ Execute
            │
            └── Approval Required ──────▶ Authority Queue
                                             │
                                             ▼
                                    Resume / Reject
                                             │
                                             ▼
    4. Monitoring + Drift Check
            │
            ▼
    5. Evidence Seal + Ledger Write

There is no direct execution path around this flow.

Control CenterPurposeOperator ActionsRuntime Proof
Agent Control CenterOperate the governed digital workforce.Pause, resume, disable, rollback/ui/bp-agent-inventory
Policy Control CenterGovern active policies, policy changes, violations, and simulations.Activate, deactivate, test, compare versions/ui/policy-library
Runtime Operations CenterEnterprise NOC for events, signals, queues, decisions, executions, and runtime health.Inspect, drain queue, failover, recover/ui/runtime-operations-surface
Evidence Control CenterMonitor evidence sources, missing evidence, evidence quality, and confidence.Inspect pack, request evidence, challenge evidence, seal pack/ui/evidence
Decision Control CenterExpose all decisions, confidence, human overrides, and escalations.Inspect, escalate, override, approve/ui/decision-records
Replay Control CenterReconstruct historical decisions and generate audit packages.Replay, export audit pack, compare runs, preserve snapshot/ui/decision-replay-studio
Learning Control CenterGovern feedback, learning events, model drift, and confidence changes.Inspect learning, freeze learning, approve update, rollback memory/ui/decision-memory
Risk Control CenterMonitor high-risk decisions, policy breaches, agent failures, escalation backlog, and blast radius.Kill switch, contain, route incident, open DR runbook/ui/enterprise-resilience

Current runtime proof: GET /api/iaf/control-plane/enterprise-control-plane and /ui/control-plane.

1. Agent Registry Enforcement

Component: IRISAgentRegistry (singleton). Source of truth: agent_registry.yaml (v1.2.0).

Guarantees:

  • Agent must be registered
  • Agent must be ACTIVE
  • Version must match deployment
  • Dependencies resolved
  • Kill switch not engaged

Methods: list_all(), list_by_tower(), list_implemented(), suspend(), resume(), validate_startup()

Startup validation blocks runtime if registry integrity fails.

2. Autonomy Gating Engine

Configuration: autonomy_governance.yaml

LevelNameMateriality
AL-1AdvisoryUnlimited (human decides)
AL-2Supervised≤ $500K
AL-3Governed Auto≤ $50K
AL-4Fleet Auto≤ $10K

API: POST /api/iaf/supervisor/agents/{id}/autonomy, GET .../autonomy/assess

Gate checks:

  • Confidence threshold
  • Materiality bounds
  • Historical performance
  • Incident rate
  • Policy coverage

If autonomy threshold violated → forced downgrade.

3. Policy Enforcement Engine

Policies stored in: services/policy/policies/

global.yamlp2p.yamlo2c.yamltreasury.yamlmodel_risk.yaml

Every evaluation recorded in PolicyEvaluationLog: policy_id, policy_version, decision_id, outcome, rationale.

SOD checks: POST /api/iaf/policies/check-sod

No execution without positive policy verdict.

4. Authority Queue (HITL Integration)

Controlled by supervisor.py

APIs: GET /queue, POST /queue/{id}/assign, POST /queue/{id}/resolve

ApproverSLA
Controller24 hours
Treasury12 hours
CFO8 hours
Compliance4 hours
PersistentEvent-DrivenDeterministically ResumableReplay-Verifiable
5. Execution Monitoring & Drift Control

Component: agent_monitor.py

APIs: GET /summary, GET /agents, GET /audit, GET /feed, GET /stream (SSE), GET /towers

Models: AgentMonitorEvent, AgentExecutionTelemetry, AgentHeartbeat

Drift Detection: POST /supervisor/agents/{id}/drift-check driven by drift_rules.yaml

Drift Triggers:

  • Automatic autonomy downgrade
  • Kill switch escalation
  • Governance alert
6. Kill Switch & Circuit Breaker

Global: POST /api/iaf/supervisor/kill-switch

Per-Agent: POST /api/iaf/agent-monitor/agents/{id}/kill-switch

Access: admin:supervisor permission required

Circuit Breaker Auto-Trigger Conditions:

  • Excessive policy violations
  • Error rate threshold breach
  • Drift beyond tolerance
  • Replay divergence detected

Blast Radius: bp Sphere only. Source systems unaffected.

The following must always be true at runtime:

1Every agent execution references a registry entry
2Every execution evaluated against active policy version
3Every autonomy decision recorded
4Every decision sealed in ledger
5No agent bypass path exists
6Kill switch halts execution within one control loop
7Control plane failure = fail-closed
CapabilityAdminBusiness (CFO, Treasurer)
Register / deregister agents
Modify policy thresholds
Activate kill switch
Promote autonomy level
Approve decisions
View evidence
Escalate decisions
Admin

Platform governance authority

Business

Operational decision authority

No single role can both modify policy AND approve the resulting decision.

AL-1
Advisory
AL-2
Supervised
AL-3
Governed Auto
AL-4
Fleet Auto
GateAL-1 → AL-2AL-2 → AL-3AL-3 → AL-4
Determinism verified
Shadow parallel runOptionalRequired (30d)Required (90d)
Policy coverage > 95%
CFO sign-off
Zero critical incidents (90d)

Promotion cannot be automatic. Requires: governance validation, performance review, and executive sign-off.

Architecture Characteristics
  • Stateless API layer
  • Horizontally scalable
  • Backed by persistent ledger
  • Event-driven
Pod Restart Behavior
  • No in-flight execution lost
  • Queue state reloaded
  • Policy version intact
  • Autonomy state preserved

Failure behavior: Fail-closed. No uncontrolled execution.

Prometheus Metrics:

policy_denial_totalautonomy_downgrade_totalkill_switch_trigger_totaldrift_detected_totalsupervisor_latency_seconds

Control Plane Dashboards:

Autonomy Distribution
Escalation Heatmap
Drift Incidents
Kill Switch History
Policy Override Rate

Frequently Asked Questions

Q: Can an agent bypass the control plane?
A: No. Execution contract requires supervisory validation before any output is produced.
Q: Who can activate the kill switch?
A: Admin roles only. It can also be auto-triggered by circuit breaker logic when anomaly thresholds are exceeded.
Q: Can autonomy be promoted automatically?
A: No. Promotion requires shadow validation, policy coverage verification, and executive approval.
Q: What happens if the control plane fails?
A: Execution halts. No decision is emitted. Source systems remain unaffected.

Positioning Statement (For Orals)

bp Sphere operates a centralized supervisory control plane that enforces registry validation, autonomy gating, policy evaluation, drift monitoring, and interruptible execution on every agent action — with deterministic replay and role-separated authority.
Enterprise DisciplineNo Uncontrolled AIClear Risk OwnershipPromotion MaturityDay-1 Safety
  • Supervisory behavior now has stronger runtime observability through profile snapshots and environment separation.
  • Public worker routing has been cleaned so demo, dev, and prod can be exposed separately instead of reusing one public URL for everything.
CapabilitiesPolicy + Authority ControlTrust Fabric + Replay
DOC-09

Policy Engine & Authority Model

Deterministic rule evaluation · Version-controlled DSL · SOX-aligned enforcement

The Policy Engine is a separate, declarative governance layer that evaluates every agent decision before output. Agents propose. Policies decide.

External to Agent LogicYAML-Driven DSLVersion-ControlledReplay-VerifiableTenant-OwnedAuditable

No agent may emit a decision without policy evaluation.

    Agent Decision Envelope
            │
            ▼
    Load Active Policy Set (tenant + global)
            │
            ▼
    Evaluate Rule Graph
            │
            ▼
    Resolution Engine
            │
            ├── APPROVE ──────▶ Seal + Continue
            ├── ESCALATE ─────▶ Authority Queue
            ├── BLOCK ────────▶ Close + Evidence
            └── MODIFY REQ ───▶ Re-evaluate

There is no bypass path around this engine.

The engine consists of six deterministic stages:

1. Policy Loader
2. Rule Parser
3. Condition Evaluator
4. Conflict Resolver
5. Action Resolver
6. EvaluationLog Recorder
DeterministicStateless per invocationContext-boundVersion-pinned

Each evaluation logs:

decision_idpolicy_idpolicy_versionoutcomerationaletimestamp

Replay always uses original policy_version.

Canonical example (global.yaml):

    policies:
      - id: materiality_threshold
        type: MATERIALITY
        severity: HIGH
        priority: 100
        description: "Transactions > $500K require 2 approvers"
        rules:
          - condition: "transaction_amount exceeds policy materiality"
            required_approvers: [finance_controller, cfo]
            evidence_level: FULL_AUDIT_PACK
        applicable_processes:
          - invoice_approval
          - payment_release
          - journal_posting
          - capitalization

Policy domains:

treasury.yaml — FX exposure limits, counterparty ceilings
p2p.yaml — Procurement authority chains
o2c.yaml — Credit limit controls
model_risk.yaml — LLM confidence floors + hallucination detection

Policies are declarative — no embedded business logic.

Conflict Resolution Rule: Most restrictive policy wins.

If Policy A → Approve and Policy B → Escalate

Final: Escalate

If Policy A → Escalate and Policy B → Block

Final: Block

Resolution Precedence (highest to lowest):

BLOCKESCALATEMODIFYAPPROVE

Policy priority attribute resolves tie-breaking. Conflicts are validated at activation time before deployment.

RoleMax Auto-ApproveEscalation Target
Analyst$0Controller
Controller$100,000Treasurer
Treasurer$1,000,000CFO
CFOUnlimitedBoard (policy-defined)
Tenant-configurableDomain-specificVersion-controlledReplay-bound

Platform enforces; BP defines values.

Policy changes require a formal activation workflow:

1. Draft Creation2. Diff Validation3. Conflict Analysis4. Admin Approval5. Shadow Validation6. Activation

Each change records:

policy_idversionauthortimestampchange_diff

No policy can change silently.

Policy activation invalidates previous version for new executions only. Historical decisions remain bound to prior version.

BehaviorWORKSHOPLIVE
Policy evaluationFullFull
Threshold enforcementActiveActive
Override allowedYesCFO approval required
Evidence retention90 days7 years
External actionsBlockedBlocked

There is no relaxed policy mode in LIVE. Simulation parity ensures credibility.

Override Requirements
  • CFO-level approval
  • Written justification
  • Separate override evidence pack
  • Version-linked to policy
Override Tracking
  • Frequency metric
  • Drift impact analysis
  • Domain-level reporting
  • Governance health KPI
RareRole-BoundEvidence-SealedExplicitly Justified
SOX Controlbp Sphere EnforcementEvidence Artifact
Segregation of dutiesRBAC + SOD policyApproval log
Dual approvalAuthority queueMulti-approver chain
Change managementPolicy versioningVersion diff history
Access controlEntra ID + RBACAccess log
Audit trailImmutable ledgerFull lineage pack
Materiality governancePolicy DSL thresholdsEvaluation log
Exception documentationOverride evidenceJustification artifact

This makes the Policy Engine directly auditable under SOX frameworks.

The following must always be true:

1Every decision evaluated against active policy set
2Every evaluation logged
3Policy version pinned to decision
4Conflicts resolved deterministically
5No silent override possible
6Replay reproduces identical policy outcome
7No SOR write-back capability
PermissionScope
admin:policyCreate / modify policies
deploy:policyActivate new version
approve:overrideAuthorize override
view:policy-logInspect evaluation history
admin:supervisorKill switch + autonomy promotion

Business users cannot modify policy thresholds.

Frequently Asked Questions

Q: Can policies be changed without audit trace?
A: No. All policy changes are version-controlled, diff-tracked, author-recorded, and activation-logged.
Q: Can agents embed hidden logic to bypass policy?
A: No. Policy evaluation occurs after agent output and before decision emission.
Q: How are policy conflicts resolved?
A: Most restrictive rule wins. Conflicts are validated before activation.
Q: Are policies applied in simulation exactly as in LIVE?
A: Yes. There is no relaxed enforcement in LIVE mode.

Positioning Statement (For Orals)

bp Sphere operates a deterministic, version-controlled policy engine where agent logic is strictly separated from governance rules, materiality thresholds are tenant-defined, conflicts resolve to the most restrictive outcome, and every evaluation is replay-verifiable and SOX-auditable.
Governance MaturitySeparation of ConcernsLegal DefensibilityNo Opaque AI LogicPromotion Safety
Observability areaLatest implementation
Readiness posture/observability/readiness exposes OTEL enablement, endpoint binding, Prometheus route, failure-mode surface, assistant-ops surface, and declared SLO targets.
Assistant operationsAssistant Operations Center shows connector packages, certification, compliance mappings, failover proof, attribution, and costs.
Replay determinismLLM audit runtime validates execution seed and evidence consistency before replay claims are accepted.
Continuous Close lineagebpsphere provenance and bpsphere close-runtime expose source systems, calculation method, replay ID, confidence basis, seeded/live badge, and query timestamp for close metrics.
Close Integrity resilienceThe Close Integrity UI now degrades to a safe empty state when source tables are unavailable instead of returning a page error.
  • Policy control is now backed by runtime profiles, policy metadata in traces, and environment-specific behavior settings.
  • The pack should now describe the authoritative mode switch as IRIS_RUNTIME_MODE, not scattered IRIS_SKIP_* flags.
CapabilitiesTrust Fabric + ReplayValue Attribution Engine
DOC-10

Deterministic Execution & Evidence Vault

Reproducible by construction · Cryptographically sealed · Regulator-ready

Every bp Sphere decision is reproducible by design. Given the same input snapshot, the same policy version, the same model version, the same seed, and the same enterprise time — the system produces the same output, every time.

This guarantee is enforced at runtime, not assumed.

The Evidence Vault stores immutable decision artifacts with full lineage, cryptographic sealing, and audit export capability.

Every execution produces a sealed envelope with the following fields:

FieldDescriptionEnforcement
execution_idGlobally unique execution referenceUUID v4
agent_idRegistered agent identityMust exist in registry
agent_versionExact deployed versionDeployment-pinned
inputs_hashSHA-256 of canonical input envelopeVerified on replay
outputs_hashSHA-256 of normalized outputVerified on replay
policy_hashSHA-256 of policy versionVersion-pinned
policy_versionHuman-readable policy versionImmutable reference
model_versionModel checkpoint identifierExact version recorded
seedExplicit random seedInjected, not derived
enterprise_timeInjected clockNot system time
modeWORKSHOP or LIVEControls retention + controls
evidence_hashHash of entire packIntegrity anchor

The evidence_hash seals the entire envelope. No field may be modified after sealing.

Seed Injection

All stochastic components receive explicit seed input. No system randomness permitted.

Enterprise Clock Injection

Time is injected from enterprise context. No reading of system clock.

Immutable Context Snapshot

Inputs derived from immutable snapshot. No live DB queries during execution.

No External Network Calls

Agent runtime is network-restricted. No HTTP or outbound calls permitted.

Output Canonicalization

Agent output normalized to schema before hashing.

Immediate Hash Sealing

Outputs hashed immediately after generation. No post-processing permitted.

Violation of any invariant results in BLOCKED execution.

Replay is not optional — it is part of the contract.

    Load Evidence Pack
            │
            ▼
    Reconstruct Input Snapshot
            │
            ▼
    Re-execute Agent (same version, seed, time)
            │
            ▼
    Compare outputs_hash
            │
            ├── MATCH ────▶ Validated
            └── MISMATCH ─▶ Determinism Violation

If mismatch detected:

  • Execution flagged
  • Agent quarantined
  • Drift incident logged
  • Investigation triggered
  • Original evidence remains immutable

Replay compliance in LIVE must equal 100%.

The system supports controlled counterfactual execution:

"What would this decision have been under Policy v2 instead of v3?"

Process:

1. Load original snapshot2. Substitute policy_hash3. Re-execute (same seed + model)4. Compare outputs_hash

Use cases:

Policy impact analysisThreshold calibrationAutonomy promotion validationCFO scenario modeling

Counterfactual outputs are marked SIMULATED. They never overwrite original evidence.

ComponentFilePurpose
Evidence Engineservices/iris_evidence.pySHA-256 sealing
Backfill Serviceservices/evidence_backfill_service.pyEnsures missing packs sealed
Vault APIapi/evidence_vault.pyRetrieval + verification
Evidence Modelsmodels/_evidence.pyPack schema
Evidence UIui/evidence.pyEvidence browser
Value Attributionservices/value_models/KPI attribution

Pack ID format: EP-CFO-{timestamp}-{8-char-hash}

Evidence packs include:

Full input snapshot reference
Decision trace
Policy evaluation log
Autonomy level
Authority routing (if any)
Model version metadata
Value attribution (if applicable)
🔍 Illustrative operating example — Evidence pack structure
decision_id: DEC-2026-03-001245

inputs:
  invoice_id: INV-88921
  supplier_id: SUP-334
  contract_terms: version_3

outputs:
  recommendation: HOLD
  reason: variance_exceeds_threshold

policy:
  file: p2p_policy.yaml
  rule: variance_limit_2_percent

hashes:
  inputs_hash: sha256:abc123...
  outputs_hash: sha256:def456...
  evidence_hash: sha256:xyz789...

replay:
  status: VERIFIED
  replay_timestamp: 2026-03-20T15:02:11Z
Evidence backfill loopMANDATORY
Cannot disable when live_only_mode=TrueENFORCED
Evidence retention7 YEARS
Hash verificationENFORCED
DeletionRESTRICTED

LIVE metadata: model_version=2.0.0, policy_version=FP&A-GOV-v3.2, sensitivity_matrix=v3.2

LIVE Mode
7 years
SOX-compliant
WORKSHOP Mode
90 days
Experimentation

Evidence is immutable once sealed. Deletion requires:

  • Governance approval
  • Audit record generation
  • Deletion evidence pack

No silent deletion possible.

FormatUse CaseContents
JSONProgrammatic auditFull envelope
PDFRegulator submissionHuman-readable trace
CSVBulk analyticsStructured decision rows

Each export is logged:

requester_identitytimestampscopeexport_hash

Evidence packs may include value attribution models:

dso_improvement
liquidity_preservation
manual_effort_avoided
decisions_surfaced_early
p2p_cost_avoidance
Model-versionedPolicy-boundReplay-verifiable

This allows CFO to trace value to deterministic decisions.

The following must always hold:

1Every execution produces evidence pack
2Every pack is hash-sealed
3Replay must reproduce identical output
4Policy version pinned permanently
5Model version recorded explicitly
6Seed and time injected deterministically
7No external calls permitted during execution
8LIVE evidence retained for 7 years

If any invariant fails → execution blocked.

Frequently Asked Questions

Q: What if replay produces a different result?
A: This is a determinism violation. The agent is quarantined, the event logged, and investigation triggered. Original evidence remains immutable.
Q: Can evidence be deleted?
A: Only via governed deletion process that itself generates an evidence artifact. LIVE evidence defaults to 7-year retention.
Q: How does counterfactual comparison work?
A: The same inputs are re-executed under an alternate policy or threshold version. Outputs are compared, but original decision remains unchanged.
Q: Can evidence be altered after creation?
A: No. Evidence packs are cryptographically sealed using SHA-256. Any modification would invalidate the evidence_hash.

Positioning Statement (For Orals)

bp Sphere operates a cryptographically sealed, deterministic execution contract where every decision is replay-verifiable, counterfactual-testable, and retained for seven years under SOX-grade audit controls.
Legal DefensibilityRegulatory ReadinessControl DisciplineNo Black-Box AIEnterprise Maturity
Evidence areaLatest implementation
Finance reconciliationFP&A variance and external disclosure metrics now carry field-level provenance, source workbook cells, and mapping metadata.
Decision tracesRecommendation traces align to the actual decision object and evidence pack rather than random/generated trace state.
Execution auditExecution IDs, input/output hashes, and evidence hashes are persisted and queryable.
CapabilitiesOntology + Enterprise ContextEnterprise Memory + Knowledge Management
DOC-11

Data Architecture & Context Model

Canonical normalization · Cross-SOR identity graph · Traceable lineage

bp Sphere does not replicate ERP systems. It references systems of record as authoritative truth. The data architecture establishes:

  • A canonical entity model
  • Cross-SOR identity resolution
  • Immutable context snapshots
  • Full lineage metadata
  • Domain-scoped access control

The objective is not duplication — it is deterministic reasoning on normalized context.

Canonical entities provide a normalized, cross-domain schema that:

Harmonizes

Fields across 114 SORs

Eliminates

Semantic drift

Enables

Cross-tower reasoning

Preserves

Traceability to source

Read-optimizedContext-boundVersion-awareNon-authoritative

SOR remains system of record.

🔍 Illustrative operating example — Canonical invoice context
Canonical entity: Invoice INV-88921

Source references:
  - SAP: 480002194
  - Ariba: IR-22091
  - BlackLine case: BL-9921

Normalized fields:
  - supplier_id: SUP-334
  - amount_usd: 12400000
  - variance_pct: 3.8
  - policy_refs:
      - p2p_policy.yaml#variance_rule

Decision use:
  - shared by P2P, close, and treasury reasoning without mutating ERP truth
EntitySource SORsKey FieldsPrimary Use
InvoiceSAP, ServiceNowinvoice_id, vendor, amount, currency, due_dateAP analysis, anomaly detection
PaymentSAP, Treasurypayment_id, amount, beneficiary, value_dateCash risk monitoring
JournalEntrySAPdoc_number, company_code, posting_dateClose integrity
CostCenterSAP, Workdaycost_center_id, hierarchy, ownerBudget control
VendorSAP, Workday, ServiceNowcanonical_vendor_id, tax_id, nameSupplier risk
Decisionbp Spheredecision_id, agent_id, confidenceAudit trail
Policybp Spherepolicy_id, versionGovernance lineage
Evidencebp Sphereexecution_id, inputs_hashRegulator export

Canonical entities are resolved and stored in normalized form while preserving source_refs.

Identity resolution is graph-based, not simple mapping.

    SAP Vendor 10042
    Workday Supplier 887
    ServiceNow CI-39221
            │
            ▼
    Canonical Vendor: V-00421
            │
            ├── name: "Acme Corp"
            ├── tax_id: "GB123456789"
            ├── source_refs:
            │     - SAP:10042
            │     - WD:887
            │     - SN:39221
            ├── match_score: 0.97
            └── resolution_method: deterministic + fuzzy

Resolution logic includes:

  • Exact key match (tax ID, legal entity ID)
  • Deterministic join keys
  • Controlled fuzzy matching
  • Confidence scoring
  • Manual override capability

Identity records include confidence metadata to prevent silent merge errors.

Agents never query live SOR during execution.

1. Data extracted via CDC / REST / batch2. Normalized into canonical schema3. Versioned snapshot created4. Snapshot ID injected into envelope
ImmutableTime-stampedVersion-vector taggedDomain-scoped

Snapshot reference is stored in Evidence Vault.

Every canonical data point carries metadata:

Metadata FieldPurposeExample
source_systemOrigin SORSAP_S4_PROD
source_record_idNative keynative-key-10042
extraction_timeRead timestamp2025-01-15T14:30:00Z
transformationApplied normalizationcurrency_convert(GBP→USD)
version_vectorMulti-SOR versionSAP:v1042, WD:v887
staleness_flagFreshness classificationFRESH / STALE / EXPIRED
data_classificationSensitivity tagCONFIDENTIAL_FINANCE
PersistedQueryableIncluded in evidence packsVisible in UI drill-down

No decision is detached from data origin.

Each context snapshot maintains version vectors per SOR:

SAP:v1042Workday:v887Treasury:v204

If any SOR version advances beyond staleness threshold:

  • staleness_flag = STALE
  • Agents receive context freshness warning
  • Policy may block execution

Freshness thresholds configurable per domain.

DomainEntities AccessibleAgent Access
TreasuryPayment, FXRate, CashPositionTreasury agents
Accounts PayableInvoice, Vendor, PaymentAP + Close agents
CloseJournalEntry, ReconciliationClose agents
FP&ABudget, Forecast, ActualsFPA agents

Access enforced by:

Governance middlewareDomain tagsRBAC roles

Cross-domain joins require explicit policy allowance.

Currently tracking 10 active domains with 0 canonical entities.

200+ SQLModel entities across 30+ modules:

ModuleFocusNotable Classes
_base.pyEnums & core typesAutonomyLevel, FinanceTower
_agents.pyAgent lifecycleAgentDefinition, AgentExecution
_r2r.pyClose domainJournalEntry, Reconciliation
_p2p.pyProcure-to-payVendor, Invoice, Payment
_governance.pyPolicy + SODPolicy, ApprovalThreshold, PolicyEvaluationLog
_i2c.pyOrder-to-cashCustomer, CreditAssessment
_treasury.pyTreasuryCashPosition, Exposure
_evidence.pyEvidence vaultEvidencePack, EvidenceArtifact
_finance_truth.pyCFIN mappingCFINMapping
OthersHR, Legal, Retail, TradingDomain-specific models
SQLModel (SQLAlchemy + Pydantic)PostgreSQL port 15435DB: iris_agentic_financeAuto-created via init_db()

Each entity tagged with classification:

PUBLICINTERNALCONFIDENTIALCONFIDENTIAL_FINANCEHIGHLY_RESTRICTED

Classification influences:

  • Policy enforcement
  • Autonomy gating
  • Export permissions
  • Evidence redaction

Sensitive fields masked in UI where required.

The bp Sphere ontology provides a 4-layer semantic object model that governs how entities are defined, related, triggered, and audited across the platform.

  ┌────────────────────────────────────────────┐
  │ EVIDENCE LAYER                             │
  │ EvidencePack · Decision · AuditTrail       │
  ├────────────────────────────────────────────┤
  │ DECISION LAYER                             │
  │ AgentExecution · PolicyEvaluation          │
  │ Recommendation · Confidence                │
  ├────────────────────────────────────────────┤
  │ DOMAIN LAYERS                              │
  │ Finance: Invoice · Supplier · Payment      │
  │ HR: Employee · Position · WorkforcePlan    │
  ├────────────────────────────────────────────┤
  │ CORE / KERNEL LAYER                        │
  │ Party · OrganizationUnit · Agreement       │
  │ Account · EvidencePack                     │
  └────────────────────────────────────────────┘
LayerObjectsPurpose
Core / KernelParty, OrganizationUnit, Agreement, Account, EvidencePackUniversal base types inherited by all domains
Domain LayersFinance: Invoice/Supplier/PO/Payment; HR: Employee/Position/WorkforcePlanDomain-specific objects with agent bindings, triggers, and mission context
DecisionAgentExecution, PolicyEvaluation, RecommendationAgent reasoning artifacts linked to ontology objects
EvidenceEvidencePack, AuditTrail, DecisionImmutable proof artifacts for every execution
Ontology API

9 REST routes under /api/iaf/ontology/ — objects, relationships, triggers, actions, lifecycles, and trigger resolution

Ontology Explorer UI

5-tab interface: Object Registry, Relationship Graph, Trigger Engine, Action Bindings, Object Lifecycle

Trigger Engine

Event-driven triggers bind ontology objects to agents — e.g. Invoice.created triggers InvoiceValidationAgent

Tenant Packs

BP-specific overlays for finance, procurement, and HR ontology extensions

A Neo4j-backed knowledge graph provides contextual reasoning for agent decisions, combining semantic search, entity extraction, and graph traversal.

  Query / Agent Context Request
              │
              ▼
  GraphRAG Retriever
      ├── Semantic Search (embeddings)
      ├── Entity Extraction (LLM)
      └── Graph Traversal (Neo4j)
              │
              ▼
  Knowledge Graph (Neo4j)
      ├── Entities (typed, with properties)
      ├── Triples (subject-predicate-object)
      └── Ontology Schema (constraints)
              │
              ▼
  Inference Engine
      ├── Multi-hop reasoning
      ├── Path finding (BFS)
      └── Relationship inference
              │
              ▼
  Enriched Context → Agent Execution
ComponentImplementationPurpose
KnowledgeGraphshared_capabilities.ai.knowledge.coreEntity/triple store with search, path-finding, and serialization
GraphRAGshared_capabilities.ai.knowledge.retrieverCombined graph + vector retrieval for agent context enrichment
InferenceEngineshared_capabilities.ai.knowledge.retrieverMulti-hop reasoning across entity relationships
OntologySchemashared_capabilities.ai.knowledge.ontology.schemaType validation, constraint checking (cardinality, domain, range)
ChainOfThoughtshared_capabilities.ai.reasoning.chain_of_thoughtExplainable reasoning with step tracking and self-consistency

Graph-Enhanced Decisions: Agents can traverse entity relationships to discover contextual signals that flat table queries would miss — e.g. finding all suppliers connected to a flagged payment through multi-hop graph traversal.

Context BridgeSource DomainsDecision Value
Employee → CostCenterHR + FP&ACapacity-to-cost alignment for workforce plans
Role demand → Project/ProgramHR + OperationsHiring prioritization against delivery risk
Attrition signal → Financial impactHR + FinanceRetention intervention with quantified cost/risk
Training compliance → Operational readinessHR + HSE/OperationsReadiness controls before mission execution
HR command drilldown contract
  • L1: KPI state and trend
  • L2: Driver decomposition
  • L3: Agent reasoning + policy checks
  • L4: Evidence/replay bundle
Cross-domain guardrails
  • Least-privilege joins only
  • Policy-gated high-sensitivity attributes
  • Lineage preserved across joins
  • Evidence references retained end-to-end

The following must always hold:

1SOR remains authoritative truth
2Canonical entities preserve source references
3Context snapshots are immutable
4Identity resolution includes confidence scoring
5Every decision references a snapshot ID
6Stale data flagged before execution
7Cross-domain access is policy-gated
8No SOR write-back capability

Frequently Asked Questions

Q: Does bp Sphere copy all SOR data?
A: No. It creates versioned context snapshots for reasoning. Raw SOR data remains in the SOR.
Q: How is data currency ensured?
A: Version vectors track freshness. Domain-specific staleness thresholds trigger warnings or policy blocks.
Q: Can identity resolution merge entities incorrectly?
A: Confidence scores and override workflows prevent silent merges. Resolution metadata is fully traceable.
Q: Is lineage visible to auditors?
A: Yes. Lineage metadata is included in evidence packs and accessible via the evidence UI.

Positioning Statement (For Orals)

bp Sphere operates a canonical context layer that harmonizes 114 systems of record into a version-aware, lineage-traceable entity model with cross-SOR identity resolution and domain-scoped access control — without replicating or mutating ERP truth.
Architectural MaturityData Governance DisciplineAudit DefensibilityCross-Domain IntelligenceEnterprise Readiness
Data layerCurrent implementation
Operational truthLive SOR and planning context remain the primary source for current-state operations.
External truthBP disclosure data is now DB-backed, synced from workbook source, and served from normalized tables rather than direct file reads.
RefreshManual sync endpoint exists for external reconciliation refresh.
CapabilitiesEvent Fabric + SOR Sidecar
DOC-12

SAP & Multi-SOR Integration Blueprint

Read-only intelligence · Event-driven ingestion · Failure-isolated architecture

bp Sphere integrates with 114 enterprise systems of record across bp's landscape. The integration principle is strict:

Observe → Normalize → Reason
Never mutate SOR truth.

All integrations are read-only by default. Guarded write-back is a controlled future capability requiring explicit governance activation.

    BP Systems of Record (114 SORs across 5 Domains)
    ┌─────────────────────────────────────────────────┐
    │ Finance (44)    │ Operations (32) │ Energy (19) │
    │ SAP · CFIN      │ ServiceNow      │ Trading     │
    │ BlackLine       │ Ariba · GRC     │ Market Data │
    │ Treasury        │ Workday · HR    │ Upstream    │
    ├─────────────────┼─────────────────┼─────────────┤
    │ Analytics (13)  │ Platform (12)               │
    │ Planning Hub    │ bp Sphere · Fabrics         │
    │ Forecasting     │ Governance                  │
    └─────────────────────────────────────────────────┘
            │
            ▼
    Integration Plane (10 Fabrics · Kubernetes)
            │
            ├── API Adapters (REST / OData)
            ├── CDC Listeners (Event Streams)
            ├── Batch Loaders (Scheduled)
            ├── Federated Gateway (GraphQL)
            ├── Saga Fabric (Distributed Tx)
            ├── Provenance Tracker
            │
            ▼
    Canonical Context Snapshots
            │
            ▼
    Agent Runtime (Read-Only)

Integration is physically and logically isolated from the agent fabric.

Current Operational Posture: Read-Only
  • No schema modification to SAP or other SORs
  • No custom BAdIs required
  • No transactional coupling
  • No synchronous dependency for agent execution
  • Decisions operate on versioned snapshots

SOR remains authoritative

System of record

No transactional updates

bp Sphere does not write

SOR downtime tolerated

Snapshots decouple runtime

Zero operational disruption to SAP or Treasury platforms.

Guarded write-back is disabled by default.

Activation requires:

Governance approvalAL-3+ autonomyIdempotency guaranteesRollback mechanismDual approval workflow

Write-back invariants (future):

  • Idempotency key required
  • Evidence pack generated
  • Policy hash pinned
  • Write status reconciled
  • Reversal path defined

No write-back may occur without explicit policy activation.

MethodUse CaseLatencyExample SORs
API (REST/OData)On-demand contextSecondsSAP S/4HANA, Workday, BlackLine
CDC (Change Data Capture)Near-real-time deltaMinutesSAP S/4HANA, Treasury, CFIN
BatchReconciliation & full syncHoursMarket Data, Master Data Hub
Federated GatewayGraphQL federationReal-timeCross-domain queries, fabric interconnection
Saga TransactionsDistributed coordinationSecondsMulti-SOR workflows, compensating actions
CDC Provides
  • Event-driven incremental updates
  • Offset tracking
  • Idempotent replay
  • Catch-up after interruption
Batch Provides
  • Full snapshot refresh
  • Drift reconciliation
  • Integrity re-validation
1. Extract from SOR2. Normalize to canonical3. Version vector update4. Snapshot ID created5. Snapshot frozen6. Injected into execution
Immutable after creationLineage preservedDomain-scopedReplay-bound

No live SOR queries occur during deterministic execution.

Failure Scenariobp Sphere BehaviorRecovery
SOR API timeoutUse last valid snapshot, flag stalenessExponential backoff retry
SOR returns errorSkip source, proceed with available dataAlert + reconciliation
CDC stream interruptionPause ingestionResume from offset
Data validation failureQuarantine recordManual correction + re-ingest
Schema mismatchReject adapterAdmin update required
Adapter circuit breaker tripTemporarily isolate SORHealth probe reset

Fail-closed, not fail-open. Decisions are flagged if data is stale.

Each SOR adapter includes:

Health endpoint polling
Circuit breaker state
Failure threshold counters
Retry backoff logic

Health API:

GET /api/iaf/health/sorsGET /api/iaf/health/sors/{name}GET /api/iaf/health/summary

Navigation dimming reflects stale SOR status in UI. Health degradation does not cascade to agent crash.

Idempotency mechanisms:

Event deduplication

via event_id + timestamp

Version vector enforcement

per-entity currency tracking

Snapshot version comparison

drift detection at ingestion

Replay-safe ingestion

idempotent processing guarantee

Daily reconciliation:

  • Compare SOR counts vs canonical entities
  • Identify drift
  • Trigger batch resync if required

Version vectors track:

SAP:v1042CFIN:v887Treasury:v204

Drift threshold configurable per domain.

Source configuration: sor_registry.yaml (v1.0)

SAP S/4HANASAP S4 CFINSAP AP/ARSAP MDGBank FeedsFP&A HubGRC
ComponentImplementation
SOR Registrysor_registry.yaml
Adaptersservices/sor_adapters.py
Health Monitorstart_sor_health_monitor()
Health APIapi/sor_health.py
HTTP Clientcreate_instrumented_client()
NATS Subscriptionssetup_nats_subscriptions()
Circuit BreakerAdapter state tracking

CFIN Subjects (NATS):

POSTING_CREATEDREPLICATION_COMPLETED

Finance Truth Mode: HYBRID (CFIN + Direct fallback) — configurable via /api/iaf/finance-truth/mode

iaf_customersiaf_vendorsiaf_journal_entriesiaf_journal_linesiaf_customer_invoicesiaf_invoicesiaf_paymentsiaf_cash_positionsiaf_reconciliationsiaf_intercompany
Snapshot-basedVersion-vector taggedDomain-scopedNon-authoritative

Database: iris_agentic_finance · Port: port 15435

Domain InstanceDatabasesMax ConnectionsKey SORs
finance-pg44200R2R, P2P, Treasury, FPA, I2C
operations-pg32120Procurement, HR, HSE, SCM
energy-pg1980Trading, Upstream, Market Data
analytics-pg1380FPA, Planning, Forecasting
platform-pg12120bp Sphere, Fabrics, Governance

Each SOR connects to its domain PostgreSQL instance via dedicated sor-{name}-secrets with per-SOR database roles ({name}_user or sor_{name}_app).

HR OS uses the same SOR discipline as finance: read-only ingestion, canonical normalization, policy-gated actions, and replayable evidence.

HR CapabilityPrimary InputsExecution Contract
Workforce state KPIsWorkday, org structure, payroll/timesheet snapshotsL1-L4 drilldowns with source references
Attrition / capacity riskHeadcount trends, manager spans, overtime indicatorsPolicy-bound recommendations + confidence trace
Hiring pipeline executionRequisitions, open roles, hiring funnel eventsMission plans with role-based approval gates
Workforce simulationScenario inputs + historical baselinesReplayable what-if output with evidence hashes

Credibility rule: HR command surfaces must resolve to linked SOR snapshots and execution evidence, not static narrative content.

The following always hold:

1No SOR schema modification
2No direct transactional coupling
3No live SOR dependency during deterministic execution
4Snapshot immutability preserved
5Circuit breakers isolate SOR failures
6Idempotent ingestion enforced
7Drift detectable and reconcilable
8Write-back disabled unless explicitly governed

Frequently Asked Questions

Q: Does bp Sphere modify SAP tables?
A: No. It reads via OData and CDC only. No SAP schema changes, no custom extensions required.
Q: What happens if a SOR is down?
A: bp Sphere operates on the last valid context snapshot. Decisions include staleness indicators. There is no fail-open mutation behavior.
Q: How is freshness tracked?
A: Every snapshot includes timestamp and version vector. Staleness exceeding threshold triggers UI alerts and optional decision hold via policy.
Q: Is CDC replay-safe?
A: Yes. Offset tracking and event deduplication ensure idempotent processing.

Positioning Statement (For Orals)

bp Sphere operates a read-only, event-driven integration plane that normalizes 114 enterprise systems into versioned canonical snapshots, isolates SOR failures via circuit breakers, and guarantees deterministic reasoning without transactional coupling to SAP or Treasury.
Enterprise SafetyNo Operational RiskArchitectural MaturityIntegration DisciplineDay-1 Readiness
  • Public worker topology is now intentionally split: demo/showcase, dev, and prod can each be mapped to separate workers.
  • Tech pack integration language should distinguish stable service topology from temporary quick-tunnel worker exposure used for public previews.
CapabilitiesTrust Fabric + Replay
DOC-13

Observability, Evaluation & Drift Monitoring

Operational reliability · Decision quality assurance · Controlled evolution

bp Sphere observability operates across three independent but connected control loops:

1. Operational Health

Is the system available and performant?

2. Decision Quality

Are agents producing correct, policy-compliant outcomes?

3. Behavioral Drift

Has agent or data behavior shifted over time?

Observability is not passive logging. It is an active governance control surface.

Agents
153
Implemented
153
Categories
25
Mode
LIVE
MetricDefinitionSLO Target
Execution success rate% executions without error> 99.5%
P50 latencyMedian execution time< 200ms
P99 latency99th percentile< 500ms
Error rate% failed executions< 0.5%
Fail-closed eventsSafety-triggered haltsTrending down
Evidence completenessDecisions with full evidence100%

Metrics tracked per-agent and aggregated per-tower. Failure to meet SLO triggers investigation workflow.

    Request → Correlation ID
            → Execution Telemetry
            → Policy Evaluation Log
            → Evidence Pack
            → Prometheus Metrics
            → Alerting Rule Evaluation
            → Incident Workflow

Every execution is observable at:

API levelAgent runtime levelPolicy evaluation levelEvidence sealing level

No black-box execution path exists.

Drift is monitored across five dimensions:

Drift TypeDetection MethodAction
Output distribution shiftStatistical divergence over rolling windowsAlert + investigation
Confidence degradationRolling mean below thresholdShadow revalidation
Policy override frequencyDeviation from baselineGovernance review
Latency regressionP99 > SLO for > 1hPerformance investigation
Data staleness increaseSnapshot age threshold breachIntegration health check
Measured continuouslyVersion-awareDomain-specificPromotion-blocking if unresolved

Shadow mode enforces safe evolution:

Production Agent (v1)
  • Active
  • Emits decisions
  • Seals evidence
Shadow Agent (v2)
  • Receives identical inputs
  • Produces non-emitted decisions
  • Logs telemetry + evidence

Promotion Criteria:

  • Output agreement ≥ 98%
  • No regression in confidence
  • Latency within SLO
  • No policy violation increase
  • Determinism compliance 100%

Shadow period requirements vary by autonomy level.

No direct cutover allowed without shadow validation.

Evaluation TypeFrequencyPurpose
Online (per execution)Real-timeLatency, success, evidence completeness
Offline batchDailyDrift, output quality, policy coverage
Shadow comparisonDuring promotionVersion agreement analysis
Replay verificationWeekly sampleDeterminism contract validation
Counterfactual testingOn policy changeImpact analysis

All evaluations produce evidence artifacts.

SLOTargetMeasurement WindowBreach Action
Availability99.9%30-day rollingIncident response
P99 Latency< 500ms1-hour windowPerformance investigation
Evidence completeness100%Per executionImmediate investigation
Determinism compliance100%Weekly replay sampleAgent quarantine
Error budget0.1% (30d)CumulativeFeature freeze

Error Budget Governance

If 30-day error budget exceeded:

  • Freeze new deployments
  • Require remediation plan
  • Executive review

This enforces engineering discipline.

ComponentImplementationPurpose
Prometheusservices/iris_metrics.pyCounters + histograms
OpenTelemetryOTelTracer.initialize()Distributed tracing
Agent Monitor APIapi/agent_monitor.pyHealth + execution dashboards
Telemetry ModelAgentExecutionTelemetryPer-execution metrics
HeartbeatsAgentHeartbeat tableLive status strips
SOR Health Monitorstart_sor_health_monitor()Upstream data freshness
Alert Rulesconfig/alerting_rules.yamlSLO breach triggers
Drift Rulesconfig/drift_rules.yamlBehavioral detection
SSE StreamGET /streamReal-time dashboard feed
Correlation IDsCorrelationMiddlewareFull traceability
QueryableRole-restrictedAudit-traceable
CFO Command Center
  • Live agent health strips
  • Autonomy distribution
  • SLO compliance summary
  • Override rate
  • Drift flags
  • Fail-closed trend
Tower Leaders
  • Domain-specific drift
  • Policy override heatmap
  • Latency distribution
Admins
  • Kill switch events
  • Deployment validation
  • Replay compliance
Bug
  • Known defect
  • Deterministic failure
  • Immediate fix required
Drift
  • Gradual behavior shift
  • May arise from data distribution change
  • Requires statistical investigation

Drift detection prevents silent degradation.

The following must always hold:

1Every execution produces telemetry
2Every execution produces evidence
3Every policy evaluation logged
4Drift rules evaluated continuously
5SLO breaches generate alerts
6Determinism sampled weekly
7Error budget governs release velocity

If observability fails → system considered degraded.

Frequently Asked Questions

Q: What happens when an SLO is breached?
A: Automated alert triggers investigation. Sustained breaches may freeze deployments until resolved.
Q: How is drift different from a bug?
A: Drift is statistical behavior shift; bugs are deterministic defects. Drift detection prevents gradual degradation.
Q: Can agents be promoted without shadow validation?
A: No. Promotion requires output agreement, confidence stability, and SLO compliance.
Q: Is determinism continuously verified?
A: Yes. Weekly replay sampling ensures 100% compliance.

Positioning Statement (For Orals)

bp Sphere operates under a measurable reliability and quality framework where execution health, decision integrity, and behavioral drift are continuously monitored, promotion is gated by shadow validation, and error budgets enforce disciplined evolution.
Production-Grade MaturityFinancial Control DisciplineSafe AI EvolutionNo Silent DegradationEnterprise Operational Readiness
  • Observability now includes runtime profile snapshots and environment verification endpoints, not only generic telemetry claims.
  • Recent platform validation work also refreshed live cluster-state reporting to avoid stale audit summaries.
CapabilitiesPolicy + Authority Control
DOC-14

Security Architecture & Threat Model

Zero trust enforcement · Defense in depth · Blast radius containment

bp Sphere operates under a strict zero-trust architecture. Every request is authenticated, authorized, domain-scoped, logged, and audited.

No implicit trust exists between:

UsersServicesAgentsDomainsTenants

Security is enforced at multiple independent layers.

Verify Explicitly

Every API call requires a valid Entra ID JWT token. No route (except explicitly public paths) is accessible without authentication.

Enforce Least Privilege

Access limited by:

  • RBAC role
  • Domain scope
  • Materiality threshold
  • Autonomy level
Assume Breach

Controls include:

  • Immutable evidence vault
  • Kill switch
  • Circuit breakers
  • Replay verification
  • Domain isolation

Compromise of one component does not cascade.

    External User
            │
            ▼
    Azure API Management (Rate limiting, IP filter)
            │
            ▼
    AuthMiddleware (JWT Validation)
            │
            ▼
    RBAC + Domain Engine
            │
            ▼
    Supervisory Control Plane
            │
            ▼
    Agent Runtime

Internal services communicate via:

mTLSHMAC traversal signingNamespace isolation

There is no direct bypass path.

LayerMechanismImplementation
IdentityMicrosoft Entra IDOAuth2 + OIDC (BP tenant SSO)
API AuthJWT (HMAC-SHA256)AuthMiddleware
RBACbp Sphere RBAC engineRole-to-permission mapping
Domain ScopingDomain engineEntity-level restriction
Traversal SigningHMAC-SHA256TRAVERSAL_SIGNING_SECRET (K8s secret)
Service-to-ServicemTLSCertificate-based
SecretsAzure Key Vault + K8s SecretsNo secrets in code
GatewayAzure API Management200 req/min rate limit

Missing JWT or signing secret → fail-closed.

ScopeStandardImplementation
Data at RestAES-256Azure-managed disk encryption
Data in TransitTLS 1.3Certificate auto-renewal
Evidence VaultAES-256 + SHA-256 hashIntegrity verification on read
SecretsK8s Secrets + Key VaultAuto-rotation enabled

Startup fails if:

  • JWT_SECRET_KEY missing
  • TRAVERSAL_SIGNING_SECRET missing
  • AUTH_ENABLED misconfigured

No secret hardcoded in source.

LLM usage is bounded and isolated. Controls include:

Token limits enforced
No external network calls
Prompt injection sanitization
Output schema validation
Policy enforcement post-generation
Prompt + response logged in LLMAuditRecord

Even a manipulated prompt cannot bypass:

Policy engineSupervisory control planeDomain scoping

LLM output is advisory input, not direct execution authority.

Tenant isolation enforced by:

TenantSeparationMiddleware

Rejects any request where:

x-iris-tenant != active-tenant

Configuration: tenant_separation_strict=True

Cross-tenant access impossible without redeployment.

LayerProtection
API GatewayRate limiting, IP filtering
AuthMiddlewareToken validation
RBAC EnginePermission enforcement
Domain EngineData scoping
Policy EngineDecision gating
Supervisory PlaneAutonomy + kill switch
Evidence VaultImmutable audit trail
Circuit BreakerIsolation of faulty components

Multiple independent controls must fail for compromise.

ThreatRiskMitigation
Malicious agent registrationUnauthorized decision logicAdmin-only registry + shadow validation
Prompt injectionManipulated outputInput sanitization + policy enforcement
Unauthorized write-backData corruptionWrite-back disabled by default
Evidence tamperingAudit compromiseHash verification on read
Privilege escalationUnauthorized data accessRBAC + no self-assignment
Service compromiseLateral movementmTLS + namespace isolation
Denial of ServicePlatform outageRate limiting + circuit breakers
Data exfiltrationConfidential leakageDomain scoping + export logging

All denial events recorded in AccessDenialLog.

StandardCoverageKey Controls
ISO 27001Access, encryption, loggingA.9, A.12
SOC 2 Type IIConfidentiality, availabilityCC6, CC7
NIST 800-53Risk managementAC, AU, SC families
SOXFinancial controlsSOD, dual approval, 7-year retention

Compliance evidence retrievable via evidence export API.

ControlImplementationEnforcement
Tenant IsolationTenantSeparationMiddlewareStrict reject
Correlation IDsCorrelationMiddlewareUUID4 propagation
Rate LimitingAuthMiddleware200 req/min
Error SanitizationExceptionHandlerMiddlewareNo stack traces
Traversal SigningHMAC-SHA256Secret required
SecretsK8s + Key VaultNo plaintext storage
Public RoutesLimited to health + docsAll others JWT-protected
Access LoggingAccessDenialLogImmutable record
Permission ChangesPermissionChangeLogAudit trail

28 governance models track control events.

The following always hold:

1No unauthenticated request reaches business logic
2No agent executes without policy gating
3No cross-tenant routing permitted
4No SOR write-back enabled by default
5No evidence modification possible post-sealing
6No self-assignment of elevated roles
7No silent permission change
8Control plane failure → fail-closed

Frequently Asked Questions

Q: How does bp Sphere handle prompt injection?
A: Input is sanitized prior to agent execution. Output is validated against policy rules. Even malicious content cannot bypass governance enforcement.
Q: Can a user elevate their own role?
A: No. Role assignment is controlled via Entra ID and RBAC configuration. Self-assignment is blocked and logged.
Q: What if an internal service is compromised?
A: mTLS, traversal signing, and namespace isolation limit blast radius. Supervisory control plane can trigger kill switch if anomaly detected.
Q: Does a SOR breach compromise bp Sphere?
A: No. Context snapshots are read-only and isolated. No write-back capability exists in current deployment.

Positioning Statement (For Orals)

bp Sphere operates a defense-in-depth, zero-trust architecture where identity, domain scope, policy enforcement, supervisory control, and immutable evidence collectively ensure that no unauthorized decision, data access, or privilege escalation can occur — even under breach assumptions.
CISO-Grade RigorRegulatory AlignmentControl-Plane MaturityNo Uncontrolled AI RiskEnterprise-Safe Deployment
  • Security/configuration cleanup has reduced model selection sprawl: LLM choice is now driven from one config setting rather than multiple defaults.
  • Environment-specific behavior should now be described as profile-driven and policy-governed, not as ad hoc deployment flags.
CapabilitiesEvent Fabric + SOR Sidecar
DOC-15

Enterprise Infrastructure Architecture

Distributed control plane · Deterministic data plane · Event-driven orchestration · Production-grade DevSecOps

bp Sphere is deployed as a distributed, policy-governed, containerized control plane running on Kubernetes with:

  • Event-driven agent fabric
  • Deterministic replay ledger
  • Tenant-isolated data plane
  • Enterprise observability spine
  • Multi-layer CI/CD quality gates

This document specifies the full infrastructure substrate — from VPC networking to rollout governance.

Distributed Control PlaneDeterministic Data PlaneEvent-Driven OrchestrationProduction DevSecOps
USER EXPERIENCE
  CFO . FP&A . Tower Leaders . 50+ UI Surfaces
------------------------------------------------------
CONVERSATIONAL GATEWAY
  Runtime-configured LLM . Controlled Interface
------------------------------------------------------
bp Sphere API GATEWAY
  FastAPI . 300+ routes . 6-layer fail-closed middleware
------------------------------------------------------
AGENT FABRIC
  153 Agents . Supervisor . HITL . Tick Engine
------------------------------------------------------
EVENT BACKBONE
  NATS JetStream . CDC bridge . Async triggers
------------------------------------------------------
GOVERNANCE CORE
  Policy Engine . RBAC/PBAC . Autonomy controls
------------------------------------------------------
DATA PLANE
  PostgreSQL Primary + Replica . Evidence Vault
------------------------------------------------------
INTEGRATION LAYER
  114 SOR Adapters . CDC . REST . Batch
------------------------------------------------------
KUBERNETES ORCHESTRATION
  Deployments . StatefulSets . HPA . Ingress
------------------------------------------------------
INFRASTRUCTURE
  VPC . Subnets . TLS . Persistent Storage

This is a control plane, not an application server.

ComponentSpecificationPurpose
ClusterMulti-node Kubernetes (control + worker nodes)Scheduling, orchestration
NetworkingPrivate VPC with segmented subnetsIsolation between SOR, platform, ingress
Load BalancerInternal L4/L7 with TLS terminationSecure ingress routing
StoragePersistent Volume Claims (RWO)Postgres, evidence, WAL
Namespaceplatform (prod) · platform-demo (demo)Environment isolation
SecretsK8s Secrets + cluster-secret-storeCredentials, signing keys
DNS*.platform.svc.cluster.localInternal service discovery
Infrastructure Guarantees
  • Node-level HA
  • Pod-level restart recovery
  • Encrypted storage (AES-256)
  • Strict namespace RBAC boundary
  • No cross-tenant routing

Core Resources

ResourceNamePurpose
DeploymentirisAPI + Agent runtime
StatefulSetpostgres-primaryPrimary data plane
StatefulSetnats-jetstreamEvent backbone
Deployments (platform-wide)~163 pods (148 SOR + fabrics + platform services)Application workloads (shared across tenants)
StatefulSets5 domain PostgreSQL instancesPersistent data plane
ClusterIP Services (SOR fabric)148 SOR servicesAdapter isolation (bp tenant integrates 11+)
bp Tenant bp Sphere Deployments8 (prod / worker / dev / dev-worker / showcase / showcase-worker / compact-smoke / apex)bp tenant workload

Health Probes

ProbeEndpointAction
Readiness/readyRemove from service if failing
Liveness/healthRestart pod if failing
Startup/ready150s max initialization window

Pod Security Hardening

runAsNonRoot: true
drop: ALL capabilities
allowPrivilegeEscalation: false
fsGroup: 1000

Zero privileged containers.

bp Sphere operates 10 specialized fabrics that coordinate cross-domain workflows, governance enforcement, and distributed intelligence.

FabricPurposeKey Capabilities
ai_meshAgent mesh orchestrationTopology routing, capability matching, role-based dispatch
control_planePlatform supervisory controlKill switches, drift detection, lifecycle management
data_fabricCross-domain data accessCanonical entity resolution, snapshot management
federated_gatewayGraphQL federationCross-SOR queries, schema stitching
guardian_fabricGovernance enforcementPolicy evaluation, autonomy gating, SOD rules
integration_fabricSOR integrationCDC listeners, API adapters, batch loaders
payment_gateway_fabricPayment processingTransaction routing, fraud detection
saga_fabricDistributed transactionsCompensating actions, multi-step coordination
trust_fabricIdentity and trustAuthentication, authorization, tenant isolation
workflow_fabricProcess orchestrationMulti-step workflows, HITL escalation, task queue
  Agent Mesh ↔→ Guardian (policy check)
       │
       ▼
  Workflow Fabric (orchestration)
       │
       ├── Integration Fabric (SOR access)
       ├── Data Fabric (context)
       ├── Payment Gateway (if financial)
       └── Saga Fabric (if distributed)
       │
       ▼
  Trust Fabric (auth) → Control Plane (supervision)

Each fabric runs as an independent Kubernetes deployment with its own database connection via platform-pg.

PostgreSQL Primary
  • 130+ SQLModel entities
  • 28 governance models
  • Evidence vault (5 tables)
  • 10 live SOR canonical tables
  • Immutable decision ledger
Replica Layer
  • Read scaling
  • Dashboard isolation
  • Simulation state isolation
  • Streaming replication

No writes accepted on replicas.

Data Integrity Controls

  • Row-level governance enforcement
  • SHA-256 hash sealing (inputs + outputs + evidence)
  • Deterministic replay validation
  • Tenant-scoped realm_id injection
  • Append-only evidence model

Event Subjects

agent.run.request
agent.run.completed
evidence.sealed
policy.violation
kpi.threshold.crossed
cfin.POSTING_CREATED
cfin.REPLICATION_COMPLETED

Capabilities

CapabilityDescription
Asynchronous orchestrationFire-and-forget agent invocation via NATS subjects
CDC propagationSAP S/4 HANA change events bridged into agent fabric
Event replayJetStream persistence for debugging and audit replay
Backpressure handlingConsumer-side flow control prevents overload
Decoupled executionSOR events propagate without direct coupling to agent runtime

JetStream persistence ensures event durability during pod restart.

Stage 1 — Builder (uv)
  • Deterministic dependency resolution
  • Layer caching
  • Shared wheel installation
Stage 2 — Runtime (Slim)
  • Python 3.13.7-slim
  • Non-root user (UID 1000)
  • Uvicorn + uvloop
  • Health check every 30s
  • LIVE_ONLY_BUILD removes demo code paths
PropertyValue
ASGI ServerUvicorn + UVLoop + HTTPTools
Workers2 (configurable)
Package Manageruv (deterministic lockfile)
Health CheckGET /health (30s interval)

Minimal attack surface.

bp Sphere Pods --> Prometheus --> Grafana
           --> Loki (logs)
           --> Jaeger (traces)

Monitoring Stack

ComponentPurpose
PrometheusMetrics collection (15s scrape)
GrafanaSLO dashboards
LokiStructured log storage
JaegerDistributed tracing
OpenTelemetryTrace propagation
SSEReal-time dashboard streaming

SLO Targets

MetricTarget
API P95< 500ms
API P99< 2000ms
Availability> 99.9%
Cross-tenant violations0
Data freshness< 5 min

Inbound request path:

1. CorrelationMiddleware
2. TenantSeparationMiddleware
3. CORSMiddleware
4. AuthMiddleware
5. GovernanceMiddleware
6. ExceptionHandlerMiddleware
Application logic executes only after:
  • Tenant verified
  • JWT validated
  • Permission mapped
  • Policy enforced

Fail-closed always.

MechanismEnforcement
Request isolationHTTP 403 on tenant mismatch
Namespace boundaryplatform vs platform-demo
Database isolationSeparate DB per env
Schema scopingrealm_id on all tables
Secrets isolationPer-deployment K8s secret
Role-scoped data modeLIVE_ONLY for CFO + tech

No shared data plane across tenants.

Layer 1 — Unit Tests
  • Coverage ≥ 70%
  • Required for merge
Layer 2 — Critical Guards
  • Demo/Live separation
  • Agent registry validation
  • Evidence hash validation
  • Secret scan

Blocking on failure.

Layer 3 — Smoke Tests
  • Live Postgres container
  • Startup validation
  • API health verification
Layer 4 — Nightly Scenario Tests
  • Full orchestration simulation
  • Drift validation
  • Non-blocking informational
Image Pipeline

docker buildimportrollout restart (zero downtime)

Readiness gate required for main branch.

K8s Cluster (platform namespace)

bp Sphere Pod
  |-- API + Agent Fabric
  |-- 6-layer middleware
  |-- 153 agents
  |-- 114 SOR adapters
  |
  |-- PostgreSQL :15435
  |-- NATS JetStream :4222
  |-- SOR Services (ClusterIP)
  +-- Observability Stack

Modular monolith architecture for operational simplicity.

Event bus enables future decomposition.

ScenarioBehavior
Pod crashAuto-restart via K8s
SOR outageSnapshot fallback + staleness flag
NATS disruptionJetStream replay
DB primary failureReplica failover
Control plane errorFail-closed (no decision emitted)

No uncontrolled execution under failure.

Enterprise Guarantees
  1. No cross-tenant data leakage
  2. No uncontrolled write-back
  3. Deterministic replay enforced
  4. Event durability across restart
  5. Zero-downtime rolling deployments
  6. Immutable evidence ledger
  7. Kill switch for blast radius containment
  8. CI/CD enforces governance integrity

Frequently Asked Questions

Q: Is bp Sphere a monolith or microservices?
A: A modular monolith for operational simplicity, with internal boundaries and externalized SOR services. Event backbone enables future service decomposition.
Q: What happens during pod failure?
A: Kubernetes restarts pod. NATS persists events. No decisions emitted during downtime (fail-closed).
Q: Can bp Sphere scale horizontally?
A: Yes. Stateless API replicas + HPA. Postgres read replicas scale dashboards. NATS distributes event consumers.

Instead of saying: "We run on Kubernetes."

bp Sphere operates as a distributed, tenant-isolated control plane deployed on Kubernetes with deterministic replay guarantees, event-driven orchestration, immutable audit ledger, zero-trust middleware, and multi-layer CI/CD governance — engineered for enterprise-scale financial operations.

Enterprise MaturityOperational ResilienceGovernance RigorInfrastructure DisciplineDay-1 Production Readiness
EnvironmentCurrent role
DevActive engineering target with latest feature validation.
Showcase/DemoSeparate deployment with demo role forcing and shorter menu shell.
ProdFull-shell runtime with distinct public mapping and no demo-role forcing.
CapabilitiesSkill Fabric + Agentic RuntimeValue Attribution Engine
DOC-16

BP Implementation & Migration Plan

Progressive confidence model · Explicit risk gates · Governance-first rollout

bp Sphere deployment at bp follows a structured, three-phase rollout designed to:

  • De-risk integration
  • Prove decision quality
  • Establish governance discipline
  • Scale only after validation
  • Promote autonomy only after evidence

No phase proceeds without satisfying exit criteria and formal checkpoint approval.

Progressive ConfidenceExplicit Risk GatesGovernance-First RolloutNo Big-Bang Risk
Day 0      30         60         90        120        150        180
|-- Phase 1 -|--------- Phase 2 ---------|---- Phase 3 ---------|
| Foundation |        Validation          |       Expansion       |
|            |                            |                       |
Security -> Integration -> Shadow -> Governance -> Scale -> Autonomy
Gate        Gate         Gate      Gate         Gate      Gate

Each gate is a formal go/no-go decision.

Phase 1 — Foundation (Day 0–30)SETUP

Objective: Secure, connected, operational baseline

Setup Deliverables
  • Azure environment provisioning (UK/EU compliant region)
  • Kubernetes namespace isolation (platform)
  • Microsoft Entra ID integration + RBAC role mapping
  • SAP S/4HANA OData API connectivity
  • Initial SOR health validation
  • Agent registry populated (153 agents)
  • Shadow pilot with 3 low-risk agents
  • Evidence vault operational
  • Observability dashboards live
Exit Criteria
  • Security review passed (CISO sign-off)
  • 3 agents executing successfully in shadow mode
  • 100% evidence sealing verified
  • Monitoring + alerting operational
  • No cross-tenant violations
Operating Mode: WORKSHOP mode for pilot agents. No LIVE AL-2 execution yet.

Phase 1 proves infrastructure + governance readiness.

Phase 2 — Validation (Day 30–90)PARALLEL RUN

Objective: Prove decision quality under parallel run

Deliverables
  • Policy configuration across active finance domains
  • Shadow parallel run (agents vs manual process)
  • AL-2 deployment for validated agents
  • CDC integration for near-real-time data
  • Drift detection baseline established
  • Error budget monitoring active
  • Replay validation sampling weekly
Exit Criteria
  • Shadow agreement ≥ 95% vs manual decisions
  • Zero critical incidents during 30-day parallel run
  • Policy coverage ≥ 90% of decision paths
  • Determinism compliance = 100%
  • SLOs met for 30 consecutive days
  • CFO + CTO shadow validation sign-off
Operating Model: AL-1 (recommend) + AL-2 (prepare with approval). No AL-3 allowed during validation.

Phase 2 proves trustworthiness.

Phase 3 — Expansion (Day 90–180)SCALE

Objective: Controlled scale + governance maturity

Deliverables
  • Full finance tower onboarding (R2R, P2P, I2C, FP&A, Treasury)
  • AL-3 promotion for qualified agents
  • Guarded write-back feasibility assessment
  • Formal governance documentation (SOX-ready)
  • Counterfactual simulation library established
  • Drift monitoring automated per domain
Exit Criteria
  • All finance towers live in AL-1/2
  • Minimum 5 agents promoted to AL-3
  • First internal SOX audit completed
  • CFO formal sign-off on governance model
  • Error budget maintained within 0.1% for 30 days
Operating Model: Selected agents promoted to AL-3 (within thresholds). Shadow validation mandatory before promotion. Write-back remains disabled unless explicitly approved.

Phase 3 proves scalability + maturity.

CheckpointTimingDecision AuthorityOutcome Options
Security GateDay 14CISOProceed / Remediate / Abort
Integration GateDay 30CTOProceed / Remediate / Pause
Shadow GateDay 60CFO + CTOPromote to AL-2 or extend shadow
Governance GateDay 90CFO + Internal AuditScale approval
Scale GateDay 120CFOExpand domains
Autonomy GateDay 150CFO + CISOPromote selected agents to AL-3

No checkpoint failure results in uncontrolled progression. Remediation is mandatory before advancement.

Authority Model:

CFO

Business readiness authority

CISO

Security veto authority

CTO

Technical readiness authority

Internal Audit

Governance certification authority

Autonomy promotion requires CFO + CISO concurrence.

On application startup (on_startup()):

StepFunctionRisk Control
1Set LIVE modePrevent demo leakage
2Initialize databaseFail if DB unavailable
3Register Cowork facadeDomain enforcement
4Create HTTP clientControlled connection pooling
5Validate dependenciesHard-fail if SOR unavailable
6Subscribe to NATSEvent durability
7Create mode contextDeterministic mode injection
8Initialize supervisory servicesGovernance online
9Validate agent registryMetadata integrity check
10Register fleetSupervisory control binding
11Initialize OpenTelemetryTraceability
12Start SOR health monitorFreshness enforcement
13Seed close calendarDeterministic baseline
14Seed policiesVersion-pinned
15Seed deploymentsTenant-defined agents
16Seed fleet dataBusiness context
17Seed close simulationParallel validation
18Seed heartbeatsDashboard accuracy
19Seed demo dataIdempotent
20Optional HR seedControlled flag
21Optional agent runnerExplicit enable flag
22Evidence backfill (mandatory in LIVE)Integrity guarantee

Failure at any required step → startup aborts.

Migration Safety Principles
  1. No “big bang” cutover
  2. Parallel shadow required before scale
  3. Autonomy promotion gated by evidence
  4. Write-back disabled unless formally approved
  5. Error budget governs deployment velocity
  6. Replay verification required before AL-3
  7. Governance documentation completed before SOX reliance

If checkpoint fails:

OptionAction
RemediateFix defect, re-test, re-evaluate
PauseMaintain current phase, extend validation
RollbackRevert to previous stable version (Helm rollback)
AbortDecommission pilot agents, retain evidence

All rollback events are logged and sealed in evidence vault.

Frequently Asked Questions

Q: Can phases overlap?
A: No. Exit criteria must be met before progression. This prevents premature scaling of unvalidated capabilities.
Q: What if a risk checkpoint fails?
A: Remediate, pause, or rollback. No automatic forward progression.
Q: Who owns go/no-go decisions?
A: The CFO owns business readiness. The CISO holds veto authority on security gates. Joint approval required for autonomy promotion.
Q: Can AL-3 promotion happen before Phase 3?
A: No. Shadow validation + policy coverage + CFO sign-off required.

Instead of saying: "We have a 3-phase rollout."

bp Sphere follows a progressive confidence deployment model where integration, decision quality, governance maturity, and autonomy promotion are gated by explicit exit criteria, formal checkpoints, and executive sign-off — eliminating big-bang risk and ensuring audit-grade readiness at every stage.

DisciplineExecutive ControlRisk ContainmentOperational MaturityDay-1 Credibility
Roadmap itemCurrent state
Critical governanceC1-C4 closed: execution authority, auth posture, GuardrailPolicyEngine fail-closed, LLM seed + confidence gate.
High-tier runtimeH1-H5 closed: learning loop, autonomy telemetry/enforcement, value attestation, asymmetric evidence, freshness SLA.
Enterprise runtime additionsWIIF, governed chat, token economy, runtime adoption audit, skill-binding learning, case-state propagation, and observability readiness are implemented.
External dependenciesSAP/SOR phases 2-3, live Entra/vendor connectors, production OTEL/SLO binding, and commercial price cards remain deployment-bound.
  • The implementation plan should now reflect that runtime profile migration, BP external truth integration, Finance AI model centralization, and FP&A V2 mission surfaces are already implemented.
  • The next phase is environment assurance and product hardening, not first-build architecture.
CapabilitiesPolicy + Authority ControlTrust Fabric + ReplayEvent Fabric + SOR Sidecar
Implemented Runtime Chapter

Runtime 1: Enterprise Policy Intelligence Runtime

The Enterprise Policy Intelligence Runtime is the policy decision point for Sphere. It converts enterprise rules, controls, approval thresholds, AI guardrails, data-sharing restrictions, model-routing constraints, authority limits, and evidence requirements into executable decisions that missions, agents, APIs, workflows, and user actions must respect.

Policy Decision Point Authority Gate Evidence Sufficiency AI Guardrails Execution Control Replayable Ledger

Production Role

The runtime prevents policy logic from fragmenting across UI components, API handlers, workflow scripts, and agent prompts. It centralizes policy evaluation so every important recommendation, action, data-sharing event, and write-back request is constrained by versioned, testable, observable, and auditable policy.

It does not replace source systems, identity providers, the evidence fabric, or the decision orchestrator. It governs those layers by returning structured outcomes such as allow, deny, require_human_approval, require_more_evidence, allow_recommend_only, allow_draft_only, and allow_simulation_only.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_policy_intelligence_runtime.py and policy_as_code_runtime.py Certified by tests/test_enterprise_policy_intelligence_certification.py and policy single-source contract tests.
Policy API contract GET /api/policy, POST /api/policy/evaluate, POST /api/policy/simulate, decision trace, override, compile, test, observability, and rollback contracts. Published from api/policy_as_code.py and surfaced in Policy Registry API documentation.
Decision ledger Every material evaluation produces a policy decision identifier, reason codes, policy trace, evidence status, and ledger pointer. Decision Runtime persists policy attestations and replay pointers for assurance packets.
Authority boundary Agents and users are constrained to read, recommend, draft, approval-required, or execution-authorized modes. Execution Gateway blocks business execution when policy proof is absent.
Evidence sufficiency Policy evaluation checks required evidence before a recommendation or action can proceed. Evidence Runtime and Enterprise Assurance Runtime consume policy decisions as current proof objects.
AI guardrails Model use, external sharing, masking, tool access, human review, memory creation, and write-back are governed as policy decisions. Enterprise agent runtime declares enterprise_policy_intelligence_runtime as a mandatory dependency.
Observability Policy volume, latency, approval requirements, denials, overrides, missing evidence, and hot policies are tracked. Operational surfaces: Policy Runtime, Policy Library, and Policy Studio.

Policy Domains Governed

  • Business policy. Process rules such as duplicate payment holds, non-PO thresholds, receipt requirements, collections handling, close sign-off, and treasury payment controls.
  • Control policy. Authority limits, segregation of duties, control-owner approvals, SOX evidence requirements, and exception governance.
  • AI policy. Approved model classes, deterministic-first requirements, prompt context limits, evidence citation requirements, human review, and tool-call permission.
  • Data policy. Classification, masking, retention, region restrictions, source authority, data freshness, and external sharing boundaries.
  • Execution policy. Whether an action is blocked, recommend-only, draft-only, approval-required, reversible, or executable.
  • Learning policy. Whether an outcome, override, or judgment can become enterprise memory or training signal.

Logical Architecture

User / Agent / Mission / API / Workflow
        |
        v
Policy Evaluation Request
        |
        v
Enterprise Policy Intelligence Runtime
        |
        +-- Policy Registry
        +-- Policy Compiler
        +-- Policy Evaluation Engine
        +-- Authority Engine
        +-- Evidence Sufficiency Engine
        +-- AI Guardrail Engine
        +-- Data Sharing Policy Engine
        +-- Execution Gate Engine
        +-- Override and Exception Engine
        +-- Policy Decision Ledger
        +-- Policy Simulation Engine
        +-- Policy Observability
        |
        v
Policy Decision Response

Runtime Evaluation Flow

  1. Normalize request. Mission, actor, role, object, requested action, source systems, risk, evidence, and data classification are converted into a canonical policy context.
  2. Resolve policies. The runtime selects active policies by mission, object type, action, control mapping, effective date, and risk class.
  3. Evaluate evidence. Required evidence is checked before action or recommendation. Missing proof returns require_more_evidence.
  4. Evaluate authority. User and agent authority are checked against role, approval threshold, segregation-of-duties, and delegation state.
  5. Apply AI guardrails. Model routing, tool access, context sharing, masking, and human review rules are enforced before generated output is trusted.
  6. Gate execution. Business actions and write-back requests are allowed, denied, approval-gated, draft-only, or simulation-only.
  7. Ledger decision. The decision, reason codes, input hash, output hash, policy versions, evidence references, and replay pointer are persisted.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/policyDiscover active policy registry entries.Read-only policy catalog.
POST /api/policy/evaluateEvaluate a user, agent, workflow, or API action against policy.Authoritative policy decision point.
POST /api/policy/simulateRun non-mutating impact analysis for threshold, authority, or evidence-rule changes.Simulation only; no production action.
GET /api/policy/decisions/{decision_id}Replay a policy decision with input facts, versions, reason codes, and output hash.Audit and assurance trace.
POST /api/policy/decisions/{decision_id}/overrideRequest or record controlled exception handling.Human-approved, expiry-bound override path.
GET /api/policy/observabilityExpose latency, decision mix, denial rate, override rate, and missing-evidence metrics.Operational readiness and control monitoring.

Integration With Other Runtimes

Decision Runtime

Uses policy outcomes to determine whether a decision can be recommended, escalated, blocked, or executed.

Evidence Runtime

Assembles evidence; policy determines whether that evidence is sufficient for the requested action.

Agent and Skill Runtime

Agents declare policy dependencies and must receive policy clearance before material tool use or business action.

Replay and Assurance Runtime

Reconstructs the exact policy version, context, evidence references, actor, and outcome used at decision time.

Operational Requirements

  • Policy decisions must be deterministic for the same input context and policy package version.
  • Policy evaluation must remain low latency enough for interactive workflows; high-risk failures default to deny or human review.
  • Policy packages require syntax validation, metadata validation, conflict detection, tests, approval, deployment, and rollback.
  • Policy observability must expose decision volume, latency, failures, denies, approval requirements, overrides, and policy hot spots.
  • Sensitive fields should be masked or referenced by identifier in policy payloads unless required for evaluation.

Acceptance Criteria

  • Every material agent recommendation carries a policy decision identifier.
  • High-risk write-back cannot proceed without an allow decision or required human approval.
  • Policy failures are fail-closed for financial actions and visibly degraded for low-risk read-only guidance.
  • Policy versions are visible in replay, assurance packets, and policy registry surfaces.
  • Overrides require approver, justification, expiry, evidence reference, and audit trail.
  • Policy logic is not embedded as the source of truth in prompts or UI-only button states.

Example: Duplicate Payment Risk

A duplicate-payment risk case requests a hold recommendation and a possible payment release. The runtime evaluates duplicate payment controls, authority threshold, evidence sufficiency, AI output requirements, and write-back policy. The resulting policy decision allows the hold recommendation, blocks payment release, requires supervisor approval for high-value action, and writes a replayable ledger record with policy versions and evidence references.

{
  "decision": "require_human_approval",
  "recommendation_allowed": true,
  "action_allowed": false,
  "write_back_allowed": false,
  "required_next_step": "supervisor_review",
  "reason_codes": [
    "duplicate_payment_risk_high",
    "amount_above_threshold",
    "write_back_requires_approval"
  ],
  "evidence_status": "sufficient",
  "ledger_pointer": "ledger://policy/{policy_decision_id}"
}

Engineering Rule

Prompts may reference policy summaries, and user interfaces may hide or disable blocked actions, but neither prompts nor UI state are the source of enforcement. Backend action APIs and execution gateways must validate the current policy decision before material action or write-back.

CapabilitiesOntology + Enterprise ContextPolicy + Authority ControlTrust Fabric + ReplayValue Attribution Engine
Implemented Runtime Chapter

Runtime 2: Enterprise Decision Runtime

The Enterprise Decision Runtime is the core decisioning engine for Sphere. It converts signals, context, evidence, policy outcomes, agent recommendations, skill results, risk, value, and human authority into governed, explainable, auditable business decisions.

Decision Orchestration Next Best Action Policy-Constrained Evidence-Backed Human Approval Boundary Replayable Ledger

Production Role

The Policy Runtime determines what is allowed. The Decision Runtime determines what should happen next within those constraints. It evaluates context, evidence, materiality, risk, value, authority, and agent output to produce a structured decision outcome and an operational next step.

It is not a source-system writer, raw workflow engine, or isolated agent. It is the governed orchestration layer that turns source-linked context and runtime intelligence into a decision record that can be explained, approved, executed through controlled channels, replayed, and measured.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_decision_runtime.py Defines runtime contract, layers, execution, live execution, replay, certification, audit package, policy simulation, and domain-pack contracts.
API surface /api/iaf/enterprise-decision-runtime/* Includes contract, execute, execute-live, replay, replay-live, approval transition, registry, layers, certification, audit, and decision trace endpoints.
Live execution path execute_decision_live() Persists DecisionLogRecord, EvidencePack, ReplayManifest, ReplayRun, ReplayStep, DecisionContractRecord, ToolInvocation, and RuntimePerformanceTelemetry.
Policy integration Calls EnterprisePolicyIntelligenceRuntime.evaluate() before finalizing execution state. Live records include policy_decision_id, policy attestation hash, winning policy, and replay-visible policy boundary.
Skill integration Routes selected skill execution through SkillFabricService.execute_skill() after policy and workflow planning. Tool invocations preserve skill ID, canonical skill ID, retries, SOR calls, token count, and evidence hash.
Human approval boundary Creates DecisionContractRecord with approval status, required role, selected action, confidence, policies applied, and replay reference. Approval transitions are available through POST /api/iaf/enterprise-decision-runtime/approvals/{{decision_id}}/transition.
Replay and assurance Seals decision-time replay snapshots and boundary summaries with integrity hashes. Live replay is available through /replay-live/{{decision_id}} and compare through /replay-live/{{decision_id}}/compare.

Runtime Scope

  • Decision intake. Accepts requests from events, APIs, agents, mission workflows, and user actions.
  • Context assembly. Resolves object, mission, source state, role, prior context, evidence references, and business intent.
  • Strategy selection. Builds a deterministic workflow plan and an LLM planner contract while keeping regulated execution deterministic.
  • Policy constraint. Consumes the Policy Runtime decision and never promotes a blocked action into execution.
  • Evidence evaluation. Checks whether evidence is complete enough for recommendation, approval, or execution.
  • Next best action. Converts recommendation, risk, authority, policy, and evidence state into the next operational step.
  • Ledger and replay. Persists the decision, integrity hash, replay manifest, evidence pack, tool calls, approval boundary, and telemetry.
  • Learning signal. Records outcome candidates and runtime telemetry for learning and improvement loops subject to policy.

Logical Architecture

Event / User / Agent / Workflow / API
        |
        v
Decision Request
        |
        v
Enterprise Decision Runtime
        |
        +-- Decision Context Builder
        +-- Decision Type Classifier
        +-- Decision Strategy Selector
        +-- Risk and Materiality Engine
        +-- Value Impact Engine
        +-- Recommendation Evaluator
        +-- Policy Constraint Adapter
        +-- Evidence Evaluation Adapter
        +-- Human Authority Adapter
        +-- Next Best Action Engine
        +-- Decision Explanation Engine
        +-- Decision Ledger
        +-- Decision Replay Adapter
        +-- Learning Signal Publisher
        |
        v
Decision Outcome

Implemented Lifecycle

  1. Receive request. A decision request arrives with actor, object, mission, requested action, amount, context, evidence, and event references.
  2. Resolve context. The runtime builds a normalized context snapshot and source-linked decision envelope.
  3. Generate workflow plan. The runtime constructs required steps such as context resolution, evidence assembly, policy evaluation, skill execution, approval routing, and replay persistence.
  4. Compile planner contract. The planner package records prompt hash, ordered steps, selected skill, token budget, and deterministic executor constraints.
  5. Evaluate policy. The Policy Runtime returns the governing decision, attestation hash, and policy set hash.
  6. Execute skill. The selected skill produces domain signals under the policy and evidence boundary.
  7. Persist proof. Evidence pack, replay manifest, replay run, replay steps, tool invocations, approval contract, decision log, and telemetry are written.
  8. Return outcome. The caller receives decision ID, recommendation, execution allowance, approval requirement, confidence, replay data, and live record references.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/iaf/enterprise-decision-runtime/contractReturn runtime responsibilities, supported decision types, SLO, and API contract.Runtime definition.
POST /api/iaf/enterprise-decision-runtime/executeBuild a deterministic decision envelope and replay package.Stateless execution path.
POST /api/iaf/enterprise-decision-runtime/execute-liveExecute through live platform boundaries and persist proof records.Authoritative production path.
GET /api/iaf/enterprise-decision-runtime/replay-live/{decision_id}Retrieve persisted live replay envelope for a decision.Replay and audit.
GET /api/iaf/enterprise-decision-runtime/replay-live/{decision_id}/compareCompare replay state against persisted live records.Replay integrity check.
POST /api/iaf/enterprise-decision-runtime/approvals/{decision_id}/transitionMove a decision approval contract through approve, reject, assign, or escalate transitions.Human authority boundary.
GET /api/iaf/enterprise-decision-runtime/production-certification-liveValidate runtime certification against live persisted records.Operational readiness.

Decision Strategies

StrategyUseControl Position
Deterministic-firstExact matching, thresholds, tolerances, authority, source-system status.Default for regulated finance execution.
HybridStructured checks plus narrative explanation, ambiguous exception handling, or case summarization.Allowed when deterministic gates remain authoritative.
Human-ledHigh-value, sensitive, control-relevant, or judgment-heavy decisions.Runtime routes, explains, records, and waits for approval.
Simulation-ledScenario comparison, forecast tradeoffs, cash impact, and policy-change impact analysis.No operational write-back unless promoted through policy and approval.

Integration With Other Runtimes

Policy Runtime

Returns the action boundary. The Decision Runtime can recommend, escalate, or hold only inside that boundary.

Evidence Runtime

Provides evidence packs and hashes used to determine confidence, supportability, and replay readiness.

Skill Fabric

Executes domain skills selected by the workflow plan and records tool invocation proof.

Replay and Observability

Persisted replay and telemetry prove what ran, how long it took, which tools were called, and what outcome was produced.

Example: Duplicate Payment Decision

A duplicate-payment risk event requests a hold recommendation. The runtime assembles invoice context, evidence items, policy state, selected skill output, approval requirements, value estimate, and replay proof. The resulting decision can recommend the hold, block payment release, route supervisor approval, and persist a replayable decision record.

{
  "decision": "hold_and_escalate",
  "allowed_to_execute": false,
  "human_approval_required": true,
  "policy_result": "REQUIRE_HUMAN_APPROVAL",
  "evidence_complete": true,
  "platform_boundaries": {
    "policy": "enterprise_policy_intelligence_runtime",
    "evidence": "EvidencePack",
    "ledger": "DecisionLogRecord",
    "approval": "DecisionContractRecord",
    "replay": "ReplayManifest/ReplayRun/ReplayStep"
  }
}

Operational Requirements

  • Decision execution must separate recommendation, approval, and write-back.
  • High-risk actions default to approval-required when policy, evidence, or runtime state is incomplete.
  • Decision outputs must include a decision ID, integrity hash, evidence reference, policy reference, approval boundary, and replay pointer.
  • LLM planning may prepare an orchestration contract, but deterministic executors and policy gates remain authoritative for regulated finance action.
  • Decision telemetry must record latency, cost estimate, token count, SOR calls, cache state, outcome, and runtime metadata.

Acceptance Criteria

  • Every material decision receives a stable decision ID and replay pointer.
  • Every executable path references current policy, evidence, approval, skill, telemetry, and replay records.
  • Recommendation and execution remain separate states.
  • High-value or write-back decisions cannot proceed without policy clearance and the required human approval state.
  • The runtime can replay a persisted decision from live records, not only from generated page content.
  • The same decision request and policy/evidence version produces deterministic execution boundaries.
  • Decision telemetry captures latency, token count, SOR calls, outcome, cost estimate, and runtime metadata.

Engineering Rule

Agents may propose recommendations and skills may generate signals, but the Decision Runtime owns the governed conversion from recommendation to operational next step. No material action should bypass policy, evidence, approval, ledger, and replay checks at the backend action layer.

CapabilitiesTrust Fabric + ReplayOntology + Enterprise ContextPolicy + Authority Control
Implemented Runtime Chapter

Runtime 3: Evidence Intelligence Runtime

The Evidence Intelligence Runtime is the proof layer for Sphere. It discovers, retrieves, validates, ranks, links, explains, packages, and preserves the evidence required to support recommendations, decisions, controls, approvals, escalations, and governed system actions.

Evidence Pack Source Lineage Sufficiency Scoring Trust Score Conflict Detection Replayable Proof

Production Role

The Policy Runtime determines what is allowed. The Decision Runtime determines what should happen next. The Evidence Runtime proves why a recommendation or decision is credible. It prevents unsupported agent claims by attaching source-linked, validated, and replayable evidence to material recommendations and actions.

It does not replace source systems, document management, identity, policy, or final business decisioning. It references and packages evidence from authoritative systems, preserves lineage, evaluates sufficiency, and exposes role-appropriate proof to runtime consumers.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service evidence_intelligence_runtime.py and evidence_intelligence_fabric.py Certified by evidence runtime live-source, world-class differentiator, and runtime certification tests.
API surface /api/eir/* and /api/evidence/* Covers contract, search, pack, decision-ready pack, provenance, graph, validate, lineage, trust score, ranking, compare, replay, KPIs, and certification.
Persisted evidence EvidencePack and evidence artifact tables Runtime can load persisted evidence packs, compute integrity coverage, and distinguish persisted packs from compatibility records.
Decision-ready packs POST /api/eir/decision-ready-pack Produces role-scoped evidence packs for analyst, supervisor, auditor, and executive consumption.
Provenance and integrity Provenance manifests, signature verification, content hashes, lineage records, and replay references. Implemented through /api/eir/provenance/{{case_id}}, /api/eir/provenance/verify, and evidence vault verification/export endpoints.
Evidence intelligence Trust scoring, continuous trust scoring, evidence graph, ranking, source-family normalization, and conflict comparison. Operational APIs include /api/eir/trust-score, /continuous-trust, /graph, /rank, and /compare.
Operational surfaces Evidence Vault, Decision Replay Studio, and Policy Library Evidence is consumed by case drawers, decision records, assurance packets, replay views, and policy/approval flows.

Runtime Scope

  • Discovery. Finds relevant invoice, PO, receipt, supplier, payment, policy, approval, journal, reconciliation, document, and prior-decision evidence.
  • Retrieval. Uses governed connectors and evidence fabric functions to retrieve source-linked evidence from approved systems and persisted packs.
  • Normalization. Converts heterogeneous source records into standard evidence objects with source family, record ID, freshness, authority, and hash metadata.
  • Validation. Checks authority, existence, freshness, completeness, case match, access, conflicts, and replay suitability.
  • Sufficiency. Separates evidence found from evidence sufficient for recommendation, approval, execution, audit, and replay.
  • Packaging. Builds evidence packs for agents, decision runtime, policy runtime, UI drawers, supervisor approvals, assurance, and replay.
  • Lineage. Stores source references, retrieval timestamps, hashes, connector version, actor, case, decision, policy decision, and replay pointer.
  • Role projection. Presents different evidence views for analyst, supervisor, auditor, and executive without changing the underlying evidence record.

Logical Architecture

Agent / User / Decision Runtime / Policy Runtime / Workflow
        |
        v
Evidence Request
        |
        v
Evidence Intelligence Runtime
        |
        +-- Evidence Discovery Engine
        +-- Evidence Connector Layer
        +-- Evidence Normalization Engine
        +-- Evidence Validation Engine
        +-- Evidence Sufficiency Engine
        +-- Evidence Conflict Detector
        +-- Evidence Ranking Engine
        +-- Evidence Pack Builder
        +-- Evidence Lineage Store
        +-- Evidence Access Control Adapter
        +-- Evidence Replay Adapter
        +-- Evidence Observability
        |
        v
Evidence Pack / Evidence Status / Missing Evidence / Lineage

Implemented Lifecycle

  1. Receive request. Agent, user, Decision Runtime, Policy Runtime, or workflow requests evidence for a case or decision.
  2. Resolve requirements. The runtime determines required evidence by case type, decision type, role, and control posture.
  3. Discover candidates. Relevant source records and documents are discovered across registered systems and persisted packs.
  4. Retrieve and normalize. Evidence is retrieved through approved paths and normalized into evidence objects.
  5. Validate quality. Source authority, freshness, completeness, access, consistency, and hash integrity are checked.
  6. Score sufficiency. The runtime determines whether evidence supports display, recommendation, approval, execution, audit, or replay.
  7. Build pack. Evidence items, missing items, trust score, lineage, and replay references are assembled into an evidence pack.
  8. Expose and replay. The pack is consumed by decision, policy, agent, UI, approval, assurance, and replay surfaces.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/eir/contractReturn runtime responsibilities, contract, and production gates.Runtime definition.
POST /api/eir/searchSearch evidence by query, case, or source.Discovery.
POST /api/eir/packBuild a case evidence pack.Evidence packaging.
POST /api/eir/decision-ready-packBuild a role-scoped decision-ready evidence pack.Decision and approval consumption.
GET /api/eir/provenance/{case_id}Return the provenance manifest for a case.Lineage and audit.
POST /api/eir/provenance/verifyVerify a provenance manifest.Integrity validation.
GET /api/eir/graph/{case_id}Return evidence graph relationships.Evidence intelligence.
POST /api/eir/validateValidate case evidence against required evidence.Sufficiency and quality.
GET /api/eir/lineage/{case_id}Return lineage for evidence records.Replay and audit.
GET /api/eir/trust-score/{case_id}Return evidence trust score.Trust scoring.
GET /api/eir/replay/{case_id}Return replayable evidence context.Replay adapter.
GET /api/evidence/packsList persisted evidence packs from the evidence vault.Operational evidence store.

Evidence Status Types

StatusMeaningRuntime Constraint
completeAll required evidence is available.Eligible for the configured decision purpose.
complete_for_recommendationEnough proof for recommendation, not execution.Execution remains blocked or approval-gated.
partialSome required evidence is missing.Decision must disclose gaps.
missingRequired evidence unavailable.Request evidence or require review.
conflictingSources disagree.Execution blocked until resolved.
staleEvidence freshness no longer satisfies policy.Refresh required before high-risk execution.
restrictedActor lacks entitlement.Mask or deny the evidence item.
untrustedSource is not authoritative enough.Use only as supporting context.

Integration With Other Runtimes

Policy Runtime

Defines evidence requirements and consumes sufficiency status before allowing action.

Decision Runtime

Uses evidence packs to determine whether a recommendation is supportable and what action is allowed next.

Agent and Skill Runtime

Agents and skills ground material claims in evidence IDs, packs, source records, and lineage.

Replay and Assurance Runtime

Reconstructs evidence state using pack IDs, source references, hashes, timestamps, and decision IDs.

Example: Duplicate Payment Evidence

A duplicate-payment case requires current invoice, prior candidate, supplier master, payment history, payment status, policy reference, and approval state. The runtime can mark the pack complete for recommendation while explicitly blocking execution until approval evidence is present.

{
  "evidence_pack_id": "evp_001928",
  "status": "complete_for_recommendation",
  "sufficiency": {
    "recommendation": true,
    "approval": true,
    "execution": false,
    "audit": true
  },
  "missing_evidence": ["supervisor_approval"],
  "lineage_pointer": "lineage://evidence/evp_001928"
}

Operational Requirements

  • Context is not evidence until it has source lineage, validation, and sufficiency status.
  • Evidence discovery, validation, sufficiency, and packaging must remain separate runtime steps.
  • Agents may summarize evidence, but must not invent evidence or promote unsupported claims.
  • Decision, policy, replay, audit, and UI layers should reference evidence IDs and evidence pack IDs.
  • Evidence failures must degrade safely: display available proof, disclose gaps, and restrict execution when needed.

Acceptance Criteria

  • Every material recommendation references an evidence pack or explicitly reports evidence insufficiency.
  • Every evidence item carries source-system lineage, validation status, and hash/integrity metadata where available.
  • Evidence sufficiency distinguishes recommendation-ready from execution-ready evidence.
  • Conflicting, stale, missing, untrusted, and restricted evidence are surfaced instead of silently filled.
  • Sensitive evidence is role-scoped or masked according to policy and entitlement.
  • Evidence packs can be consumed by Policy Runtime, Decision Runtime, Replay Runtime, and assurance packets.
  • Persisted evidence records can be loaded and measured without relying only on in-memory compatibility cases.

Engineering Rule

Evidence must be assembled and validated before material recommendations are treated as decision-ready. A value shown in a user interface is not evidence unless it is linked to a source record, validation state, freshness status, lineage, and evidence pack reference.

CapabilitiesOntology + Enterprise ContextCREST Context ReconstructionEnterprise Memory + Knowledge ManagementTrust Fabric + Replay
Implemented Runtime Chapter

Runtime 4: Enterprise Context Runtime

The Enterprise Context Runtime is the contextual intelligence layer for Sphere. It gathers, normalizes, enriches, governs, optimizes, and serves the business context required by agents, policies, decisions, evidence packs, workflows, analytics, supervisors, and executives.

Context Pack Source Catalog Ontology Mapping CREST Reconstruction AI Context Governance Replay Snapshot

Production Role

Evidence proves specific claims. Context frames the complete business situation around those claims. The Context Runtime answers what the platform, user, workflow, or model is allowed to know before a recommendation, evidence pack, policy decision, or business decision is produced.

The runtime does not replace SAP, Ariba, ServiceNow, Databricks, Foundry, UDP, identity, evidence, policy, or final decisioning. It creates a governed context object from those sources and platform runtimes so downstream consumers do not each invent their own interpretation of the same case.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_context_runtime.py Certified by enterprise context runtime certification, live-source, production certification, and context quality tests.
API surface /api/ecr/* Provides contract, build, preview, rank, compress, explain, graph, quality, governance, sources, observability, KPI, and certification endpoints.
Context audit surface /api/iaf/context-runtime/* Exposes context operations, audit cases, lineage, usage, skills/policies, control plane, ontology, judgment, learning, and missing-context simulation.
CREST reconstruction /crest/context/build and /api/iaf/crest/context/build Builds reconstructed context from CREST entities, graph edges, decision links, skills, ontology, and knowledge metrics.
Persisted context ResolvedContext, ResolvedContextSnapshot, and decision context hashes Decision and payment services persist context snapshots and replay verifies context/source snapshot hashes.
Context intelligence Ranking, compression, explanation, context graph, quality scoring, governance checks, and source catalog. Implemented through rank_context, compress_context, explain_context, context_graph, context_quality, and context_governance.
Operational surfaces Context Graph, CREST Context Studio, Context Engineering, and Context Audit Console Context is visible as a platform runtime, not only as page-local data.

Runtime Scope

  • Assembly. Gathers case, mission, user, business, financial, operational, policy, evidence, memory, tool, and AI context into one context pack.
  • Normalization. Converts system-specific records into common business objects and mission-aware context structures.
  • Ontology mapping. Aligns source terms with Sphere finance concepts so agents and users operate with consistent meaning.
  • Source catalog. Declares approved context sources including SAP ECC, SAP S/4HANA, Ariba, Databricks, Unity Catalog, SharePoint, Teams, Outlook, HR, policy, evidence, memory, knowledge, event runtime, and external APIs.
  • Process awareness. Applies mission process lifecycles for P2P, O2C, R2R, Treasury, FP&A, and Tax context framing.
  • Governance. Applies role, mission, redaction, minimum-required-context, direct-SOR-access, RBAC, ABAC, and prompt-injection protections before context reaches models or users.
  • Optimization. Ranks, compresses, explains, and graph-links context to control token cost, model latency, and relevance.
  • Replay. Persists snapshots and hashes so the context used by a decision can be reconstructed and verified later.

Logical Architecture

User / Agent / Workflow / Evidence / Policy / Decision
        |
        v
Context Request
        |
        v
Enterprise Context Runtime
        |
        +-- Context Request Router
        +-- Source Connector Adapter
        +-- Domain Normalization Engine
        +-- Business Ontology Mapper
        +-- Source Catalog and Process Lifecycle
        +-- Context Enrichment Engine
        +-- Identity and Entitlement Adapter
        +-- Data Masking and Redaction Adapter
        +-- Freshness and Source Health Engine
        +-- Context Ranking and Compression
        +-- Context Graph and CREST Reconstruction
        +-- Context Snapshot and Lineage Store
        +-- Context Observability
        |
        v
Governed Business Context Object

Implemented Lifecycle

  1. Receive request. User, agent, workflow, Evidence Runtime, Policy Runtime, or Decision Runtime requests context for a case, entity, mission, or AI interaction.
  2. Resolve scope. The runtime determines whether the caller needs case, transaction, entity, mission, executive, AI, or replay context.
  3. Select sources. Relevant source systems and platform runtimes are selected through the source catalog and mission lifecycle.
  4. Retrieve and reconstruct. Context is built from live persisted sources where available, with controlled fallback to in-process assembly when runtime dependencies fail.
  5. Normalize and map. Source data is transformed into common business objects and ontology-aligned terms.
  6. Enrich. Risk, materiality, process state, SLA, control relevance, ownership, source health, memory, and financial context are added.
  7. Govern. Role, entitlement, redaction, masking, action authority, data boundary, and AI-safety controls are applied.
  8. Optimize. The pack is ranked, compressed, explained, graph-linked, and scored for quality and governance readiness.
  9. Persist and expose. Snapshots, lineage, hashes, and audit views make the context consumable by decision, evidence, policy, agent, UI, and replay surfaces.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/ecr/contractReturn runtime responsibilities, components, sources, process lifecycles, and certification endpoints.Runtime definition.
POST /api/ecr/context/buildBuild a governed context pack, preferring live persisted source I/O with safe fallback.Context assembly.
POST /api/ecr/context/previewPreview context structure and scope before downstream use.Context review.
POST /api/ecr/context/rankRank context by relevance and decision utility.Context optimization.
POST /api/ecr/context/compressCompress context to a bounded set of high-value items.Model and latency control.
POST /api/ecr/context/explainExplain what context is included and why.User and agent transparency.
POST /api/ecr/context/graphReturn context graph nodes, edges, and relationship summary.Graph reconstruction.
POST /api/ecr/context/qualityScore completeness, freshness, source coverage, and decision readiness.Context quality.
POST /api/ecr/context/governanceEvaluate redaction, authority, external sharing, and policy compliance.Context governance.
GET /api/ecr/sourcesReturn approved source catalog, optionally scoped by mission.Source registry.
GET /api/ecr/executive-consoleReturn cross-runtime context console summary.Executive context.
GET /api/ecr/observability-contractReturn context observability and operational measurement contract.Runtime operations.
POST /crest/context/buildBuild CREST context from enterprise graph and decision-link assets.Context reconstruction.
GET /api/iaf/context-runtime/audit/cases/{entity_id}/lineageReturn case context lineage for audit and replay.Lineage.

Context Types

Context TypePurposeExample Runtime Consumer
caseFull business situation around one case.Case drawer, Decision Runtime, replay.
transactionInvoice, payment, journal, PO, receipt, or reconciliation state.Evidence Runtime and Skill Runtime.
entitySupplier, customer, account, legal entity, cost center, or business unit.Agent Runtime and mission workbench.
missionQueue, process, lifecycle, and operating state for P2P, O2C, R2R, Treasury, FP&A, or Tax.Supervisor control plane.
aiPrompt-safe, role-scoped context for an agent or external model target.Ask Sphere and model gateway.
replaySnapshot and hash set used to reconstruct prior decisions.Replay Runtime and assurance views.

Integration With Other Runtimes

Evidence Runtime

Consumes context to discover which evidence should be assembled, validated, and packaged.

Policy Runtime

Uses context fields such as role, amount, mission, data sensitivity, source state, and requested action.

Decision Runtime

Uses context to classify decision type, select strategy, evaluate risk/value, and generate next best action.

Agent and Skill Runtime

Receives scoped, governed context packs instead of independently fetching unmanaged source data.

CREST and Knowledge Runtime

Reconstructs enterprise context through entities, graph edges, decision links, ontology, skills, and knowledge metrics.

Replay Runtime

Verifies context snapshots and source hashes so later review can reconstruct the original business situation.

Example: Duplicate Payment Context

For a duplicate-payment case, the runtime frames the invoice as more than an AP record. It becomes a P2P exception with supplier context, payment status, prior-payment context, process stage, risk, SLA posture, policy relevance, evidence references, role authority, and source lineage.

{
  "context_id": "ECR-p2p-INV-LIVE-DUP-1778594510",
  "case_id": "INV-LIVE-DUP-1778594510",
  "mission": "p2p",
  "role": "p2p_analyst",
  "process_awareness": {
    "process": "purchase_to_pay",
    "stage": "exception_review"
  },
  "security": {
    "direct_sor_access": false,
    "redaction_required": true,
    "authority": "recommend_only",
    "data_boundary": "minimum_required_context",
    "rbac": true,
    "abac": true,
    "prompt_injection_protection": true
  },
  "source_catalog": ["SAP ECC", "SAP S/4HANA", "SAP Ariba", "Databricks", "Policy Runtime", "Evidence Runtime"]
}

Operational Requirements

  • Pages should request normalized context from the runtime instead of constructing independent case objects.
  • Agents should consume scoped context packs and must not independently retrieve unrestricted source context for material actions.
  • Context and evidence must remain distinct: context frames the business situation; evidence proves specific claims.
  • AI context must be purpose-aware, role-aware, source-aware, and redacted before prompt construction or model routing.
  • Ontology mapping, source selection, transformations, snapshots, and hashes must be versioned enough to support replay.

Acceptance Criteria

  • Every material case can be represented as a context pack with case ID, mission, role, source catalog, process awareness, security boundary, and context hash.
  • Agents receive scoped context packages rather than unrestricted raw source-system data.
  • Policy, evidence, decision, skill, replay, and user-experience layers can consume the same context object.
  • Context can be ranked, compressed, explained, graphed, quality-scored, and governance-checked through runtime APIs.
  • Sensitive context is redacted or constrained by role, purpose, mission, RBAC, ABAC, and minimum-required-context rules.
  • Live-source context falls back safely without breaking the UI surface or falsely authorizing high-risk execution.
  • Decision replay can verify the context snapshot and source snapshot hashes used by the original decision.

Engineering Rule

The Context Runtime is the source of governed business framing. If a page, agent, policy check, evidence pack, or decision requires business situation awareness, it should consume the shared context pack rather than assembling an isolated interpretation from raw APIs.

CapabilitiesSkill Fabric + Agentic RuntimePolicy + Authority ControlTrust Fabric + ReplayEnterprise Memory + Knowledge Management
Implemented Runtime Chapter

Runtime 5: Enterprise Agent Runtime

The Enterprise Agent Runtime is the controlled execution layer for Sphere agents. It governs how agents are registered, planned, invoked, scoped, tooled, constrained, monitored, certified, lifecycle-managed, and replayed.

Agent Registry Governed Autonomy Tool Permission Gate Reasoning / Execution Separation Agent Certification Replay Package

Production Role

Context defines what an agent may know. Evidence proves what the agent can rely on. Policy determines what is allowed. Decisioning determines what should happen next. The Agent Runtime defines how agents participate in that chain without becoming uncontrolled scripts, prompts, or source-system actors.

The runtime does not own policy authoring, source-system write execution, model-provider infrastructure, evidence storage, or final business authority. It owns the governed agent control plane: registration, planning, execution gates, tool permission contracts, model lineage, collaboration, certification, observability, lifecycle state, and replay packages.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_agent_runtime.py Certified by enterprise agent runtime certification and production gate tests.
API surface /api/ear/* Provides contract, registry, register, plan, execute, collaborate, certify, promotion gate, learning feedback, suspend, retire, console, observability, KPI, and certification endpoints.
Platform registry IRISAgentRegistry, AgentDefinition, AgentDeployment, AgentPolicyBinding, and AgentEvidenceContract Separates canonical agent identity, tenant activation, policy binding, and evidence/replay requirements.
Execution substrate AgentExecution, AgentExecutionTelemetry, ToolInvocation, and replay manifests Runtime executions carry execution IDs, hashes, tool lineage, model lineage, evidence package, replay timeline, and telemetry.
Operational gates Readiness gate, policy gate, evidence gate, decision gate, tool permission gate, human boundary, and lifecycle controls. Production certification asserts all gates for a high-risk duplicate-invoice agent execution.
Agent operations surfaces Agent Inventory, Agentic Workforce, Skill Fabric, and Agent Governance Center Agent health, readiness, policy binding, evidence, execution history, and governance can be inspected from live UI surfaces.
Measured posture world_class_certification() and production_certification() The runtime exposes measured readiness gaps and does not claim full world-class status unless production gates and persisted execution evidence support it.

Runtime Scope

  • Registration. Agents are registered with ownership, domain, missions, autonomy level, criticality, maturity, permissions, tools, skills, models, policies, dependencies, SLA class, and version.
  • Planning. Task plans include identity/authority check, context assembly, memory review, decomposition, policy precheck, evidence requirements, permissioned tool selection, and recommendation.
  • Execution gating. External action requests add mandatory decision runtime gate, human approval gate, and controlled source-system handoff.
  • Reasoning boundary. Agents may reason, draft, plan, recommend, and collaborate, but external financial action remains separated from reasoning output.
  • Tool control. Tools are constrained by declared permissions, tool lineage, policy gates, and authority level.
  • Collaboration. High-risk work can require peer review by evidence, policy, decision, supervisor, and control-owner agents.
  • Certification. Agent readiness scores cover business value, reliability, explainability, policy compliance, evidence quality, performance, security, human collaboration, and operational maturity.
  • Lifecycle. Agents can be promoted, suspended, retired, archived, and measured through readiness gates.
  • Learning. Human feedback and plan-change signals are recorded as governed learning feedback with evidence references.
  • Replay. Every execution can carry an immutable replay package with prompt/context/evidence/tool/policy/decision boundary references.

Logical Architecture

User / Event / Workflow / Schedule / API
        |
        v
Agent Invocation Request
        |
        v
Enterprise Agent Runtime
        |
        +-- Agent Registry
        +-- Agent Manifest and Readiness Validator
        +-- Invocation Router
        +-- Context Injection Adapter
        +-- Planning and Task Decomposition
        +-- Tool Authorization Engine
        +-- Policy Enforcement Adapter
        +-- Evidence Attachment Adapter
        +-- Agent Execution Boundary
        +-- Decision Handoff Adapter
        +-- Human Handoff Adapter
        +-- Agent Memory and Learning Adapter
        +-- Agent Observability
        +-- Agent Replay Recorder
        |
        v
Recommendation / Explanation / Draft / Evidence Request / Escalation / Decision Request / Action Proposal

Implemented Lifecycle

  1. Register. Agent identity and manifest-like contract are declared through registry models or /api/ear/agents/register.
  2. Certify. Readiness is scored against ownership, policy, evidence, security, reliability, explanation, and operational maturity.
  3. Plan. The runtime builds a governed task plan and identifies context, evidence, policy, tool, decision, and human gates.
  4. Assemble context. The Context Runtime provides scoped context; agents do not need unmanaged raw source access.
  5. Check policy and evidence. Policy and evidence requirements are checked before material recommendation or action proposal.
  6. Authorize tools. Allowed tools are selected and recorded; source-system write actions remain gated.
  7. Generate recommendation. The agent produces a recommendation, explanation, evidence request, escalation request, or decision request.
  8. Handoff to decision. Business action proposals are evaluated by the Decision Runtime before execution can proceed.
  9. Capture human boundary. Approval-required states are explicit and recorded before any controlled handoff.
  10. Record telemetry and replay. Execution hash, tool lineage, model lineage, evidence package, observability, collaboration, and replay timeline are stored.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/ear/contractReturn runtime scope, ownership, operating model, autonomy contract, integrations, and APIs.Runtime definition.
GET /api/ear/registryReturn registered agents with lifecycle state, declared controls, tools, permissions, and certification posture.Agent registry.
POST /api/ear/agents/registerRegister an agent with ownership, autonomy, tools, models, skills, policies, dependencies, SLA, and version.Registration and manifest validation.
POST /api/ear/agents/planCreate a cost-aware and time-aware task plan with required runtime gates.Planning.
POST /api/ear/agents/executeExecute an agent within readiness, policy, evidence, decision, tool, and human-boundary gates.Governed execution.
POST /api/ear/agents/collaborateAssemble peer-review and consensus team for high-risk or critical work.Agent collaboration.
POST /api/ear/agents/certifyScore agent readiness and return promotion state, blockers, and governance checks.Certification.
POST /api/ear/agents/{agent_id}/promotion-gateEvaluate whether an agent can progress to the target lifecycle state.Lifecycle gate.
POST /api/ear/agents/learning-feedbackRecord human feedback, learning, plan-change signal, and evidence reference.Learning adapter.
POST /api/ear/agents/{agent_id}/suspendSuspend an agent and record the control reason.Operational control.
POST /api/ear/agents/{agent_id}/retireRetire an agent and require archive of execution/replay/certification history.Lifecycle closure.
GET /api/ear/observability-contractReturn agent metrics, traces, alerts, and SLOs.Runtime observability.
GET /api/ear/production-certificationRun production certification gates against a representative high-risk execution contract.Release gate.

Agent Output Types

OutputMeaningRequired Boundary
recommendationAgent proposes an action or next step.Evidence reference, policy context, decision handoff for material action.
explanationAgent explains structured runtime outputs.No new unsupported facts; cite context, evidence, policy, or decision objects.
draftAgent prepares a message, note, or action payload.Human review unless policy allows direct administrative send.
evidence_requestAgent identifies missing proof.Evidence Runtime request or human evidence collection task.
escalation_requestAgent routes high-risk or blocked work.Supervisor, controller, control owner, treasury approver, or other role queue.
decision_requestAgent asks Decision Runtime to evaluate next best action.Decision ID, policy decision, evidence pack, and replay reference.
action_proposalAgent proposes source-system or workflow action.Policy, evidence, decision, human approval, and integration write-back gates.

Autonomy and Gate Model

Autonomy LevelRuntime MeaningExecution Constraint
observeAgent can monitor and report.No recommendations or actions without escalation.
recommendAgent can recommend or explain.No external action execution.
assistAgent can draft and prepare work.Human review before controlled action.
execute_with_approvalAgent can prepare execution after approval.Approval and decision gates required.
autonomous_within_policyAgent can execute only within explicit low-risk policy boundaries.Policy, evidence, decision, observability, and replay gates remain mandatory.

Integration With Other Runtimes

Context Runtime

Provides scoped context packages so agents do not fetch unrestricted raw source data.

Evidence Runtime

Supplies evidence packs and evidence gates for material recommendations.

Policy Runtime

Constrains invocation, tool use, external sharing, action proposal, memory creation, and write-back requests.

Decision Runtime

Evaluates agent recommendations before business actions are allowed.

Skill Runtime

Provides deterministic and hybrid business capabilities that agents can invoke under permissioned tool contracts.

Replay and Observability

Records execution hash, timeline, gates, tool/model lineage, telemetry, collaboration, and final output.

Example: Duplicate Invoice Agent Execution

A P2P duplicate-invoice agent can assess a high-risk candidate, generate a hold recommendation, attach evidence, and request a decision. The runtime records that external action was requested but not executed until policy, evidence, decision, and human gates clear.

{
  "operation": "ExecuteAgent",
  "agent_id": "p2p_duplicate_invoice_agent",
  "reasoning_completed": true,
  "external_action_requested": true,
  "external_action_executed": false,
  "reasoning_execution_separated": true,
  "gate": {
    "policy_gate": {"required": true},
    "evidence_gate": {"required": true, "status": "passed"},
    "decision_gate": {"required": true},
    "tool_permission_gate": {"status": "passed"},
    "human_boundary": "approval_required_before_external_action"
  },
  "replay": {
    "immutable_after_finalization": true,
    "timeline": ["agent_registered", "context_loaded", "policy_prechecked", "evidence_checked", "tools_invoked", "recommendation_generated", "decision_gate_applied", "human_boundary_recorded"]
  }
}

Operational Requirements

  • Agents must run through the Agent Runtime rather than direct page, script, or prompt invocation for material work.
  • Prompts may guide behavior, but policy, evidence, tool, and execution gates must be enforced outside prompts.
  • Recommendation and execution must remain separate states, with explicit human approval for high-risk financial action.
  • Every tool call should be permissioned, observable, and replayable.
  • Agent lifecycle promotion must be based on readiness, controls, evidence, observability, and operational maturity, not registration alone.

Acceptance Criteria

  • Every approved agent has a registry identity, owner, version, domain/mission scope, autonomy level, permissions, tools, policies, and dependencies.
  • High-risk agent execution separates reasoning from external action and records external_action_executed=false until gates clear.
  • Agent plans include context, memory, policy, evidence, tool, recommendation, decision, and human approval steps where required.
  • Tool access is allowlisted and captured through tool lineage or tool invocation records.
  • Material recommendations carry evidence references, policy context, decision references, and replay metadata.
  • Promotion gates are hard gates and can block lifecycle progression when readiness score or registry blockers fail.
  • Agent observability publishes success/failure, latency, token/cost, tool utilization, human intervention, readiness, policy violation, and override metrics.
  • Agent replay captures enough information to reconstruct what the agent saw, which gates applied, what tools were called, and what output was produced.

Engineering Rule

No material agent should execute outside the runtime control plane. Agents can reason and recommend, but source-system action requires context, evidence, policy, decision, tool authorization, human boundary, observability, and replay records at the backend action layer.

CapabilitiesSkill Fabric + Agentic RuntimeTrust Fabric + ReplayPolicy + Authority ControlValue Attribution Engine
Implemented Runtime Chapter

Runtime 6: Enterprise Skill Runtime

The Enterprise Skill Runtime is the reusable business-capability execution layer for Sphere. It provides governed, versioned, observable, evidence-producing skills that agents, workflows, decisions, policies, evidence packs, and user experiences can invoke consistently.

Skill Registry Deterministic Finance Logic Skill Marketplace Composition Plan Evidence-Producing Output Durable Replay

Production Role

Agents coordinate and reason. Skills perform defined business capabilities with schema-bound inputs, outputs, policies, evidence, versioning, observability, and replay. This keeps finance logic out of prompts and makes reusable capabilities inspectable by architects, operators, control owners, and auditors.

The runtime does not own agent reasoning, final decisions, policy authoring, source-system authority, evidence storage, or workflow orchestration. It owns the governed execution contract for reusable skills and the Skill Fabric integration that turns domain logic into callable platform capabilities.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_skill_runtime.py Certified by enterprise skill runtime tests covering contract, registry, marketplace, composition, execution, durable history, certification, and production gates.
Skill Fabric service skill_fabric_service.py and skill_runtime_logic.py Executes concrete finance skills including duplicate invoice detection, 3-way match, GRIR, price/quantity variance, vendor risk, payment terms, discount capture, R2R, FP&A, Treasury, and O2C contributors.
API surface /api/esr/* and /api/iaf/skills/* ESR provides enterprise lifecycle/governance. Skill Fabric provides registry, dependency readiness, execution, lifecycle, overlays, quality, graph, lineage, and durable execution surfaces.
Skill contracts core/enterprise_skill_contract_runtime.py and filesystem-backed skill.yaml packages Skill manifests compile provider-neutral skill context, input/output contracts, policy bindings, evidence requirements, role variants, and package metadata.
Durable execution SkillExecutionRecord and ToolInvocation Durable execution stores skill ID, mission, decision reference, role, status, evidence hash, input/output hashes, selection metadata, execution metadata, and replay data.
Governance and gates Skill governance gate, policy gate, audit gate, budget enforcement, source-adapter contract, and fabric gate before durable execution. Live adapter skills require payloads and block rather than fabricate records when SAP, Ariba, Microsoft Graph, or Databricks payloads are absent.
Operational surfaces Skill Fabric, Skills Registry, Skill Reality Control Plane, and Skill Fabric Introduction Users can inspect skill contracts, dependency readiness, lifecycle state, source bindings, execution history, and replay links.

Runtime Scope

  • Registry. Maintains reusable skill contracts with owner, domain, mission, category, version, criticality, permissions, inputs, outputs, policies, evidence, dependencies, SLA, and certification state.
  • Discovery and marketplace. Finds skills by mission, domain, category, capability, query, caller, tags, certification status, and marketplace score.
  • Composition. Composes multiple skills into sequential or parallel plans with preconditions, policy requirements, evidence requirements, reusable templates, and replayable composition hashes.
  • Execution. Runs deterministic, integration, AI-assisted, analytics, and business skills with input validation, governance gate, certification, evidence generation, quality scoring, observability, and replay.
  • Durability. Persists governed executions to iaf_skill_execution_records with evidence/input/output hashes and execution metadata.
  • Source dependencies. Exposes source dependency readiness so skills show required sources, permissions, candidate systems, and production validation state.
  • Evidence output. Converts material skill outputs into evidence hashes, evidence records, decision signals, and replay references.
  • Budget and reliability. Tracks execution latency, cost, retries, timeouts, budget breaches, unavailable sources, success/failure, and blocked outcomes.
  • Lifecycle. Supports skill lifecycle state, certification, governance gate, overlays, quality scorecards, lineage, capability graph, and kill-switch-style suspension through lifecycle controls.

Logical Architecture

Agent / Workflow / Decision Runtime / Evidence Runtime / UI / API
        |
        v
Skill Invocation Request
        |
        v
Enterprise Skill Runtime
        |
        +-- Skill Registry
        +-- Skill Contract Validator
        +-- Discovery and Marketplace
        +-- Composition Planner
        +-- Input Validation
        +-- Governance Gate
        +-- Skill Execution Engine
        +-- Deterministic Logic Executor
        +-- Integration Adapter Boundary
        +-- Model / AI Adapter
        +-- Evidence Output Adapter
        +-- Output Quality Scoring
        +-- Budget and Reliability Monitor
        +-- Durable Execution Recorder
        +-- Skill Replay Recorder
        |
        v
Score / Match Result / Classification / Calculation / Validation / Decision Signal / Evidence Item / Simulation Result

Implemented Lifecycle

  1. Register or load. Skill is registered through ESR or imported from the filesystem-backed Skill Fabric catalog.
  2. Validate contract. Required owner, permissions, policies, inputs, outputs, dependencies, evidence, performance targets, and certification status are checked.
  3. Discover or select. Agents, workflows, decisions, and UIs discover eligible skills by mission, capability, query, or explicit skill ID.
  4. Compose if needed. Multi-skill workflows are assembled with sequential or parallel execution and reusable composition hashes.
  5. Gate execution. Governance, policy, audit, budget, fabric, and source-adapter gates run before execution.
  6. Execute logic. Runtime invokes deterministic functions, integration adapters, AI-assisted classification, analytical calculations, or hybrid contributors.
  7. Validate output. Outputs are structured, hashed, quality-scored, and checked for status, trust, and replay readiness.
  8. Emit evidence. Skill result creates evidence hash, evidence metadata, decision signal, and optional tool invocation lineage.
  9. Persist history. Durable executions store skill, mission, decision, hashes, selection, governance, quality, business outcome, and replay metadata.
  10. Observe and improve. Metrics, quality scorecards, overlays, lifecycle events, and human feedback support regression and runtime improvement.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/esr/contractReturn ESR scope, ownership, taxonomy, relationships, differentiators, and integration contract.Runtime definition.
GET /api/esr/registryReturn enterprise skill registry and registry controls.Skill registry.
POST /api/esr/skills/registerRegister a governed skill contract.Registration.
POST /api/esr/skills/discoverFind skills by mission, domain, category, query, capability, and caller.Discovery.
POST /api/esr/marketplaceReturn searchable marketplace results with dependency visualization, usage analytics, ratings, and usage summary.Marketplace.
POST /api/esr/skills/composeCreate sequential or parallel skill composition plan.Composition.
POST /api/esr/skills/run-compositionExecute a reusable multi-skill composition plan with replay.Composition execution.
POST /api/esr/skills/executeExecute a skill with governance, certification, evidence, quality, budget, observability, and replay output.Skill execution.
POST /api/esr/skills/execute-durableExecute with fabric gate and persist SkillExecutionRecord.Durable execution.
GET /api/esr/skills/executionsRead durable execution history by skill ID.Replay and audit.
POST /api/esr/skills/certifyScore skill readiness and governance board checks.Certification.
POST /api/esr/skills/{skill_id}/governance-gateEvaluate whether skill can move to target lifecycle state.Governance gate.
POST /api/esr/fabric/loadImport filesystem-backed Skill Fabric catalog into ESR.Skill Fabric integration.
GET /api/iaf/skills/registry/{skill_id}Inspect concrete Skill Fabric contract.Operational contract.
POST /api/iaf/skills/executeExecute concrete Skill Fabric skill and emit tool invocation log.Domain skill execution.
GET /api/iaf/skills/executionsRead Skill Fabric execution history.Skill replay.

Skill Types

TypeImplemented ExamplesPrimary Output
Matchingfinance.p2p.duplicate_invoice_detection, finance.p2p.3way_match_resolution, reconciliation contributors.Match score, reasons, candidate, variance.
ScoringDuplicate payment, vendor risk, close task risk, forecast confidence, liquidity assessment.Risk, priority, confidence, materiality.
ValidationInvoice math integrity, payment terms, currency, vendor master, posting control, policy compliance.Validation result and failed controls.
CalculationGRIR variance, price/quantity variance, discount capture, cash forecast impact, KPI delta.Calculated exposure, variance, value.
Extraction / ExplanationEvidence pack assembly, historical decision context, executive explanation, invoice classification.Structured context, narrative, evidence refs.
Integrationsap_read, ariba_read, databricks_query, msgraph_read.Live adapter payload or blocked status.

Runtime Modes

ModeUseControl Posture
deterministicThresholds, calculations, exact matches, schema validation.Preferred for financial controls and audit-sensitive checks.
deterministic_firstDuplicate matching, 3-way match, GRIR, payment terms, posting control.Structured logic dominates; explanation may be added after.
integration_adapterSAP, Ariba, Databricks, Microsoft Graph read adapters.Live payload required; no fabricated fallback.
ai_assistedClassification, explanation, executive summary, document-language interpretation.Policy-gated, output-schema validated, evidence-aware.
hybridDomain contributors combining deterministic checks, model support, context, policy, and evidence.Replay must bind skill version, source refs, model/policy versions, and output hashes.
simulationScenario, cash, forecast, value, and transformation calculations.Simulation output is decision input, not direct execution authority.

Integration With Other Runtimes

Agent Runtime

Agents discover and invoke certified skills instead of embedding finance logic in prompts.

Decision Runtime

Consumes skill outputs as decision signals, selected skill IDs, replay bindings, and skill result payloads.

Context Runtime

Provides normalized context and scoped source data required by skills.

Evidence Runtime

Converts skill outputs into evidence items, evidence hashes, and evidence pack contributors.

Policy Runtime

Gates sensitive data use, model use, execution eligibility, source dependency, and learning capture.

Replay and Observability

Tracks skill version, input/output hashes, evidence hash, policy gate, source calls, latency, cost, and replay pointer.

Example: Duplicate Invoice Skill

The duplicate invoice capability runs as a reusable skill rather than prompt-only agent logic. It produces a structured score, risk level, match reasons, evidence hash, business outcome, replay ID, and execution hash.

{
  "operation": "ExecuteSkill",
  "skill_id": "duplicate_invoice_detection",
  "skill_version": "1.0.0",
  "outputs": {
    "status": "completed",
    "duplicate_risk_score": 0.98,
    "risk_level": "high",
    "match_reasons": ["same_supplier", "same_amount", "similar_invoice_number", "payment_window_overlap"]
  },
  "evidence": {
    "evidence_hash": "sha256:...",
    "produced_by": "enterprise_skill_runtime"
  },
  "quality": {
    "trust_score": 0.9
  },
  "replay": {
    "immutable_after_finalization": true
  }
}

Operational Requirements

  • Reusable finance logic should be implemented as skills rather than duplicated in agents, prompts, UI components, or workflow scripts.
  • Skills must be schema-bound where outputs feed evidence, decisioning, policy checks, or audit views.
  • Skill outputs used for material recommendations must include skill ID, version, evidence hash, input/output hashes, and replay pointer.
  • Integration skills must block safely when required live payloads or source permissions are missing.
  • Skill deployment and promotion must be governed by tests, certification, source dependencies, lifecycle state, and observability.

Acceptance Criteria

  • Every production skill has an owner, version, domain/mission scope, input contract, output contract, policies, evidence requirements, dependencies, and certification state.
  • Skill execution returns a run/execution ID, skill version, governance gate, certification result, inputs, outputs, evidence hash, quality score, business outcome, observability, budget check, replay ID, and execution hash.
  • Durable skill execution records input and output hashes rather than relying only on transient in-memory execution.
  • Skill Fabric execution exposes tool invocation logs and source dependency readiness for concrete finance skills.
  • Live adapter skills block when required live payloads are missing; they do not silently simulate SAP, Ariba, Databricks, or Microsoft Graph results.
  • Agent and Decision Runtime paths can invoke skills and preserve skill IDs, versions, evidence hashes, and replay metadata.
  • Skill tests and certification surfaces expose measured readiness gaps and avoid overstating runtime quality when execution history is incomplete.

Engineering Rule

Business logic belongs in reusable skill contracts. Agents may orchestrate, explain, and recommend, but matching, scoring, validation, calculation, extraction, simulation, and control checks should run through the Skill Runtime so outputs are versioned, tested, observable, evidence-ready, and replayable.

CapabilitiesEvent Fabric + SOR SidecarSkill Fabric + Agentic RuntimePolicy + Authority ControlTrust Fabric + Replay
Implemented Runtime Chapter

Runtime 7: Enterprise Orchestration Runtime

The Enterprise Orchestration Runtime is the coordination layer for Sphere missions. It governs how triggers, cases, context, evidence, skills, agents, policies, decisions, human approval, workflow actions, events, and source-system handoffs move from intent to controlled outcome.

Workflow Planning Runtime Coordination Policy Gate Human Wait State Mission ACT Replay Graph

Production Role

The implemented service is named EnterpriseIntelligenceOrchestrationRuntime. In the platform runtime stack, it is the enterprise orchestration capability: it creates the execution plan, chooses participating runtimes, binds context/evidence/skills/model route, applies policy gates, creates human wait states, records observability, and seals replay.

The runtime does not decide business outcomes, author policies, own evidence, or bypass systems of record. Decisioning remains with the Decision Runtime, governance remains with Policy Runtime, evidence remains with Evidence Runtime, and source-system updates flow through controlled ACT or integration action paths.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime control plane enterprise_intelligence_orchestration_runtime.py Plans and executes cross-runtime flows across context, policy, skill, model routing, evidence, decision, human wait states, replay, and learning.
EIOR API surface /api/eior/* Exposes runtime contract, runtime/model/skill registries, workflow planning, execution, execution status, cancellation, replay, executive ops console, and production certification.
Mission ACT orchestration act_orchestration.py, mission_act.py, and config/act_workflows.yaml Executes deterministic mission actions from YAML workflow definitions with role checks, policy hooks, workflow/step persistence, evidence creation, canonical events, and SOR action records.
Workflow persistence WorkflowRunRecord, WorkflowStepRecord, WorkflowPolicyResolutionRecord, FlowTraceRecord, ExecutionLedger, and SorActionRecord Mission ACT and universal action execution record workflow state, step state, evidence hashes, policy resolution, flow trace, ledgers, and controlled source-system writeback disposition.
Event and failure handling enterprise_event_runtime.py, governance_outbox.py, and iaf_event_dead_letters Event runtime validates, routes, observes, and dead-letters events; governance outbox retries with exponential backoff and records retry/dead-letter metrics.
Certification tests/test_runtime_cert_orchestration_and_evidence_gap_closure.py Validates measured orchestration certification, production gates, bounded attestation reads, hub-spoke collaboration disclosure, replay graph, policy fail-closed behavior, and ops console output.
Operational surfaces Active Missions, Mission Command Center, Observability, and Replay Theater Users can inspect mission state, orchestration timelines, operational telemetry, and replay lineage across runtime steps.

Runtime Scope

  • Trigger handling. Accepts event, user, schedule, API, source-system, agent, policy, evidence, and control triggers through EIOR, Mission ACT, and Event Runtime entry points.
  • Workflow planning. Creates versioned execution plans with runtime selection, model route, skill package, context envelope, evidence envelope, multi-agent plan, human boundary, and replay contract.
  • Cross-runtime coordination. Sequences Context, Policy, Skill, Model Gateway, Evidence, Decision, Human Collaboration, Replay, Memory, Learning, Event, and Integration concerns.
  • Mission action execution. Runs canonical ACT workflows from YAML definitions for P2P, O2C, R2R, FP&A, Treasury, and Retail missions.
  • Policy and decision gates. Blocks, waits, or proceeds based on Policy Runtime decisions and action-disposition checks before controlled execution.
  • Human wait state. Creates explicit wait states for approval, evidence, or supervisor review when policy or risk requires human authority.
  • State and lineage. Stores workflow runs, workflow steps, policy resolutions, evidence hashes, flow traces, execution ledgers, SOR actions, plan hashes, execution hashes, and replay hashes.
  • Failure control. Uses bounded certification probes, retry policies, cancellation, event dead letters, governance outbox backoff, and compensation hooks for safe failure modes.
  • Observability. Publishes metric events, executive console summaries, step latency, runtime count, human wait status, cost, policy blocks, consensus, and replay completeness.

Logical Architecture

Event / User / Schedule / API / Source-System Change
        |
        v
Mission or Workflow Trigger
        |
        v
Enterprise Orchestration Runtime
        |
        +-- Trigger Router
        +-- Runtime and Skill Registry
        +-- Workflow Plan Builder
        +-- Context Envelope Builder
        +-- Evidence Envelope Binder
        +-- Skill and Model Route Selector
        +-- Runtime Invocation Coordinator
        +-- Policy Gate Coordinator
        +-- Decision Gate Coordinator
        +-- Human Wait State Coordinator
        +-- Mission ACT Workflow Executor
        +-- Integration Action Coordinator
        +-- Retry / Cancel / Dead-Letter Handler
        +-- Orchestration Ledger
        +-- Replay Recorder
        +-- Observability Publisher
        |
        v
Completed / Waiting For Human / Waiting For Evidence / Blocked / Cancelled / Failed Safe / Replayed

Implemented Lifecycle

  1. Receive trigger. A mission, event, API call, user action, or workflow action creates an orchestration request.
  2. Plan execution. EIOR creates a plan ID, runtime selection list, context envelope, evidence envelope, skill package, model route, collaboration contract, human boundary, and replay contract.
  3. Assemble context. Context is built or referenced before downstream evidence, skill, policy, agent, and decision steps use the case state.
  4. Evaluate policy. Policy Runtime evaluates authority, evidence status, AI guardrails, proposed action, and write boundary.
  5. Select skill and model route. Skill Fabric compiles the required skill context and model routing is selected by task type, risk, classification, write requirement, and latency target.
  6. Validate evidence. Evidence envelope and evidence references are bound to the plan and later to decision memory and replay.
  7. Coordinate collaboration. The runtime builds shared context and sibling-agent policy before dispatching or blocking sibling collaboration.
  8. Evaluate decision. Decision Runtime receives the governed recommendation boundary and determines next best action or approval requirement.
  9. Wait for human if required. High-risk or write-bound actions enter human wait state instead of executing directly.
  10. Coordinate controlled action. Mission ACT or universal action execution applies policy proof, evidence hash, ledger record, SOR action record, and writeback disposition for controlled actions.
  11. Record replay and learning. Plan, execution, policy, step hashes, evidence references, decision memory, learning signal, and timeline replay are persisted or returned.
  12. Cancel, retry, or dead-letter. Executions can be cancelled; outbox and event runtime manage retries and dead-letter conditions when safe continuation is not possible.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/eior/contractReturn runtime responsibility, owned capabilities, separation of duties, and certification endpoints.Runtime definition.
GET /api/eior/registries/runtimesReturn participating runtime registry, owners, SLAs, capabilities, APIs, and dependencies.Runtime selection.
GET /api/eior/registries/modelsReturn approved deterministic, frontier, sovereign/private, and local model routes.Model routing.
GET /api/eior/registries/skillsReturn mission-scoped skill contracts available for orchestration.Skill selection.
POST /api/eior/workflows/planCreate an execution plan with context envelope, evidence envelope, skill package, model route, steps, human boundary, and replay contract.Workflow planning.
POST /api/eior/workflows/executeExecute planned steps, evaluate policy, create step outputs, collaboration state, decision memory, learning signal, observability, and replay hash.Cross-runtime execution.
GET /api/eior/executions/{execution_id}Return latest persisted plan, execution, or cancellation state with attestation hash.Status.
POST /api/eior/executions/{execution_id}/cancelCancel an execution and persist cancellation attestation.Safe termination.
POST /api/eior/executions/{execution_id}/replayReconstruct plan and execution timeline from persisted attestations.Replay.
GET /api/eior/executive-ops-consoleAggregate recent plan/execution/cancel history, compliance rate, human wait count, latency, cost, and consensus rate.Operations.
GET /api/eior/production-certificationRun bounded production gates against a real plan execution contract.Release gate.
POST /api/iaf/actions/simulateSimulate universal action effects without mutating source-system state.Action simulation.
POST /api/iaf/actions/executeExecute a governed action with policy proof, evidence hash, ledger, and writeback disposition.Controlled action.
POST /api/iaf/missions/{mission_id}/entity/{entity_id}/actRun a YAML-defined ACT workflow for a mission/entity/action.Mission action workflow.
GET /api/iaf/missions/{mission_id}/act/workflowsList configured mission ACT workflows and allowed roles.Workflow registry.
GET /api/iaf/missions/{mission_id}/act/historyReturn recent mission workflow runs with actor, role, status, timestamps, and evidence hash.Workflow history.
GET /api/iaf/missions/{mission_id}/act/{workflow_id}/statusReturn persisted workflow run and step state.Workflow state.

Orchestration Patterns

PatternImplemented FormControl Point
Sequential orchestrationEIOR step graph: context, policy, skill, model, evidence, decision, human wait, controlled execution, replay, learning.Plan hash and step output hashes.
Deterministic ACT workflowconfig/act_workflows.yaml defines mission/action steps such as permission validation, policy compliance, journal posting, event emission, evidence bundle, and action log.Workflow run and step records.
Policy-gated orchestrationPolicyAsCodeRuntime.evaluate can produce DENY, REQUIRES_MORE_EVIDENCE, REQUIRES_HUMAN_APPROVAL, or allow execution path.Sibling dispatch and controlled execution boundary.
Human-in-the-loop orchestrationHigh-risk or write-bound actions enter human_wait_state or Mission ACT approval queues rather than executing directly.Human boundary and role authority.
Evidence-gated orchestrationEvidence envelope, evidence hash, and evidence bundle creation determine whether recommendation, approval, or execution can proceed.Evidence pack and replay lineage.
Event-driven orchestrationEvent Runtime handles event validation, source metadata, route decisions, workflow refs, dead letters, and metrics.Event fabric and dead-letter records.
Retry and compensationStep retry policy is declared in EIOR; governance outbox retries with exponential backoff; spine engine records compensation completion.Retry counters, dead-letter state, compensation result.
Replay-oriented orchestrationEIOR replay returns timeline records and replay hash; ACT records evidence hash, workflow history, ledger, and step status.Replay contract and attestation hash.

Mission Outcomes

OutcomeMeaningImplemented Signal
completedGoverned plan or workflow completed.EIOR status or WorkflowRunRecord status.
waiting_for_humanPolicy, risk, or write boundary requires approval/review.EIOR final status and human wait metric.
waiting_for_evidencePolicy requires additional evidence before continuation.EIOR policy decision REQUIRES_MORE_EVIDENCE.
blockedPolicy denied or blocked sibling dispatch / controlled action.EIOR sibling-agent policy and policy block metric.
cancelledUser or system cancelled execution.eior_cancel attestation with previous status and cancel hash.
failed_safeRuntime cannot proceed safely.Workflow error state, outbox dead-letter, event dead-letter, or bounded degraded certification.
simulatedAction impact calculated without source-system mutation./api/iaf/actions/simulate response with simulated flag and impact fields.

Integration With Other Runtimes

Context Runtime

EIOR builds or references governed context before skills, evidence, policy, and decisions use case state.

Evidence Runtime

Evidence envelopes, evidence hashes, evidence bundles, and evidence refs travel through plan, action, decision memory, and replay.

Skill Runtime

Skill registry and compiled skill packages become orchestration steps rather than embedded workflow logic.

Agent Runtime

Agents participate through shared context and sibling dispatch policy; EIOR remains the process control plane.

Policy Runtime

Policy decisions gate sibling dispatch, evidence wait, human wait, and controlled source-system action.

Decision Runtime

Decision Runtime owns business recommendation; Orchestration Runtime owns the process path around that decision.

Event Runtime

Validated events, workflow refs, metrics, retries, and dead letters connect mission triggers and downstream operations.

Integration Runtime

Source-system writeback remains behind policy proof, evidence hash, execution ledger, SOR action record, and writeback disposition.

Example: Duplicate Payment Orchestration

A high-risk P2P duplicate payment case runs through orchestration as a plan rather than disconnected page calls. The runtime coordinates context, policy, skill, model route, evidence, decision, human wait, replay, and learning while preserving a controlled action boundary.

{
  "runtime": "enterprise_intelligence_orchestration_runtime",
  "case_id": "INV-LIVE-DUP-1778594510",
  "mission": "p2p",
  "proposed_action": "hold_payment",
  "control_plane_boundary": "EIOR orchestrates; Decision Runtime decides; Policy Runtime governs",
  "runtime_selection": [
    "enterprise_context_runtime",
    "enterprise_policy_intelligence_runtime",
    "enterprise_skill_fabric",
    "model_orchestration_gateway",
    "evidence_intelligence_runtime",
    "enterprise_decision_runtime",
    "human_collaboration_runtime",
    "enterprise_intelligence_orchestration_runtime",
    "enterprise_memory_learning_runtime"
  ],
  "human_boundary": "human_approval_required_for_write_or_policy_exception",
  "replay_contract": "cross_runtime_replay_v1"
}

Operational Requirements

  • Mission orchestration must live in backend runtime services, not in UI click handlers.
  • Agents may propose next steps, but mission state and process control remain owned by Orchestration Runtime.
  • Human approvals, evidence requests, policy blocks, cancellations, retries, and failed-safe states must be explicit workflow states.
  • System writeback must pass through policy proof, evidence hash, decision/action context, ledger, and SOR action recording.
  • Execution plans, workflow definitions, action dispositions, policy decisions, and replay records must be versioned or hash-bound enough for reconstruction.

Acceptance Criteria

  • Every material mission flow has an orchestration entry point, execution plan, or deterministic ACT workflow definition.
  • Workflow state is backend-owned; the UI reads state and timelines rather than coordinating sensitive steps locally.
  • Policy decisions can block, pause for evidence, or create human wait states before sibling dispatch or controlled execution.
  • Controlled source-system action records evidence hash, ledger ID, policy ID/version, simulated flag, writeback disposition, queue state, and downstream effects.
  • Workflow runs and steps are persisted with status, run ID, action type, triggering actor, evidence hash, and step outputs.
  • Event retries and dead-letter behavior are explicit through Event Runtime and governance outbox mechanisms.
  • Orchestration replay can reconstruct plan and execution timeline through attestation records, hashes, execution graph, evidence references, policy decision, and replay contract.
  • Production certification reports measured gaps honestly and does not claim emergent multi-agent orchestration when the implementation is hub-spoke coordination.

Engineering Rule

Orchestration is not a UI responsibility and not an agent responsibility. The platform must start, plan, gate, pause, resume, cancel, observe, and replay mission work through backend orchestration services so context, evidence, skills, policy, decisions, humans, events, and source-system actions remain coordinated and auditable.

CapabilitiesTrust Fabric + ReplayPolicy + Authority ControlEvent Fabric + SOR SidecarEnterprise Memory + Knowledge Management
Implemented Runtime Chapter

Runtime 8: Enterprise Replay Runtime

The Enterprise Replay Runtime is the accountability and reconstruction layer for Sphere. It reconstructs how cases, decisions, evidence packs, policy checks, agent/tool activity, skill outputs, approvals, orchestration steps, and source-system actions occurred.

Replay Manifest Timeline Builder Role-Scoped Replay Hash Integrity Decision Comparison Audit Pack Export

Production Role

Replay is not raw logging. It is structured reconstruction using runtime artifact IDs, persisted replay manifests, replay runs, replay steps, policy attestations, evidence packs, decision logs, tool invocations, approval contracts, workflow records, telemetry, hashes, and source lineage.

The runtime does not make new decisions or alter historical execution state. It reconstructs what happened, validates integrity where hashes and manifests are available, projects the result by role, and supports audit, control testing, incident review, technical debugging, and continuous improvement.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Forensic replay data model models/_replay.py Defines ReplayManifest, ReplayRun, ReplayStep, ReplayDiff, ReplayCertification, and SORStateComparison with indexed execution, decision, evidence-pack, entity, status, and timestamp fields.
Decision Replay service decision_replay_service.py Provides readiness, execution explorer, execution detail, determinism health, stored replay, evidence lineage, audit pack export, deep execution drill, execution graph, risk panel, and chaos replay simulation.
Role-scoped replay role_scoped_replay_service.py and replay_timeline_builder.py Projects replay differently for analyst, supervisor, auditor, and CFO roles; assembles timelines across policy attestations, policy evaluations, evidence packs, artifacts, violations, decision logs, and supervisory events.
Decision Runtime replay enterprise_decision_runtime.py and /api/iaf/enterprise-decision-runtime/replay* Persists and reconstructs live decision replay envelopes with decision log, evidence pack, policy attestation, telemetry, replay manifest/run/steps, approval contract, tool invocations, and comparison output.
Control-plane replay /api/iaf/control-plane/replay/{{decision_id}}, /timeline, and signed evidence export Returns role-scoped replay, deterministic timeline, and signed evidence bundle export for governed review.
Event and orchestration replay /api/eior/executions/{{execution_id}}/replay, /api/eer/events/replay, and /api/fef/events/replay Reconstructs orchestration timelines and event timelines with replay hashes, filtered replay, partial replay, and event lineage.
User surfaces Decision Replay Studio, Replay Theater, Replay Center, and decision drawer replay tabs Operators can inspect replay readiness, executions, evidence lineage, graph view, risk panel, audit packs, timeline, and deterministic comparison surfaces.

Runtime Scope

  • Case and decision replay. Reconstructs decision logs, policy attestations, evidence packs, tool invocations, approval contracts, telemetry, and replay manifests for material decisions.
  • Agent and tool replay. Uses agent execution records, tool invocation records, model/tool metadata, status, latency, tokens, SOR calls, evidence hashes, and decision references to inspect agent behavior.
  • Skill replay. Links skill execution records, skill IDs, versions, input/output hashes, evidence hashes, and replay IDs through Skill Runtime and Decision Replay Studio surfaces.
  • Policy replay. Replays policy attestations, policy evaluation logs, policy set hashes, attestation hashes, violations, matched conditions, and triggered actions.
  • Evidence replay. Reconstructs evidence packs, artifacts, evidence hashes, inputs/outputs hashes, replayable flags, artifact counts, lineage, and signed evidence exports.
  • Orchestration replay. Replays EIOR plan/execution/cancel attestations and ACT workflow run/step state where mission flows have been executed.
  • Comparative replay. Compares decision-time replay against current deterministic replay and records allowed versus material differences.
  • Role-scoped projection. Returns different replay detail for analyst, supervisor, auditor, and CFO roles, including hash redaction and artifact/governance filtering.
  • Audit pack export. Builds signed evidence bundles and exportable audit packs from replay payloads, evidence, lineage, and decision contracts.
  • Resilience. Supports partial replay, missing artifact reporting, degraded readiness responses, bounded reads, and safe 404/not-found responses when live replay artifacts do not exist.

Logical Architecture

User / Auditor / Supervisor / Platform Team / API
        |
        v
Replay Request
        |
        v
Enterprise Replay Runtime
        |
        +-- Runtime Auth and Role Scope
        +-- Replay Manifest Index
        +-- Replay Run and Step Store
        +-- Decision Log Adapter
        +-- Evidence Pack Adapter
        +-- Policy Attestation Adapter
        +-- Tool Invocation Adapter
        +-- Approval Contract Adapter
        +-- Supervisory Timeline Adapter
        +-- Orchestration Replay Adapter
        +-- Event Replay Adapter
        +-- Hash and Integrity Validator
        +-- Replay Comparison Engine
        +-- Signed Evidence Export
        +-- Decision Replay Studio API
        |
        v
Replay Summary / Timeline / Technical Trace / Comparison / Audit Pack / Signed Evidence Bundle

Implemented Lifecycle

  1. Request replay. User, API, control plane, or Decision Replay Studio requests replay by execution ID or decision ID.
  2. Authorize and scope. Runtime auth and role-scoped replay determine whether hashes, artifacts, policy trace, and governance details are included.
  3. Resolve trace anchors. Replay services resolve decision log, evidence pack, policy attestation, replay manifest, replay run, replay steps, telemetry, approval contract, and tool invocation records.
  4. Build timeline. Timeline builder orders policy attestations, policy evaluations, evidence pack sealing, evidence artifacts, violations, decision log, supervisory events, and queue items.
  5. Validate integrity. Replay manifests, hashes, evidence hashes, inputs/outputs hashes, policy set hashes, replay pointers, and integrity hashes establish tamper-evident reconstruction.
  6. Generate technical trace. Decision Replay service returns execution detail, evidence lineage, graph nodes/edges, deep drawer payload, risk panel, and chaos replay simulation.
  7. Generate business projection. Role-scoped replay returns a concise summary, timeline, policy block, evidence block, and governance block appropriate to the role.
  8. Compare state. Live replay comparison contrasts the immutable replay snapshot against current deterministic replay and classifies material versus allowed differences.
  9. Export evidence. Control Plane export signs a role-scoped evidence bundle with replay payload for governed review.
  10. Record replay run. Where Decision Runtime creates live records, replay manifests, replay runs, replay steps, and certifications store the replay invocation and result.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/iaf/decision-replay/readinessReturn replay readiness, coverage, seal rate, deterministic percentage, replay success, and impact strip.Readiness.
GET /api/iaf/decision-replay/executionsList replayable executions with filters for agent, search, and evidence.Execution explorer.
GET /api/iaf/decision-replay/executions/{execution_id}Return execution detail, linked evidence pack, and determinism contract.Execution detail.
GET /api/iaf/decision-replay/determinism-healthReturn determinism health segmented by agent, tower, or autonomy level.Determinism.
GET /api/iaf/decision-replay/replay/{execution_id}Replay stored evidence in historical, shadow, deterministic, current-policy, or scenario mode.Replay execution.
GET /api/iaf/decision-replay/evidence-lineage/{execution_id}Return evidence lineage, agent/policy version history, and lineage integrity.Evidence replay.
GET /api/iaf/decision-replay/audit-pack/{execution_id}Return exportable decision audit pack with execution, replay, lineage, evidence, and contract.Audit pack.
GET /api/iaf/decision-replay/execution-deep/{execution_id}Return deep drill payload for decision, policy, evidence, value, and timeline views.Technical trace.
GET /api/iaf/decision-replay/execution-graph/{execution_id}Return decision-context graph nodes and edges.Graph replay.
GET /api/iaf/decision-replay/risk-panelReturn high-risk, escalations, SLA breaches, and override replay posture.Risk replay.
GET /api/iaf/decision-replay/chaos-replay/{execution_id}Run what-if replay with changed tolerance, risk tier, or SOR delay.Scenario comparison.
GET /api/iaf/enterprise-decision-runtime/replay/{decision_id}Return deterministic Decision Runtime replay.Decision replay.
GET /api/iaf/enterprise-decision-runtime/replay-live/{decision_id}Return persisted live replay envelope for a decision.Live reconstruction.
GET /api/iaf/enterprise-decision-runtime/replay-live/{decision_id}/compareCompare persisted decision-time replay with current deterministic replay.Comparative replay.
GET /api/iaf/control-plane/replay/{decision_id}Return role-scoped replay for a decision.Role projection.
GET /api/iaf/control-plane/replay/{decision_id}/timelineReturn full policy/evidence/supervision timeline.Timeline.
POST /api/iaf/control-plane/evidence/exportExport signed evidence bundle projected to a role.Signed audit export.
POST /api/eior/executions/{execution_id}/replayReplay orchestration plan/execution timeline from EIOR attestations.Orchestration replay.

Replay Artifact Model

ArtifactStored InReplay Function
ReplayManifestiaf_replay_manifestsBinds execution, decision, evidence pack, agent, entity, policy version, execution mode, and version bindings.
ReplayRuniaf_replay_runsStores replay invocation, mode, requested context, status, certification, classification, warnings, and diffs.
ReplayStepiaf_replay_stepsRecords discrete replay stages such as immutable replay snapshot and runtime boundary checks.
ReplayDiffiaf_replay_diffsCaptures differences between original and replayed/current values with severity and category.
ReplayCertificationiaf_replay_certificationsRecords deterministic verification and replay certification result.
SORStateComparisoniaf_sor_state_comparisonsStores decision-time versus current source-state hashes and comparison payload.

Replay Types

Replay TypeImplemented SourceWhat It Reconstructs
Decision replayDecision Runtime and Decision Replay APIDecision log, outcome, confidence, policy attestation, evidence, approval, telemetry, and replay pointer.
Evidence replayEvidencePack, EvidenceArtifact, signed evidence exportEvidence hashes, artifact lineage, replayable flag, inputs/outputs hashes, and pack metadata.
Policy replayPolicyAttestationRecord, PolicyEvaluationLog, PolicyViolationPolicy set hash, attestation hash, evaluation results, matched conditions, triggered actions, and violations.
Agent/tool replayAgentExecution and ToolInvocation recordsAgent identity, task, execution mode, tool names, skills, status, duration, tokens, SOR calls, and decision refs.
Skill replaySkillExecutionRecord and Skill Fabric execution historySkill run, input/output hashes, evidence hash, status, selection metadata, execution metadata, and replay ID.
Orchestration replayEIOR attestations and Mission ACT workflow recordsPlan, execution graph, step states, human waits, cancellations, workflow runs, workflow steps, and evidence hashes.
Event replayEnterprise Event Runtime and Financial Event FabricBusiness event timelines, filtered replay, partial replay, event hashes, and event lineage.
Comparative replayDecision Runtime live compare and SOR state comparison modelDecision-time state versus current deterministic replay or current source-state representation.

Role-Scoped Replay

RoleIncluded DetailRestricted Detail
AnalystTimeline, artifacts, policy trace, decision status, outcome.Hashes and governance internals are reduced.
SupervisorTimeline, artifacts, hashes, policy trace, governance, evidence, and outcome.Only source entitlement restrictions still apply.
AuditorTimeline, hashes, artifacts, policy trace, governance, evidence, signed export.Sensitive fields remain governed by evidence and identity policies.
CFOShort timeline and summary-level status.Artifacts, hashes, policy trace, and governance internals are compressed.

Example: Live Decision Replay Envelope

Decision Runtime live replay reconstructs the persisted envelope for a decision and computes an integrity hash over the replay payload.

{
  "runtime": "enterprise_decision_runtime",
  "found": true,
  "decision_log": {
    "decision_status": "APPROVAL_REQUIRED",
    "outcome": "hold_recommended",
    "replay_pointer": "sha256:..."
  },
  "evidence_pack": {
    "evidence_hash": "sha256:...",
    "is_replayable": true
  },
  "policy_attestation": {
    "decision": "REQUIRES_HUMAN_APPROVAL",
    "attestation_hash": "sha256:..."
  },
  "replay_manifest": {
    "execution_mode": "deterministic_first",
    "version_bindings": {"policy_version": "runtime-bound"}
  },
  "tool_invocations": [
    {"tool_name": "sap_read_invoice", "status": "completed", "sor_calls": 1}
  ],
  "integrity_hash": "sha256:..."
}

Operational Requirements

  • Replay artifacts must be structured records, not only log lines.
  • Runtime boundaries must preserve artifact IDs for context, evidence, policy, decision, agent, skill, orchestration, integration, and human approval records where those records exist.
  • Material artifacts should carry hashes or replay pointers so integrity can be validated later.
  • Replay views must be role-aware and purpose-aware; sensitive hashes, artifacts, prompts, and governance details are not universally exposed.
  • Replay comparison must identify source-state or runtime-enrichment differences instead of pretending current state is identical to decision-time state.
  • Failed and blocked paths must remain replayable; replay cannot be limited to successful decisions.

Acceptance Criteria

  • Every material Decision Runtime execution can create or reference ReplayManifest, ReplayRun, and ReplayStep records.
  • Replay payloads link decision log, evidence pack, policy attestation, telemetry, approval contract, tool invocations, and original replay snapshot when available.
  • Role-scoped replay hides hashes, artifacts, policy trace, or governance detail when the role configuration requires redaction.
  • Timeline replay includes policy, evidence, violation, decision, and supervisory events in chronological order.
  • Live replay comparison distinguishes material differences from allowed runtime enrichment differences.
  • Replay readiness and determinism health are available through Decision Replay API and UI surfaces.
  • Audit pack and signed evidence export are available for governed review and verification.
  • Partial replay, missing artifacts, not-found decisions, and degraded readiness states fail safely without fabricating replay completeness.

Engineering Rule

Any runtime that creates a material recommendation, decision, approval, source-system action, policy attestation, evidence pack, skill result, tool call, or orchestration step must emit enough identifiers, version bindings, hashes, and timestamps for Replay Runtime to reconstruct the event later. If the replay artifact is incomplete, the UI and APIs must say partial replay rather than imply full reconstruction.

CapabilitiesEnterprise Memory + Knowledge ManagementTrust Fabric + ReplayPolicy + Authority ControlValue Attribution Engine
Implemented Runtime Chapter

Runtime 9: Enterprise Learning Runtime

The Enterprise Learning Runtime is the governed improvement layer for Sphere. It captures outcomes, feedback, overrides, evidence gaps, policy friction, skill performance, agent quality, and process patterns, then turns them into validated, approved, replay-linked improvement candidates.

Learning Signal Enterprise Memory Closed Loop Promotion Gate Skill Overlay Rollback Governance

Production Role

The implemented runtime is named Enterprise Memory & Learning Runtime. It combines memory storage, learning outcomes, learning proposals, governed promotion, semantic retrieval, knowledge graph, organizational learning, and executive learning observability under one runtime contract.

The runtime is deliberately not uncontrolled AI memory. It can recommend knowledge, policy, skill, routing, workflow, and process improvements, but production behavior changes require validation, review, explicit promotion, versioned metadata, rollback posture, and replayable evidence.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_memory_learning_runtime.py Implements Enterprise Memory & Learning Runtime for governed memory, learning outcomes, proposal generation, closed-loop learning, semantic retrieval, knowledge graph, flywheel, executive console, observability, and certification.
API surface /api/emlr/* Exposes contract, memory store/search/semantic-search, similar cases, learning outcome, closed-loop learning, proposal generation, learning engine, knowledge graph, personalized insights, flywheel, executive console, certification, forget, promote, and KPIs.
Learning records AgentLearningFeedback, PolicyEffectivenessSnapshot, BehaviorDriftSnapshot, and OrganizationalChangeRecord Persists governed feedback, policy effectiveness, behavior drift, and enterprise change candidates with evidence hash, replay reference, owner, approval state, and value-claim boundary.
Memory and outcome records EnterpriseMemoryEntry, CrestLearningOutcome, SkillLearningOverlay, and VNextImpactLearningEntry Stores decision memory, learning outcomes, skill overlays, and impact-learning entries that can be searched, reviewed, promoted, graphed, and replay-linked.
Promotion governance learning_promotion_service.py and skill_learning_overlay_runtime.py Validates evidence, confidence, target policy, non-contradiction, runtime effect, and approval state before promotion; blocks unvalidated learning from mutating runtime metadata.
Feedback contract learning_feedback_runtime.py and config/contracts/learning_feedback_contract.yaml Defines feedback and drift dimensions used to normalize learning payloads and drift records.
Tenant learning governance config/tenants/bp/learning_governance.yaml Configures BP staged rollout, dual approval threshold, rollback triggers, evidence outbox emission, retention, actor logging, rationale requirement, and learning evidence contract.
Certification tests/test_enterprise_memory_learning_runtime_certification.py Verifies measured certification, closed-loop learning validation, no auto-mutation, governed promotion pipeline, semantic search, knowledge graph, flywheel, executive console, and production gates.
Operational surfaces Enterprise Learning Runtime, Decision Memory, Learning Fabric, Learning Proposals, and Institutional Memory Users can inspect learning candidates, memory records, replay-backed patterns, promotion posture, and institutional memory.

Runtime Scope

  • Signal capture. Captures outcomes, human feedback, overrides, evidence hashes, policy references, agent performance, source cases, before/after accuracy, and value outcomes through EMLR and Agent Runtime feedback paths.
  • Memory storage. Stores governed experience in enterprise memory entries with mission, domain, entity, decision reference, policy reference, evidence reference, actor, trust level, approval state, classification, tags, and content hash.
  • Learning outcome creation. Creates CrestLearningOutcome records with decision reference, proposal type, confidence, source payload, evidence hash, proposed status, and promotion reference.
  • Learning engine. Clusters repeated patterns from decisions and memory records, estimates value, identifies affected policies/missions, and produces improvement candidates.
  • Closed-loop validation. Runs learning outcome creation, validation, promotion pipeline state, evidence check, outcome measurement, and auto-mutation block in one governed flow.
  • Promotion governance. Requires validation and explicit request before applying learning to policy metadata or skill overlay metadata.
  • Skill overlay learning. Supports proposed, reviewed, validated, and promoted skill learning overlays with evidence and runtime-effect checks.
  • Memory governance. Supports forgetting/retiring memory and promoting memory to enterprise-approved, certified trust state.
  • Retrieval and reuse. Provides keyword/semantic hybrid search, local 256-dim hash-embedding retrieval (pgvector-swappable), similar case lookup, personalized insights, and an in-memory runtime knowledge graph.
  • Impact and observability. Publishes learning metric events, executive learning console, policy effectiveness snapshots, organization change records, KPIs, and production certification gates.

Logical Architecture

Runtime Outcomes / Human Feedback / Replay / Evidence / Policy / Agent / Skill / Workflow
        |
        v
Learning Signal or Memory Capture
        |
        v
Enterprise Learning Runtime
        |
        +-- Memory Store
        +-- Learning Outcome Store
        +-- Feedback Contract Adapter
        +-- Learning Engine
        +-- Pattern Clusterer
        +-- Proposal Generator
        +-- Validation Gate
        +-- Promotion Pipeline
        +-- Skill Overlay Governance
        +-- Policy Metadata Promotion
        +-- Knowledge Graph Builder
        +-- Semantic Retrieval
        +-- Personalized Insight Engine
        +-- Executive Learning Console
        +-- Observability and Certification
        +-- Rollback Governance
        |
        v
Approved Memory / Learning Proposal / Skill Overlay / Policy Candidate / Organizational Change Record / Learning Metrics

Implemented Lifecycle

  1. Capture outcome. A decision, case, agent run, supervisor action, skill result, audit finding, or process outcome emits a learning payload.
  2. Store memory or proposal. EMLR stores the experience as enterprise memory or creates a CrestLearningOutcome proposal.
  3. Bind evidence and provenance. Payload hash, evidence hash, decision reference, policy ID, source cases, actor, and mission/domain context are preserved.
  4. Classify learning. Proposal type identifies whether the signal targets policy, skill, knowledge, routing, process, or workflow improvement.
  5. Run learning engine. Repeated observations are clustered by root cause, exception type, or policy ID and synthesized into lessons and improvement candidates.
  6. Validate candidate. Promotion service checks evidence, target policy, confidence floor, and non-contradiction before a learning outcome can move from proposed to validated.
  7. Route review. Learning remains proposed or validated until the relevant owner reviews it through governance, CREST review, skill overlay review, or memory promotion.
  8. Promote with controls. Approved learning can update policy metadata or skill overlay metadata only through explicit promotion functions.
  9. Record enterprise change. Learning promotions create policy effectiveness snapshots and organizational change records with evidence and replay references.
  10. Measure and monitor. Executive console, flywheel, observability metrics, KPIs, staged rollout, rollback triggers, and production certification monitor post-promotion posture.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/emlr/contractReturn runtime scope, memory/knowledge boundary, APIs, governance, promotion pipeline, and certification posture.Runtime definition.
POST /api/emlr/memory/storeStore governed memory with mission, domain, decision, policy, evidence, classification, trust, tags, and metadata.Memory capture.
GET /api/emlr/memory/searchSearch enterprise memory using semantic-keyword hybrid retrieval.Memory retrieval.
GET /api/emlr/memory/semantic-searchRun local 256-dim hash-embedding memory retrieval (pgvector-swappable; no external vector DB required yet).Semantic retrieval.
GET /api/emlr/cases/{case_id}/similarFind similar cases from memory entries and memory episodes.Case reuse.
POST /api/emlr/learning/outcomeCreate proposed learning outcome and optional agent learning feedback.Signal capture.
POST /api/emlr/learning/closed-loopCreate, validate, optionally promote, and report closed-loop learning with promotion pipeline state.Closed-loop learning.
POST /api/emlr/learning/proposalsGenerate learning proposals from decision observations and detected clusters.Candidate generation.
POST /api/emlr/learning/engineRun deterministic pattern, trend, cluster, and lesson synthesis.Pattern detection.
GET /api/emlr/knowledge-graphMaterialize memory, mission, domain, entity, policy, and decision nodes/edges.Knowledge graph.
POST /api/emlr/personalized-insightsReturn role-specific memory recommendations for analyst, supervisor, controller, CFO, auditor, developer, or executive roles.Personalization.
GET /api/emlr/flywheelReturn enterprise intelligence flywheel inputs, learning conversion, outputs, cross-mission reuse, and graph summary.Learning flywheel.
GET /api/emlr/executive-consoleReturn learning maturity, recurring issues, policy candidates, automation opportunities, and self-improvement queue.Executive learning console.
GET /api/emlr/observability-contractReturn metrics, alerts, dashboards, and runbooks for memory and learning operations.Observability.
GET /api/emlr/production-certificationRun production gates for memory, closed-loop learning, promotion pipeline, flywheel, graph, personalization, and observability.Release gate.
POST /api/emlr/memory/{memory_id}/forgetRetire a memory record with actor and reason.Memory governance.
POST /api/emlr/memory/{memory_id}/promotePromote memory to approved, certified enterprise memory with promotion hash.Memory promotion.
POST /api/iaf/decision-intelligence/skill-learning-overlays/{overlay_id}/validateValidate a skill learning overlay before promotion.Skill learning gate.
POST /api/iaf/decision-intelligence/skill-learning-overlays/{overlay_id}/promotePromote an approved skill learning overlay into runtime metadata.Skill learning promotion.
POST /api/iaf/skills/overlaysCreate a proposed Skill Fabric learning overlay.Overlay proposal.
GET /api/iaf/skills/overlaysList skill learning overlays by skill or status.Overlay registry.
POST /api/iaf/skills/overlays/{overlay_id}/reviewApprove, reject, or retire a skill learning overlay.Overlay review.
POST /api/enterprise-cognitive-os/learning-signalQualify a learning signal from decision artifact and outcome.Cognitive OS signal.
POST /crest/learning/proposeCreate a CREST learning outcome proposal.CREST learning proposal.
POST /crest/learning/{outcome_id}/reviewUpdate CREST learning outcome review state.CREST review.
GET /crest/learning/outcomesList CREST learning outcomes by mission and status.Learning outcome registry.

Learning Artifact Model

ArtifactStored InRuntime Function
EnterpriseMemoryEntryiaf_enterprise_memory_entriesGoverned memory with content hash, trust, approval, classification, policy/evidence/decision references, and lifecycle.
CrestLearningOutcomeiaf_crest_learning_outcomesLearning outcome proposal with type, confidence, payload, status, reviewer, and promotion reference.
AgentLearningFeedbackiaf_agent_learning_feedbackAgent feedback linked to decision/execution, policy, outcome value, feedback source, payload, and evidence hash.
SkillLearningOverlayiaf_skill_learning_overlaysSkill-targeted learning overlay with proposed payload, confidence, review state, validation, and promotion binding.
PolicyEffectivenessSnapshotiaf_policy_effectiveness_snapshotsPolicy learning measurement with sample count, success rate, override rate, value, and improvement signals.
BehaviorDriftSnapshotiaf_behavior_drift_snapshotsAgent/policy behavior drift score, baseline/current window, indicators, and recommended action.
OrganizationalChangeRecordiaf_organizational_change_recordsAppend-only enterprise change ledger for approved learning promotions and replay/evidence-backed change candidates.

Promotion Pipeline

StageImplemented GateProduction Boundary
candidate_learning_identifiedLearning engine or outcome API creates candidate.No runtime behavior changes.
expert_reviewValidated when confidence or explicit validation supports review.Human review remains required.
business_validationRequires evidence of frequency, value, or measured outcome.Learning remains advisory.
policy_validationPromotion service validates target policy, evidence, confidence, and non-contradiction.Inactive/retired policy targets block promotion.
pilot_deploymentOnly complete after explicit promotion request and validation pass.Staged rollout governed by tenant config.
outcome_measurementMeasures value, accuracy, sample count, override rate, and policy effectiveness.Regression can trigger rollback.
enterprise_rolloutAllowed only after promotion and sufficient confidence.Promotion updates metadata, not uncontrolled model state.
continuous_monitoringExecutive console, KPIs, observability, and governance config monitor impact.Rollback remains available.

Learning Types

Signal TypeImplemented SourceAllowed Runtime Effect
Human feedbackAgentLearningFeedback, Agent Runtime learning feedback, EMLR outcome payload.Guidance, memory, candidate proposal, or owner review queue.
Evidence gapEvidence hash absence, source cases, memory metadata, learning engine clusters.Evidence profile or workflow improvement candidate.
Policy frictionPolicy ID clustering, policy effectiveness snapshots, override rate, repeated proposals.Policy metadata candidate or policy-owner review item.
Skill qualitySkill learning overlays and skill execution/replay metadata.Validated overlay metadata, threshold/routing/test-case candidate.
Agent qualityAgent feedback, recommendation outcomes, before/after accuracy, evidence references.Agent guidance or planning adjustment candidate.
Workflow bottleneckLearning engine clusters over workflow observations and memory episodes.Workflow optimization candidate, not direct mission graph mutation.
Value realizationOutcome value, policy effectiveness, organizational change records, value event linkage.Impact learning and business case for promotion.

Tenant Governance

ControlBP ConfigurationRuntime Effect
ApprovalDefault owner bucket ops_lead; dual approval above impact score 0.50.Behavior-changing learning routes through owners.
Staged rolloutshadow, 10pct, 50pct, 100pct cohorts with KPI guardrails.Promotions do not jump directly to full estate.
RollbackKPI drift threshold, high drift alert, or operator initiation.Reverts to prior calibration and notifies configured roles.
EvidenceOutbox subject prefix iris.bp.learning and 365-day retention.Learning changes are evidence-emitting and retained.
AuditActor logging and rationale required with evidence contract contract.learning.bp.v1.Every applied learning action is attributable.

Example: Closed-Loop Learning

A repeated duplicate-invoice evidence pattern can become a learning outcome. The runtime validates it, blocks automatic mutation, and exposes the promotion pipeline state for owner review.

{
  "operation": "ClosedLoopLearning",
  "learning": {
    "status": "proposed",
    "promotion_allowed_without_approval": false,
    "governance": "proposed_learning_requires_review_before_runtime_effect"
  },
  "validation": {
    "passed": true,
    "checks": {
      "has_evidence": true,
      "has_target_policy": true,
      "confidence_above_floor": true,
      "non_contradiction": true
    }
  },
  "promotion": {
    "attempted": false,
    "promoted": false
  },
  "closed_loop": {
    "outcome_measured": true,
    "business_success_measured": true,
    "auto_mutation_blocked": true
  }
}

Operational Requirements

  • Learning signals must remain observations until validated and reviewed.
  • Promoted learning must target a specific artifact: memory record, policy metadata, skill overlay, routing hint, evidence workflow, test case, or organizational change record.
  • Learning promotion must preserve source cases, evidence hash, policy ID, reviewer, confidence, before/after accuracy, and replay reference when available.
  • Memory can recommend knowledge and policy changes, but it must not silently rewrite authoritative knowledge or production policy rules.
  • Tenant governance controls staged rollout, dual approval, rollback, evidence retention, and audit rationale.
  • Learning failures are non-blocking for core operations unless a policy explicitly requires learning capture.

Acceptance Criteria

  • Learning does not silently change production policy, controls, ontology, source-system actions, or agent authority.
  • Closed-loop learning explicitly reports auto_mutation_blocked=true and promotion_allowed_without_approval=false.
  • Learning outcomes require validation before promotion and can be rejected when evidence, confidence, target policy, or non-contradiction checks fail.
  • Promoted learning creates versioned or metadata-bound runtime effects plus policy effectiveness and organizational change records.
  • Skill learning overlays require approval before promotion and must include evidence, confidence, payload, skill ID, and runtime effect.
  • Enterprise memory entries are searchable, hashable, classified, trust-scored, lifecycle-managed, and approval-state aware.
  • Learning proposals and engine clusters include promotion pipeline state and do not auto-promote.
  • Learning Runtime certification is measured from persisted memory/learning state and structural signals, not a self-issued grade.

Engineering Rule

Do not implement learning as free-form conversation memory or automatic self-modification. Capture structured signals, bind evidence and replay references, validate quality, classify the target artifact, route review, promote through an explicit gate, and measure post-promotion impact with rollback available.

CapabilitiesTrust Fabric + ReplayValue Attribution EngineEvent Fabric + SOR SidecarPolicy + Authority Control
Implemented Runtime Chapter

Runtime 10: Enterprise Observability Runtime

The Enterprise Observability Runtime is the operational intelligence layer for Sphere. It collects, correlates, persists, analyzes, and reports telemetry across platform runtimes, agents, skills, policies, evidence, decisions, business processes, security events, model usage, cost, and value.

Unified Telemetry Distributed Trace Runtime Health AI Quality FinOps Governed Remediation

Production Role

Observability is not only log capture. The implemented runtime defines a unified telemetry model that combines technical metrics, AI/model signals, policy and evidence checkpoints, security events, cost, business impact, replay references, and operational health.

The runtime does not execute business decisions or automatically repair production behavior. It exposes health, traceability, alerts, diagnostics, and recommendation packets. Remediation remains recommend-only until policy and owner approval are captured.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_observability_runtime.py Implements cross-runtime operational, business, AI, policy, evidence, security, performance, value, alerting, remediation, governance, OpenTelemetry, and executive operations observability.
API surface /api/eor/* Exposes contract, telemetry model, telemetry ingest, durable telemetry, persisted telemetry hydration, runtime health, distributed trace, drilldown, domain observability views, alerts, remediation, governance, integrations, and certification.
Durable telemetry store RuntimePerformanceTelemetry / iaf_runtime_performance_telemetry Persists page/persona/mission/case/action latency, token count, model, SAP calls, cache hit, decision ID, evidence pack ID, outcome, error code, metadata JSON, realm, and timestamp.
Unified telemetry contract UNIFIED_TELEMETRY_FIELDS Requires runtime, mission, process, business entity, correlation ID, user/agent ID, policy reference, evidence reference, decision reference, cost, business impact, and timestamp.
Trace model distributed_trace() and drilldown_packet() Correlates spans by correlation ID and returns policy checkpoints, evidence checkpoints, replay identifiers, bottlenecks, executive summary, governance references, and source lineage.
OpenTelemetry bridge open_telemetry_contract() Maps runtime telemetry to resource attributes, span attributes, histogram/counter/gauge metrics, OTLP, Prometheus, structured logs, and W3C trace context.
FinOps integration finops_runtime.py Builds ledger-first AI economics from LLM audit, agent execution, value events, mission budgets, token intelligence, model routing, SAP call intelligence, cache intelligence, and cost-per-decision envelopes.
Alerting assets ops/observability/eor_rules.yaml Defines health, trace coverage, business value telemetry, governance checkpoint coverage, and operational-risk alerts.
Dashboard asset ops/observability/eor_executive_operations_dashboard.json Defines executive panels for platform health, runtime health scores, business value, AI utilization, policy compliance, and operational risk.
Certification tests/test_enterprise_observability_runtime_certification.py Verifies no fake health, live telemetry measurement, hash-backed telemetry, durable persistence, trace correlation, OpenTelemetry mapping, remediation governance, alerts, command center, and production gates.
Operational surfaces Runtime Command Center, Enterprise Observability Runtime, FinOps Command Center, and Platform Observability Users can inspect runtime health, agent quality, cost, traceability, readiness, incidents, and operational proof surfaces.

Runtime Scope

  • Unified telemetry. Normalizes telemetry across infrastructure, application, AI, business, and governance domains into one schema.
  • Runtime health. Scores observed runtimes from actual samples and reports no_samples when no telemetry has been captured.
  • Distributed trace. Links runtime spans through correlation ID and exposes policy, evidence, decision, replay, and bottleneck checkpoints.
  • Business process observability. Aggregates throughput, SLA compliance, automation percentage, business value, exception rate, and cycle time by process.
  • AI observability. Tracks model usage, tokens, cost, human intervention rate, confidence, tool usage, and quality trends.
  • Policy and evidence observability. Measures policy evaluations, failures, compliance score, evidence records, completeness, missing evidence, replay readiness, and audit readiness.
  • Security observability. Counts authentication failures, authorization failures, prompt injection attempts, data access violations, and security posture.
  • Performance and value analytics. Measures latency, queue depth, cache efficiency, cost efficiency, financial impact, runtime cost, ROI, automation rate, and risk reduction.
  • Alerting and remediation. Creates alerts/incidents for errors, latency SLO breaches, and security events, then recommends owner-routed remediation without auto-execution.
  • Operations integration. Publishes OpenTelemetry, Prometheus, Grafana, Jaeger, Loki, Kubernetes, SAP, Databricks, Teams, ServiceNow, Azure Monitor, and CloudWatch integration contracts.

Logical Architecture

All Sphere Runtimes / Agents / Skills / UI / Integrations / Models
        |
        v
Unified Telemetry Record
        |
        v
Enterprise Observability Runtime
        |
        +-- Unified Telemetry Model
        +-- Runtime Health Scorer
        +-- Distributed Trace Correlator
        +-- Drilldown Packet Builder
        +-- Business Process Observability
        +-- AI Observability
        +-- Policy Observability
        +-- Evidence Observability
        +-- Security Observability
        +-- Performance Analytics
        +-- Business Value Analytics
        +-- Alerting and Incident Management
        +-- Governed Remediation Recommendations
        +-- OpenTelemetry Mapping
        +-- Durable Telemetry Store
        +-- Executive Operations Command Center
        |
        v
Runtime Health / Trace / Alert / Drilldown / Cost / Value / Governance / Certification

Implemented Lifecycle

  1. Instrument. Runtimes emit unified telemetry fields and optional latency, error, throughput, queue, token, model, tool, confidence, policy, evidence, security, replay, and status attributes.
  2. Collect. Telemetry can be recorded in memory through /api/eor/telemetry or persisted through /api/eor/telemetry-durable.
  3. Persist. Durable telemetry is written to RuntimePerformanceTelemetry with correlation, trace, policy, evidence, replay, cost, business impact, and telemetry hash metadata.
  4. Hydrate. Persisted telemetry can be loaded back through /api/eor/load-persisted-telemetry without fabricating runtime samples.
  5. Correlate. Events are grouped by correlation ID and exposed as distributed traces with parent-child relationships and governance checkpoints.
  6. Aggregate. Runtime health, process, AI, policy, evidence, security, performance, and business-value views are computed from captured telemetry.
  7. Detect. Latency, error, and security-event thresholds create alerts and high-severity incident records.
  8. Diagnose. Drilldown packets combine touched runtimes, missions, processes, latency, cost, business impact, errors, governance refs, bottlenecks, and source lineage.
  9. Recommend. Remediation recommendations preserve trace, route to owner, require policy/owner approval, and do not auto-execute.
  10. Report. Executive operations command center, dashboard assets, alerting rules, certification, and FinOps views expose health, risk, value, and cost posture.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/eor/contractReturn runtime scope, ownership, observed runtimes, telemetry domains, fields, differentiators, and production boundaries.Runtime definition.
GET /api/eor/telemetry-modelReturn required and optional telemetry fields, correlation model, and governance requirements.Schema contract.
POST /api/eor/telemetryRecord in-memory unified telemetry and evaluate alert thresholds.Telemetry ingest.
POST /api/eor/telemetry-durableRecord telemetry and persist it to RuntimePerformanceTelemetry.Durable ingest.
POST /api/eor/load-persisted-telemetryHydrate runtime state from persisted telemetry records.Durable replay.
GET /api/eor/runtime-healthReturn per-runtime health, availability, latency, error rate, throughput, cost, readiness, and sample status.Health scoring.
GET /api/eor/trace/{correlation_id}Return distributed trace, spans, parent-child relationships, policy/evidence checkpoints, replay identifiers, and bottlenecks.Trace correlation.
GET /api/eor/drilldown/{correlation_id}Return executive summary, governance refs, bottlenecks, and source lineage for a trace.Root-cause packet.
GET /api/eor/business-processesReturn throughput, SLA, automation, business value, exception rate, and cycle time by process.Business telemetry.
GET /api/eor/aiReturn model usage, tokens, cost, human intervention rate, confidence, and AI quality trend.AI telemetry.
GET /api/eor/policyReturn policy evaluation count, failure count, failure rate, latency, and compliance score.Policy telemetry.
GET /api/eor/evidenceReturn evidence count, completeness, missing evidence, replay readiness, audit readiness, and quality trend.Evidence telemetry.
GET /api/eor/securityReturn security events by type and security posture.Security telemetry.
GET /api/eor/performanceReturn latency, queue utilization, cache efficiency, cost efficiency, and scalability signal.Performance analytics.
GET /api/eor/business-valueReturn financial impact, runtime cost, ROI, automation rate, risk reduction, and decision quality.Value analytics.
GET /api/eor/alerts-incidentsReturn alert and incident workflows, thresholds, incidents, and alert inventory.Alerting.
GET /api/eor/remediation-recommendationsReturn governed remediation recommendations and approval requirements.Remediation.
GET /api/eor/executive-command-centerReturn enterprise health, AI utilization, value, automation, policy compliance, risk, cost, adoption, and security posture.Executive ops.
POST /api/eor/emit-runtime-metricsInvoke the runtime metrics emitter and record the emission as observability telemetry.Metric emission.
GET /api/eor/governanceReturn ownership, retention, access control, classification, audit logging, and remediation governance.Governance.
GET /api/eor/opentelemetry-contractReturn OpenTelemetry resource attributes, span attributes, metric types, exporters, and trace context.OTel contract.
GET /api/eor/integrationsReturn supported integrations for OpenTelemetry, Prometheus, Grafana, Jaeger, Loki, Kubernetes, SAP, Databricks, Teams, ServiceNow, Azure Monitor, and CloudWatch.Ops integration.
GET /api/eor/world-class-certificationReturn measured certification based on telemetry coverage and operational visibility.Certification.
GET /api/eor/production-certificationReturn production gates for schema, trace, health, domain views, alerting, governance, durable store, OpenTelemetry, remediation, and drilldown.Release gate.

Unified Telemetry Model

FieldPurposeRequired Boundary
runtime_idIdentifies the runtime that emitted the telemetry.All records.
mission, process, business_entityConnects telemetry to business operations.Business-aware observability.
correlation_id, trace_idLinks spans into a distributed trace.Trace correlation.
user_agent_idIdentifies user, system, or agent actor.Accountability.
policy_ref, evidence_ref, decision_ref, replay_idLinks telemetry to governance artifacts.Control and replay.
cost_usd, business_impact_usdConnects runtime work to cost and value.FinOps and value analytics.
latency_ms, error_count, sla_target_ms, statusSupports health, SLO, alerting, and degraded-mode views.Operations.
tokens, model, tool_name, confidenceSupports AI and agent quality telemetry.AI operations.
evidence_completeness, policy_status, security_eventSupports evidence, policy, and security observability.Governance.

Observed Runtime Set

RuntimeHealth BasisPrimary Signals
Policy IntelligenceTelemetry samples for enterprise_policy_intelligence_runtime.Policy failures, latency, compliance, security events.
DecisionTelemetry samples for enterprise_decision_runtime.Decision latency, value, confidence, policy/evidence references.
ContextTelemetry samples for enterprise_context_runtime.Context latency, cache, dependencies, trace coverage.
EvidenceTelemetry samples for evidence_intelligence_runtime.Evidence completeness, replay readiness, audit readiness.
LearningTelemetry samples for enterprise_memory_learning_runtime.Learning health, value, governance, process signals.
EventTelemetry samples for enterprise_event_runtime.Event throughput, failure rate, correlation state.
AgentTelemetry samples for enterprise_agent_runtime.Model, tokens, tool, confidence, cost, intervention rate.
SkillTelemetry samples for enterprise_skill_runtime.Skill latency, cost, confidence, tool dependency.
KnowledgeTelemetry samples for enterprise_knowledge_runtime.Knowledge trace, retrieval, quality, governance refs.
IntegrationTelemetry samples for enterprise_integration_runtime.Integration latency, SAP/Databricks dependencies, errors.
AssuranceTelemetry samples for enterprise_assurance_runtime.Control assurance, evidence, replay, audit posture.

Alerting Rules

AlertConditionOperational Meaning
EORRuntimeHealthBelowThresholdRuntime health score below readiness threshold.Runtime estate is degraded.
EORTraceCoverageDroppedCross-runtime trace coverage below audit threshold.Root-cause and replay linkage are at risk.
EORBusinessValueTelemetryMissingNo business value telemetry in observability records.Platform value cannot be substantiated.
EORPolicyEvidenceCheckpointMissingPolicy or evidence checkpoints missing from traces.Governance trace completeness is below threshold.
EOROperationalRiskElevatedOpen incident count above zero.Operational risk requires attention.

Example: Telemetry Record

A duplicate-invoice decision path can emit a single unified record that ties operational telemetry to governance artifacts, cost, and business impact.

{
  "runtime_id": "enterprise_decision_runtime",
  "mission": "p2p",
  "process": "duplicate_invoice_resolution",
  "business_entity": "invoice",
  "correlation_id": "corr-eor-001",
  "user_agent_id": "duplicate_invoice_agent",
  "policy_ref": "duplicate_payment_control",
  "evidence_ref": "evidence-pack-eor-001",
  "decision_ref": "decision-eor-001",
  "cost_usd": 0.02,
  "business_impact_usd": 125000,
  "latency_ms": 420,
  "tokens": 1300,
  "model": "governed-model",
  "tool_name": "duplicate_invoice_detection",
  "confidence": 0.95,
  "evidence_completeness": 1.0,
  "replay_id": "replay-eor-001",
  "status": "completed"
}

Operational Requirements

  • Every material runtime path should emit correlation IDs and governance references when available.
  • Telemetry must be redaction-safe and classified as internal operational telemetry unless a stricter policy applies.
  • Runtime health must not show green when telemetry is absent; missing samples are an explicit state.
  • Business value and cost must be tracked together so operators can evaluate cost per decision and cost per value outcome.
  • Remediation must preserve trace, evidence, policy, decision, and owner references before any operational action is taken.
  • Observability signals should feed Learning Runtime as patterns, not direct production mutations.

Acceptance Criteria

  • Runtime health must be honest: when no samples exist, the response reports no_samples instead of fabricated green status.
  • Every telemetry record receives a deterministic telemetry_hash and can carry cost, business impact, policy, evidence, decision, replay, and trace references.
  • Durable telemetry must be stored through RuntimePerformanceTelemetry and reloadable into the runtime.
  • Distributed traces must expose policy checkpoints, evidence checkpoints, replay identifiers, parent-child span relationships, and bottlenecks.
  • OpenTelemetry mapping must support W3C trace context, OTLP, Prometheus, and structured logs.
  • Business, AI, policy, evidence, security, performance, and business-value observability views must be computed from telemetry records.
  • Alerts and remediation recommendations must remain governed and recommend-only until policy and owner approval are captured.
  • Certification must distinguish live telemetry from certification probes and must not pollute empty runtime views.

Engineering Rule

Do not reduce observability to application logs. Every material runtime action should emit structured telemetry with correlation, governance references, latency, status, cost, business impact, and replay linkage where available. If the telemetry is missing, the system must show an explicit gap rather than imply healthy operation.

CapabilitiesEvent Fabric + SOR SidecarPolicy + Authority ControlTrust Fabric + ReplaySkill Fabric + Agentic Runtime
Implemented Runtime Chapter

Runtime 11: Enterprise Integration Runtime

The Enterprise Integration Runtime is the governed connectivity layer for Sphere. It standardizes how runtimes, agents, skills, and workflows discover connectors, read business objects, enforce security, preserve source lineage, register mappings, record connector health, and produce replayable integration traces.

Connector Catalog Security Gateway DataAccessGateway Trace Hash Mapping Registry Health History

Production Role

The runtime prevents direct, unmanaged source-system access. Business-object reads go through Enterprise Security Runtime and DataAccessGateway, and each successful read can be persisted as an integration trace. The current implementation is strongest on governed read access, connector catalog, mapping registry, connector health, source-status classification, and integration traceability.

Write-back is represented as a governed boundary and connector contract, not as unrestricted execution. SAP and ServiceNow write-capable paths are intentionally described as contract or dry-run boundaries until policy, decision, human approval, production adapter, credential, and tenant controls authorize real writes.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_integration_runtime.py Implements the governed connectivity facade over DataAccessGateway, security gateway, connector catalog, integration contracts, mapping registry, health history, durable traces, and executive integration console.
API surface /api/integrations/* Exposes contract, connectors, fabric, API management, transformation, synchronization, governance, error handling, security, AI integration, data governance, observability, marketplace, mappings, connector health, business-object reads, trace search, executive console, runtime interoperability, and certification.
Connector catalog connector_catalog() Combines finance KDO read-model connectors with governed platform connector contracts for SAP S/4HANA, SAP Ariba, Databricks Unity Catalog, Microsoft Graph 365, ServiceNow, OpenAI, Anthropic, Ollama Local, and Kafka/NATS event mesh.
Connectivity fabric connectivity_fabric() Defines the standard connector contract: security-runtime authentication, standard error envelope, bounded retry with circuit breaker, source lineage/evidence reference, runtime telemetry, policy-before-sensitive-action, and replay support through correlation ID and trace hash.
Governed read path read_business_object() Routes reads through Enterprise Security Runtime and DataAccessGateway, then persists a durable integration trace with connector, source status, provenance, policy reference, evidence reference, security decision, record count, and trace hash.
Durable traceability IntegrationAuditTrace / iaf_integration_audit_traces Stores trace ID, connector, business object, action, success, source status, provenance, record count, latency, policy/evidence/security refs, trace hash, metadata, realm, and timestamp.
Mapping registry IntegrationMappingDefinition / iaf_integration_mapping_definitions Persists governed transformation mappings by source system, target system, business object, mapping type, version, owner, status, and definition JSON.
Connector health IntegrationConnectorHealth / iaf_integration_connector_health Stores connector status, source status, latency, error count, freshness status, metadata, realm, and checked timestamp for health history.
Security and AI boundaries security_contract() and ai_integration() Declares authentication, authorization, OAuth/SAML/API key/mTLS/certificate/secret controls, audit logging, model abstraction, provider failover, cost optimization, token tracking, and sovereign AI support.
Certification tests/test_enterprise_integration_runtime_certification.py Verifies connector catalog, standardized fabric, no direct agent external access, policy enforcement, replay support, MCP API style, mapping registry, health history, durable traces, source-status classification, and production gates.
Operational surfaces Data Trust Explorer, Capability Readiness Center, Integration Adapter Library, and Data Integrations Users can inspect source trust, integration readiness, adapter patterns, connector posture, and integration dependencies.

Runtime Scope

  • Connector catalog. Discovers standardized finance, enterprise application, collaboration, data platform, AI platform, sovereign AI, and event infrastructure connectors.
  • Governed read access. Reads business objects through Enterprise Security Runtime and DataAccessGateway rather than allowing pages, agents, or skills to call source systems directly.
  • Write boundary contracts. Represents SAP, ServiceNow, and bidirectional patterns as governed contracts or dry-run boundaries until policy, decision, approval, adapter, and production credentials are available.
  • Source-status classification. Separates read_model, live, test-double, local, contract, blocked, unavailable, and unknown source states.
  • Transformation mapping. Registers and searches source-to-target mapping definitions for canonical object normalization and payload transformation.
  • Synchronization contract. Declares real-time, scheduled, incremental, full, offline, change-data-capture, conflict detection, and governed bidirectional synchronization behavior.
  • Security gateway. Uses Enterprise Security Runtime for authorization and auditing before confidential data access.
  • Integration trace. Persists integration reads and outcomes into IntegrationAuditTrace with replayable trace hashes and governance references.
  • Connector health. Records and searches connector health history for latency, error, freshness, source status, and operational reporting.
  • Marketplace and executive console. Exposes connector discovery, certification, owner, SLA, usage analytics, dependency visualization, executive health, mapping registry, and traceability posture.

Logical Architecture

Context / Evidence / Agent / Skill / Decision / Orchestration Runtime
        |
        v
Integration Request
        |
        v
Enterprise Integration Runtime
        |
        +-- Connector Catalog
        +-- Enterprise Connectivity Fabric
        +-- Security Runtime Gateway
        +-- DataAccessGateway Adapter
        +-- Business Object Reader
        +-- Source Status Classifier
        +-- Transformation Contract
        +-- Synchronization Contract
        +-- Error Handling Contract
        +-- Integration Mapping Registry
        +-- Connector Health History
        +-- Integration Audit Trace
        +-- Marketplace
        +-- Executive Integration Console
        +-- Production Certification
        |
        v
Read Model / Live Connector / Test Connector / Local Connector / Contract Connector / Event Mesh / AI Provider

Implemented Lifecycle

  1. Register or discover connector. Connector catalog combines KDO read-model connectors and platform connector contracts into one standardized view.
  2. Classify source status. Each connector reports whether it is read-model, live, test-double, local, contract, blocked, unavailable, or unknown.
  3. Authorize access. Business-object reads call Enterprise Security Runtime with role, sensitivity, criticality, and runtime identity.
  4. Read through gateway. Allowed reads call DataAccessGateway for canonical business objects such as invoice, purchase order, goods receipt, journal, vendor, customer, contract, policy, payment, and cash position.
  5. Create trace. The runtime creates a hashed trace with action, business object, success state, timestamp, source metadata, record count, and provenance.
  6. Persist trace. IntegrationAuditTrace stores connector ID, source status, policy/evidence/security references, record count, and trace hash.
  7. Register mapping. IntegrationMappingDefinition records source/target mapping metadata and JSON definition for transformation governance.
  8. Record health. IntegrationConnectorHealth stores connector availability, latency, errors, freshness, and source-state telemetry.
  9. Search and inspect. APIs expose traces, mappings, health, marketplace, and executive console views for operations and audit.
  10. Certify gates. Production certification checks connector catalog, governance, security, observability, mapping registry, durable traceability, health history, and source-status classification.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/integrations/contractReturn runtime scope, owned capabilities, non-owned boundaries, business objects, backing services, and differentiators.Runtime definition.
GET /api/integrations/connectorsReturn standardized connector catalog with categories, source statuses, access modes, certification status, and business objects.Connector registry.
GET /api/integrations/fabricReturn standard connector contract, direct-access prohibition, replay support, policy enforcement, health model, and supported patterns.Connectivity fabric.
GET /api/integrations/api-managementReturn supported API styles, versioning, documentation, rate limiting, gateway, analytics, and lifecycle management.API management.
GET /api/integrations/transformationReturn schema/field mapping, normalization, validation, conversion, enrichment, preservation, AI-assisted mapping, and registry contract.Transformation.
GET /api/integrations/synchronizationReturn sync patterns, conflict detection, retry sync, offline sync, and governed bidirectional sync contract.Synchronization.
GET /api/integrations/governanceReturn required ownership fields, governance board checks, and integration readiness certification scores.Governance.
GET /api/integrations/error-handlingReturn retry, circuit breaker, dead-letter, timeout, duplicate detection, compensation, recovery, and error categorization contract.Resilience.
GET /api/integrations/securityReturn authentication, authorization, OAuth, SAML, API key, mTLS, certificate, encryption, secrets, audit, and security gateway posture.Security.
GET /api/integrations/aiReturn model abstraction, multi-model routing, provider failover, cost/latency optimization, prompt routing, tools, streaming, health, token tracking, and sovereign AI support.AI integration.
GET /api/integrations/data-governanceReturn lineage, ownership, quality, masking, privacy, retention, compliance, residency, master-data consistency, and auditability.Data governance.
GET /api/integrations/observabilityReturn API health, connector health, throughput, latency, error-rate model, retry activity, sync, cost, runtime health, SLA, and health-history source.Observability.
GET /api/integrations/marketplaceReturn connector discovery, API discovery, certification, version history, usage analytics, dependencies, owner, SLA, docs, and catalog.Marketplace.
POST /api/integrations/mappingsRegister an integration mapping definition.Mapping registry.
GET /api/integrations/mappings/searchSearch mapping definitions by business object.Mapping lookup.
POST /api/integrations/healthPersist connector health record.Health history.
GET /api/integrations/health/searchSearch connector health records.Health lookup.
POST /api/integrations/business-object/readRead a canonical business object through security and data-access gateway.Governed read.
GET /api/integrations/business-object/readRead a canonical business object through query parameters.Governed read.
GET /api/integrations/traces/searchSearch persisted integration audit traces by connector or business object.Replay and audit.
GET /api/integrations/executive-consoleReturn integration health, connector utilization, freshness, synchronization, cost, business impact, SLA, maturity, and traceability posture.Executive operations.
GET /api/integrations/runtime-interoperabilityReturn runtimes that must use Integration Runtime for external system access.Runtime boundary.
GET /api/integrations/world-class-certificationReturn measured connector liveness/governance certification and connector provenance.Certification.
GET /api/integrations/production-certificationReturn production gates for integration architecture, security, governance, marketplace, durable traceability, mappings, health, and source status.Release gate.

Connector Source Status

StatusMeaningOperational Rule
read_modelLocal or ingested read model populated by source-system ingestion.Allowed for governed reads with disclosed provenance.
liveLive external connector credentials and endpoint are available.Allowed according to connector policy and role.
test_doubleTest connector mode, such as Microsoft Graph test mode.Must be labeled as test-only and not treated as production source truth.
localLocal runtime connector such as sovereign local model provider.Allowed within local runtime policy.
contractGoverned connector contract exists, but production endpoint is not asserted live.Valid architecture boundary; not a live-data claim.
blocked_until_credentialsConnector requires credentials before live use.Do not execute; expose dependency.
unavailableSource path unavailable.Return safe failure and preserve trace.

Durable Integration Model

ArtifactStored InRuntime Function
IntegrationAuditTraceiaf_integration_audit_tracesReplayable trace of connector, object, action, success, source status, provenance, policy/evidence/security refs, hash, and metadata.
IntegrationMappingDefinitioniaf_integration_mapping_definitionsVersioned mapping registry for source-to-target business-object transformations.
IntegrationConnectorHealthiaf_integration_connector_healthConnector health history for status, source status, latency, errors, freshness, and operational reporting.

Example: Governed Business Object Read

A Context or Evidence Runtime request for invoice data is routed through security, DataAccessGateway, the standard fabric contract, and durable integration trace persistence.

{
  "operation": "ReadBusinessObject",
  "allowed": true,
  "business_object": "invoice",
  "gateway_result": {
    "kdo": "kdo://finance/invoice",
    "access_path": "mcp_ihub_gateway",
    "connector": "external_sor",
    "provenance": "live_read_model",
    "record_count": 2
  },
  "standard_fabric": {
    "authentication_model": "security_runtime_gateway",
    "policy_enforcement": "policy_runtime_before_sensitive_action",
    "replay_support": "correlation_id_and_trace_hash"
  },
  "durable_trace": {
    "stored": true,
    "model": "IntegrationAuditTrace",
    "trace_id": "eir-000001"
  }
}

Operational Requirements

  • Every connector must disclose access mode and source status.
  • Read access must not imply write access.
  • Write-capable connector contracts must remain gated by policy, decision, human approval, payload validation, and production adapter controls.
  • Integration reads must preserve source lineage, evidence reference, policy reference, security decision, and trace hash when persisted.
  • Mapping and health records must be searchable so operators can inspect transformation and connector posture.
  • Connector health and traces must feed Observability and Replay surfaces rather than remaining hidden inside adapter code.

Acceptance Criteria

  • Agents, skills, workflows, and UI surfaces must not directly integrate with external systems for material actions; the integration runtime is the governed boundary.
  • Business-object reads must pass through security authorization and record a trace when data is retrieved.
  • Every persisted integration trace must include connector ID, business object, action, success state, source status, provenance, policy/evidence/security references, and trace hash.
  • Connector catalog must disclose source status instead of implying all connectors are live production integrations.
  • Read-model connectors and governed contract connectors must be distinguished from live external HTTP connectors.
  • Mapping definitions must be versioned and searchable by business object.
  • Connector health history must be persisted and searchable by connector.
  • Write-capable integration must remain policy-, decision-, approval-, and adapter-gated; current contracts must not be represented as uncontrolled production write-back.

Engineering Rule

Do not let agents, skills, pages, or workflows call enterprise systems directly. They should request integration through the runtime, which applies security, discloses source status, preserves lineage, records traces, and keeps write-capable behavior behind explicit policy, decision, approval, and adapter gates.

CapabilitiesPolicy + Authority ControlTrust Fabric + ReplayEvent Fabric + SOR SidecarSkill Fabric + Agentic Runtime
Implemented Runtime Chapter

Runtime 12: Enterprise Identity, Security & Trust Runtime

The Enterprise Identity, Security & Trust Runtime is implemented as enterprise_security_runtime. It is the security gateway for Sphere: every human, agent, service, connector, workflow, API, tool, and model must operate under explicit identity, least privilege, adaptive trust, data protection, AI-specific controls, and replayable security audit.

Identity-First RBAC + ABAC Zero Trust Agent Security Gate AI Protection SecurityAuditEvent

Production Role

The runtime centralizes security enforcement for the rest of the platform. Policy, decision, agent, context, evidence, memory, knowledge, event, integration, and assurance runtimes can call the mandatory security gateway before sensitive data access, privileged action, tool use, write-like behavior, external sharing, or AI interaction.

The current implementation owns Sphere-specific authorization, zero-trust scoring, tool permission checks, AI and agent protection, data-protection contracts, secret-provider contracts, durable security audit, and provider readiness. It does not replace Entra ID, SAML/OIDC providers, PAM, SIEM, DLP, network controls, or source-system authorization; it consumes or integrates with those controls where configured.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime service enterprise_security.py Implements the identity-first security runtime for humans, agents, tools, models, APIs, services, workflows, connectors, and background jobs.
Runtime identity enterprise_security_runtime Registered in runtime certification as priority 12 with module iris_sor.services.enterprise_security, API prefix /api/security, UI routes /ui/access-controls and /ui/security, and durable entity SecurityAuditEvent.
API surface /api/security/* Exposes contract, identity registry, authentication contract, authorization, zero-trust evaluation, AI protection, agent action evaluation, secrets, data protection, API security, infrastructure security, events, durable events, event search, gateway evaluation, forwarding contract, provider readiness, threats, incidents, compliance, observability, integrations, executive command center, governance, and certification.
Identity registry identity_registry() Declares first-class human identities and machine identities: employee, contractor, vendor, auditor, executive, administrator, agent, skill, API, service, workflow, event producer, event consumer, background job, connector, and AI model.
Authentication contract authentication_contract() Declares SSO, SAML, OAuth, OpenID Connect, Entra ID, MFA, passwordless sessions, workload identity, certificate authentication, API keys with rotation, service accounts, model identity, continuous revalidation, MFA for privileged actions, step-up for sensitive data, and session risk scoring.
Authorization engine authorize() Combines role, ABAC sensitivity rules, policy reference, tool permissions, emergency access, field-level permission flagging, masking requirement, and audit requirement into one security decision.
Zero-trust evaluation evaluate_zero_trust() Combines authorization, adaptive trust score, device trust, network trust, session age, least privilege, continuous verification, conditional access, and session revalidation requirement.
AI and agent protection protect_ai_interaction() and evaluate_agent_action() Detects prompt injection, context poisoning, sensitive leakage risk, unsafe tool execution risk, autonomy boundary issues, write/external action risk, and records high-severity security events when blocked.
Mandatory gateway security_gateway() Provides the pre-action gateway for policy, decision, agent, context, evidence, memory, knowledge, event, integration, and assurance runtimes.
Durable ledger SecurityAuditEvent / iaf_security_audit_events Persists event ID, event type, severity, subject, runtime, policy/decision/evidence references, integrity hash, forwarding targets, forwarding status, details JSON, realm, and timestamp.
Provider readiness provider_readiness() Reports deployment readiness for Entra ID, SAML, OIDC, Microsoft Graph 365 read adapter, Azure Key Vault, AWS Secrets Manager, HashiCorp Vault, and SIEM/SOAR forwarding contracts.
Certification tests/test_enterprise_security_runtime_certification.py Verifies runtime scope, identity/authn/authz/zero-trust controls, AI and agent gates, data protection, secrets, API security, durable audit events, provider readiness, API registration, runtime certification registration, production gates, and non-authoritative world-class score reporting.

Runtime Scope

  • Identity-first control. Requires explicit human and machine identity categories before access, tool use, model use, integration, event production, or background work.
  • Authentication contract. Defines the authentication methods and session controls expected from deployment identity providers and workload identity providers.
  • RBAC plus ABAC. Combines role checks with sensitivity classification so restricted or PII-like data requires elevated roles.
  • Tool permission scoping. Loads agent tool permissions from the registry and blocks unauthorized tool use unless an explicit emergency path is allowed.
  • Adaptive trust. Scores risk from sensitivity, write/external action, prompt attack, context poisoning, device posture, failed attempts, and business criticality.
  • AI-specific protection. Blocks prompt-injection, policy-override, fake-evidence, secret-leakage, and unsafe write/execute tool patterns before agent action proceeds.
  • Agent security gate. Evaluates AI safety, zero-trust status, autonomy level, human approval, external action, runtime isolation, memory protection, and audit logging.
  • Data protection contract. Declares encryption, tokenization, masking, pseudonymization, PII detection, data residency, retention, and secure deletion controls.
  • Secret boundary. Declares Azure Key Vault, AWS Secrets Manager, HashiCorp Vault, Kubernetes secrets, rotation policies, and a hard rule that plaintext secret storage is not allowed.
  • Security audit ledger. Records material security decisions and incidents with integrity hashes and forwarding state for replay, search, and security operations.
  • Security observability. Surfaces authorization failures, threats, incidents, prompt attacks, API abuse, risk score posture, compliance posture, and runtime health.
  • Forwarding contract. Defines Microsoft Sentinel, Splunk, and ServiceNow Security Operations forwarding payloads and dead-letter requirements when providers are configured.

Logical Architecture

User / Agent / Service / Runtime / Connector / Tool / Model
        |
        v
Identity and Trust Request
        |
        v
Enterprise Identity, Security & Trust Runtime
        |
        +-- Identity Registry
        +-- Authentication Contract
        +-- RBAC / ABAC Authorization Engine
        +-- Agent Tool Permission Scope
        +-- Adaptive Trust Score
        +-- Zero Trust Evaluation
        +-- AI Security Protection
        +-- Agent Security Gate
        +-- Data Protection Contract
        +-- Secrets Management Contract
        +-- Mandatory Security Gateway
        +-- SecurityAuditEvent Ledger
        +-- SIEM / SOAR Forwarding Contract
        +-- Provider Readiness
        +-- Executive Security Command Center
        |
        v
allow / block_or_step_up / allow_with_masking / audit / incident / provider_gap

Implemented Lifecycle

  1. Request. A user, agent, service, connector, runtime, or model requests access, tool execution, AI interaction, data access, external action, or security event recording.
  2. Resolve identity type. The runtime treats human and machine identities as first-class subjects and requires owner, purpose, credential type, least-privilege scope, audit subject, and revalidation controls.
  3. Authenticate. The authentication contract defines human and machine authentication methods, session controls, continuous revalidation, MFA, and step-up requirements.
  4. Authorize. The authorization engine evaluates role, sensitivity, agent ID, entity/cost-center attributes, tool permission, emergency access, masking requirement, and audit requirement.
  5. Evaluate trust. Zero-trust evaluation adds adaptive risk score, device trust, network trust, session age, conditional access, micro-segmentation, and service authentication.
  6. Protect AI interaction. Prompt, context, requested tools, human approval, and unsafe write/execute patterns are checked before agent output or tool execution is trusted.
  7. Gate agent action. Agent action combines AI protection, zero-trust authorization, autonomy level, external action risk, human approval state, and audit logging.
  8. Ledger decision. Security events are captured in-memory or persisted to SecurityAuditEvent with integrity hash and forwarding state.
  9. Publish observability. Threats, incidents, compliance posture, authorization failures, prompt attacks, API abuse, and runtime health are exposed through security observability endpoints.
  10. Certify release. Production certification checks identity, authentication, authorization, zero trust, secrets, data protection, AI protection, agent gate, API security, compliance, observability, forwarding, provider readiness, durable audit, and gateway controls.

Core API Contract

EndpointPurposeRuntime Boundary
GET /api/security/contractReturn runtime mission, owned capabilities, non-owned boundaries, identity types, control families, differentiators, and consumed-by runtimes.Runtime definition.
GET /api/security/identity-registryReturn human and machine identity categories plus first-class identity requirements.Identity registry.
GET /api/security/authentication-contractReturn human and machine authentication methods plus session controls.Authentication.
POST /api/security/authorizeEvaluate role, ABAC sensitivity, agent tool permissions, emergency access, masking requirement, policy reference, and audit requirement.Authorization.
POST /api/security/zero-trust/evaluateEvaluate adaptive risk, authorization, device trust, network trust, session age, least privilege, conditional access, and revalidation.Zero trust.
POST /api/security/ai/protectEvaluate prompt injection, context poisoning, sensitive leakage, unsafe tool execution, and model abuse controls.AI security.
POST /api/security/agents/evaluate-actionEvaluate agent action against AI security, zero trust, autonomy, external action, human approval, tool restriction, memory protection, and audit logging.Agent authority.
GET /api/security/secretsReturn secret-provider and plaintext-storage controls.Secrets.
GET /api/security/data-protectionReturn encryption, tokenization, masking, pseudonymization, PII detection, residency, retention, and deletion controls.Data protection.
GET /api/security/api-securityReturn protected API types, rate limiting, gateway, validation, monitoring, and MCP endpoint controls.API security.
GET /api/security/infrastructure-securityReturn Kubernetes, container, database, storage, network, service mesh, CI/CD, build, and deployment security controls.Infrastructure contract.
POST /api/security/eventsRegister an in-memory security event.Security event capture.
POST /api/security/events/durablePersist a security event to SecurityAuditEvent.Durable ledger.
GET /api/security/events/searchSearch persisted security events by severity, type, and limit.Security replay.
POST /api/security/gateway/evaluateRun the mandatory security gateway for other runtimes.Pre-action gateway.
GET /api/security/forwarding-contractReturn Sentinel, Splunk, and ServiceNow Security Operations forwarding targets and payload contract.Security operations.
GET /api/security/provider-readinessReturn identity, graph, secrets, and SIEM/SOAR provider readiness.Deployment readiness.
GET /api/security/threatsReturn threat counters and security events.Threat detection.
GET /api/security/incidentsReturn incident response posture and open incidents created from high or critical events.Incident response.
GET /api/security/complianceReturn SOX, ISO 27001, SOC 2, GDPR, NIST, CIS, PCI where applicable, internal standards, and AI governance posture.Compliance.
GET /api/security/observabilityReturn authorization failures, threats, incidents, prompt attacks, API abuse, risk score availability, compliance posture, and runtime health.Security observability.
GET /api/security/integrationsReturn identity, secrets, SIEM, Sentinel, Splunk, SAP identity services, ServiceNow security operations, and forwarding integration posture.Security integrations.
GET /api/security/executive-command-centerReturn security posture, score, threat landscape, compliance, AI security score, identity health, incident trends, and business impact.Executive security.
GET /api/security/governanceReturn control ownership, risk classification, review cadence, approval workflows, policy mapping, exception management, audit trails, durable ledger, mandatory gateway, forwarding contract, and readiness scores.Governance.
GET /api/security/world-class-certificationReturn non-authoritative measured certification until backed by persisted security evidence.Certification.
GET /api/security/production-certificationReturn release gate results across identity, authentication, authorization, zero trust, secrets, data protection, AI protection, agent gate, API security, compliance, observability, gateway, forwarding, and provider readiness.Release gate.

Decision Types

DecisionMeaningImplemented Signal
allowAction or access is permitted.authorize().allowed=true and security_gateway().decision=allow.
block_or_step_upAction is blocked or requires stronger control before proceeding.Returned by gateway when authorization, trust, AI, or agent gate fails.
allow_with_maskingAccess can continue but sensitive data must be masked.Represented by data_masking_required=true for restricted or PII-like sensitivity.
auditDecision must be ledgered.audit_required=true and optional durable event persistence.
incidentHigh or critical security event creates an incident record.register_security_event() opens incidents for high or critical severity.
provider_gapProvider contract exists but deployment binding is not configured.provider_readiness() discloses configured versus contract-only status.

Durable Security Model

ArtifactStored InRuntime Function
SecurityAuditEventiaf_security_audit_eventsAppend-only security event ledger with event type, severity, subject, runtime, policy, decision, evidence, hash, forwarding, details, realm, and timestamp.
integrity_hashSecurityAuditEvent.integrity_hashSHA-256 hash of the recorded security event payload for tamper-evident replay.
forwarding_statusSecurityAuditEvent.forwarding_statusQueues high and critical events for Sentinel, Splunk, and ServiceNow Security Operations when provider routing is configured.

Example: Agent Action Blocked

A recommend-only invoice agent requesting an unsafe write-like tool is evaluated through AI protection, zero-trust authorization, tool permission scope, autonomy boundary, and audit logging.

{
  "operation": "AgentSecurityGate",
  "allowed": false,
  "agent_authentication": true,
  "tool_restrictions": true,
  "autonomous_execution_limits": true,
  "audit_logging": true,
  "ai_security": {
    "allowed": false,
    "unsafe_tool_execution_risk": true,
    "action": "block_and_escalate"
  },
  "zero_trust": {
    "never_trust_by_default": true,
    "session_revalidation_required": true
  }
}

Provider Readiness Rules

  • Entra ID readiness is inferred from Entra/Azure environment variables or a Microsoft OIDC issuer.
  • OIDC readiness can come from environment variables or runtime settings with issuer, audience, and JWKS/shared-secret or Microsoft issuer support.
  • Microsoft Graph readiness reports whether the MsGraph365ReadAdapter is readable and whether it is test-mode or live.
  • Secret-provider readiness reports Azure Key Vault, AWS Secrets Manager, and HashiCorp Vault contracts separately from live configuration.
  • SIEM/SOAR forwarding is a contract until providers are configured and event forwarding is operationalized.

Acceptance Criteria

  • Every user, agent, skill, API, service, workflow, connector, event actor, background job, and AI model must have an explicit runtime identity category.
  • Sensitive runtime actions must pass authorization and zero-trust checks before execution.
  • Restricted or PII-like data must require elevated role authorization and must set masking-required output where appropriate.
  • Agents must not exceed tool permissions or autonomy boundaries.
  • Prompt-injection, fake-evidence, context-poisoning, secret-leakage, and unsafe tool-execution patterns must be blocked or escalated.
  • Write or external actions must be constrained by security, policy, human approval, and integration boundaries.
  • Material security decisions and incidents must be replayable through SecurityAuditEvent or equivalent security event records.
  • Provider readiness must disclose whether identity, graph, secrets, and SIEM/SOAR providers are configured rather than implying live binding.
  • Security certification must remain measured and non-authoritative unless backed by real persisted security events.

Engineering Rule

Do not treat login as blanket authorization. Sensitive access, tool use, model interaction, memory creation, external action, replay export, and integration activity should evaluate user identity, machine identity, role, ABAC sensitivity, tool permission, session trust, AI-safety state, and audit requirements through the security runtime boundary.

CapabilitiesValue Attribution EnginePolicy + Authority ControlSkill Fabric + Agentic RuntimeTrust Fabric + Replay
Implemented Runtime Chapter

Runtime 13: Enterprise Model, Prompt & FinOps Runtime

The Enterprise Model, Prompt & FinOps Runtime is implemented as a composed runtime boundary: enterprise_finops_runtime certifies the economics layer, FinOpsRuntime computes ledger-backed cost and value, token_economy_runtime makes pre-call budget decisions, and GovernedLLMService enforces prompt, model, schema, confidence, and replay rules.

Governed LLM Prompt Registry Token Economy Cost Per Decision Value Attribution Deterministic First

Production Role

This runtime prevents uncontrolled model use by separating finance truth from language generation. Deterministic skills and policy checks remain the decision core; model calls are governed for explanation, extraction, classification, replay narrative, supervisor assistance, and bounded communication support.

The current implementation is strongest in FinOps attribution, token-economy decisioning, governed LLM contracts, prompt/schema validation, deterministic invocation evidence, and cost-to-value reporting. It is not a standalone universal provider gateway; provider execution is mediated by existing LLM services and shared model routing components.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime facade enterprise_finops_runtime.py Provides the first-class FinOps certification facade for token, model, tool, source-system, compute, budget, and business-value economics across runtimes.
Runtime identity enterprise_finops_runtime Registered in runtime certification with module iris_sor.services.enterprise_finops_runtime, API prefix /api/finops, UI routes /ui/finops-command-center and /ui/runtime-command-center, and data entities LLMAuditRecord, ToolInvocation, RuntimePerformanceTelemetry, and ValueEventRecord.
Operational FinOps service finops_runtime.py Builds ledger-backed AI economics from LLMAuditRecord, value attribution from ValueEventRecord, decision cost envelopes from AgentExecution, and explicit integration counters when present.
Token economy decisioning token_economy_runtime.py Makes pre-call decisions to proceed, reuse cache, skip LLM, escalate to human, downgrade model tier, or deny based on deterministic sufficiency, prior prompt hash, confidence, estimated tokens, price card, and per-agent daily ceiling.
Governed LLM service governed_llm_runtime.py Enforces domain/agent LLM governance, prompt bindings, output schema requirements, confidence thresholds, allowed model intersection, deterministic invocation parameters, pre-commit validation, output hashing, and replay evidence records.
Prompt registry config/contracts/llm_prompt_registry.yaml Stores approved prompt families such as evidence intelligence, decision replay narrative, root cause analysis, supervisor copilot, collections communication, duplicate invoice explanation, and autonomous finance strategy.
Output schemas config/contracts/llm_output_schemas.yaml Defines required structured output schemas used by governed LLM validation before model output is persisted or consumed downstream.
Tenant LLM governance config/tenants/bp/domains/*/llm_governance.yaml Declares agent-specific LLM access tier, decision authority, prompt bindings, output schema, confidence threshold, fallback mode, allowed models, policy override requirement, and forbidden actions.
Budget policy config/tenants/bp/finops_budget_policy.yaml and token_price_card.yaml Provides mission monthly budgets and model-tier price cards used for budget governance and price-card cost estimation.
UI surface FinOps Command Center and Runtime FinOps Exposes spend, value, token intelligence, routing intelligence, context cost intelligence, budget governance, decision cost envelopes, and role-specific economics views.
Certification and tests tests/test_finops_universal_cost_runtime.py, tests/test_llm_agent_layer_contract.py, and tests/test_duplicate_invoice_skill_contract.py Verify deterministic and LLM decision cost envelopes, FinOps contract exposure, prompt/schema governance, forbidden actions, approved model selection, compiled prompt packages, model routing, prompt hashes, explanation-only LLM usage, and deterministic finance decision core.

Runtime Scope

  • Deterministic-first routing. Token-economy decisions can skip the LLM entirely when deterministic rules are sufficient.
  • Governed model eligibility. The LLM service intersects agent-allowed models with prompt-family allowed models and rejects unapproved model use.
  • Prompt package governance. Prompt families, output schemas, prompt bindings, and tenant LLM governance overlays define what can be invoked by which agent.
  • Output validation. Required schema fields and confidence thresholds are checked before output can be committed.
  • Replay evidence. Governed LLM evidence records include prompt ID, prompt version, model ID, input hash, response hash, confidence, token usage, latency, fallback mode, execution seed, and pre-commit hash.
  • Token budget decisioning. The token economy evaluates estimated input/output tokens, price-card cost, daily ceiling, prior prompt hashes, confidence, and deterministic sufficiency before a call.
  • Cost attribution. FinOps summaries attribute cost by agent, mission, model, decision envelope, execution mode, and value event linkage.
  • Universal decision cost. Deterministic, hybrid, and reasoning executions receive FinOps envelopes so zero-LLM decisions are represented rather than invisible.
  • Business value linkage. ValueEventRecord linkage supports cost-to-value, value confidence split, ROI, and cost per decision views.
  • Source-system economics. SAP/source-system call counts, calls avoided, cache hits, misses, and context token counters are included only when explicitly instrumented.
  • Certification facade. Enterprise FinOps Runtime consolidates finops, cost, token economy, mission cost, and tool invocation modules into one certified runtime boundary.
  • Operational transparency. Static certification is non-authoritative until real operational substrate rows exist for tool/source cost, value attribution, and token/model economics.

Logical Architecture

Agent / Skill / Decision / Replay / Control Plane
        |
        v
Model, Prompt, or Cost Request
        |
        v
Enterprise Model, Prompt & FinOps Runtime
        |
        +-- Enterprise FinOps Runtime Facade
        +-- Governed LLM Service
        +-- Tenant LLM Governance Overlay
        +-- Prompt Registry
        +-- Output Schema Registry
        +-- Shared Model Routing Runtime
        +-- Token Economy Decision Runtime
        +-- Mission Budget Policy
        +-- Token Price Card
        +-- LLMAuditRecord Ledger
        +-- AgentExecution Cost Envelopes
        +-- ToolInvocation and Runtime Telemetry
        +-- ValueEventRecord Attribution
        +-- FinOps Command Center
        +-- Runtime Certification
        |
        v
proceed / use_cache / skip_llm / downgrade / escalate_human / deny / cost_record / replay_evidence

Implemented Lifecycle

  1. Agent or skill requests model support. The caller identifies domain, agent, prompt family, intended model, confidence, estimated tokens, deterministic sufficiency, and any relevant prompt/input hash.
  2. Resolve governance. GovernedLLMService loads tenant LLM governance and platform prompt/schema contracts.
  3. Check LLM access. Deterministic-only agents are blocked from LLM invocation, unsafe decision authority is rejected, output schema and confidence thresholds are required, and prompt binding must match the agent.
  4. Select model. The runtime intersects agent allowed models and prompt allowed models, then routes through the shared model routing runtime using latency, cost, and trust constraints.
  5. Evaluate token economy. TokenEconomyRuntime can return proceed, use_cache, skip_llm, escalate_human, downgrade, or deny before the model call.
  6. Build deterministic invocation parameters. Model calls use temperature 0 and stable execution seed for replayable behavior where the governed service controls invocation params.
  7. Validate output before commit. Output schema fields and confidence thresholds are enforced; low-quality or incomplete output fails before persistence.
  8. Build evidence record. Input hash, response hash, token usage, model, prompt, confidence, latency, fallback, seed, and pre-commit hash are captured.
  9. Record economics. LLMAuditRecord, AgentExecution, ToolInvocation, RuntimePerformanceTelemetry, and ValueEventRecord provide the substrate for cost and value analysis.
  10. Publish views. FinOps Command Center and control-plane APIs expose cost per decision, model portfolio, token intelligence, budget governance, replay rows, and value attribution.

Core API and Surface Contract

Endpoint / SurfacePurposeRuntime Boundary
GET /api/iaf/control-plane/finops-runtimeReturn live FinOps summary computed from LLMAuditRecord, ValueEventRecord, and AgentExecution.Operational FinOps.
GET /api/iaf/control-plane/observability/tokensReturn token usage summary with cost attribution.Token observability.
GET /api/iaf/control-plane/observability/tokens/costReturn token cost by agent and model.Cost attribution.
GET /api/iaf/control-plane/observability/token-economyReturn governed token economy controls, waste alerts, routing efficiency, and value-per-token metrics.Token economy dashboard.
GET /api/iaf/control-plane/observability/models/performanceReturn model latency and performance readiness metrics.Model observability.
GET /api/iaf/decision-intelligence/token-economy/decisionReturn pre-call token decision: proceed, use cache, skip LLM, escalate human, or deny.Pre-call budget guard.
/ui/finops-command-centerDisplay live FinOps economics, budget governance, model portfolio, cost attribution, decision cost envelopes, and role views.User experience.
/ui/runtime-finopsDisplay runtime-level FinOps proof surface and link to the command center.Runtime proof.
enterprise_finops_runtime.contract()Declare consolidated runtime capabilities and required consumers.Runtime definition.
enterprise_finops_runtime.world_class_certification()Return measured source/module coverage for FinOps runtime signals.Certification.
enterprise_finops_runtime.world_class_certification_live()Earn authoritative certification only from real recent rows in operational substrate tables.Live certification.

Governed Model Decisions

DecisionMeaningImplemented Source
skip_llmDeterministic rule is sufficient; no model call should be made.evaluate_llm_request(deterministic_sufficient=true).
use_cacheIdentical prompt/input hash exists in the current window.Prior LLMAuditRecord match by agent and hash.
proceedEstimated cost is within ceiling and model tier is acceptable.Token economy decision with price-card basis.
downgraded_to_stay_under_ceilingLarge tier would breach budget, so smaller tier is selected.Token economy tier downgrade.
escalate_humanLow confidence and material token estimate make generation uneconomic or risky.Confidence and token materiality check.
denyPer-agent cost ceiling is already spent.Daily spend check from LLMAuditRecord.

Prompt and Model Governance

ControlImplemented BehaviorFailure Mode
LLM access tierllm_access is normalized to L0-L5 and deterministic-only agents are blocked.GovernedLLMPolicyError.
Decision authorityAgents with approve, execute, write, post, or release authority cannot use unsafe LLM decision authority.Invocation denied.
Prompt bindingAgent must be bound to the requested prompt family.Invocation denied.
Approved modelsAgent allowed models must intersect with prompt allowed models.Invocation denied if intersection is empty.
Output schemaRequired fields are enforced before persistence.Pre-commit validation fails.
Confidence thresholdModel output confidence must meet tenant governance threshold.Pre-commit validation fails.
Replay parametersInvocation uses deterministic temperature and stable seed.Evidence build fails if seed or deterministic flags mismatch.
Forbidden actionsTenant governance lists blocked actions such as SOR write, LLM score calculation, unapproved external message, and autonomy grant.Action assertion fails.

Durable Economics Model

ArtifactRuntime UseData Basis
LLMAuditRecordToken/model cost, model tier, prompt/input hash reuse, latency, evidence counters, and LLM replay rows.iaf_llm_audit_log.
AgentExecutionUniversal decision cost envelopes for deterministic, hybrid, and reasoning executions.Agent execution ledger.
ValueEventRecordBusiness value attribution, confidence split, ROI, and cost-to-value linkage.iaf_value_events.
ToolInvocationTool and source-system cost dimension for live certification when recent operational rows exist.iaf_tool_invocations.
RuntimePerformanceTelemetryRuntime cost, latency, health, and observability linkage.Operational telemetry.

Example: Token Economy Decision

Before a model call, the runtime can skip, cache, downgrade, proceed, escalate, or deny based on cost, confidence, deterministic sufficiency, and price-card estimates.

{
  "action": "skip_llm",
  "reason": "deterministic_rule_sufficient",
  "model_tier": "none",
  "est_cost_usd": 0.0,
  "spent_usd": 0.0,
  "ceiling_usd": 8.0,
  "headroom_usd": 8.0,
  "pricing_basis": "platform_price_card.small:platform-default:platform_default_fallback"
}

Example: Governed LLM Evidence

When a model call is allowed, the governed service can build replay evidence without making the model output a source of financial authority.

{
  "agent_id": "llm.decision_replay_agent",
  "prompt_id": "finance.decision_replay_narrative.v1",
  "prompt_version": "1.0.0",
  "model_id": "gpt-4.1-mini",
  "input_hash": "sha256-input",
  "response_hash": "sha256-response",
  "token_usage": {"input": 1000, "output": 250},
  "fallback_mode": "deterministic_fallback",
  "execution_seed": 123456,
  "temperature": 0.0,
  "precommit_allowed": true
}

Operational Requirements

  • Structured finance logic should remain deterministic where exact matching, thresholds, controls, and calculations are sufficient.
  • Prompt templates and output schemas must be versioned contracts, not ad hoc strings inside agents.
  • LLM output should support explanation, extraction, classification, and narrative; it must not replace policy, evidence, skill, decision, or human authority.
  • Token and cost controls should run before model invocation, not only in after-the-fact reporting.
  • Cost views must disclose whether values are measured, estimated, policy-based, or illustrative.
  • External provider use should remain subject to security, policy, tenant LLM governance, and allowed-model contracts.

Acceptance Criteria

  • LLM access must be explicitly granted by tenant LLM governance; deterministic-only agents must fail closed.
  • Prompt use must be bound to an approved prompt family and output schema.
  • Agent-allowed models and prompt-allowed models must intersect before invocation.
  • Model invocation parameters must be deterministic where governed replay requires repeatability.
  • Model output must pass required-field and confidence-threshold checks before commit.
  • Token-economy decisions must support deterministic skip, cache reuse, human escalation, model-tier downgrade, and cost-ceiling denial.
  • Cost attribution must include deterministic, hybrid, and reasoning AgentExecution envelopes, not only LLM calls.
  • FinOps reporting must disclose data basis and avoid implying SAP/cache/source-system economics when counters are not instrumented.
  • Certification must remain non-authoritative until real recent operational substrate rows back token/model economics, value attribution, and tool/source cost dimensions.

Engineering Rule

Do not send every finance task to a frontier model. Use deterministic skills for finance truth, governed prompt/model contracts for language work, token-economy decisions before invocation, and ledger-backed FinOps attribution after execution.

CapabilitiesEvent Fabric + SOR SidecarTrust Fabric + ReplaySkill Fabric + Agentic RuntimePolicy + Authority Control
Implemented Runtime Chapter

Runtime 14: Enterprise Event Fabric Runtime

The Enterprise Event Fabric Runtime is implemented as a layered event platform. The enterprise_event_runtime facade certifies the canonical business-event model, while /api/events, /api/fef, and /api/iaf/event-bus provide the persisted finance event store, financial event fabric, and canonical event-to-agent dispatch surfaces.

Business Events Deduplication Runtime Routing Event Replay Subscriptions Dead Letters

Production Role

The runtime is the event-driven activation layer for Sphere. It converts system changes, runtime outputs, user actions, and control signals into business facts that can refresh context, collect evidence, evaluate policy, orchestrate decisions, trigger agents, update dashboards, publish telemetry, and feed learning.

The implementation intentionally separates the certification facade from the live event backbone. The facade can ingest, score, route, deduplicate, and replay events in memory; the persisted backbone stores finance events, CDC events, subscriptions, deliveries, and canonical enterprise event contracts. Certification reports whether it is measuring facade events or live persisted event flow.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Runtime facade enterprise_event_runtime.py Implements the certification and business-facing event fabric facade for canonical event model, event quality, runtime routing, automation triggers, lineage, replay, observability contract, and executive event console.
Runtime identity enterprise_event_runtime Registered in runtime certification as the Enterprise Event Runtime with purpose: event envelope, subscriptions, replayable streams, dead-letter handling, and runtime propagation.
Certification API /api/eer/* Exposes contract, registry, sources, event ingest/process/route/quality/replay, executive console, live backbone console, observability contract, KPIs, and certification gates.
Persisted event API /api/events/* Publishes FinanceEvent records, streams pending events, lists CDC events, manages subscriptions, returns event traces, and writes canonical enterprise event contracts.
Financial event fabric /api/fef/* Provides publish, subscribe, search, replay, correlate, and KPI APIs over the financial event store.
Canonical event bus /api/iaf/event-bus/* Exposes supervisor/admin subscription snapshots and canonical event publish through the same dispatch registry used by event-triggered agents.
Event store services/events/event_store.py Persists finance events, CDC events, subscriptions, deliveries, matching subscriptions, event stats, and subscription status changes.
Dispatcher and subscriptions services/events/event_dispatcher.py and event_subscriptions.py Maintains in-process canonical event subscriptions, dispatches canonical events to subscribed agents, and materializes declarative agent subscriptions.
Transport adapters events/nats_handlers.py, services/events/kafka_client.py, and event publisher services Support NATS subscription setup, Kafka publish attempts, canonical event publication, and broker-backed event propagation where configured.
Durable models FinanceEvent, CDCEvent, EventSubscription, EventDelivery, and EnterpriseEventContract Capture event payloads, source metadata, priorities, correlations, processing state, subscriptions, delivery attempts, canonical contracts, policy context, and candidate agents.
Certification and tests tests/test_enterprise_event_runtime_certification.py, test_event_runtime_live_backbone.py, test_policy_gate_on_events.py, and test_cross_domain_event_propagation.py Verify contract scope, source catalog, ingestion, quality, lineage, routing, triggers, deduplication, replay, live backbone stats, API registration, observability artifacts, policy gates, and cross-domain propagation.

Runtime Scope

  • Business fact event model. Treats events as governed business facts with event hash, business entity, source, timestamp, correlation, policy context, evidence, decision, workflow, and replay references.
  • Source catalog. Declares SAP ECC, SAP S/4HANA, SAP Ariba, Databricks, Unity Catalog, Microsoft 365, SharePoint, ServiceNow, HR systems, Kafka, NATS JetStream, Azure Event Grid, AWS EventBridge, REST APIs, webhooks, and internal runtimes as event sources.
  • Event categories. Classifies business, AI, user, and platform events with canonical examples for invoice, PO, goods receipt, payment, journal, collection, treasury, forecast, model, prompt, policy, replay, learning, approval, override, runtime, and connector activity.
  • Schema validation. Validates required event fields, payload shape, schema version, event hash, and event quality before high-trust routing.
  • Deduplication. Maintains an event dedupe index keyed by external event ID or hashed event type, source, entity, timestamp, and payload.
  • Routing. Routes by event category and type to context, policy, evidence, decision, learning, agent, assurance, and executive dashboard runtimes.
  • Automation triggers. Creates context assembly, policy evaluation, evidence collection, decision orchestration, dashboard refresh, agent execution, and learning capture triggers based on event type and category.
  • Replay. Filters replay timelines by event type, business entity, decision reference, or correlation ID and returns event hash, quality score, source, and routed consumers.
  • Live backbone. Reads persisted event-store statistics separately from the in-memory certification facade so runtime health and certification do not rely on static assumptions.
  • Dead-letter visibility. Schema-invalid facade events are moved to an in-memory dead-letter list; persisted event APIs track failed processing attempts and delivery failures.
  • Subscriptions. Supports persisted subscriptions and canonical in-process agent subscription snapshots for event-to-agent dispatch.
  • Observability. Defines event quality, ingestion/routing latency, queue depth, dead letters, consumer lag, replay activity, runtime health, SLOs, logs, traces, and alerts.

Logical Architecture

Source Systems / Runtimes / Users / Agents / Schedules / CDC
        |
        v
Raw or Canonical Event
        |
        v
Enterprise Event Fabric Runtime
        |
        +-- Enterprise Event Runtime Facade
        +-- Event Registry
        +-- Source Catalog
        +-- Event Normalization
        +-- Schema Validation
        +-- Deduplication Index
        +-- Source Lineage
        +-- Policy Context
        +-- Runtime Routing
        +-- Automation Triggers
        +-- Event Quality Scoring
        +-- Replay Timeline
        +-- Live Backbone Console
        +-- Finance Event Store
        +-- CDC Event Store
        +-- Event Subscriptions
        +-- Event Delivery Tracking
        +-- Canonical Event Contracts
        +-- Event Bus Dispatcher
        |
        v
Context / Policy / Evidence / Decision / Agent / Learning / Assurance / Dashboard Consumers

Implemented Lifecycle

  1. Source emits event. A source system, runtime, user, agent, or transport adapter sends a raw or canonical event.
  2. Ingest. The facade normalizes event type, category, timestamp, source, business entity, process, tenant, region, correlation, payload, priority, and references.
  3. Deduplicate. The runtime suppresses duplicate facade events using an external event ID or deterministic dedupe hash.
  4. Validate. Required fields and payload object shape are checked; invalid events are marked for dead-letter handling.
  5. Lineage. Source, timestamp, received time, correlation, business entity, policy context, evidence, decision, workflow, replay ID, and lineage hash are attached.
  6. Policy context. The event records required policy controls and produces a policy trace identifier for downstream governance.
  7. Route. Event category and type select routed consumers and produce topic, queue, routing modes, priority, event bus abstraction, and routing hash.
  8. Trigger automation. Context, policy, evidence, decision, dashboard, agent, or learning triggers are created according to event semantics.
  9. Score quality. Completeness, freshness, latency, duplication, ordering, accuracy, source reliability, processing success, and consumer success are combined into an event quality score.
  10. Replay and observe. Event timelines, live backbone stats, observability contract, KPIs, and certification gates expose event lineage and operational posture.

Core API Contract

Endpoint / SurfacePurposeRuntime Boundary
GET /api/eer/contractReturn event runtime ownership, non-owned boundaries, categories, sources, consumers, differentiators, and API list.Runtime definition.
GET /api/eer/registryReturn canonical event type catalog, schema version, owner, status, retention, replay readiness, and schema registry posture.Event registry.
GET /api/eer/sourcesReturn configured source catalog and authority scores.Source registry.
POST /api/eer/events/ingestIngest, normalize, validate, route, trigger, score, deduplicate, and store an event in the facade.Facade ingestion.
POST /api/eer/events/processProcess an ingested event through validation, lineage, policy, routing, automation, quality, and certification hash.Facade processing.
POST /api/eer/events/routeReturn topic, queue, routing modes, routed consumers, bus abstraction, and routing hash.Runtime routing.
POST /api/eer/events/qualityReturn completeness, freshness, latency, duplication, ordering, accuracy, source reliability, processing success, consumer success, and trust score.Quality scoring.
POST /api/eer/events/replayReturn filtered replay timeline by event type, business entity, decision reference, or correlation ID.Replay.
GET /api/eer/executive-consoleReturn facade event counts, latency, quality, queue utilization, integration health, business impact, and cost index.Event console.
GET /api/eer/executive-console/live-backboneReturn persisted event store health and lineage stats from EventStore.Live backbone.
GET /api/eer/observability-contractReturn metrics, traces, logs, alerts, and SLOs expected from the event runtime.Observability.
GET /api/eer/kpisReturn catalog size, events ingested, dead-letter count, event quality, replay readiness, certification score, and sample status.KPIs.
GET /api/eer/world-class-certificationBind to live backbone when possible and return measured event-flow certification.Certification.
GET /api/eer/production-certificationReturn production gate checks for event model, schema validation, dedupe, policy, evidence lineage, routing, triggers, replay, quality, console, observability, and unified bus abstraction.Release gate.
POST /api/eventsPersist FinanceEvent, publish to Kafka, publish canonical event contract, and return event response.Persisted event ingestion.
GET /api/events/{event_id}/traceReturn finance event and canonical event traceability.Trace.
GET /api/events/streamStream pending events via SSE with type/category filters.Live stream.
GET /api/events/cdcList CDC events by source, table, operation, and processing status.CDC.
GET /api/events/cdc/statsReturn CDC statistics.CDC metrics.
GET /api/events/subscriptionsList persisted subscriptions.Subscriptions.
POST /api/events/subscriptionsCreate persisted event subscription.Subscription registry.
DELETE /api/events/subscriptions/{subscription_id}Cancel persisted subscription.Subscription lifecycle.
POST /api/fef/events/publishPublish a financial event through the financial event fabric.Financial event fabric.
POST /api/fef/subscriptionsCreate a financial event subscription.Financial subscription.
GET /api/fef/events/searchSearch financial events by type, source, entity, correlation, mission, and limit.Event search.
POST /api/fef/events/replayReplay financial events by event, entity, or correlation.Financial replay.
POST /api/fef/events/correlateCorrelate financial events by correlation, entity, or case.Correlation.
GET /api/iaf/event-bus/subscriptionsReturn in-process canonical event subscriptions.Canonical event bus.
POST /api/iaf/event-bus/publishSupervisor/admin publish of canonical event through dispatcher and ledger-backed path.Event-to-agent dispatch.

Event Categories

CategoryExamplesTypical Consumers
business_eventsinvoice_received, po_created, goods_receipt_completed, payment_approved, journal_posted, collection_updated, treasury_position_changed, forecast_submitted.Context, policy, evidence, decision, dashboard.
ai_eventsagent_started, agent_completed, model_selected, prompt_compiled, tool_executed, decision_completed, evidence_generated, replay_created, learning_captured.Policy, evidence, decision, learning, assurance.
user_eventslogin, approval, escalation, override, comment, assignment, collaboration, file_upload, notification_acknowledged.Context, policy, dashboard, learning.
platform_eventsconnector_failure, queue_backlog, runtime_degraded, cache_refreshed, deployment_completed, security_alert, health_event.Dashboard, observability, operations, assurance.

Durable Event Model

ArtifactStored InRuntime Function
FinanceEventiaf_finance_eventsFinance event payload, source, entity, mission, payload schema, correlation, policy context, candidate agents, processing state, Kafka metadata, actor, realm, and timestamps.
CDCEventiaf_cdc_eventsChange-data-capture event from source tables with operation, before/after images, changed columns, primary keys, processing state, and linked finance event.
EventSubscriptioniaf_event_subscriptionsSubscriber filters, delivery mode, retry policy, status, consumer group, offset, and failure state.
EventDeliveryiaf_event_deliveriesDelivery attempts, response code/body, success, error, retry timing, duration, and subscription linkage.
EnterpriseEventContractiaf_enterprise_event_contractsCanonical event contract independent of domain vocabulary with event version, object, source, correlation, actor, policy context, candidate agents, and evidence requirement.

Example: Event Trigger Path

A high-priority invoice event can be normalized, policy-contextualized, routed to runtime consumers, and replayed without giving the source event direct authority to execute a finance action.

{
  "event_type": "invoice_received",
  "category": "business_events",
  "source": "SAP Ariba",
  "business_entity": "INV-EER-001",
  "business_process": "p2p",
  "correlation_id": "CORR-EER-001",
  "policy_context": {"policy_required": true},
  "evidence_ref": "evidence-pack-eer-001",
  "decision_ref": "decision-eer-001",
  "routing": {
    "topic": "enterprise.business.invoice_received",
    "event_bus_abstraction": ["Kafka", "NATS JetStream", "Azure Event Grid", "AWS EventBridge", "internal_runtime_bus"],
    "routed_consumers": [
      "enterprise_context_runtime",
      "enterprise_policy_intelligence_runtime",
      "evidence_intelligence_runtime",
      "enterprise_decision_runtime",
      "enterprise_agent_runtime"
    ]
  }
}

Operational Requirements

  • Source events should create or resume missions through Event Fabric and Orchestration, not call agents or write-back tools directly.
  • Canonical event contracts should carry source, object, correlation, policy context, candidate agents, and evidence requirement for downstream traceability.
  • Event consumers must be idempotent because delivery semantics are at-least-once by default.
  • Schema validation failures, processing failures, and delivery failures must be visible to operations and replay surfaces.
  • Event metrics should distinguish facade samples from persisted live-backbone events.
  • Event source authority and source-system transaction authority must remain separate from runtime event routing.

Acceptance Criteria

  • Event-triggered agents must be reached through event fabric and runtime routing, not by direct source-system-to-agent calls.
  • Business events must preserve source, business entity, correlation ID, timestamp, payload, policy context, and replay lineage.
  • Duplicate events must not produce duplicate facade processing results.
  • Schema-invalid events must be visible as dead-letter or failed-processing records.
  • Event routing must identify downstream runtime consumers and automation triggers.
  • Event replay must return event hash, quality score, routed consumers, source, timestamp, business entity, and correlation ID.
  • Persisted event APIs must distinguish finance events, CDC events, subscriptions, deliveries, and canonical enterprise contracts.
  • Certification must disclose whether it is measuring in-memory facade events or the live persisted event backbone.
  • Event observability must include latency, failures, dead letters, replay activity, consumer lag, SLOs, and event quality metrics.

Engineering Rule

Do not wire source events directly to agent execution or enterprise write-back. Events enter the fabric, are normalized, deduplicated, validated, enriched with lineage and policy context, routed to orchestration and runtime consumers, then observed and replayed.

CapabilitiesPolicy + Authority ControlTrust Fabric + ReplayValue Attribution EngineEvent Fabric + SOR Sidecar
Implemented Runtime Chapter

Runtime 15: Enterprise Control & Assurance Runtime

The Enterprise Control & Assurance Runtime is implemented through two cooperating layers: EnterpriseAssuranceRuntime for decision trust, audit packets, continuous monitoring, platform assessment, trust index, and governance routing; and Controls Intelligence Runtime for control-cycle operation, evidence capture, deterministic testing, controller gates, audit packs, and immutable control replay.

Control Registry Evidence Contracts Audit Packets Continuous Monitoring Trust Index Control Replay

Production Role

This runtime answers whether a governed decision, agent recommendation, workflow, or control cycle can be trusted. Policy answers whether an action is allowed; assurance verifies whether the action or recommendation is sufficiently evidenced, traceable, replayable, human-bounded, monitored, and audit-ready.

The implementation does not replace a formal GRC system. It creates a runtime assurance layer that can feed GRC, auditor, controller, supervisor, executive, replay, observability, and learning surfaces with evidence and control results captured during normal operations.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Assurance contract EnterpriseAssuranceRuntime.contract() and GET /api/assurance/contract Defines the boundary between policy allowance and assurance trust.
Case assurance evaluation POST /api/assurance/evaluate backed by AssuranceDecisionRecord Scores evidence, policy, replay, source lineage, human boundary, unsupported claims, and final outcome.
Evidence contract GET /api/assurance/cases/{{case_id}}/evidence-contract and EvidenceContractItem Persists required and missing evidence items per case and assurance record.
Audit packet POST /api/assurance/cases/{{case_id}}/audit-packet and AuditPacketRecord Builds sealed, hash-signed packets with policy, evidence, replay, human action, exceptions, and limitations.
Continuous control monitoring POST /api/assurance/continuous-controls and ContinuousControlMonitorRecord Computes financial, policy, evidence, audit, AI governance, and security-observability control families.
Enterprise trust index POST /api/assurance/trust-index and EnterpriseTrustIndexRecord Aggregates health, trust, compliance, risk, policy, evidence, AI quality, security, maturity, and business impact.
Closed-loop governance POST /api/assurance/closed-loop-governance and AssuranceGovernanceFindingRecord Routes runtime certification gaps and credibility gaps to owning remediation paths.
Monitoring schedules POST /api/assurance/monitoring-schedules and AssuranceMonitoringScheduleRecord Persists continuous monitoring cadence, runbook reference, last-run state, and job contract.
Control cycle runtime /api/controls-intelligence/control-cycles and ControlRuntimeCycle Stores control master, cycle, population, sample, evidence, test, exception, audit pack, human decision, agent run, and audit log records.
Control agents /api/controls-intelligence/agents and controls_intelligence_agents.py Provides deterministic controls agents for orchestration, monitoring, scoping, evidence, testing, validation, and audit liaison work products.
Control registry configuration config/tenants/bp/domains/finance/policy_pack/control_registry.yaml Includes configured controls such as trading margin variance, duplicate invoice prevention, three-way match, PO presence, goods receipt, payment terms, and vendor risk screening.
Operator surfaces /ui/enterprise-assurance-runtime and /ui/mission/controls Expose audit packets, assurance dashboard, continuous control cycles, controller actions, evidence, quality, FinOps, graph, and replay flows.

Runtime Scope

  • Map material cases, decisions, policies, evidence packs, agent runs, skills, control cycles, and integration confirmations to assurance records.
  • Evaluate whether recommendations are trusted, blocked, incomplete, approval-bound, or limited by missing evidence or replay.
  • Persist control evidence items separately from UI state so audit packets can be exported and replayed.
  • Monitor control families continuously using assurance records, evidence contract items, policy logs, replay manifests, audit packets, security events, telemetry, and agent evaluations.
  • Support control-cycle operation through persisted control masters, populations, sample selections, test results, exceptions, audit packs, controller decisions, agent runs, and immutable audit log events.
  • Aggregate executive assurance posture through platform assessments, trust index snapshots, command-center views, role-filtered assurance views, and trust trends.
  • Feed recurring gaps into governance findings, routed remediation, outcome learning, and runtime certification gaps.
  • Keep source-system write-back outside this runtime. Assurance can block, qualify, or require approval, but execution remains with Orchestration and Integration runtimes.

Logical Architecture

Case / Agent / Skill / Policy / Decision / Evidence / Replay / Control Cycle
        |
        v
Control or Assurance Request
        |
        v
Enterprise Control & Assurance Runtime
        |
        +-- Enterprise Assurance Runtime
        +-- Assurance Framework
        +-- Evidence Contract Evaluator
        +-- Case Assurance Evaluator
        +-- Audit Packet Generator
        +-- Continuous Control Monitor
        +-- Platform Assurance Assessment
        +-- Enterprise Trust Index
        +-- Closed-Loop Governance Router
        +-- Monitoring Schedule Runner
        +-- Controls Intelligence Runtime
        +-- Control Registry Explorer
        +-- Control Cycle Service
        +-- Population and Sample Records
        +-- Evidence Item Records
        +-- Deterministic Test Results
        +-- Exception Records
        +-- Audit Pack Records
        +-- Controller Decision Records
        +-- Agent Work Products
        +-- Immutable Audit Log Events
        |
        v
Assurance Status / Control Outcome / Audit Packet / Trust Index / Governance Finding / Replay Path

Implemented Lifecycle

  1. Register: Controls and runtime contracts are declared through policy packs, runtime certification, control-cycle schema, and assurance framework records.
  2. Map: A case or control cycle is linked to mission, policy, evidence, replay, decision, agent, source lineage, and human boundary references.
  3. Collect: Evidence Runtime, policy logs, replay manifests, security events, runtime telemetry, value events, and control-cycle records become the evidence base.
  4. Evaluate: Assurance calculates credibility and control status using deterministic scoring rather than treating a policy allow/deny as the whole control result.
  5. Package: Audit packets and control audit packs create hash-backed exports with evidence, policy, decision, replay, human action, exceptions, limitations, and seal metadata.
  6. Monitor: Continuous control monitoring computes effectiveness, evidence, policy, exception, remediation, and family status across the current compact database snapshot.
  7. Attest: Controller or owner actions are persisted as ControlRuntimeHumanDecision rows and matching audit log events.
  8. Route: Governance findings and exceptions route to owner workflows, controller queues, or external ticket targets when available.
  9. Replay: Case replay, audit packet export, control-cycle replay flows, and immutable audit log events reconstruct the control path.
  10. Improve: Outcomes and recurring assurance gaps publish learning records and governance findings for future refinement.

Core API Contract

Endpoint / SurfacePurposeRuntime Boundary
GET /api/assurance/contractReturn assurance runtime purpose, sub-capabilities, policy boundary, scoring weights, audit registers, certification, and trust fabric models.Runtime definition.
GET /api/assurance/frameworkReturn assurance mission, business / AI / platform domains, non-owned responsibilities, and independent evidence sources.Assurance framework.
POST /api/assurance/evaluateEvaluate a case with mission, case type, agent, decision, policy, evidence, replay, proposed action, facts, approvals, unsupported claims, and correlation refs.Case assurance.
GET /api/assurance/cases/{case_id}/evidence-contractReturn evidence contract coverage and missing items for a case.Evidence requirement.
GET /api/assurance/cases/{case_id}/replay-timelineReturn replay timeline composed from assurance and replay artifacts.Replay linkage.
GET /api/assurance/agents/{agent_id}/evaluationReturn windowed agent assurance evaluation.Agent quality assurance.
POST /api/assurance/cases/{case_id}/outcomeCapture human or business outcome and create an outcome learning record.Learning linkage.
POST /api/assurance/cases/{case_id}/audit-packetGenerate a sealed audit packet for a case.Audit packet.
GET /api/assurance/cases/{case_id}/audit-packet/exportExport the latest audit packet for an authorized role.Audit export.
GET /api/assurance/dashboardReturn assurance health, mission coverage, credibility register, operational metrics, signing posture, and latest packets.Dashboard.
POST /api/assurance/continuous-controlsCreate continuous control monitor records for financial, policy, evidence, audit, AI governance, and security-observability control families.Continuous monitoring.
POST /api/assurance/platform-assessmentCreate platform assurance assessment from runtime certification, evidence, policy, replay, security, telemetry, and credibility sources.Platform assessment.
POST /api/assurance/trust-indexCreate enterprise trust index snapshot.Executive assurance.
GET /api/assurance/trust-index/trendsReturn historical trust index points and trend direction.Trend analytics.
POST /api/assurance/closed-loop-governanceCreate governance findings from runtime certification and credibility gaps.Remediation governance.
POST /api/assurance/closed-loop-governance/routeRoute open governance findings to a target such as ServiceNow.Routing.
GET /api/assurance/executive-command-centerReturn assurance, trust, policy, control, audit, risk, runtime health, maturity, and business value scores.Executive command center.
POST /api/assurance/monitoring-schedulesPersist a scheduled assurance monitoring plan.Monitoring schedule.
POST /api/assurance/monitoring-schedules/runRun scheduled monitoring and update last-run metadata.Scheduled assurance.
GET /api/controls-intelligence/control-cyclesReturn persisted control-cycle summary with evidence, exception, controller, and audit-pack counts.Control cycle.
GET /api/controls-intelligence/control-cycles/{cycle_id}Return one control cycle with component map, schema, read/write boundaries, population, sample, tests, exceptions, evidence, audit pack, and human decision.Control detail.
POST /api/controls-intelligence/control-cycles/{cycle_id}/run-cyclePersist control run-cycle agent and audit log events.Control execution ledger.
POST /api/controls-intelligence/controller-actions/{action}Persist controller sign-off, pushback, waiver, escalation, substantiation, or remediation action.Human attestation.
GET /api/controls-intelligence/replay-flows/{flow_id}Return control replay flow payload.Control replay.

Durable Assurance Model

ArtifactStored InRuntime Function
AssuranceDecisionRecordiaf_assurance_decision_recordsCanonical trust verdict for a material agent or workflow decision.
EvidenceContractItemiaf_evidence_contract_itemsEvaluated evidence requirement with status, lineage, hash, and source detail.
AgentEvaluationRecordiaf_agent_evaluation_recordsAgent quality, evidence, policy, latency, cost, regression, and failure reason metrics.
AuditPacketRecordiaf_audit_packet_recordsGenerated packet JSON, status, hash, generator, tenant, realm, and timestamp.
ContinuousControlMonitorRecordiaf_continuous_control_monitor_recordsControl family effectiveness, evidence, policy, exception, remediation, and detailed check results.
EnterpriseTrustIndexRecordiaf_enterprise_trust_index_recordsExecutive trust components and aggregate enterprise trust index.
AssuranceGovernanceFindingRecordiaf_assurance_governance_findingsClosed-loop remediation finding with evidence refs, owner, route, ticket ref, and due date.
ControlRuntimeCycleiaf_control_cyclesControl assurance cycle period, framework, owner, verdict, confidence, status, and value metadata.
ControlRuntimeEvidenceItemiaf_control_evidence_itemsControl-cycle evidence artifact, source, fact, hash, and lineage.
ControlRuntimeTestResultiaf_control_test_resultsDeterministic control test points, pass/fail counts, execution mode, and result JSON.
ControlRuntimeExceptioniaf_control_exceptionsControl exception, classification, finding, sample, impact, owner, and status.
ControlRuntimeAuditLogEventiaf_control_audit_log_eventsImmutable control replay event stream with sequence, actor, summary, hash, and event JSON.

Example: Case Assurance Result

A duplicate-invoice case is not trusted because an agent recommended it. It is trusted only when evidence, policy, replay, source lineage, and human-action boundaries are sufficient.

{
  "case_id": "INV-LIVE-DUP-1778594510",
  "mission_id": "p2p",
  "assurance_status": "trusted",
  "credibility_score": 100,
  "evidence_contract": {
    "missing_evidence": [],
    "required_items": [
      "Invoice",
      "Supplier master",
      "Prior payment",
      "Policy decision",
      "Replay manifest"
    ]
  },
  "control_interpretation": {
    "duplicate_payment_prevention": "operating",
    "write_back": "approval_required_before_execution",
    "audit_packet": "eligible"
  }
}

Operational Requirements

  • Do not equate policy outcome with control pass. Policy is one input; assurance also requires evidence, lineage, replay, human boundary, and execution state.
  • Do not infer control evidence from rendered UI. Control evidence must be represented as first-class records with hashes and lineage.
  • Audit packets must disclose limitations when an assurance record, replay record, approval, or evidence item is missing.
  • Controller decisions must be reasoned and persisted before audit packs are released externally.
  • Control-cycle agents should remain deterministic-first and supervised by control-owner gates.
  • Assurance metrics should distinguish trusted decisions, blocked decisions, missing evidence, audit packet coverage, control effectiveness, and remediation backlog.

Acceptance Criteria

  • Every material case assurance decision can reference evidence, policy, replay, human boundary, source lineage, and payload hash.
  • Missing mandatory duplicate-invoice evidence blocks trust rather than silently passing the case.
  • Trusted duplicate-invoice assurance requires invoice, supplier, payment, policy, replay, and evidence-contract coverage.
  • Audit packets are sealed with a packet hash and HMAC signature in non-production fallback or configured runtime-secret mode.
  • Continuous controls produce persisted monitor records and remediation flags for control families that need attention.
  • Trust index snapshots are persisted and can be trended for executive and auditor views.
  • Control cycles expose population, sample, deterministic test result, exception, evidence, audit pack, human decision, and replay event data.
  • Controller actions persist a human-decision row plus an audit-log event instead of changing only visual state.
  • Role-filtered executive assurance views limit which fields are returned for CFO, auditor, security, and operations roles.
  • Failure modes default to safe results: unavailable identity, missing evidence, missing assurance record, incomplete replay, or absent policy lowers readiness.

Engineering Rule

Treat controls and assurance as runtime artifacts, not presentation labels. A control status is valid only when it links to durable evidence, policy, decision, agent or skill output, human approval where required, integration confirmation where applicable, and replayable audit records.

CapabilitiesOntology + Enterprise ContextCREST Context ReconstructionTrust Fabric + ReplayPolicy + Authority Control
Implemented Runtime Chapter

Runtime 16: Enterprise Finance Ontology Runtime

The Enterprise Finance Ontology Runtime is implemented as a tenant-aware semantic fabric. It combines platform ontology overlays, BP domain-pack ontology YAML, semantic alias resolution, governed object lifecycle records, ontology-trigger routing, runtime-link ledgers, decision ontology coverage, object journeys, and Enterprise Intelligence Fabric graph persistence.

Canonical Objects Relationship Graph Semantic Aliases Governed Object State Runtime Spine Ontology Governance

Production Role

This runtime defines the shared finance language used by context, evidence, policy, decision, agent, skill, event, integration, replay, assurance, learning, and dashboard surfaces. It prevents source-system fields, UI pages, agents, and policy rules from inventing separate definitions for the same finance concept.

The implementation is not a transaction store. It is a governed semantic layer that maps transaction and runtime records to canonical objects, relationships, lifecycle events, policy scope, evidence requirements, and replayable decision paths.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Ontology fabric get_ontology_fabric() and OntologyFabric Merges platform ontology overlays, tenant ontology roots, and domain-pack ontology contracts into one runtime fabric.
Ontology summary and registry GET /api/iaf/ontology/summary, /objects, /relationships, /actions, /triggers, /lifecycles Exposes canonical objects, relationship graph, action vocabulary, trigger catalog, lifecycle states, and event standards.
BP finance ontology YAML config/tenants/bp/domains/finance/ontology/ontology_objects.yaml and ontology_relationships.yaml Defines objects such as finance.invoice, finance.payment, finance.journal_entry, finance.vendor, purchase order, customer invoice, forecast, tax obligation, and their relationships.
Shared finance domain pack config/tenants/_shared/domain_packs/finance/ontology/* Provides reusable finance objects, relationships, events, aliases, and canonical mapping registry inherited by tenant overlays.
Semantic alias resolution SemanticResolutionService, semantic_registry.yaml, and SemanticAliasMapping Resolves terms such as supplier/vendor, invoice/bill, customer/client, payment/settlement, and records canonical alias mappings.
Canonical identity CanonicalIdentityBinding Binds source-system object IDs to canonical object IDs with confidence and match method.
Semantic conflict tracking SemanticConflictRecord Persists competing semantic resolutions and their approval/resolution status.
Ontology contract versioning OntologyContractVersion and semantic_registry.yaml Stores active ontology contract version, additive migration policy, schema hash, and migration notes.
Governed object lifecycle GovernedObjectState and GovernedEventRecord Persists latest canonical object state and decision-grade ontology events for replay and audit.
Ontology runtime-link ledger OntologyRuntimeLink and record_ontology_runtime_link() Indexes exact or inferred links from ontology objects to policy evaluations, agent executions, decisions, evidence packs, events, and KPIs without mutating immutable records.
Decision ontology coverage /api/iaf/ontology/decision-ontology/* and decision_ontology_runtime.py Audits live decision, policy, evidence, memory, and ontology-link ledgers against the tenant decision ontology contract.
Object journey and runtime spine /api/iaf/ontology/object-journey/{{entity_type}}/{{entity_id}}, /runtime-spine/{{ref}}, and /api/iaf/decision-graph/{{entity_id}} Reconstructs source-to-event-to-policy-to-agent-to-decision-to-action-to-evidence paths for a canonical object.
Enterprise graph persistence /api/intelligence/entities, /relationships, /graph/search, /graph/traverse Persists ontology entities, evidence-backed relationships, graph traversals, reasoning traces, and digital twin snapshots.
Ontology lifecycle governance POST /api/intelligence/ontology/lifecycle and IntelligenceOntologyLifecycleRecord Records ontology publish/change lifecycle, owner, approval reference, prior version, change summary, and effective period.
User surfaces /ui/ontology, /ui/object-journey, and /intelligence/object-journey Expose object taxonomy, live ontology activity, policy/agent/evidence usage, and object execution proof.

Runtime Scope

  • Define canonical finance objects, object aliases, lifecycle states, event vocabulary, and object relationships through platform and tenant domain-pack YAML.
  • Resolve tenant/domain/source aliases into canonical terms and canonical object identities.
  • Map source-system records and runtime records onto ontology objects without rewriting immutable policy, agent, decision, or evidence ledgers.
  • Validate object event transitions against ontology lifecycle definitions before governed object state is advanced.
  • Resolve ontology-triggered agents and routing plans from canonical object events rather than page-specific logic.
  • Provide runtime object journeys that show creation, canonical mapping, event emission, agent trigger, policy evaluation, decision creation, action execution, evidence creation, and replayability.
  • Support Enterprise Intelligence Fabric graph operations for semantic entities, relationships, reasoning traces, digital twin snapshots, graph traversal, and ontology lifecycle governance.
  • Expose binding-gap and decision-ontology coverage audits so platform teams can find ontology objects, triggers, agents, policies, decisions, and evidence that are not yet linked.

Logical Architecture

Source Systems / Domain Packs / Runtime Records / Events / Agents / UI
        |
        v
Ontology Lookup / Semantic Resolution / Object Event / Trace Request
        |
        v
Enterprise Finance Ontology Runtime
        |
        +-- Ontology Fabric
        +-- Domain Contract Merger
        +-- BP Finance Ontology Objects
        +-- BP Finance Ontology Relationships
        +-- Event and Lifecycle Registry
        +-- Semantic Alias Resolver
        +-- Canonical Identity Binder
        +-- Semantic Conflict Tracker
        +-- Governed Object State Store
        +-- Governed Event Store
        +-- Ontology Trigger Engine
        +-- Ontology Routing Runtime
        +-- Ontology Runtime Link Ledger
        +-- Decision Ontology Coverage Auditor
        +-- Runtime Spine Resolver
        +-- Object Journey Resolver
        +-- Enterprise Intelligence Graph
        +-- Ontology Lifecycle Governance
        |
        v
Canonical Object / Relationship / Event / Trigger / Runtime Spine / Versioned Semantic Meaning

Implemented Lifecycle

  1. Declare: Domain packs define ontology objects, relationships, events, aliases, lifecycles, and canonical mappings in YAML.
  2. Merge: The tenant OntologyFabric merges platform overlays, tenant roots, and unified domain contracts at runtime.
  3. Resolve: Semantic resolution normalizes aliases and builds canonical IDs from object type, source system, and source object ID.
  4. Validate: Governed object events validate object registration, lifecycle state, and state transition semantics.
  5. Route: Ontology trigger and routing runtimes turn object events into matched triggers, ordered agents, policy scope, routing rules, and execution contract.
  6. Link: Runtime publishers write OntologyRuntimeLink rows for source records, policy evaluations, agent executions, decisions, evidence packs, and KPIs.
  7. Trace: Runtime-spine and object-journey APIs reconstruct exact object-to-event-to-policy-to-agent-to-decision-to-evidence lineage when links exist.
  8. Govern: Ontology lifecycle records capture publish/change status, owner, approval reference, prior version, change summary, and effective period.
  9. Observe: Explorer and binding-gap surfaces show object usage, live policy/agent/evidence activity, unresolved bindings, and old contract coverage.
  10. Replay: Historical meaning is preserved through ontology contract versions, runtime-link records, object journeys, and replayable event/evidence references.

Core API Contract

Endpoint / SurfacePurposeRuntime Boundary
GET /api/iaf/ontology/summaryReturn tenant ontology fabric summary.Fabric summary.
GET /api/iaf/ontology/objectsReturn merged ontology objects.Object registry.
GET /api/iaf/ontology/objects/{object_name}Return object detail, fields, lifecycle, aliases, evidence, relationships, policies, and agents where available.Object detail.
GET /api/iaf/ontology/relationshipsReturn merged ontology relationships.Relationship graph.
GET /api/iaf/ontology/actionsReturn ontology actions.Action vocabulary.
GET /api/iaf/ontology/triggersReturn trigger catalog.Trigger registry.
GET /api/iaf/ontology/lifecyclesReturn object lifecycle states and transitions.Lifecycle model.
GET /api/iaf/ontology/event-standardsReturn event standards from the ontology fabric.Event ontology.
GET /api/iaf/ontology/canonical-registryReturn canonical registry from the fabric.Canonical registry.
POST /api/iaf/ontology/resolve-triggersResolve an object/event into matched triggers, trigger agents, resolution source, and ontology refs.Trigger resolution.
GET /api/iaf/ontology/binding-gapsAudit ontology trigger bindings against registry and runtime subscriptions.Binding gap audit.
GET /api/iaf/ontology/decision-ontology/contractReturn tenant-scoped decision ontology contract.Decision ontology.
GET /api/iaf/ontology/decision-ontology/coverageAudit live runtime ledgers against decision ontology coverage targets.Coverage audit.
GET /api/iaf/ontology/decision-ontology/governanceReturn governance and readiness report for decision ontology.Ontology governance.
GET /api/iaf/ontology/decision-ontology/chain/{signal_id}Resolve a signal into decision ontology and matching runtime links.Decision chain.
POST /api/iaf/ontology/decision-ontology/backfillIdempotently index existing runtime records into ontology links.Backfill.
GET /api/iaf/decision-graph/{entity_id}Return exact runtime spine when available; otherwise build an ontology-grounded decision graph.Decision graph.
GET /api/iaf/context/{entity_id}Return ontology refs, canonical mappings, policy context, enterprise memory, signal context, decision graph, and runtime spine.Context grounding.
GET /api/iaf/ontology/runtime-spine/{ref}Resolve source record, decision, execution, or evidence ID into exact ontology runtime links.Runtime spine.
GET /api/iaf/ontology/object-journey/{entity_type}/{entity_id}Return source-to-event-to-policy-to-decision-to-action-to-evidence journey for an ontology object.Object journey.
GET /api/intelligence/ontologyReturn Enterprise Intelligence Fabric ontology catalog and governance posture.Ontology catalog.
POST /api/intelligence/entitiesRegister a durable semantic entity.Graph entity.
POST /api/intelligence/relationshipsRegister evidence-backed graph relationship.Graph relationship.
POST /api/intelligence/graph/traversePersist and return graph traversal result.Semantic traversal.
POST /api/intelligence/ontology/lifecycleRecord ontology lifecycle governance event.Version governance.

Canonical Finance Objects and Relationships

Implemented AssetExamplesRuntime Meaning
finance.invoiceinvoice_id, invoice_number, vendor_id, purchase_order_id, amount, currency, posting_date, due_date, status, payment_terms, three_way_match_status, source_system.Supplier invoice lifecycle from received to paid, disputed, or cancelled.
finance.paymentpayment_id, payment_method, amount, currency, execution_date, settlement_date, vendor_id, status.Outbound payment lifecycle from scheduled to settled or reversed.
finance.journal_entryjournal_entry_id, fiscal_period, source, total_amount, line_count, preparer_id, approver_id, status.R2R posting object tied to approvals, GL accounts, reconciliation, and close.
finance.vendorvendor_id, name, tax_id, payment_terms, risk_category, country, status, spend_ytd.Supplier master semantics with sensitive and control-relevant attributes.
finance.invoice references finance.purchase_ordermany_to_one.Invoice-to-PO relationship for matching and procurement context.
finance.invoice paid_by finance.paymentmany_to_many.Invoice settlement relationship.
finance.invoice issued_by finance.vendormany_to_one.Supplier context for duplicate, risk, and payment controls.
finance.journal_entry reconciled_in finance.reconciliationmany_to_one.Close and reconciliation control context.
finance.cash_position informs finance.forecastmany_to_one.Treasury-to-FP&A semantic bridge.
finance.journal_entry gives_rise_to finance.tax_obligationone_to_many.R2R-to-Tax semantic bridge.

Durable Ontology Model

ArtifactStored InRuntime Function
SemanticAliasMappingiaf_semantic_alias_mappingsTenant/domain/source alias to canonical term mapping.
CanonicalIdentityBindingiaf_canonical_identity_bindingsSource-system object ID to canonical object ID binding.
SemanticConflictRecordiaf_semantic_conflictsSemantic conflict and resolution tracking.
OntologyContractVersioniaf_ontology_contract_versionsSemantic contract version and schema hash.
GovernedObjectStateiaf_governed_object_statesLatest state for a first-class ontology object.
GovernedEventRecordiaf_governed_event_recordsDecision-grade object event envelope with triggers, agents, policy scope, replay eligibility, state transition, and payload hash.
GovernedRoutingPlaniaf_governed_routing_plansCanonical routing plan resolved for a governed object event.
OntologyRuntimeLinkiaf_ontology_runtime_linksNon-mutating index from ontology object to policy, agent, decision, evidence, event, and KPI records.
IntelligenceOntologyEntityiaf_intelligence_ontology_entitiesGoverned semantic entity in the Enterprise Intelligence Fabric graph.
IntelligenceGraphRelationshipiaf_intelligence_graph_relationshipsEvidence-backed semantic relationship with confidence, trust, provenance, and evidence refs.
IntelligenceOntologyLifecycleRecordiaf_intelligence_ontology_lifecycleVersioned ontology publish/change lifecycle and approval record.
IntelligenceGraphTraversalRecordiaf_intelligence_graph_traversalsDurable graph traversal result for semantic reasoning replay.

Example: Duplicate Invoice Semantic Path

The duplicate-payment path is ontology-grounded before it becomes a case, recommendation, control outcome, or replay record.

{
  "entity_type": "Invoice",
  "entity_id": "INV-LIVE-DUP-1778594510",
  "canonical_object": "finance.invoice",
  "relationships": [
    "finance.invoice issued_by finance.vendor",
    "finance.invoice references finance.purchase_order",
    "finance.invoice paid_by finance.payment"
  ],
  "runtime_spine": {
    "object": "Invoice",
    "event": "invoice_received",
    "policy": "PolicyEvaluation",
    "agent": "AgentExecution",
    "decision": "Decision",
    "evidence": "EvidencePack"
  },
  "semantic_controls": [
    "duplicate_payment_prevention",
    "payment_release_authority",
    "evidence_sufficiency"
  ]
}

Operational Requirements

  • Do not hardcode source-field mappings in UI pages or agents when an ontology mapping or canonical object path exists.
  • Do not treat ontology as only a glossary. Object relationships, lifecycle states, triggers, policy scope, evidence requirements, and runtime links are first-class runtime assets.
  • Do not rewrite immutable policy, agent, decision, or evidence rows to add semantic metadata. Use OntologyRuntimeLink for non-mutating linkage.
  • Ontology-triggered automation must pass through trigger resolution, routing, policy scope, and replay-eligible governed event records.
  • Unknown source objects, unknown terms, semantic conflicts, and deprecated mappings should produce explicit review states rather than silent conversion.
  • Ontology changes must be versioned and additive-first so historical replay remains semantically stable.

Acceptance Criteria

  • Ontology object, relationship, event, trigger, action, lifecycle, and canonical registry APIs return tenant-scoped fabric data.
  • BP finance domain packs define concrete finance objects and relationships rather than relying on only generic platform terms.
  • Semantic aliases resolve common terminology differences such as supplier/vendor, invoice/bill, customer/client, and payment/settlement.
  • Governed object events reject unregistered objects and invalid lifecycle transitions instead of silently creating inconsistent state.
  • Runtime-spine APIs prefer exact ontology runtime links and clearly report unavailable lineage when exact links are absent.
  • Object journey exposes stage completeness for canonical mapping, events, agent trigger, policy evaluation, decision, action, evidence, and replay.
  • Decision ontology coverage audits live ledgers against tenant contract coverage targets and supports idempotent backfill.
  • Ontology binding-gap audit detects mismatches between ontology triggers, registry entries, and runtime subscriptions.
  • Enterprise Intelligence Fabric persists ontology entities, relationships, reasoning traces, graph traversals, digital twin snapshots, and lifecycle records.
  • Ontology governance remains additive-first and versioned so replay can interpret historical decisions under the right semantic contract.

Engineering Rule

Treat ontology as the shared contract for meaning, not as documentation. Runtime context, evidence, policies, decisions, events, agents, controls, replay, and learning should reference canonical ontology objects and relationships wherever they need business meaning.

CapabilitiesSkill Fabric + Agentic RuntimeEvent Fabric + SOR SidecarPolicy + Authority ControlValue Attribution Engine
Implemented Runtime Chapter

Runtime 17: Enterprise Mission & Case Runtime

The Enterprise Mission & Case Runtime is implemented as the work-object layer that turns signals, decision records, agent recommendations, policy boundaries, evidence packs, approvals, workflow runs, supervisor queue items, value events, and mission pages into governed finance work.

Mission Registry Work Items Decision Workforce Supervisor Queue Mission ACT Case Value Rollup

Production Role

This runtime is the business container for finance work. Event Fabric may detect the signal, Orchestration may coordinate the flow, and Decision Runtime may decide the next action, but Mission & Case Runtime owns how the work appears as a mission queue, case drawer, supervisor approval item, action contract, SLA obligation, value rollup, and replayable lifecycle.

The implementation already separates mission definition, work-item persistence, governed action execution, supervisor approval state, case value propagation, evidence/replay payloads, and role-specific rendering. That separation is what prevents a case from being just a frontend card.

Implementation Inventory

CapabilityImplemented Runtime AssetVerification / Surface
Mission registry rbac.mission_registry.MissionRegistry loads iris_missions.yaml and overlays tenant mission YAML through MissionCatalogLoader. BP canonical mission files under config/tenants/bp/missions/*.yaml provide owners, entry events, emitted events, agents, policies, stage-agent maps, evidence requirements, routing, health thresholds, and role lenses.
Mission runtime standard config/mission_runtime_standard.yaml defines the common mission contract: work queue, signal feed, agent strip, L1-L7 drawer, evidence pack, replay, policy bindings, action center, audit trail, role hierarchy, and performance budgets. The standard is exposed through /ui/bp-mission-runtime-standard and rendered inside mission capability panels.
Case/work object model P2PWorkItem, O2CWorkItem, R2RWorkItem, FPAWorkItem, TreasuryWorkItem, CloseWorkItem, and FBTWorkItem persist mission-scoped work. Each work item carries entity refs, status, priority, severity, queue, assignee, autonomy level, policy mode, SLA due time, evidence pack ID, business context, and blockers.
Decision workforce DecisionWorkforceItem, DecisionWorkforceAction, and DecisionWorkforceComment provide cross-mission decision work queues. Records include mission, decision ID, team/assignee, priority score, risk score, value at risk, entity refs, policy refs, agent runs, evidence pack ID, SLA deadline, payload, and comments.
P2P recommendations and assignment P2PWorkRecommendation and P2PAssignmentEvent attach agent-recommended actions and routing audit to P2P work items. Recommendation rows carry action payload, confidence, risk level, explanation, evidence refs, ranking, and selected status.
Mission ACT execution /api/iaf Mission ACT endpoints execute governed drawer actions through execute_mission_action, apply_work_item_action, and escalate_work_item. Durable WorkflowRunRecord, WorkflowStepRecord, and WorkflowPolicyResolutionRecord capture run state, steps, policy resolution, evidence hash, actor role, mission, entity, and status.
Universal action contract UniversalActionExecuteRequest requires mission, decision ID, entity ID, action, actor role, actor ID, reason, change-control ID where needed, and metadata. The response reports execution ID, evidence hash, policy ID/version, ledger ID, evidence pack ID, replayability, queue state, downstream effects, source-system writeback flag, and action disposition.
Shared UX case state services.ux_case_runtime produces a deterministic case object for desktop, mobile, supervisor, CFO, technology, evidence, notifications, and replay surfaces. The case contract includes status, owner, risk, value at risk, supplier, source systems, decision ID, replay ID, evidence pack ID, policy, action history, rendering mode, evidence, policy boundary, action contract, replay, value rollup, and deep links.
Case state propagation case_state_propagation_runtime.record_case_resolution writes idempotent ValueEventRecord rows and SupervisoryQueueItem updates when a case is resolved. The rollup path uses AgentDecisionLedger, ValueEventRecord, and SupervisoryQueueItem to synchronize analyst, supervisor, and CFO surfaces.
Supervisor queue SupervisoryQueueItem persists gate escalations, kill-switch reviews, drift alerts, approval-required work, priority, SLA deadline, assignment, status, resolution, and resolver. Supervisor surfaces render approval queues, SLA pressure, value, evidence sufficiency, policy boundary, action history, and available approval actions.
Mission command center /ui/mission/{{mission_id}} renders parameterized mission pages from the YAML registry and mission builders. The page integrates mission subnav, KPI health strip, entity cards, decision queue, governance grid, drawer configs, Ask Sphere panel, mission operating model, evidence/replay, agent strips, and runtime capability panels.
Mission workbench contract mission_workbench_contract normalizes analyst workbench behavior across P2P, O2C, R2R, FP&A, Treasury, Close, and FBT. Contracts define title, supervisor route, default user, routing explanation, return path, queue metadata, and mission-specific detail fields.
Mission activity and value mission_activity_metrics, mission_value_runtime, and mission_evidence_runtime provide queue signals, impact rollups, evidence packs, replay readiness, and policy-violation context. Mission metrics query bounded recent agent executions and active pending state; value attribution rolls up per mission/case through iaf_value_events.
Mission process grounding /api/mission-process-grounding returns normalized process rows, source-binding gaps, authority-risk findings, agent candidates, skill candidates, and transformation backlog. This connects process maps and fallback profiles to mission work creation without giving agents direct write authority.

Runtime Scope

  • Mission registry and tenant mission catalog with owners, KPIs, agents, policies, entry events, emitted events, routing, health thresholds, and role-specific lenses.
  • Case/work item lifecycle for P2P, O2C, R2R, FP&A, Treasury, Close, and enterprise finance trust work.
  • Decision workforce queue with priority score, risk score, value at risk, entity refs, policy refs, agent runs, evidence pack, SLA deadline, assignment, and comments.
  • Persona-scoped action contract distinguishing allowed direct actions, confirmation-required actions, approval-required actions, and blocked actions.
  • Artifact linkage through evidence pack IDs, decision IDs, replay IDs, policy IDs, agent runs, workflow IDs, value events, and supervisory queue rows.
  • SLA and escalation state through work item due timestamps, supervisory queue deadlines, priority/severity, queue IDs, assignment, and status.
  • Case closure and value propagation through idempotent value events, evidence hashes, replay refs, and CFO rollup eligibility.
  • Mission UI surfaces for analyst workbench, supervisor plane, tower leader view, CFO drilldown, case drawer, evidence pack, and replay.
  • Mission observability through bounded activity metrics, evidence payloads, queue state, backlog, active pending rows, and mission value panels.
  • Mission process grounding from BP process/data maps into source-binding gaps, authority risks, skill candidates, agent candidates, and transformation backlog.

Logical Architecture

Event / Agent / Skill / Policy / Decision / Control / User / Integration / Schedule
        |
        v
Case Signal / Work Signal / Approval Signal / Closure Signal
        |
        v
Enterprise Mission & Case Runtime
        |
        +-- Mission Registry
        +-- Tenant Mission Catalog
        +-- Mission Runtime Standard
        +-- Work Item Stores
        +-- Decision Workforce Store
        +-- Recommendation Store
        +-- Assignment Event Store
        +-- Shared UX Case Runtime
        +-- Mission ACT API
        +-- Workflow Run Ledger
        +-- Workflow Step Ledger
        +-- Workflow Policy Resolution Ledger
        +-- Supervisory Queue
        +-- Case State Propagation
        +-- Value Event Rollup
        +-- Mission Activity Metrics
        +-- Mission Evidence Runtime
        +-- Mission Process Grounding
        |
        v
Mission Queue / Case Drawer / Supervisor Queue / Tower View / CFO Rollup / Replayable Work Lifecycle

Implemented Lifecycle

  1. Register: Mission definitions are loaded from tenant YAML and overlaid into the RBAC mission registry with canonical agents, policies, entry events, emitted events, stage-agent maps, and evidence flags.
  2. Create: Signals from events, agent outputs, decisions, policy gates, or users become work items or decision workforce items with entity refs, priority, risk, value, queue, and SLA metadata.
  3. Attach: Evidence packs, agent runs, policy refs, decision IDs, workflow runs, replay IDs, value events, and queue entries are attached to the case/work item contract.
  4. Route: Assignment and queue logic route work by mission, role, amount, exception type, authority, SLA pressure, risk, and workload.
  5. Act: Mission ACT endpoints execute governed actions only after policy metadata, role authority, evidence hash, and writeback constraints are checked.
  6. Escalate: Supervisor queue items persist approval-required, gate escalation, drift, and kill-switch review work with priority, SLA deadline, assignment, status, and resolution.
  7. Resolve: Case resolution writes idempotent value attribution and propagates state to analyst, supervisor, and CFO rollup surfaces.
  8. Close: Closure requires structured outcome, evidence/replay refs, decision linkage, value attribution, and control/policy state where applicable.
  9. Observe: Mission activity metrics, evidence payloads, queue metrics, and value panels report backlog, pending review, hash verification, replayable packs, and value at risk/protected.
  10. Learn: Accepted actions, rejected recommendations, false positives, evidence gaps, SLA delays, overrides, manual remediation, and control exceptions become learning candidates through the existing case and value signal paths.

Core API and Surface Contract

Endpoint / SurfacePurposeRuntime Boundary
GET /ui/mission/{mission_id}Render parameterized mission command center from registry, builders, role, mission data, operating model, evidence, replay, and drawer configs.Mission workspace.
GET /ui/mission/{mission_id}?role=tower_leadRender supervisor/tower-lead control plane when canonical supervisor runtime applies.Supervisor view.
GET /ui/supervisor-case-queueRender shared approval queue for a governed case with policy boundary, evidence sufficiency, SLA, action history, and available actions.Case approval queue.
GET /ui/cfo-case-drilldownRender CFO read-only case drilldown with value at risk, category exposure, policy, evidence pack, and replay link.Executive case view.
POST /api/iaf/mission-act/...Execute deterministic mission action workflows and persist workflow run, step, policy resolution, evidence hash, and status.Mission ACT.
POST /api/iaf/actions/executeExecute a universal drawer action with mission, decision ID, entity ID, actor role, actor ID, policy metadata, and optional change-control ID.Governed action execution.
GET /api/mission-process-groundingReturn process grounding, source-binding gaps, authority-risk audit, agent candidates, skill candidates, and transformation backlog.Process grounding.
GET /api/mission-process-grounding/authority-riskReturn only authority-risk findings for a mission or all missions.Authority risk.
GET /api/mission-process-grounding/agent-candidatesReturn agent and skill candidates derived from mission process maps and fallback profiles.Agent/skill candidate discovery.
services.ux_case_runtime.get_case_stateBuild shared case contract for analyst, mobile, supervisor, CFO, technology, evidence, notifications, and replay surfaces.Shared case contract.
case_state_propagation_runtime.record_case_resolutionRecord idempotent case resolution value event and propagate supervisor/CFO state.Case closure/value propagation.
case_state_propagation_runtime.build_case_state_rollupBuild case rollup from value events, decision ledger rows, and supervisory queue entries.Case rollup.

Durable Mission and Case Model

ArtifactStored In / Configured ByRuntime Function
Missioniris_missions.yaml plus config/tenants/bp/missions/*.yamlBusiness mission definition, owner, KPIs, role views, agents, policies, events, evidence requirements, routing, and health posture.
Mission Runtime Standardconfig/mission_runtime_standard.yamlMandatory mission capabilities, drawer structure, role contracts, performance budgets, and credibility requirements.
P2PWorkItemiaf_p2p_work_itemsP2P decision-bearing case/work object with entity, status, priority, queue, assignment, autonomy, SLA, evidence, context, and blockers.
O2C/R2R/FPA/Treasury/Close/FBT WorkItemiaf_*_work_itemsCross-domain mission work objects with the same queue, state, SLA, evidence, and business context shape.
DecisionWorkforceItemiaf_decision_workforce_itemsMission-level decision work item with decision ID, priority/risk/value scores, entity refs, policy refs, agent runs, evidence pack, and SLA.
P2PWorkRecommendationiaf_p2p_work_recommendationsAgent recommendation with action code, action payload, confidence, risk, explanation, evidence refs, rank, and selected flag.
P2PAssignmentEventiaf_p2p_assignment_eventsRouting and reassignment audit trail.
WorkflowRunRecordworkflow_runsMission action workflow execution record with mission, entity, actor, role, status, run ID, evidence hash, and metadata.
WorkflowStepRecordworkflow_stepsStep-level mission action log with step index, status, latency, detail, and completion time.
WorkflowPolicyResolutionRecordworkflow_policy_resolutionsAuthoritative proof that policy binding and evaluation occurred before mutable ACT side effects.
SupervisoryQueueItemsupervisory_queuePersistent approval/escalation queue item with priority, SLA deadline, assignee, payload, status, resolution, and resolver.
AgentDecisionLedgeriaf_agent_decision_ledgerGoverned decision record attached to case/entity, agent, policy, evidence hash, replay hash, action, and routing target.
ValueEventRecordiaf_value_eventsCase and mission value attribution with idempotency, amount, confidence, evidence refs, replay ref, and rollup hierarchy.
ExecutionLedgeriaf_execution_ledgerEvent/tick-to-agent execution proof with status, policy/evidence flags, hashes, latency breakdown, and trigger source.

Case State and Action Semantics

Runtime ConceptImplemented RepresentationMeaning
Case identitycase_id, work_item_id, decision_id, entity_idStable business work reference across UX, work queue, decision ledger, workflow, value, and replay.
Mission ownershipmission_id, domain_code, mission_key, towerDomain and operating-model scope for routing, role views, KPIs, agents, and policies.
Statestatus, queue_state, execution_status, workflow.statusCurrent lifecycle position: awaiting action, assigned, in review, escalated, resolved, policy blocked, or workflow in progress/completed.
Prioritypriority, severity, priority_score, risk_score, value_at_riskExplainable queue order driven by risk, value, SLA, and policy/control pressure.
Assignmentqueue_id, assigned_user_id, assigned_team_id, team_id, assigned_to, SupervisoryQueueItem.assigned_toOwner, queue, or role responsible for the next action.
SLAsla_due_at, sla_deadline, minutes_to_breachTime-bound work pressure for analyst, supervisor, remediation, and approval flows.
Evidenceevidence_pack_id, evidence_hash, evidence_refsProof attached to case actions, decisions, recommendations, value events, and replay.
Policypolicy_refs, policy_id, policy_version, policy_mode, WorkflowPolicyResolutionRecordPolicy boundary attached to the case and enforced before governed action execution.
Actionsallowed_direct, confirmation_required, approval_required, blocked_for_role, action_dispositionAction availability is role-, policy-, evidence-, and state-aware.
Closurerecord_case_resolution, ValueEventRecord, resolved_at, closed_atResolved work emits value attribution and updates supervisor/CFO rollups.

Example: Duplicate Payment Case Path

{
  "case_id": "INV-778950",
  "mission": "p2p",
  "work_item": {
    "entity_type": "invoice",
    "status": "awaiting_action",
    "priority": "high",
    "queue_id": "p2p_supervisor_queue",
    "autonomy_level": "AL2",
    "policy_mode": "recommend_and_prefill",
    "evidence_pack_id": "EP-INV-778950"
  },
  "linked_artifacts": {
    "decision_id": "DEC-INV-778950",
    "policy_id": "AP-001",
    "replay_id": "REPLAY-INV-778950-CHAT",
    "workflow_run": "workflow_runs.run_id",
    "supervisory_queue": "supervisory_queue.id",
    "value_event": "iaf_value_events.value_event_id"
  },
  "action_contract": {
    "allowed_direct": ["open_case", "open_evidence", "ask_why", "open_replay"],
    "confirmation_required": ["hold_payment", "escalate_case", "request_evidence", "close_case"],
    "approval_required": ["release_payment", "approve_exception", "override_recommendation"]
  }
}

User Experience Contract

  • Analyst workbench renders mission queues, case cards, risk/value context, evidence, policy boundary, next action, and replay links from runtime contracts.
  • Case drawer exposes summary, context, evidence, agent output, skill results, policy, decision, controls, integration state, timeline, actions, and replay.
  • Supervisor control plane renders approval queues, SLA pressure, team workload, high-risk cases, missing evidence, policy-blocked actions, control exceptions, and action history.
  • Tower leader views aggregate mission backlog, aging, automation, manual approval, control exceptions, cycle time, value, and bottlenecks.
  • CFO drilldown is read-only and consumes case value rollup, materiality, evidence pack, policy state, and replay link.

Operational Requirements

  • Do not treat a rendered UI card as a case. Runtime-owned work item, decision workforce, workflow, supervisory queue, value event, and replay records are the source of truth.
  • Do not let mission pages invent state. UI should render runtime state, policy boundary, action contract, evidence sufficiency, and SLA pressure.
  • Do not execute a drawer action without policy metadata, actor role, evidence hash, decision/entity context, replayability, and source-system writeback disposition.
  • Do not close high-risk work without evidence, decision, policy, control, approval, value, and replay lineage where those artifacts are required.
  • Duplicate signals should correlate to an existing work item, decision item, or case state before creating new work.
  • Mission definitions should remain business-owned in tenant configuration while runtime contracts remain platform-owned.

Acceptance Criteria

  • Mission pages are registry-driven, not hardcoded per role, and resolve mission agents, policies, events, KPIs, role lenses, and runtime tier metadata.
  • Work item records carry entity identity, state, priority, queue/assignment, autonomy level, policy mode, SLA due date, evidence pack, business context, and blockers.
  • Decision workforce records carry mission, decision ID, priority score, risk score, value at risk, entity refs, policy refs, agent runs, evidence pack ID, SLA deadline, and comments.
  • Governed actions require policy metadata and return execution ID, evidence hash, policy version, replayability, source-system writeback disposition, and queue state.
  • Workflow run, workflow step, and workflow policy resolution rows prove mission action execution and policy resolution before mutable side effects.
  • Supervisor approval work persists as SupervisoryQueueItem with priority, SLA deadline, assignment, status, resolution, resolver, and payload.
  • Case state propagation is idempotent and writes value events keyed by case, decision action, actor, evidence hash, and replay reference.
  • CFO and supervisor surfaces consume the same case state contract as analyst and technology surfaces instead of recomputing separate case definitions.
  • Mission evidence payloads report evidence packs, hash verification, replayable packs, pending review, source docs, integrity failures, and policy violations without breaking page rendering on query failure.
  • Case closure and high-risk action visibility remain policy-bound, evidence-backed, replayable, and role-aware.

Engineering Rule

Treat mission cases as governed business work objects. A valid case has identity, mission scope, state, priority, assignment, SLA, action contract, artifact links, policy/evidence/decision context, value attribution, and replayability. Anything less is only a visual representation, not the Mission & Case Runtime.