bp Sphere Technology, Data & Governance Deep Dive

bp Sphere · enterprise execution runtime, governance, replay, evidence, resilience, and operating model
Standalone deep-dive route

bp Sphere Enterprise Execution Runtime

bp Sphere is being positioned here as a governed enterprise orchestration and evidence runtime operating across bp’s existing AWS, Azure, SAP, Ariba, Databricks, UDP, and operational control landscape. This page is intentionally structured as a slide-style solution narrative: each section answers why bp needs the capability, what problem it solves, how it works, how it would be implemented, and what questions bp still needs to answer before execution planning.

Audience
Architecture, Security, Data, Cloud, Finance, Integration, Ops
Stance
Governed enterprise execution, not AI theater
Phase 1 motion
Read-only observe and assist
Control principle
Policy-bound, replayable, human-supervised
Outcome
Architecture and governance alignment with next-step decisions
Introduction

What this deep dive must answer for bp

This is not an AI demo. It is the technology, data, governance, cyber, resilience, digital workplace, and controllership deep dive that must show how bp Sphere fits into bp’s real architecture and operating model without forcing ERP replacement or uncontrolled autonomy.

Purpose

  • Validate integration into bp technology architecture.
  • Align on data quality dependencies and the data governance model.
  • Evaluate supplier approaches to cyber security, access management, disaster recovery, resilience, and digital workplace change.

Expected outputs

  • Validated integration approach into bp technology architecture.
  • Clarity on additional inputs suppliers need to refine their solution.
  • Clarity on what additional details bp still needs from suppliers.

Control questions

  • How does governance scale safely across humans and agents?
  • How are replay, evidence, and controllership preserved?
  • What are the key risks, assumptions, and dependencies for pilot execution?

P2P analyst mission lens

This page also needs to explain what this architecture means for a real operator. For a P2P analyst, bp Sphere should not feel like another dashboard. It should feel like a governed decision workspace where live queue priorities, agent recommendations, escalation context, evidence, replay, and enterprise impact stay in one mission surface.

  • Priority queue replaces static worklists with governed decisions that need action now.
  • Human-agent collaboration stays inline through recommendations, evidence requests, and supervisor escalation.
  • Replay and proof stay attached to each case so analysts can trust why something was flagged.
  • Business impact remains visible so analysts act as financial flow operators, not AP clerks clearing tickets.
Platform Deep-Dive

The five planes that complete the architecture story

Beyond the orchestration runtime, replay, and controlled execution that the solution covers as a whole, five architectural planes determine whether bp Sphere holds up under real enterprise scrutiny: how agents are authored, how their reasoning is intermediated, how they are grounded and remembered, how their behaviour is continuously evaluated, and how they coordinate across domains. Each plane is unpacked in the slide it most naturally belongs to.

Slide 04

Author plane & Model Gateway

Skill Studio plus signed registry, paired with an LLM gateway that redacts, routes, moderates, and meters every reasoning call.

Slide 05

Engineering trace & evaluation

Per-step reasoning trace plus offline and online evaluation, sitting beneath business replay and sharing the same evidence chain.

Slide 06

Multi-agent coordination

Mediator plus devil's-advocate pattern, quorum rules, and collusion defences for any cross-domain decision.

Slide 08

Authored, owned, governable

Every Phase 1 agent is a signed artefact in the bp Sphere registry with a named owner, opening the path for ops teams to become creators of agentic capability.

Slide 10

Knowledge plane & agent memory

Unstructured grounding from policy and runbook documents, plus three memory stores agents use across runs without writing back to systems of record.

Cross-cutting

One evidence chain

Author signature, model gateway log, trace span, eval verdict, multi-agent dissent, and replay entry stitch into a single audit-grade chain.

Slide 01 · Executive context

What bp is really asking for and what this deep dive must prove

Credibility first No ERP replacement narrative

bp is not evaluating AI in the abstract. bp is evaluating whether AI can operate safely inside a regulated enterprise. This deep dive must show that bp Sphere can operate inside bp’s environment without forcing replatforming, without bypassing governance, and without pretending that generic copilots equal enterprise execution.

What we heard from bp

  • bp does not want ERP replacement, platform lock-in, or another disconnected AI estate.
  • The runtime must align to existing AWS, Azure, IAM, Entra ID, and cloud guardrail standards.
  • Productivity must be realized with operational safety, bounded writes, and named accountability.
  • The enterprise is federated: fragmented ownership, mixed system maturity, regional process variants, and uneven data quality are the reality.
  • Governance, controllership, evidence, and operational trust matter more than novelty.
  • Transformation must be phased, measurable, human-supervised, and politically adoptable.

What this deep dive must establish

Enterprise credibility

Show that the orchestration model fits a real bp estate rather than a toy greenfield runtime.

Governance trust

Make policy, approval boundaries, replay, evidence, and rogue-agent prevention visually explicit.

Operational realism

Show human-agent collaboration, escalation, and business impact through finance operating moments.

Architecture defensibility

Explain what bp already has, what AWS/Azure provide, and what bp Sphere uniquely adds.

Transformation practicality

Demonstrate a progressive roadmap starting with read-only agents rather than uncontrolled writes.

Differentiation

Make it clear why this is not 'just copilots', 'just workflows', or 'just cloud AI primitives'.

Boardroom message: the challenge is no longer software. The challenge is fragmented decisions, fragmented visibility, and fragmented accountability across a strong but disconnected estate.

Day 2 deep-dive purpose

  • Validate integration into bp technology architecture rather than propose a parallel stack.
  • Understand dependency on data quality and align the data governance model needed for credible operationalization.
  • Evaluate supplier approaches to cyber security and access management, disaster recovery and resilience, and digital workplace change.
  • Surface the risks, assumptions, and dependencies that shape pilot scope, control design, and rollout readiness.
  • Clarify additional inputs suppliers need from bp and what additional questions bp needs answered before solution refinement.
  • Align on the controllership operating model and the human-governance expectations behind future enterprise autonomy.

Expected outputs from this deep dive

  • Validated integration approach into bp technology architecture.
  • Clarity on data quality dependencies and the aligned data governance model.
  • Validated supplier approaches to cyber security and access management, disaster recovery and resilience, and digital workplace transformation.
  • A concrete list of risks, assumptions, dependencies, and architecture decisions requiring follow-up.
  • Clarity on additional inputs suppliers need to refine their solution and what additional details bp still needs from suppliers.
  • A credible path to align the controllership operating model with replay, evidence, policy runtime, and human-supervised execution.

Questions to resolve with bp across data, security, runtime, and operating model

Data readinessIdentity + accessRuntime standardsResilience expectationsDigital workplaceControllership model

Each later slide expands one of these domains so the deep dive can end with concrete follow-up actions instead of generic agreement.

Credibility test for the page

  • Can bp architecture teams see how the runtime integrates without ERP replacement?
  • Can data teams see the dependency on governed products, lineage, and quality stewardship?
  • Can security teams see IAM, workload identity, cyber telemetry, and controlled write boundaries?
  • Can controllership leaders see how replay, evidence, approvals, and policy runtime preserve trust?
  • Can both sides leave with explicit decisions, open questions, assumptions, and next-step inputs?
Slide 02 · Current enterprise environment

Strong foundation. Fragmented execution.

Federated enterprise reality

bp already has major systems, cloud, and data investments. Systems are not broken. Execution is fragmented. The gap is enterprise coordination, operational trust, and the ability to turn fragmented signals into governed, replayable execution across domains.

Enterprise estate in scope

SAP ECC / S4HANAAribaBlackLineTreasury platformsUDPDatabricksPalantirAWSAzureLaunchPad / YalaRegional ERP variantsFinance modernization

The gap

  • Approvals, policies, workflows, AI pilots, and analytics remain functionally or technically siloed.
  • Operational intelligence is fragmented across dashboards and system-specific views instead of one supervised runtime.
  • Logging exists, but enterprise replay, evidence lineage, and governance reconstruction are not yet first-class operating assets.
  • Cross-domain orchestration across finance, procurement, treasury, and operations is still limited.

Current state → transitional state → future state

Current

  • Fragmented approvals
  • Siloed automation
  • Limited enterprise replay
  • Manual exception triage
  • Cross-domain visibility gaps

Transitional

  • Read-only intelligence agents
  • Replay-backed visibility
  • Supervisor escalation surfaces
  • Policy-aware prioritization
  • Governed exception operations

Future

  • Cross-domain coordination
  • Bounded autonomy
  • Enterprise decision runtime
  • Institutional operational memory
  • Digital twin and strategic intelligence
Problem statement: analysts still search across systems, controllers still reconstruct evidence manually, auditors still rebuild history, and leaders still lack one enterprise view of decision posture.
Slide 03 · Strategic positioning

bp Sphere is the enterprise execution runtime, not another dashboard or copilot

bp Sphere throughout

The strongest message is not “we have features.” It is “traditional dashboards tell you what happened; bp Sphere governs what happens next.” This slide must visually separate static analytics, copilots, workflow engines, and governed operational orchestration.

What this is not

  • Not an ERP replacement
  • Not a separate cloud platform
  • Not a menu of disconnected copilots
  • Not another BI surface with AI labels
  • Not uncontrolled autonomous AI

What AWS / Azure already provide

  • Compute, model runtimes, workflow primitives, logging, eventing, and infra observability
  • Useful AI and automation building blocks
  • Strong cloud services, but not enterprise semantic orchestration or supervisory governance

What bp Sphere adds

  • Enterprise orchestration runtime
  • Policy-bound execution boundaries
  • Replay and evidence contracts
  • Supervisory governance and escalation
  • Cross-domain semantic context and operational intelligence
Positioning statement: Claude can reason. Azure can host. AWS can orchestrate. bp Sphere governs. bp already owns the systems of record, the cloud, the IAM estate, and the data programs. bp Sphere becomes the connective intelligence and governance fabric that makes those assets operate together safely.
Slide 04 · Enterprise orchestration architecture

How bp Sphere operates across the enterprise: from hybrid signals to governed execution and replay

Architecture visual

Every transaction follows the same lifecycle: supplier invoice, treasury payment, journal entry, or forecast revision. bp Sphere uses one runtime for every business process, every agent, and every decision: signal, context, evidence, policy, decision, execution, and replay.

Enterprise sources
Systems of record
SAP, S/4HANA, Ariba, BlackLine, treasury, procurement, finance close, supplier and operational events.
Existing fabric
bp cloud + integration
AWS, Azure, API gateways, event buses, middleware, IAM, UDP, Databricks, Palantir, regional ownership boundaries.
bp Sphere
Signal and semantic intake
Canonical event overlay, ontology resolution, context normalization, cross-system correlation, workload identity propagation.
bp Sphere
Agentic orchestration runtime
Specialized operational, governance, and intelligence agents coordinate through events, policy, and bounded autonomy tiers.
Governance core
Policy + execution gates
Approval boundaries, SoD enforcement, authority thresholds, risk classification, execution gateway, human override, fail-safe throttles.
Trust layer
Evidence + replay
Lineage, hashes, manifests, approval trace, agent participation chain, replay reconstruction, audit reconstruction, forensic evidence.
Operating surfaces
Human + agent runtime
Analyst decision theaters, supervisor command surfaces, runtime control plane, treasury/close impact views, operational intelligence surfaces.
Outcome
Governed enterprise execution
Controlled writes, measurable productivity, working-capital intelligence, supplier trust protection, replayable operations, institutional learning.

Author plane and Model Gateway

How agents enter the runtime, and how their reasoning is intermediated

Two planes sit either side of the orchestration runtime above and determine whether the agent estate remains governable as it grows. On the inbound side, the Author plane turns business intent into signed, registered agent artefacts. On the outbound side of every reasoning step, the Model Gateway intercepts the call to any LLM for redaction, routing, moderation, and metering. Together they make the difference between a controlled estate and a sprawl of unaccountable agents.

Why this matters for bp

The model behind any agent is interchangeable without re-certifying the agent itself because the boundary is the gateway, not the model. Every new agent enters the estate through a single bp-owned author surface, which means no shadow deployments, no untracked prompts, and no off-fabric model calls.

Model Gateway

LLM intermediation

  • PII redaction, prompt-injection and jailbreak scanning, model routing across Azure OpenAI, Bedrock, and private models.
  • Output moderation, citation checks, and schema validation before responses re-enter the runtime.

Cost and trace emission

Every reasoning call emits tokens used, latency, model selected, redaction statistics, and a trace span, feeding FinOps, SOC monitoring, and the engineering trace surface introduced on Slide 05.

Author plane

Skill Studio

Declarative agent specs in Markdown or YAML with tool binding, knowledge attachment, prompt template, and an eval set. No-code for domain experts, with pair-authoring alongside engineering for tool makers.

Registry & promotion pipeline

Signed artefacts for skills, tools, MCP servers, models, prompts, and knowledge packs. Promotion runs through sandbox, shadow, canary, bounded, and general, and each gate requires a passing eval and council sign-off.

Position in the architecture flow
Author plane
Skill Studio plus Registry. Specs flow into the orchestration runtime as signed, versioned agents with no bespoke deployments.
bp Sphere
Agentic orchestration runtime
Coordinates the registered agents arriving from the author plane.
Model Gateway
LLM intermediation
Every reasoning step routes through the bp-controlled gateway before reaching any model.
Governance core
Policy + execution gates
Also receives prompt-injection and jailbreak signals from the gateway alongside its existing controls.

Local-to-bp production integration mapping

Current demo proof bp production target Integration pattern Governance model
Local SAP invoice simulation SAP S/4HANA / CFIN OData, BAPI mediation, SAP event mesh, canonical event overlay bp IAM, approval audit, replay-backed write attribution
Local procurement and supplier state Ariba and supplier workflows APIs, webhooks, supplier event normalization, policy-bound gateway Supplier stewardship, policy-bound workflow mutation, named owner
Local replay queue and runtime events Splunk, SIEM, SOC, and runtime ops tooling OpenTelemetry spans, event forwarding, replay event export Immutable retention, security forwarding, privileged traceability
Local evidence manifests UDP / Databricks governed products Lineage-aware evidence linkage, governed product references, metadata joins Steward-owned lineage, retention policy, evidence schema contract
Local runtime-health and deployment proof CloudWatch, Azure Monitor, LaunchPad / Yala ops views Health APIs, runtime control tower aggregation, deployment-state federation bp runbooks, on-call ownership, change and incident controls

Critical production message

bp Sphere operates inside bp’s governance and cloud boundaries rather than introducing a parallel operating model.

  • The demo surfaces prove orchestration behavior; the production mapping above proves how those behaviors swap onto bp-approved systems and controls.
  • No shadow IAM, no parallel data regime, and no hidden direct-write path are introduced in order to move from the local proof estate into production.
  • Every local proof path must point to a named bp production target, a concrete integration contract, and a clear governance owner before pilot approval.
Slide 05 · Signal to replay lifecycle

Signal-to-replay lifecycle: every operational moment must flow to replay and evidence

Signature trust visual

This is the page’s most important trust diagram. It shows how operational signals become governed decisions, controlled execution, evidence packs, and replayable enterprise history. It should be the visual anchor for duplicate payment, policy breach, and close-pressure demos.

01
Signal
Invoice, policy, treasury, close, supplier, or data-quality event enters the runtime through API or event observation.
02
Context
bp Sphere resolves enterprise context: supplier, business unit, authority, working-capital posture, and prior incidents.
03
Policy
Execution boundaries, approval thresholds, SoD rules, and risk-tier guardrails are evaluated before any action.
04
Orchestration
Specialized agents coordinate classification, matching, recommendation, escalation, and evidence capture.
05
Decision
A recommendation, hold, escalation, or governed automation path is produced with confidence and alternatives.
06
Controlled execution
Writes flow only through approval gates and execution gateways; no rogue direct write path exists.
07
Evidence
Lineage, hashes, approvals, and agent participation are packaged into an immutable evidence contract.
08
Replay
The decision is reconstructable end to end for audit, forensic review, operator trust, and runtime learning.
09
Learning
Patterns from replay, overrides, and outcomes improve prioritization and bounded autonomy over time.
Why this matters: cloud logs are not enterprise replay. Replay here means reconstructing the operational chain: signal, context, policy checks, agent coordination, human approvals, execution outcome, and the evidence manifest that proves what happened.

Engineering trace and evaluation harness

The surfaces beneath replay

Replay is the controller-facing record of what the enterprise did. Beneath it, bp Sphere runs two further surfaces: a per-step engineering trace that explains why an agent reasoned the way it did, and a continuous evaluation harness that proves the agent estate is not getting quietly worse over time. The three views share a single evidence chain: replay for controllers, trace for engineering and risk, eval for early-warning drift.

Why this matters

Replay reconstructs the operational decision for controllers. Trace reconstructs the reasoning path for engineers. Eval continuously verifies that the reasoning path still matches the gold standard, turning drift from a post-incident discovery into a leading indicator.

Trace surface

Per-step LLM and tool observability

  • LLM span: prompt, completion, tokens, latency, model, and cost.
  • Tool span: input, output, retries, and side-effects.
  • Retrieval span: query, chunks fetched, and citation overlap.
  • Hallucination markers when output diverges from citations.
  • Searchable by agent, user, supplier, and time window.
  • Diffable across model versions and prompt versions.

Evaluation harness

Offline and online quality gating

  • Golden suite of 100-500 labelled cases per agent, owned by the agent author.
  • Regression gate: any change to prompt, model, tool, or knowledge re-runs the evals.
  • Online probes: sampled live decisions checked against an oracle.
  • Faithfulness scoring: output measured against retrieved citations.
  • Drift dashboards covering confidence, override rate, and escalation rate.
  • Quarterly adversarial red-team probes for every high-risk agent.

View 1

Business replay

Signal to context to policy to orchestration to decision to execution to evidence to replay to learning. Controller and auditor surface.

View 2

Engineering trace

Per-step LLM and tool spans, retrieval chunks, and faithfulness scores. Engineer and risk surface, linked to the same evidence pack.

View 3

Evaluation harness

Golden-set regression gates and online drift dashboards. Early-warning surface for the governance council.

Live runtime proof

This deep dive must not stop at diagrams. These live surfaces prove the runtime is operating, governed, and replayable inside the current estate.

What each surface proves

  • High-risk runtime demo: one operational moment from signal to policy to escalation to evidence to replay.
  • Control tower: live agents, engineering trace, eval status, deployment state, policy gates, and evidence integrity.
  • Operator console: production swap sequence, solution choreography, and the presenter-facing proof path across runtime, resilience, and assurance.

Runtime visibility expectations

  • Active agents and policy evaluations must be visible as operational state, not hidden inside backend logs.
  • Evidence completeness, replay backlog, deployment state, and AWS/Azure runtime location must be inspectable during the deep dive.
  • The page should leave no doubt that bp Sphere has a control plane, not just a slide narrative.
Slide 06 · Rogue agent prevention

Controlled execution and rogue-agent prevention architecture

Governance-safe execution

This slide must remove fear. The point is to show that phase 1 is read-only, that controlled execution is explicit, and that no agent can write into SAP, Ariba, or treasury systems outside governed pathways.

Agent recommendation zone

  • Read-only intake and semantic correlation
  • Risk scoring, prioritization, and recommendation
  • Evidence generation before execution
  • No direct write permission to core systems

Policy + approval gateway

  • Approval matrix and authority threshold checks
  • SoD enforcement and supplier exception rules
  • Escalation gate when risk, value, or evidence completeness demands it
  • Named human accountability with override and abort path

Controlled execution gateway

  • Execution only through bounded adapters
  • Replay and evidence package attached to every action
  • Post-action observability, fail-safe, and recovery hooks
  • No uncontrolled write, no hidden mutation, no black-box action path

Critical first-phase rule

  • Phase 1 agents observe, classify, correlate, prioritize, escalate, and replay.
  • Phase 1 agents do not post journals, execute payments, mutate workflows, or change master data.
  • This reduces governance resistance while creating immediate productivity and visibility.

What bp stakeholders should hear

  • This is not autonomous AI executing in production without oversight.
  • This is a controlled enterprise intelligence and orchestration layer with explicit write boundaries.
  • Replay and evidence are non-negotiable control features, not optional diagnostics.

Multi-agent coordination

Conflicts, looping handoffs, and collusion

The three zones above govern a single agent's path to action. As soon as two or more agents share authority over the same supplier, invoice, or treasury decision, a distinct set of failure modes appears: conflicting recommendations, looping handoffs, capability creep, and most dangerously, two weak agents quietly reinforcing each other's bad call. bp Sphere addresses this with an explicit coordination plane that sits on top of the gateways above.

Most dangerous failure mode

The most dangerous failure mode in a multi-agent estate is not a single rogue agent. It is two agents quietly agreeing with each other when they should not. The mediator and devil's-advocate pattern makes that structurally hard, and the evidence pack captures the disagreement so controllers can see exactly how consensus was reached.

Coordination primitives

  • Mediator agent required for any cross-domain decision and never optional.
  • Typed message contract covering claim, evidence, confidence, and dissent.
  • Conflict policy that tie-breaks by risk class and escalates by default.
  • Deadlock breaker with bounded hop count and wall-time before automatic human handover.
  • Quorum rules so high-risk decisions require independent agents in agreement.

Collusion and sycophancy defences

  • Independent reasoning: agents see each other's evidence, never each other's drafts.
  • Diverse models: high-risk multi-agent flows route across at least two model families.
  • Devil's-advocate agent argues against the emerging consensus and its output is preserved in evidence.
  • Anti-sycophancy probes inject false premises to test whether agents push back.
  • Collusion detection flags cases where agents agree faster than baseline on a high-risk case.

Step 1

Specialist agents

Invoice Risk, Treasury Exposure, and Supplier Risk each produce an evidence pack independently.

Step 2

Mediator agent

Reconciles claims, surfaces dissent, and applies quorum and risk-class rules.

Step 3

Devil's-advocate

Required for higher-tier decisions, and its argument is captured in evidence regardless of the final call.

Step 4

Policy gateway

Receives a single mediated recommendation together with the full dissenting trace.

Cross-domain decision pattern

The coordination plane turns parallel specialist outputs into one governed recommendation with explicit dissent, bounded escalation, and a preserved audit trail for controllers and risk teams.

Slide 06A · Agent execution runtime

Agent execution runtime: one runtime for many agents

Execution discipline

This slide answers a common boardroom question directly: if there are many agents, who controls them? The answer is that there are not many separate runtimes. There is one governed runtime through which all agents operate.

Business event
Signal intake
Invoice, trade event, journal request, forecast revision, treasury alert.
Context layer
Situation model
Business, process, policy, historical, user, and financial context.
Evidence layer
Evidence assembly
Documents, approvals, records, communications, and prior decisions.
Policy runtime
Guardrail evaluation
Authority matrix, SoD, thresholds, and risk policy.
Agent runtime
Specialized reasoning
Domain agents reason, score, recommend, and coordinate.
Human governance
Approve or escalate
Analyst, supervisor, controller, or executive review path.
Execution gateway
Controlled action
Only approved actions reach enterprise systems.
Replay runtime
Enterprise memory
Every step remains reconstructable and attributable.
Key message: 130 agents do not mean 130 operating models. They operate through one runtime with one set of evidence, policy, approval, replay, and telemetry contracts.
Slide 06B · Enterprise context layer

Enterprise context layer: the difference between answering questions and making decisions

Context as operating asset

Without context, AI answers prompts. With enterprise context, AI makes governed decisions. This is where bp Sphere becomes different from standalone copilots or document search tools.

Business context

Legal entity, business unit, desk, portfolio, supplier, customer, asset, cost center, and materiality.

Process context

Current workflow stage, prior approvals, open exceptions, unresolved reconciliations, and case age.

Financial context

Exposure, liquidity, margin, budget, actuals, plan, working capital, and close pressure.

Policy context

Authority matrix, SoD, treasury policy, procurement policy, forecast policy, and execution boundaries.

Operational context

Queue state, source freshness, runtime status, upstream failures, and control-plane conditions.

Historical context

Similar cases, accepted overrides, prior outcomes, dispute patterns, supplier behavior, and decision history.

Why bp needs this

  • Because the current environment stores facts but does not assemble operating context across systems.
  • Because analysts still spend time collecting context before they can apply judgment.
  • Because without context, AI becomes generic assistance instead of enterprise decision support.

Questions for bp

  • Which context dimensions are already trusted and production-ready?
  • Which domains suffer most from missing historical or policy context today?
  • Who owns freshness and stewardship for each context class?
Slide 06C · Evidence intelligence fabric

Evidence intelligence fabric: the trust layer reused across every mission

Differentiator

Evidence intelligence is one of the clearest reasons bp Sphere is not just another AI layer. It assembles the proof package behind every recommendation and makes that package reusable across domains.

Evidence discovery

Find the records, documents, approvals, communications, and prior decisions relevant to the case.

Evidence resolution

Resolve duplicates, conflicting versions, and related entities across systems and repositories.

Evidence extraction

Extract clauses, milestones, invoice fields, approval facts, and policy-relevant attributes.

Evidence lineage

Preserve source references, timestamps, hashes, authorship, and transformation history.

Evidence packaging

Package the decision trail into one controller-facing and auditor-facing proof set.

Evidence viewer

Expose that package inline for analysts, supervisors, controllers, and approvers.

Reuse story: the same evidence fabric supports journal validation, contract validation, credit risk, maintenance verification, forecasting, and close intelligence. That is platform leverage, not use-case sprawl.
Slide 07 · Progressive transformation roadmap

Progressive transformation roadmap: current state to strategic intelligence enterprise

Transformation slide

This roadmap positions bp Sphere as a phased modernization journey rather than a big-bang replacement. It emphasizes coexistence, progressive orchestration, bounded autonomy, and the evolution of the human operating model.

Phase 0
0-3 months

Foundation alignment

Target architecture, IAM, runtime certification, integration strategy, semantic model, replay/evidence standards, and AI governance council alignment.

Architecture

Define orchestration runtime positioning across AWS, Azure, LaunchPad/Yala, and enterprise integration standards.

Security

Align workload identity, privileged execution policy, telemetry forwarding, and approval requirements.

Data

Identify governed data products, stewardship, quality gaps, and semantic normalization strategy.

Phase 1
3-9 months

Observe and assist

Read-only operational intelligence agents, replay, evidence, runtime observability, analyst copilots, and supervisor visibility without uncontrolled writes.
Phase 2
9-18 months

Governed automation

Supervised orchestration, policy-bound automation, controlled execution gateways, and human oversight of bounded write-back actions.
Phase 3
18-30 months

Cross-domain coordination

Finance, procurement, treasury, supplier, and close operations coordinate through one enterprise semantic and event fabric.
Phase 4-5
30+ months

Enterprise decision runtime to strategic intelligence

Bounded autonomy, digital twin simulation, institutional learning, dynamic policy optimization, and predictive enterprise operations.
Slide 08 · Phase 1 read-only agents

Initial read-only agents and their role in building trust

Safest starting point

The first wave should not focus on autonomous execution. It should focus on operational intelligence, replayability, policy visibility, and supervisor confidence. These agents observe, correlate, classify, explain, prioritize, and escalate without mutating enterprise systems.

Recommended initial agent set

Invoice Risk Intelligence AgentReconciliation Intelligence AgentTreasury Exposure Intelligence AgentPolicy Violation Intelligence AgentOperational Replay AgentRuntime Observability AgentClose Confidence AgentSupplier Risk Correlation AgentException Prioritization AgentData Quality Intelligence Agent

What they prove

  • Enterprise visibility and prioritization without governance fear
  • Replay, evidence, and supervisor surfaces as first-class capabilities
  • Operational productivity gains before write automation
  • Semantic intelligence and cross-system correlation without replatforming

Human + agent operating model for phase 1

Analysts
Exception governance
Review high-value or high-risk exceptions, evaluate evidence, and approve or escalate decisions.
Supervisors
Threshold management
Manage escalations, capacity, approval routing, and early governance signals.
Read-only agents
Observe and assist
Continuously monitor signals, classify risk, surface bottlenecks, and prepare replay/evidence packages.
Governance teams
Assure and certify
Define runtime controls, review drift, and certify policy boundaries before broader automation.
Finance leaders
Operational intelligence
See close confidence, liquidity exposure, supplier posture, and controllership impact from one surface.
Outcome
Trust accumulation
Enterprise confidence in AI grows through visibility, replay, and bounded operating behavior.

How these Phase 1 agents are authored and owned

Each of the ten agents above is a signed artefact in the bp Sphere registry introduced on Slide 04, with a named bp owner. The owner is the author, the on-call, and the person who approves promotion or retirement. That ownership model makes the Phase 1 estate governable from day one, and it opens the path for ops folks across AP, treasury, controllership, and supplier ops to become creators of agentic capability rather than only consumers of it.

Named owner for every agent

Each agent is owned by a senior ops SME in the relevant domain. The owner authors the spec, owns the eval suite, signs off on promotion, and decides when to retire. Accountability is structural, not procedural.

Citizen authoring open from Day 1

Four of these ten agents, Invoice Risk, Reconciliation, Policy Violation, and Data Quality, are explicitly designed to be authored or refined by ops folks through Skill Studio, with platform engineering pairing only on tool-making. This is how bp grows the agent estate without bottlenecking on engineering capacity.

Author = named owner

Operational accountability and runtime accountability live with the same person.

Spec + evals + sign-off

The authored spec, owned eval suite, and promotion decision form the evidence root.

Registry-backed

Every agent is signed, versioned, and promoted through the registry rather than bespoke deployment paths.

4 / 10 citizen-authorable on Day 1

The initial estate is deliberately designed so a meaningful subset can be created and refined by domain operators, not only by engineering teams.

First entries in bp's Skill Studio registry

These ten agents are the first signed entries in bp's Skill Studio registry. Each one carries a named owner, an eval suite, and an explicit retire-or-renew decision path. The estate grows from this registry, not from one-off bespoke deployments.

Slide 09 · Capability matrix

bp vs AWS/Azure vs bp Sphere capability heatmap

Architecture ownership clarity

This matrix is one of the most important artifacts in the narrative. It clarifies what bp already owns, what cloud providers supply, what bp Sphere uniquely contributes, and where the enterprise gaps exist today.

Capability domain bp internal AWS / Azure bp Sphere Gap / implication
Identity and security Entra ID, IAM, cloud guardrails, existing access standards Workload identity primitives, IAM roles, managed identities Agent identity governance, named accountability, policy-aware permissioning Major gap is AI/runtime-specific access governance rather than raw IAM absence.
Runtime infrastructure Hybrid AWS/Azure compute and container runtime AI runtimes, workflow services, autoscaling, eventing primitives Governed orchestration runtime, cross-cloud coordination, supervisory control logic Clouds provide primitives; bp Sphere provides enterprise orchestration semantics.
Integration and eventing Fragmented APIs, middleware, regional patterns EventBridge, Event Grid, API management, workflow steps Canonical context, semantic event overlay, cross-system normalization Major gap is enterprise semantic coordination, not API connectivity alone.
Data and semantics UDP, Databricks, governed data initiatives, stewardship programs Cloud data tooling and lineage utilities Operational context layer, ontology runtime, digital twin context Semantics and enterprise operational context remain immature today.
Agentic orchestration Minimal enterprise lifecycle governance for agents Agent frameworks and workflow helpers Multi-agent coordination, escalation runtime, bounded autonomy, institutional learning This is one of the biggest differentiators versus native cloud AI.
Governance and policy Existing enterprise policies and fragmented approvals Cloud guardrails, basic AI guardrails, not enterprise operations governance Runtime risk classification, approval governance, controlled execution, autonomy governance Major gap today is policy-aware operational AI, not policy absence.
Replay and evidence Logging and audit, but limited operational replay or cryptographic evidence Logs and monitoring, not enterprise replay Decision replay, evidence manifests, lineage, governance replay, trust runtime This is a primary strategic differentiator and skeptic-killer capability.
Human supervisory runtime Analyst workflows and fragmented escalation No native supervisor coordination model Decision theaters, supervisor command surfaces, dynamic thresholds, runtime interventions Needed to make AI augmentation culturally and operationally credible.
Observability and control plane Infrastructure monitoring Cloud telemetry and logs Operational control plane, semantic observability, escalation visibility, replay status Infrastructure telemetry does not equal enterprise operational observability.
Slide 09A · Data readiness service

Data readiness service: trust must be measurable before autonomy is discussable

Data trust

A credible enterprise platform needs a visible answer to data trust. The question is not only whether data exists. The question is whether it is complete, fresh, owned, traceable, and reliable enough for governed decisions.

Completeness

Required fields, relationships, attachments, and event coverage present for the decision type.

Freshness

Current enough for the operating decision and within an agreed SLA by source.

Lineage

Source, transformation, and stewardship path visible to operators and auditors.

Ownership

Named steward and operational team accountable for quality and remediation.

Trust

Quality signals aggregated into one business-facing confidence posture.

Criticality

Material sources receive stronger monitoring and lower tolerance for degradation.

Illustrative readiness index

Domain / sourceReadinessImplication
SAP finance events92%Suitable for governed observation and decision support.
Databricks planning products88%Strong for FP&A context, but needs freshness governance on critical cycles.
Supplier master61%High governance risk; Sphere should expose this weakness rather than hide it.

Why bp needs this

  • Because trust concerns are often really data-trust concerns expressed in business language.
  • Because the runtime should degrade safely when readiness falls below policy thresholds.
  • Because bp leadership needs to see where autonomy is blocked by data quality, not just by policy caution.
Slide 10 · Federated data and cloud federation

Federated data strategy and AWS + Azure federation model

Data and cloud alignment

bp does not need another giant centralized platform pitch. It needs a coherent explanation of how bp Sphere overlays a federated data and cloud estate, leverages existing UDP and Databricks investments, and coordinates signals across AWS and Azure without pretending the enterprise will be homogeneous.

AWS domain
Operational and integration workloads
Event sources, containers, step primitives, APIs, runtime services, and domain systems already operating inside AWS-aligned controls.
Azure domain
Identity and AI estate
Entra ID, Azure AI, analytics workloads, enterprise governance services, and additional runtime estates already approved inside Azure boundaries.

bp Sphere federation spine

  • Canonical enterprise event overlay across AWS and Azure
  • Semantic context and ontology resolution over federated data products
  • Policy and identity propagation from bp-approved IAM models
  • Replay and evidence runtime independent of any single cloud primitive
  • Cross-cloud orchestration with local execution boundaries preserved
Data foundation
UDP + Databricks + governed products
Use existing data products, lineage, stewardship, and access patterns as the substrate for context and intelligence rather than forcing a new data monopoly.
Enterprise result
Federated operational intelligence
The runtime sees enterprise motion without requiring every source to be centralized or every process to be standardized first.

Data strategy expectations from bp

  • Clarify which finance and operational datasets are already governed and production-ready.
  • Map stewardship, ownership, lineage, and metadata standards that the runtime must honor.
  • Treat semantic normalization as a progressive program, not a prerequisite big-bang cleanup exercise.
  • Use federated data intelligence rather than forced centralization where not politically or technically realistic.

Questions that belong on the slide

  • Which datasets are already approved for operational AI observation?
  • What lineage and metadata standards are already required?
  • Where are the biggest current data quality pain points in finance operations?
  • How does bp want long-term semantic federation to evolve across UDP, Databricks, and domain teams?

Knowledge plane and agent memory

Grounding for accurate, citation-backed reasoning

The federation spine above carries bp's structured data estate: UDP, Databricks, SAP, and Ariba. Agents operate on a parallel plane of unstructured knowledge, policy PDFs, supplier MSAs, controllership runbooks, and SOX documentation, and on distinct memory stores that allow them to be useful across runs without writing back into systems of record. The plane below reuses the same stewardship, lineage, and quality standards bp already applies to its data products, extended to documents and agent memory.

Governance continuity

Ingest, classification, freshness SLA, and steward attribution mirror the governance principles bp already applies to UDP and Databricks. No new governance regime is introduced; the existing one is extended to documents and agent memory so the data stewardship community can adopt the knowledge plane without learning a separate model.

Stage 1

Ingest

  • Connectors to SharePoint, Confluence, and controllership repositories.
  • Document classification and sensitivity tagging.
  • Versioning and freshness SLA per source.
  • Author and steward attribution captured at ingest.

Stage 2

Process and index

  • Semantic chunking that respects clauses and tables.
  • Hybrid vector and keyword index.
  • Entity linking into the canonical ontology used by structured data.
  • Quality scoring per chunk for retrieval ranking.

Stage 3

Serve to agents

  • Grounding-only retrieval so no agent answer lands without a cited chunk.
  • Citation identifiers flow into the evidence pack alongside structured signals.
  • Per-agent knowledge scopes enforce least privilege.
  • Stale-content quarantine triggers on freshness-SLA breach.

Memory plane

Three stores agents draw on across runs

  • <strong>Episodic memory:</strong> per-run journal reusable as case memory, such as a supplier flagged 11 days ago with a known resolution. Bounded retention and fully replayable.
  • <strong>Institutional memory:</strong> approved override patterns, accepted lessons, and near-miss incidents, promoted only by the Governance Council and never auto-learned.
  • <strong>Working memory:</strong> short-lived session context and intermediate state carried across a bounded task without becoming a system-of-record write-back path.

Agentic compounding asset

Institutional memory is bp's compounding agentic asset: the durable layer where approved lessons and accepted patterns accumulate under governance instead of being re-learned from scratch every run.

Slide 10A · bp AI standards and governance contract

Platform-first standards alignment: Yalla, Nexus, LaunchPad, MCP, and A2A

Mandatory alignment

The reviewed bp AI standards make one thing explicit: suppliers are not being judged only on runtime vision. They are being judged on whether their agent estate can operate inside bp's approved platforms, open standards, governance gates, and accountability model. This slide makes that contract explicit.

What bp standards require

  • Platform-first delivery through bp internal developer pathways such as LaunchPad / Yalla, not vendor-owned infrastructure.
  • Centralized discovery, reuse, and registration through Nexus before new agents or MCP servers are created.
  • Mandatory use of MCP for tool and data integration and A2A for multi-agent collaboration.
  • AI use-case registration, risk triage, and production gating through AI LaunchPad with named human accountability.
  • No direct or ad-hoc enterprise data access outside approved MCP, API, or iHub-style pathways.
  • Three Lines of Defence, human oversight, and auditable telemetry for every production AI capability.

What bp Sphere must prove

  • Agents are discoverable, owned, and registered in a governance model that aligns to Nexus concepts.
  • Every capability has an accountable owner, policy bindings, promotion gates, and retirement path.
  • Multi-agent behavior is explicit, orchestrated, and replayable rather than hidden in prompt chains.
  • The runtime can show AI inventory, risk class, and LaunchPad-style certification posture for each production use case.
  • The operating model stays inside bp's cloud, identity, and governance boundaries instead of introducing a parallel estate.

Platform-first

All AI runs on bp-approved cloud and runtime pathways. The narrative should explicitly reject vendor-owned infrastructure and shadow identity models.

Discoverable and reusable

The agent estate should expose registry metadata, reusable MCP servers, and A2A coordination contracts so capabilities built once can be composed many times.

Governed for production

Unregistered agents, unclassified risks, or non-compliant delivery paths should be impossible to present as “production-ready.”

Slide 10B · Golden Path delivery and assurance

Golden Path delivery, evaluation assurance, and agentic UX compliance

Engineering credibility

The standards material raises the bar beyond runtime observability. bp expects Golden Path delivery, bp-controlled repositories and pipelines, formal design review, real evaluation discipline, and a consistent agentic UX with confidence, source attribution, override, and accessibility built in.

Golden Path delivery

  • bp-controlled repositories and Azure DevOps pipelines rather than vendor-hosted source control.
  • Pre-approved build, security, performance, and deployment gates in the standard CI/CD path.
  • Technical Design Document and Technical Design Review evidence before production certification.
  • Runbooks, monitoring, rollback, and transition documentation as part of delivery, not post-go-live cleanup.

Evaluation and assurance

  • Ground-truth evaluation metrics for AI quality; UAT alone is not enough.
  • Component and end-to-end agent evaluation, including retrieval quality, reasoning relevance, tool behavior, and workflow completion.
  • Drift thresholds, regression history, and promotion blocking when quality degrades.
  • Adversarial or red-team testing for high-risk AI implementations.

Agentic UX compliance

  • Confidence signals, source attribution, and human override controls in every AI-facing mission surface.
  • Consistent role-based UX so analysts, supervisors, and leaders do not relearn each agent.
  • Accessibility and design-system compliance as part of AI readiness, not visual polish.
  • Trust, control, and transparency treated as UX requirements rather than hidden engineering details.

What the live system should show

  • Registry metadata that looks like a governed production asset, not a demo list of agents.
  • Trace drawers with prompt, model, retrieval, latency, citations, confidence, retries, and policy context.
  • Evaluation surfaces with golden datasets, pass/fail trends, drift warnings, and promotion gates.
  • Mission surfaces where confidence, evidence, and override remain inline for operators.

Proof language for the page

The strongest claim is not “the AI is smart.” It is “the AI estate is governable, testable, discoverable, and auditable inside bp's existing standards.” This is the gap that many suppliers will not close.

Slide 11 · Cyber security and access management

Cyber security, IAM, and controlled access model

Security validation

This deep dive needs to show that bp Sphere integrates into bp’s security architecture rather than creating a shadow runtime. The emphasis is zero-trust alignment, managed identities, controlled privilege, named accountability, and SOC-visible telemetry.

Identity and access architecture

Human identity
Entra ID + enterprise IAM
Named human users remain the authority for analysts, supervisors, approvers, and privileged intervention paths.
Workload identity
Managed service identities
Agents run with governed workload identities, least privilege, scoped data access, and no hidden local account model.
Privilege boundary
Policy + approval gateway
Privileged actions route through approval, policy checks, execution boundaries, and replay-backed attribution before write adapters are invoked.
Cyber oversight
SOC / SIEM integration
Security-relevant telemetry, unusual execution paths, and privileged decision events feed existing detection and monitoring patterns.

What this validates for bp

  • No separate shadow IAM model is introduced; bp identity remains the source of truth.
  • Privileged execution can be bounded, audited, and escalated explicitly.
  • Named accountability exists for human and agent actions through replay-backed attribution.
  • Runtime telemetry can align to enterprise cyber monitoring and compliance expectations.

Questions to resolve with bp security

  • What workload identity, JIT access, and PAM patterns are approved for production runtime components?
  • What telemetry classes must be forwarded into SOC/SIEM and cyber audit workflows?
  • What runtime isolation and cross-cloud segmentation standards must the pilot satisfy?
  • What is the required model for break-glass, revocation, and privileged approval logging?

Operational IAM detail

  • OIDC-based federation into bp identity with no separate local identity store for the runtime.
  • SCIM-aligned lifecycle integration so role and entitlement changes propagate into mission surfaces and approvals.
  • Managed workload identities for agents, tools, and connectors, each scoped to least privilege and bounded data domains.
  • JIT privileged execution and PAM-compatible break-glass for sensitive actions, with replay-backed privileged traceability.
  • Model access routed through the gateway so prompt, tool, and retrieval spans inherit the same identity and policy context.

Security-team message

Security review should be able to map the runtime to existing enterprise controls in operational language: federation, lifecycle, privileged access, segmentation, and traceability. “Supports Entra ID” is not enough; the page needs to show how access is provisioned, approved, elevated, traced, and revoked in practice.

Slide 12 · Disaster recovery and resilience

Disaster recovery, resilience, and safe failure model

Operational resilience

This needs to show how the runtime behaves when dependencies fail, how human fallback works, and how replay supports recovery. The goal is to prove graceful degradation, not a brittle orchestration engine.

Detect

  • Runtime anomalies, queue stalls, evidence lag, replay gaps, and degraded upstream integrations are surfaced early.
  • Cross-cloud dependencies and source-system disruption are visible in the control plane instead of being hidden in infrastructure logs.

Contain

  • Autonomy can be throttled or downgraded by policy tier when trust conditions degrade.
  • Execution pathways can fall back to read-only or human-only modes rather than silently failing or continuing unsafely.
  • Supervisor visibility remains available even when automation is reduced.

Recover

  • Replay helps reconstruct the incident path, the impacted decisions, and the recovery sequence.
  • Human fallback, queue continuity, and evidence integrity remain central to operational restart decisions.
  • Post-incident lessons feed governance tuning, resilience certification, and phase progression decisions.

What this answers

  • How does bp Sphere fail safely without creating uncontrolled writes or hidden decisions?
  • How are incidents replayed and reconstructed for cyber, audit, and operations teams?
  • How do human fallback and bounded autonomy work under runtime degradation?
  • What is the continuity story when one cloud plane or source system is impaired?

Questions to resolve with bp

  • What RTO/RPO and SLA expectations apply to the first pilot domain?
  • What DR and failover patterns are already approved across AWS and Azure for this runtime class?
  • What crisis-management hooks must integrate with existing resilience workflows?
  • What level of degraded operation is acceptable before manual-only fallback becomes mandatory?

Failure scenario that must be demoed

01
Azure runtime unavailable
A model-hosting or orchestration dependency becomes impaired in one cloud plane during hypercare.
02
Policy runtime downgrades autonomy
Write-capable paths are automatically reduced to supervised or read-only mode by policy tier.
03
Execution shifts to human-governed fallback
Analysts and supervisors keep operating from the same mission surface while automation is throttled.
04
Replay continuity preserved
Replay, evidence, and incident context continue to accumulate so the failure can be reconstructed end to end.
05
Incident bridge activated
Control tower and resilience surfaces expose state, owners, and recovery sequence until service returns.

Failure proof surfaces

  • The page should show one explicit degraded-mode path instead of implying that resilience exists somewhere in the platform.
  • This is where bp sees that replay continuity, human fallback, and policy downgrade are designed into the runtime rather than added after incidents.
Slide 12A · Runtime operations center

Runtime operations center: who runs bp Sphere at 2am?

Operate the platform

A boardroom-ready platform story needs an operations answer, not just an architecture answer. This slide explains how the runtime is run, monitored, escalated, and corrected when something goes wrong.

Monitor

  • Agent health
  • Replay health
  • Evidence health
  • Queue health
  • Policy violations
  • Data quality
  • Escalation backlog
  • Cost posture

Operate

  • On-call ownership
  • Incident bridge
  • Runbooks
  • Kill switches
  • Fallback activation
  • Threshold tuning

Improve

  • Root cause patterns
  • Override trends
  • Failure trends
  • Policy friction
  • Cost anomalies
  • Remediation backlog
Operating answer: bp Sphere is not self-running magic. It needs a named runtime operating team, clear escalation ownership, and one control surface that combines trust, policy, operations, and cost signals.
Slide 13 · Digital workplace and adoption

Digital workplace, adoption, and workforce operating model

Change management

bp will want to understand how work changes, not just how systems connect. This slide should make it credible that analysts, supervisors, and finance leaders move into mission surfaces, inline replay, and human-agent collaboration without adding tool sprawl.

Digital workplace shift

Today

  • Inboxes, reports, queues, and system hopping
  • Manual prioritization and limited context
  • After-the-fact audit reconstruction
  • Fragmented collaboration and escalation

Transitional

  • Read-only agent assistance
  • Replay-backed work queues
  • Supervisor-visible escalations
  • Policy-aware prioritization and explainability

Future

  • Decision theaters instead of dashboard sprawl
  • Human-agent collaborative execution
  • Enterprise-impact-aware work surfaces
  • Institutional learning embedded in daily operations

Adoption and change model

  • Start with visible productivity and trust gains rather than autonomous execution.
  • Use role-specific mission surfaces for analysts, supervisors, and finance leaders.
  • Treat replay, evidence, and explainability as adoption enablers, not technical extras.
  • Measure productivity, confidence, override patterns, and operational load reduction as part of rollout governance.

Questions to resolve with bp

  • Which user groups should experience the first digital workplace transition surface?
  • What training, adoption, and operating-model risks exist for the first rollout?
  • What KPIs define meaningful workforce productivity improvement for the pilot?
  • How should this align to broader digital workplace and transformation programs already underway?
Slide 13A · Hypercare intelligence

Hypercare intelligence: how the first production wave is stabilized

Adoption + control

The first production wave should not rely on generic project hypercare. It should use runtime intelligence to detect adoption issues, data breakdowns, agent failures, and escalation trends before they damage confidence.

Detect

Adoption issues, data issues, agent failures, policy friction, backlog spikes, and unusual override patterns.

Predict

Which queues will breach, which data sources are becoming unreliable, and which teams are heading toward failure or overload.

Recommend

Corrective actions: retraining, source remediation, policy tuning, threshold adjustment, or fallback activation.

Why this matters

  • Because the first 90 days determine whether the platform is trusted or treated as another fragile initiative.
  • Because hypercare needs to be intelligence-driven, not only meeting-driven.
  • Because bp will judge the platform by how it behaves when adoption, data, or process reality becomes messy.

Questions for bp

  • Who owns hypercare decisions across operations, architecture, security, and controllership?
  • Which hypercare signals should trigger executive escalation?
  • What remediation actions can be automated versus requiring human approval?
Slide 14 · Controllership operating model

Controllership operating model and finance governance alignment

Finance trust

This must explicitly answer how controllership works in the target model. Replay, evidence, policy, approval traceability, close confidence, and treasury exposure need to be framed as part of finance governance, not just runtime plumbing.

Controllership model

Control objective
Policy-aligned execution
Thresholds, SoD, approval matrices, and exception pathways remain explicit and reviewable.
Audit objective
Replay + evidence
Every critical decision path is attributable, reconstructable, and ready for controller or auditor challenge.
Close objective
Confidence and readiness
Backlog, approvals, variances, and unresolved exceptions inform close confidence and escalation early.
Treasury objective
Liquidity-aware posture
Working-capital and payment decisions propagate visibly into exposure and release decisions.

What controllers should be able to say

  • We can see why a recommendation was made, what policy governed it, and who approved it.
  • We can replay how a high-risk decision emerged and what evidence existed at the time.
  • We can distinguish read-only intelligence, supervised automation, and controlled write execution clearly.
  • We can connect transaction-level interventions to close readiness, exposure, and financial stewardship outcomes.

Questions to resolve with bp controllership

  • What are the non-negotiable auditability and replay requirements for finance operations?
  • How should close confidence, treasury exposure, and exception posture be surfaced to controllers?
  • What dual authorization, evidence retention, and sign-off standards apply?
  • What controllership KPIs most clearly prove value in the first phase?

FBT control reality this must map to

  • SOX and non-SOX control posture across payables, cash and banking, local close, receivables, tax, and treasury-related processes.
  • Unauthorized approval, SOD access, fraud monitoring, and access review concerns already operated by FBT teams.
  • Reconciliation, audit assurance, and evidence sufficiency expectations for finance-critical interventions.
  • Liquidity, counterparty, and treasury risk implications when a payment, hold, or escalation changes enterprise posture.

What the runtime should eventually expose

  • Control-mapping views that show which FBT control families a mission or agentic decision supports.
  • Explicit SOD, unauthorized approval, fraud-risk, and reconciliation signals on high-value cases.
  • Controller-facing evidence packs that map replay to policy, approval, and audit obligations.
  • Treasury and controllership consequence previews attached to sensitive P2P and close decisions.
Slide 14A · Next-wave mission agents

Next-wave mission agents and enterprise coordination build queue

Policy-bound enterprise execution agents

The next enhancement wave should not be framed as more bots. It should be framed as policy-bound enterprise execution agents operating inside the bp Sphere intelligence fabric. The emphasis is cross-mission orchestration, evidence generation, live operational intelligence, and enterprise decision coordination.

Phase 1 priority agents

  • P2P Control Tower Supervisor Agent to coordinate queues, escalations, workload balancing, and SLA jeopardy across AP, procurement, and treasury.
  • Enterprise Signal Correlation Agent to connect P2P, Treasury, O2C, FP&A, R2R, and ST&S into replayable causal chains.
  • Close Confidence Agent to predict period-end delay, control gaps, and missing evidence for controllership and CFO audiences.
  • Enterprise Liquidity Optimization Agent to coordinate cash posture, payment timing, working capital, and treasury actions.
  • Evidence and Replay Agent as the signature cross-mission differentiator for immutable evidence packs and provenance.

Why these agents matter

  • They elevate the estate from process automation toward enterprise decision operations.
  • They strengthen supervisory intelligence, not just task execution.
  • They create visible cross-mission coordination rather than isolated domain wins.
  • They directly reinforce the control-tower, replay, and policy-runtime claims already shown in the solution narrative.
Cross-mission orchestrationEvidence generationEnterprise prioritizationHuman-supervised executionLeadership decision queuesReplayable coordination

P2P

Control tower supervision, duplicate intelligence, vendor health and friction, procurement policy drift, and working-capital optimization.

O2C and R2R

Revenue leakage, customer risk and exposure, autonomous dispute resolution, close confidence, journal risk, and reconciliation intelligence.

Treasury, FP&A, and ST&S

Enterprise liquidity optimization, forecast drift, scenario orchestration, supply disruption prediction, margin optimization, and trade exposure intelligence.

ST&S runtime proof now available

The ST&S architecture is no longer only a roadmap claim. The live trading runtime now exposes an autonomy matrix plus a context-graph and decision-ledger surface grounded in Endur-led trade context, shipping state, and source freshness.

How to use it in the room

  • Use the matrix to show progressive autonomy and safe first-release scope across trading and shipping.
  • Use the ST&S graph to show ontology plus runtime context rather than generic RAG or prompt-only AI.
  • Use the decision ledger to answer who recommended what, under which approval boundary, and from which systems.
  • Use context quality to address bp concerns about trust, freshness, provenance, and operational safety.
Critical runtime requirement. Every new agent must follow one standard enterprise contract: business purpose, trigger events, inputs, decision logic, HITL, policy controls, evidence and replay, actions, outputs, KPIs, integrations, security, observability, failure handling, and learning loop.

First-wave implementation matrix

The first three agents now need a delivery artifact with agent ID, owner, LaunchPad ID, policy pack, event inputs, replay contract, UI surface, evaluation gate, and delivery phase.

P2P Control Tower Supervisor Agent

LaunchPad ID LP-P2P-CTRL-001. Runtime proof should be live queue reprioritization, replayable escalation rationale, and bounded workload coordination on the analyst mission surface.

Enterprise Signal Correlation and Close Confidence

LaunchPad IDs LP-ENT-CORR-001 and LP-R2R-CLOSE-001. Runtime proof should show cross-mission causal chains, close-confidence change, and controller-facing evidence sufficiency.

Slide 15 · Human + agent operating model and control plane

Multi-actor operating model and enterprise runtime control plane

Human + agent coexistence

This slide must answer how humans and agents coexist, who owns what, and how governance works at runtime. It should make the future-state operating model emotionally and operationally credible rather than abstract.

Human + agent operating model

Analysts
Exception governors
Validate high-risk recommendations, review evidence, resolve escalations, and maintain accountability.
Supervisors
Runtime managers
Balance workload, tune thresholds, approve sensitive actions, and manage drift or overload signals.
Agents
Specialized coworkers
Observe signals, correlate context, score risk, recommend actions, assemble evidence, and route escalations.
Governance teams
Assurance stewards
Certify policies, approve autonomy boundaries, review replay, and assure operational safety across domains.

Enterprise runtime control plane

  • Runtime health, event throughput, policy evaluations, escalations, evidence status, and replay availability must be visible as operational intelligence, not just infra telemetry.
  • The control plane should support threshold tuning, drift visibility, approval bottlenecks, failure injection, and supervisor intervention from one place.
  • This is how bp Sphere becomes a governed enterprise runtime instead of a collection of agents hidden behind screens.
Agent healthPolicy evaluationsEscalation visibilityReplay queuesEvidence completenessAutonomy levelsFailure injectionDrift monitoring

Messier enterprise reality scenarios

  • Duplicate supplier master and region-specific workflow variance causing the same operational event to look different across systems.
  • Delayed reconciliation and partial ERP coexistence creating close-pressure hotspots that do not map neatly to one source system.
  • Conflicting policy versions, hypercare instability, and rollout inconsistency across regions or business units.
  • Treasury delay propagation where a blocked approval changes liquidity posture faster than the source teams can see manually.

How the demos should handle reality

  • Show one or two messy enterprise scenarios directly, rather than only clean happy-path flows.
  • Tie every messy scenario to a specific proof surface: mission runtime, control tower, resilience, replay, or operator console.
  • Demonstrate that bp Sphere is useful precisely because the landscape is inconsistent, not because the page pretends it is clean.
Slide 16 · Digital twin and strategic intelligence

Enterprise finance digital twin and strategic intelligence architecture

Future-state enterprise vision

The narrative should not stop at transaction automation. It should show the long-term direction: invoice, payment, liquidity, supplier trust, and close operations feeding a governed finance digital twin and strategic intelligence runtime.

Operational inputs
P2P, R2R, treasury, supplier, close
Live signals, approvals, exposures, backlog, disputes, policy events, and replay history form the factual substrate.
Digital twin core
Enterprise finance twin
Working-capital posture, close confidence, treasury exposure, supplier concentration, and scenario baselines are modeled from live enterprise flow.
Strategic intelligence
Simulation and optimization
Liquidity scenarios, discount optimization, DPO tradeoffs, policy-impact simulation, and cross-domain prioritization become governed decisions.
Governed outcome
Predictive but bounded enterprise operations
Human-supervised autonomy evolves only after trust, evidence, and control-plane maturity are proven in earlier phases.
Strategic positioning: the first successful milestone is not “AI executing actions.” It is “AI creating enterprise visibility, replayability, prioritization, and governed operational intelligence.” The digital twin arrives on top of that trust foundation.
Slide 16A · Local runtime to bp production runtime

Local runtime to bp production runtime: the path from proof to controlled enterprise rollout

Execution path

This is the implementation bridge executives and architects will ask for. It explains how bp Sphere moves from local proof surfaces to controlled enterprise runtime without hand-waving the hard parts away.

Stage 1
Sandbox

Proof runtime

Validate context assembly, evidence packaging, replay, and operator surfaces with controlled datasets and seeded scenarios.
Stage 2
Pilot

Read-only production observation

Connect to approved production-like feeds, observe events, build replay, and prove value without uncontrolled writes.
Stage 3
Controlled automation

Human-approved execution

Enable bounded write actions through explicit gateways, approval chains, and policy controls.
Stage 4
Enterprise rollout

Scaled operating model

Extend to additional business units, regions, and functional domains with formal runtime operations and assurance.
Stage 5
Cross-domain coordination

Enterprise execution layer

Move from single-domain value to coordinated finance, procurement, treasury, ST&S, and operations decisioning.
Slide 17 · Questions, demo moments, and close

How the deep dive should run and what it should close with

Decision and alignment closeout

Session agenda

00:00-00:15
Welcome, context & objectives
Why bp Sphere exists, what bp is really asking for, and what this deep dive must prove.
00:15-00:35
bp landscape, Quantum transformation & key gaps
Strong foundations, fragmented execution, and where coordination, trust, and replay are still missing.
00:35-01:00
bp Sphere vision & architecture
Positioning, enterprise execution runtime, and the common lifecycle across invoices, journals, payments, and forecasts.
01:00-01:25
AWS + Azure federated orchestration model
How the runtime operates across the existing cloud estate without creating a parallel platform.
01:25-01:45
SAP, Ariba, UDP & Databricks integration strategy
Systems remain distributed while context, evidence, and policy become unified.
01:45-02:00
Break
Reset before moving into trust, replay, policy, and execution control.
02:00-02:25
Signal -> replay -> evidence lifecycle
How operational signals become governed decisions, evidence contracts, and enterprise memory.
02:25-02:50
Controllership, policy runtime & human governance
How controls, approvals, and replay become operating features rather than after-the-fact audit activity.
02:50-03:10
Cyber security, IAM & operational trust
Identity, telemetry, privileged execution, and runtime security boundaries.
03:10-03:30
Resilience, safe failure & hypercare intelligence
Degraded mode, recovery, runtime operations, and first-wave production stabilization.
03:30-03:45
Digital workplace & human + agent operating model
How analysts, supervisors, controllers, and leaders work through one mission surface.
03:45-04:00
Risks, dependencies, open questions & next steps
What bp still needs to answer across architecture, data, security, and operations.

Best demo moments from one mission surface

Duplicate payment prevention

Signal, orchestration, policy, supervisor escalation, replay, evidence, and measurable protected value.

Early payment discount optimization

A CFO-relevant example showing working-capital intelligence and liquidity-aware decisioning.

Supplier trust escalation

A strong emotional scenario showing cross-functional impact, team routing, and operational criticality.

GR/IR mismatch auto-resolution

The best adaptive-autonomy demonstration: high-confidence recommendation, human approval, and learning loop.

Quarter-close pressure

Shows unresolved invoices, reconciliation backlog, treasury dependency, and close confidence movement.

Policy breach prevention

A governance-first scenario for auditors, controllership, and skeptical architecture leaders.

Questions to resolve with bp

  • Which domains and data products are approved for pilot observation first?
  • What are the enterprise standards for AI runtime isolation, SIEM forwarding, and privileged execution?
  • Who owns runtime certification, policy approval, and autonomy boundary decisions?
  • What metrics define a successful phase 1 observe-and-assist pilot?

Risks, assumptions, and dependencies

  • Integration readiness and data-product maturity will vary by domain and region.
  • Pilot credibility depends on agreed identity, telemetry, and approval patterns up front.
  • Governance adoption requires controller, cyber, architecture, and operations sponsorship, not just technology approval.
  • Read-only phase 1 scope is the safest commercial and operational entry point; write-back assumptions should stay explicitly bounded.

What we need from bp next

Architecture

Approved integration patterns, target systems for pilot connection, runtime hosting boundaries, and production environment sequencing.

Data

Named data products, stewardship owners, readiness posture, lineage standards, and access approval for the first domains.

Security

Identity model, PAM/JIT expectations, telemetry forwarding requirements, segmentation standards, and privileged execution rules.

Operations

On-call ownership, hypercare model, resilience expectations, pilot KPIs, and governance forums for certification and escalation.

Visible proof 1

Live runtime proof

Signal, orchestration, policy, evidence, replay, and human escalation running on live bp Sphere surfaces.

Visible proof 2

Production swap proof

Explicit mapping from local proof sources to bp production targets, integration contracts, and governance owners.

Visible proof 3

Failure and governance proof

Policy blocking, degraded mode, replay continuity, and human fallback shown as operational behavior rather than slide claims.

The future is not isolated copilots or uncontrolled agents.
The future is governed enterprise execution: replayable, evidence-backed, policy-aware, human-supervised, and credible inside bp’s federated cloud, data, and operational ecosystem. bp Sphere should be presented as the orchestration, trust, and governance fabric that helps bp progressively evolve from fragmented execution into a controllable enterprise decision runtime.