bp Sphere Technology, Data & Governance Deep Dive
bp Sphere Enterprise Execution Runtime
bp Sphere is being positioned here as a governed enterprise orchestration and evidence runtime operating across bp’s existing AWS, Azure, SAP, Ariba, Databricks, UDP, and operational control landscape. This page is intentionally structured as a slide-style solution narrative: each section answers why bp needs the capability, what problem it solves, how it works, how it would be implemented, and what questions bp still needs to answer before execution planning.
What this deep dive must answer for bp
This is not an AI demo. It is the technology, data, governance, cyber, resilience, digital workplace, and controllership deep dive that must show how bp Sphere fits into bp’s real architecture and operating model without forcing ERP replacement or uncontrolled autonomy.
Purpose
- Validate integration into bp technology architecture.
- Align on data quality dependencies and the data governance model.
- Evaluate supplier approaches to cyber security, access management, disaster recovery, resilience, and digital workplace change.
Expected outputs
- Validated integration approach into bp technology architecture.
- Clarity on additional inputs suppliers need to refine their solution.
- Clarity on what additional details bp still needs from suppliers.
Control questions
- How does governance scale safely across humans and agents?
- How are replay, evidence, and controllership preserved?
- What are the key risks, assumptions, and dependencies for pilot execution?
P2P analyst mission lens
This page also needs to explain what this architecture means for a real operator. For a P2P analyst, bp Sphere should not feel like another dashboard. It should feel like a governed decision workspace where live queue priorities, agent recommendations, escalation context, evidence, replay, and enterprise impact stay in one mission surface.
- Priority queue replaces static worklists with governed decisions that need action now.
- Human-agent collaboration stays inline through recommendations, evidence requests, and supervisor escalation.
- Replay and proof stay attached to each case so analysts can trust why something was flagged.
- Business impact remains visible so analysts act as financial flow operators, not AP clerks clearing tickets.
The five planes that complete the architecture story
Beyond the orchestration runtime, replay, and controlled execution that the solution covers as a whole, five architectural planes determine whether bp Sphere holds up under real enterprise scrutiny: how agents are authored, how their reasoning is intermediated, how they are grounded and remembered, how their behaviour is continuously evaluated, and how they coordinate across domains. Each plane is unpacked in the slide it most naturally belongs to.
Slide 04
Author plane & Model Gateway
Skill Studio plus signed registry, paired with an LLM gateway that redacts, routes, moderates, and meters every reasoning call.
Slide 05
Engineering trace & evaluation
Per-step reasoning trace plus offline and online evaluation, sitting beneath business replay and sharing the same evidence chain.
Slide 06
Multi-agent coordination
Mediator plus devil's-advocate pattern, quorum rules, and collusion defences for any cross-domain decision.
Slide 08
Authored, owned, governable
Every Phase 1 agent is a signed artefact in the bp Sphere registry with a named owner, opening the path for ops teams to become creators of agentic capability.
Slide 10
Knowledge plane & agent memory
Unstructured grounding from policy and runbook documents, plus three memory stores agents use across runs without writing back to systems of record.
Cross-cutting
One evidence chain
Author signature, model gateway log, trace span, eval verdict, multi-agent dissent, and replay entry stitch into a single audit-grade chain.
What bp is really asking for and what this deep dive must prove
bp is not evaluating AI in the abstract. bp is evaluating whether AI can operate safely inside a regulated enterprise. This deep dive must show that bp Sphere can operate inside bp’s environment without forcing replatforming, without bypassing governance, and without pretending that generic copilots equal enterprise execution.
What we heard from bp
- bp does not want ERP replacement, platform lock-in, or another disconnected AI estate.
- The runtime must align to existing AWS, Azure, IAM, Entra ID, and cloud guardrail standards.
- Productivity must be realized with operational safety, bounded writes, and named accountability.
- The enterprise is federated: fragmented ownership, mixed system maturity, regional process variants, and uneven data quality are the reality.
- Governance, controllership, evidence, and operational trust matter more than novelty.
- Transformation must be phased, measurable, human-supervised, and politically adoptable.
What this deep dive must establish
Enterprise credibility
Show that the orchestration model fits a real bp estate rather than a toy greenfield runtime.
Governance trust
Make policy, approval boundaries, replay, evidence, and rogue-agent prevention visually explicit.
Operational realism
Show human-agent collaboration, escalation, and business impact through finance operating moments.
Architecture defensibility
Explain what bp already has, what AWS/Azure provide, and what bp Sphere uniquely adds.
Transformation practicality
Demonstrate a progressive roadmap starting with read-only agents rather than uncontrolled writes.
Differentiation
Make it clear why this is not 'just copilots', 'just workflows', or 'just cloud AI primitives'.
Day 2 deep-dive purpose
- Validate integration into bp technology architecture rather than propose a parallel stack.
- Understand dependency on data quality and align the data governance model needed for credible operationalization.
- Evaluate supplier approaches to cyber security and access management, disaster recovery and resilience, and digital workplace change.
- Surface the risks, assumptions, and dependencies that shape pilot scope, control design, and rollout readiness.
- Clarify additional inputs suppliers need from bp and what additional questions bp needs answered before solution refinement.
- Align on the controllership operating model and the human-governance expectations behind future enterprise autonomy.
Expected outputs from this deep dive
- Validated integration approach into bp technology architecture.
- Clarity on data quality dependencies and the aligned data governance model.
- Validated supplier approaches to cyber security and access management, disaster recovery and resilience, and digital workplace transformation.
- A concrete list of risks, assumptions, dependencies, and architecture decisions requiring follow-up.
- Clarity on additional inputs suppliers need to refine their solution and what additional details bp still needs from suppliers.
- A credible path to align the controllership operating model with replay, evidence, policy runtime, and human-supervised execution.
Questions to resolve with bp across data, security, runtime, and operating model
Each later slide expands one of these domains so the deep dive can end with concrete follow-up actions instead of generic agreement.
Credibility test for the page
- Can bp architecture teams see how the runtime integrates without ERP replacement?
- Can data teams see the dependency on governed products, lineage, and quality stewardship?
- Can security teams see IAM, workload identity, cyber telemetry, and controlled write boundaries?
- Can controllership leaders see how replay, evidence, approvals, and policy runtime preserve trust?
- Can both sides leave with explicit decisions, open questions, assumptions, and next-step inputs?
Strong foundation. Fragmented execution.
bp already has major systems, cloud, and data investments. Systems are not broken. Execution is fragmented. The gap is enterprise coordination, operational trust, and the ability to turn fragmented signals into governed, replayable execution across domains.
Enterprise estate in scope
The gap
- Approvals, policies, workflows, AI pilots, and analytics remain functionally or technically siloed.
- Operational intelligence is fragmented across dashboards and system-specific views instead of one supervised runtime.
- Logging exists, but enterprise replay, evidence lineage, and governance reconstruction are not yet first-class operating assets.
- Cross-domain orchestration across finance, procurement, treasury, and operations is still limited.
Current state → transitional state → future state
Current
- Fragmented approvals
- Siloed automation
- Limited enterprise replay
- Manual exception triage
- Cross-domain visibility gaps
Transitional
- Read-only intelligence agents
- Replay-backed visibility
- Supervisor escalation surfaces
- Policy-aware prioritization
- Governed exception operations
Future
- Cross-domain coordination
- Bounded autonomy
- Enterprise decision runtime
- Institutional operational memory
- Digital twin and strategic intelligence
bp Sphere is the enterprise execution runtime, not another dashboard or copilot
The strongest message is not “we have features.” It is “traditional dashboards tell you what happened; bp Sphere governs what happens next.” This slide must visually separate static analytics, copilots, workflow engines, and governed operational orchestration.
What this is not
- Not an ERP replacement
- Not a separate cloud platform
- Not a menu of disconnected copilots
- Not another BI surface with AI labels
- Not uncontrolled autonomous AI
What AWS / Azure already provide
- Compute, model runtimes, workflow primitives, logging, eventing, and infra observability
- Useful AI and automation building blocks
- Strong cloud services, but not enterprise semantic orchestration or supervisory governance
What bp Sphere adds
- Enterprise orchestration runtime
- Policy-bound execution boundaries
- Replay and evidence contracts
- Supervisory governance and escalation
- Cross-domain semantic context and operational intelligence
How bp Sphere operates across the enterprise: from hybrid signals to governed execution and replay
Every transaction follows the same lifecycle: supplier invoice, treasury payment, journal entry, or forecast revision. bp Sphere uses one runtime for every business process, every agent, and every decision: signal, context, evidence, policy, decision, execution, and replay.
Author plane and Model Gateway
How agents enter the runtime, and how their reasoning is intermediated
Two planes sit either side of the orchestration runtime above and determine whether the agent estate remains governable as it grows. On the inbound side, the Author plane turns business intent into signed, registered agent artefacts. On the outbound side of every reasoning step, the Model Gateway intercepts the call to any LLM for redaction, routing, moderation, and metering. Together they make the difference between a controlled estate and a sprawl of unaccountable agents.
Why this matters for bp
The model behind any agent is interchangeable without re-certifying the agent itself because the boundary is the gateway, not the model. Every new agent enters the estate through a single bp-owned author surface, which means no shadow deployments, no untracked prompts, and no off-fabric model calls.
Model Gateway
LLM intermediation
- PII redaction, prompt-injection and jailbreak scanning, model routing across Azure OpenAI, Bedrock, and private models.
- Output moderation, citation checks, and schema validation before responses re-enter the runtime.
Cost and trace emission
Every reasoning call emits tokens used, latency, model selected, redaction statistics, and a trace span, feeding FinOps, SOC monitoring, and the engineering trace surface introduced on Slide 05.
Author plane
Skill Studio
Declarative agent specs in Markdown or YAML with tool binding, knowledge attachment, prompt template, and an eval set. No-code for domain experts, with pair-authoring alongside engineering for tool makers.
Registry & promotion pipeline
Signed artefacts for skills, tools, MCP servers, models, prompts, and knowledge packs. Promotion runs through sandbox, shadow, canary, bounded, and general, and each gate requires a passing eval and council sign-off.
Local-to-bp production integration mapping
| Current demo proof | bp production target | Integration pattern | Governance model |
|---|---|---|---|
| Local SAP invoice simulation | SAP S/4HANA / CFIN | OData, BAPI mediation, SAP event mesh, canonical event overlay | bp IAM, approval audit, replay-backed write attribution |
| Local procurement and supplier state | Ariba and supplier workflows | APIs, webhooks, supplier event normalization, policy-bound gateway | Supplier stewardship, policy-bound workflow mutation, named owner |
| Local replay queue and runtime events | Splunk, SIEM, SOC, and runtime ops tooling | OpenTelemetry spans, event forwarding, replay event export | Immutable retention, security forwarding, privileged traceability |
| Local evidence manifests | UDP / Databricks governed products | Lineage-aware evidence linkage, governed product references, metadata joins | Steward-owned lineage, retention policy, evidence schema contract |
| Local runtime-health and deployment proof | CloudWatch, Azure Monitor, LaunchPad / Yala ops views | Health APIs, runtime control tower aggregation, deployment-state federation | bp runbooks, on-call ownership, change and incident controls |
Critical production message
bp Sphere operates inside bp’s governance and cloud boundaries rather than introducing a parallel operating model.
- The demo surfaces prove orchestration behavior; the production mapping above proves how those behaviors swap onto bp-approved systems and controls.
- No shadow IAM, no parallel data regime, and no hidden direct-write path are introduced in order to move from the local proof estate into production.
- Every local proof path must point to a named bp production target, a concrete integration contract, and a clear governance owner before pilot approval.
Signal-to-replay lifecycle: every operational moment must flow to replay and evidence
This is the page’s most important trust diagram. It shows how operational signals become governed decisions, controlled execution, evidence packs, and replayable enterprise history. It should be the visual anchor for duplicate payment, policy breach, and close-pressure demos.
Engineering trace and evaluation harness
The surfaces beneath replay
Replay is the controller-facing record of what the enterprise did. Beneath it, bp Sphere runs two further surfaces: a per-step engineering trace that explains why an agent reasoned the way it did, and a continuous evaluation harness that proves the agent estate is not getting quietly worse over time. The three views share a single evidence chain: replay for controllers, trace for engineering and risk, eval for early-warning drift.
Why this matters
Replay reconstructs the operational decision for controllers. Trace reconstructs the reasoning path for engineers. Eval continuously verifies that the reasoning path still matches the gold standard, turning drift from a post-incident discovery into a leading indicator.
Trace surface
Per-step LLM and tool observability
- LLM span: prompt, completion, tokens, latency, model, and cost.
- Tool span: input, output, retries, and side-effects.
- Retrieval span: query, chunks fetched, and citation overlap.
- Hallucination markers when output diverges from citations.
- Searchable by agent, user, supplier, and time window.
- Diffable across model versions and prompt versions.
Evaluation harness
Offline and online quality gating
- Golden suite of 100-500 labelled cases per agent, owned by the agent author.
- Regression gate: any change to prompt, model, tool, or knowledge re-runs the evals.
- Online probes: sampled live decisions checked against an oracle.
- Faithfulness scoring: output measured against retrieved citations.
- Drift dashboards covering confidence, override rate, and escalation rate.
- Quarterly adversarial red-team probes for every high-risk agent.
View 1
Business replay
Signal to context to policy to orchestration to decision to execution to evidence to replay to learning. Controller and auditor surface.
View 2
Engineering trace
Per-step LLM and tool spans, retrieval chunks, and faithfulness scores. Engineer and risk surface, linked to the same evidence pack.
View 3
Evaluation harness
Golden-set regression gates and online drift dashboards. Early-warning surface for the governance council.
Live runtime proof
This deep dive must not stop at diagrams. These live surfaces prove the runtime is operating, governed, and replayable inside the current estate.
What each surface proves
- High-risk runtime demo: one operational moment from signal to policy to escalation to evidence to replay.
- Control tower: live agents, engineering trace, eval status, deployment state, policy gates, and evidence integrity.
- Operator console: production swap sequence, solution choreography, and the presenter-facing proof path across runtime, resilience, and assurance.
Runtime visibility expectations
- Active agents and policy evaluations must be visible as operational state, not hidden inside backend logs.
- Evidence completeness, replay backlog, deployment state, and AWS/Azure runtime location must be inspectable during the deep dive.
- The page should leave no doubt that bp Sphere has a control plane, not just a slide narrative.
Controlled execution and rogue-agent prevention architecture
This slide must remove fear. The point is to show that phase 1 is read-only, that controlled execution is explicit, and that no agent can write into SAP, Ariba, or treasury systems outside governed pathways.
Agent recommendation zone
- Read-only intake and semantic correlation
- Risk scoring, prioritization, and recommendation
- Evidence generation before execution
- No direct write permission to core systems
Policy + approval gateway
- Approval matrix and authority threshold checks
- SoD enforcement and supplier exception rules
- Escalation gate when risk, value, or evidence completeness demands it
- Named human accountability with override and abort path
Controlled execution gateway
- Execution only through bounded adapters
- Replay and evidence package attached to every action
- Post-action observability, fail-safe, and recovery hooks
- No uncontrolled write, no hidden mutation, no black-box action path
Critical first-phase rule
- Phase 1 agents observe, classify, correlate, prioritize, escalate, and replay.
- Phase 1 agents do not post journals, execute payments, mutate workflows, or change master data.
- This reduces governance resistance while creating immediate productivity and visibility.
What bp stakeholders should hear
- This is not autonomous AI executing in production without oversight.
- This is a controlled enterprise intelligence and orchestration layer with explicit write boundaries.
- Replay and evidence are non-negotiable control features, not optional diagnostics.
Multi-agent coordination
Conflicts, looping handoffs, and collusion
The three zones above govern a single agent's path to action. As soon as two or more agents share authority over the same supplier, invoice, or treasury decision, a distinct set of failure modes appears: conflicting recommendations, looping handoffs, capability creep, and most dangerously, two weak agents quietly reinforcing each other's bad call. bp Sphere addresses this with an explicit coordination plane that sits on top of the gateways above.
Most dangerous failure mode
The most dangerous failure mode in a multi-agent estate is not a single rogue agent. It is two agents quietly agreeing with each other when they should not. The mediator and devil's-advocate pattern makes that structurally hard, and the evidence pack captures the disagreement so controllers can see exactly how consensus was reached.
Coordination primitives
- Mediator agent required for any cross-domain decision and never optional.
- Typed message contract covering claim, evidence, confidence, and dissent.
- Conflict policy that tie-breaks by risk class and escalates by default.
- Deadlock breaker with bounded hop count and wall-time before automatic human handover.
- Quorum rules so high-risk decisions require independent agents in agreement.
Collusion and sycophancy defences
- Independent reasoning: agents see each other's evidence, never each other's drafts.
- Diverse models: high-risk multi-agent flows route across at least two model families.
- Devil's-advocate agent argues against the emerging consensus and its output is preserved in evidence.
- Anti-sycophancy probes inject false premises to test whether agents push back.
- Collusion detection flags cases where agents agree faster than baseline on a high-risk case.
Step 1
Specialist agents
Invoice Risk, Treasury Exposure, and Supplier Risk each produce an evidence pack independently.
Step 2
Mediator agent
Reconciles claims, surfaces dissent, and applies quorum and risk-class rules.
Step 3
Devil's-advocate
Required for higher-tier decisions, and its argument is captured in evidence regardless of the final call.
Step 4
Policy gateway
Receives a single mediated recommendation together with the full dissenting trace.
Cross-domain decision pattern
The coordination plane turns parallel specialist outputs into one governed recommendation with explicit dissent, bounded escalation, and a preserved audit trail for controllers and risk teams.
Agent execution runtime: one runtime for many agents
This slide answers a common boardroom question directly: if there are many agents, who controls them? The answer is that there are not many separate runtimes. There is one governed runtime through which all agents operate.
Enterprise context layer: the difference between answering questions and making decisions
Without context, AI answers prompts. With enterprise context, AI makes governed decisions. This is where bp Sphere becomes different from standalone copilots or document search tools.
Business context
Legal entity, business unit, desk, portfolio, supplier, customer, asset, cost center, and materiality.
Process context
Current workflow stage, prior approvals, open exceptions, unresolved reconciliations, and case age.
Financial context
Exposure, liquidity, margin, budget, actuals, plan, working capital, and close pressure.
Policy context
Authority matrix, SoD, treasury policy, procurement policy, forecast policy, and execution boundaries.
Operational context
Queue state, source freshness, runtime status, upstream failures, and control-plane conditions.
Historical context
Similar cases, accepted overrides, prior outcomes, dispute patterns, supplier behavior, and decision history.
Why bp needs this
- Because the current environment stores facts but does not assemble operating context across systems.
- Because analysts still spend time collecting context before they can apply judgment.
- Because without context, AI becomes generic assistance instead of enterprise decision support.
Questions for bp
- Which context dimensions are already trusted and production-ready?
- Which domains suffer most from missing historical or policy context today?
- Who owns freshness and stewardship for each context class?
Evidence intelligence fabric: the trust layer reused across every mission
Evidence intelligence is one of the clearest reasons bp Sphere is not just another AI layer. It assembles the proof package behind every recommendation and makes that package reusable across domains.
Evidence discovery
Find the records, documents, approvals, communications, and prior decisions relevant to the case.
Evidence resolution
Resolve duplicates, conflicting versions, and related entities across systems and repositories.
Evidence extraction
Extract clauses, milestones, invoice fields, approval facts, and policy-relevant attributes.
Evidence lineage
Preserve source references, timestamps, hashes, authorship, and transformation history.
Evidence packaging
Package the decision trail into one controller-facing and auditor-facing proof set.
Evidence viewer
Expose that package inline for analysts, supervisors, controllers, and approvers.
Progressive transformation roadmap: current state to strategic intelligence enterprise
This roadmap positions bp Sphere as a phased modernization journey rather than a big-bang replacement. It emphasizes coexistence, progressive orchestration, bounded autonomy, and the evolution of the human operating model.
Foundation alignment
Architecture
Define orchestration runtime positioning across AWS, Azure, LaunchPad/Yala, and enterprise integration standards.
Security
Align workload identity, privileged execution policy, telemetry forwarding, and approval requirements.
Data
Identify governed data products, stewardship, quality gaps, and semantic normalization strategy.
Observe and assist
Governed automation
Cross-domain coordination
Enterprise decision runtime to strategic intelligence
Initial read-only agents and their role in building trust
The first wave should not focus on autonomous execution. It should focus on operational intelligence, replayability, policy visibility, and supervisor confidence. These agents observe, correlate, classify, explain, prioritize, and escalate without mutating enterprise systems.
Recommended initial agent set
What they prove
- Enterprise visibility and prioritization without governance fear
- Replay, evidence, and supervisor surfaces as first-class capabilities
- Operational productivity gains before write automation
- Semantic intelligence and cross-system correlation without replatforming
Human + agent operating model for phase 1
How these Phase 1 agents are authored and owned
Each of the ten agents above is a signed artefact in the bp Sphere registry introduced on Slide 04, with a named bp owner. The owner is the author, the on-call, and the person who approves promotion or retirement. That ownership model makes the Phase 1 estate governable from day one, and it opens the path for ops folks across AP, treasury, controllership, and supplier ops to become creators of agentic capability rather than only consumers of it.
Named owner for every agent
Each agent is owned by a senior ops SME in the relevant domain. The owner authors the spec, owns the eval suite, signs off on promotion, and decides when to retire. Accountability is structural, not procedural.
Citizen authoring open from Day 1
Four of these ten agents, Invoice Risk, Reconciliation, Policy Violation, and Data Quality, are explicitly designed to be authored or refined by ops folks through Skill Studio, with platform engineering pairing only on tool-making. This is how bp grows the agent estate without bottlenecking on engineering capacity.
Author = named owner
Operational accountability and runtime accountability live with the same person.
Spec + evals + sign-off
The authored spec, owned eval suite, and promotion decision form the evidence root.
Registry-backed
Every agent is signed, versioned, and promoted through the registry rather than bespoke deployment paths.
4 / 10 citizen-authorable on Day 1
The initial estate is deliberately designed so a meaningful subset can be created and refined by domain operators, not only by engineering teams.
First entries in bp's Skill Studio registry
These ten agents are the first signed entries in bp's Skill Studio registry. Each one carries a named owner, an eval suite, and an explicit retire-or-renew decision path. The estate grows from this registry, not from one-off bespoke deployments.
bp vs AWS/Azure vs bp Sphere capability heatmap
This matrix is one of the most important artifacts in the narrative. It clarifies what bp already owns, what cloud providers supply, what bp Sphere uniquely contributes, and where the enterprise gaps exist today.
| Capability domain | bp internal | AWS / Azure | bp Sphere | Gap / implication |
|---|---|---|---|---|
| Identity and security | Entra ID, IAM, cloud guardrails, existing access standards | Workload identity primitives, IAM roles, managed identities | Agent identity governance, named accountability, policy-aware permissioning | Major gap is AI/runtime-specific access governance rather than raw IAM absence. |
| Runtime infrastructure | Hybrid AWS/Azure compute and container runtime | AI runtimes, workflow services, autoscaling, eventing primitives | Governed orchestration runtime, cross-cloud coordination, supervisory control logic | Clouds provide primitives; bp Sphere provides enterprise orchestration semantics. |
| Integration and eventing | Fragmented APIs, middleware, regional patterns | EventBridge, Event Grid, API management, workflow steps | Canonical context, semantic event overlay, cross-system normalization | Major gap is enterprise semantic coordination, not API connectivity alone. |
| Data and semantics | UDP, Databricks, governed data initiatives, stewardship programs | Cloud data tooling and lineage utilities | Operational context layer, ontology runtime, digital twin context | Semantics and enterprise operational context remain immature today. |
| Agentic orchestration | Minimal enterprise lifecycle governance for agents | Agent frameworks and workflow helpers | Multi-agent coordination, escalation runtime, bounded autonomy, institutional learning | This is one of the biggest differentiators versus native cloud AI. |
| Governance and policy | Existing enterprise policies and fragmented approvals | Cloud guardrails, basic AI guardrails, not enterprise operations governance | Runtime risk classification, approval governance, controlled execution, autonomy governance | Major gap today is policy-aware operational AI, not policy absence. |
| Replay and evidence | Logging and audit, but limited operational replay or cryptographic evidence | Logs and monitoring, not enterprise replay | Decision replay, evidence manifests, lineage, governance replay, trust runtime | This is a primary strategic differentiator and skeptic-killer capability. |
| Human supervisory runtime | Analyst workflows and fragmented escalation | No native supervisor coordination model | Decision theaters, supervisor command surfaces, dynamic thresholds, runtime interventions | Needed to make AI augmentation culturally and operationally credible. |
| Observability and control plane | Infrastructure monitoring | Cloud telemetry and logs | Operational control plane, semantic observability, escalation visibility, replay status | Infrastructure telemetry does not equal enterprise operational observability. |
Data readiness service: trust must be measurable before autonomy is discussable
A credible enterprise platform needs a visible answer to data trust. The question is not only whether data exists. The question is whether it is complete, fresh, owned, traceable, and reliable enough for governed decisions.
Completeness
Required fields, relationships, attachments, and event coverage present for the decision type.
Freshness
Current enough for the operating decision and within an agreed SLA by source.
Lineage
Source, transformation, and stewardship path visible to operators and auditors.
Ownership
Named steward and operational team accountable for quality and remediation.
Trust
Quality signals aggregated into one business-facing confidence posture.
Criticality
Material sources receive stronger monitoring and lower tolerance for degradation.
Illustrative readiness index
| Domain / source | Readiness | Implication |
|---|---|---|
| SAP finance events | 92% | Suitable for governed observation and decision support. |
| Databricks planning products | 88% | Strong for FP&A context, but needs freshness governance on critical cycles. |
| Supplier master | 61% | High governance risk; Sphere should expose this weakness rather than hide it. |
Why bp needs this
- Because trust concerns are often really data-trust concerns expressed in business language.
- Because the runtime should degrade safely when readiness falls below policy thresholds.
- Because bp leadership needs to see where autonomy is blocked by data quality, not just by policy caution.
Federated data strategy and AWS + Azure federation model
bp does not need another giant centralized platform pitch. It needs a coherent explanation of how bp Sphere overlays a federated data and cloud estate, leverages existing UDP and Databricks investments, and coordinates signals across AWS and Azure without pretending the enterprise will be homogeneous.
bp Sphere federation spine
- Canonical enterprise event overlay across AWS and Azure
- Semantic context and ontology resolution over federated data products
- Policy and identity propagation from bp-approved IAM models
- Replay and evidence runtime independent of any single cloud primitive
- Cross-cloud orchestration with local execution boundaries preserved
Data strategy expectations from bp
- Clarify which finance and operational datasets are already governed and production-ready.
- Map stewardship, ownership, lineage, and metadata standards that the runtime must honor.
- Treat semantic normalization as a progressive program, not a prerequisite big-bang cleanup exercise.
- Use federated data intelligence rather than forced centralization where not politically or technically realistic.
Questions that belong on the slide
- Which datasets are already approved for operational AI observation?
- What lineage and metadata standards are already required?
- Where are the biggest current data quality pain points in finance operations?
- How does bp want long-term semantic federation to evolve across UDP, Databricks, and domain teams?
Knowledge plane and agent memory
Grounding for accurate, citation-backed reasoning
The federation spine above carries bp's structured data estate: UDP, Databricks, SAP, and Ariba. Agents operate on a parallel plane of unstructured knowledge, policy PDFs, supplier MSAs, controllership runbooks, and SOX documentation, and on distinct memory stores that allow them to be useful across runs without writing back into systems of record. The plane below reuses the same stewardship, lineage, and quality standards bp already applies to its data products, extended to documents and agent memory.
Governance continuity
Ingest, classification, freshness SLA, and steward attribution mirror the governance principles bp already applies to UDP and Databricks. No new governance regime is introduced; the existing one is extended to documents and agent memory so the data stewardship community can adopt the knowledge plane without learning a separate model.
Stage 1
Ingest
- Connectors to SharePoint, Confluence, and controllership repositories.
- Document classification and sensitivity tagging.
- Versioning and freshness SLA per source.
- Author and steward attribution captured at ingest.
Stage 2
Process and index
- Semantic chunking that respects clauses and tables.
- Hybrid vector and keyword index.
- Entity linking into the canonical ontology used by structured data.
- Quality scoring per chunk for retrieval ranking.
Stage 3
Serve to agents
- Grounding-only retrieval so no agent answer lands without a cited chunk.
- Citation identifiers flow into the evidence pack alongside structured signals.
- Per-agent knowledge scopes enforce least privilege.
- Stale-content quarantine triggers on freshness-SLA breach.
Memory plane
Three stores agents draw on across runs
- <strong>Episodic memory:</strong> per-run journal reusable as case memory, such as a supplier flagged 11 days ago with a known resolution. Bounded retention and fully replayable.
- <strong>Institutional memory:</strong> approved override patterns, accepted lessons, and near-miss incidents, promoted only by the Governance Council and never auto-learned.
- <strong>Working memory:</strong> short-lived session context and intermediate state carried across a bounded task without becoming a system-of-record write-back path.
Agentic compounding asset
Institutional memory is bp's compounding agentic asset: the durable layer where approved lessons and accepted patterns accumulate under governance instead of being re-learned from scratch every run.
Platform-first standards alignment: Yalla, Nexus, LaunchPad, MCP, and A2A
The reviewed bp AI standards make one thing explicit: suppliers are not being judged only on runtime vision. They are being judged on whether their agent estate can operate inside bp's approved platforms, open standards, governance gates, and accountability model. This slide makes that contract explicit.
What bp standards require
- Platform-first delivery through bp internal developer pathways such as LaunchPad / Yalla, not vendor-owned infrastructure.
- Centralized discovery, reuse, and registration through Nexus before new agents or MCP servers are created.
- Mandatory use of MCP for tool and data integration and A2A for multi-agent collaboration.
- AI use-case registration, risk triage, and production gating through AI LaunchPad with named human accountability.
- No direct or ad-hoc enterprise data access outside approved MCP, API, or iHub-style pathways.
- Three Lines of Defence, human oversight, and auditable telemetry for every production AI capability.
What bp Sphere must prove
- Agents are discoverable, owned, and registered in a governance model that aligns to Nexus concepts.
- Every capability has an accountable owner, policy bindings, promotion gates, and retirement path.
- Multi-agent behavior is explicit, orchestrated, and replayable rather than hidden in prompt chains.
- The runtime can show AI inventory, risk class, and LaunchPad-style certification posture for each production use case.
- The operating model stays inside bp's cloud, identity, and governance boundaries instead of introducing a parallel estate.
Platform-first
All AI runs on bp-approved cloud and runtime pathways. The narrative should explicitly reject vendor-owned infrastructure and shadow identity models.
Discoverable and reusable
The agent estate should expose registry metadata, reusable MCP servers, and A2A coordination contracts so capabilities built once can be composed many times.
Governed for production
Unregistered agents, unclassified risks, or non-compliant delivery paths should be impossible to present as “production-ready.”
Golden Path delivery, evaluation assurance, and agentic UX compliance
The standards material raises the bar beyond runtime observability. bp expects Golden Path delivery, bp-controlled repositories and pipelines, formal design review, real evaluation discipline, and a consistent agentic UX with confidence, source attribution, override, and accessibility built in.
Golden Path delivery
- bp-controlled repositories and Azure DevOps pipelines rather than vendor-hosted source control.
- Pre-approved build, security, performance, and deployment gates in the standard CI/CD path.
- Technical Design Document and Technical Design Review evidence before production certification.
- Runbooks, monitoring, rollback, and transition documentation as part of delivery, not post-go-live cleanup.
Evaluation and assurance
- Ground-truth evaluation metrics for AI quality; UAT alone is not enough.
- Component and end-to-end agent evaluation, including retrieval quality, reasoning relevance, tool behavior, and workflow completion.
- Drift thresholds, regression history, and promotion blocking when quality degrades.
- Adversarial or red-team testing for high-risk AI implementations.
Agentic UX compliance
- Confidence signals, source attribution, and human override controls in every AI-facing mission surface.
- Consistent role-based UX so analysts, supervisors, and leaders do not relearn each agent.
- Accessibility and design-system compliance as part of AI readiness, not visual polish.
- Trust, control, and transparency treated as UX requirements rather than hidden engineering details.
What the live system should show
- Registry metadata that looks like a governed production asset, not a demo list of agents.
- Trace drawers with prompt, model, retrieval, latency, citations, confidence, retries, and policy context.
- Evaluation surfaces with golden datasets, pass/fail trends, drift warnings, and promotion gates.
- Mission surfaces where confidence, evidence, and override remain inline for operators.
Proof language for the page
The strongest claim is not “the AI is smart.” It is “the AI estate is governable, testable, discoverable, and auditable inside bp's existing standards.” This is the gap that many suppliers will not close.
Cyber security, IAM, and controlled access model
This deep dive needs to show that bp Sphere integrates into bp’s security architecture rather than creating a shadow runtime. The emphasis is zero-trust alignment, managed identities, controlled privilege, named accountability, and SOC-visible telemetry.
Identity and access architecture
What this validates for bp
- No separate shadow IAM model is introduced; bp identity remains the source of truth.
- Privileged execution can be bounded, audited, and escalated explicitly.
- Named accountability exists for human and agent actions through replay-backed attribution.
- Runtime telemetry can align to enterprise cyber monitoring and compliance expectations.
Questions to resolve with bp security
- What workload identity, JIT access, and PAM patterns are approved for production runtime components?
- What telemetry classes must be forwarded into SOC/SIEM and cyber audit workflows?
- What runtime isolation and cross-cloud segmentation standards must the pilot satisfy?
- What is the required model for break-glass, revocation, and privileged approval logging?
Operational IAM detail
- OIDC-based federation into bp identity with no separate local identity store for the runtime.
- SCIM-aligned lifecycle integration so role and entitlement changes propagate into mission surfaces and approvals.
- Managed workload identities for agents, tools, and connectors, each scoped to least privilege and bounded data domains.
- JIT privileged execution and PAM-compatible break-glass for sensitive actions, with replay-backed privileged traceability.
- Model access routed through the gateway so prompt, tool, and retrieval spans inherit the same identity and policy context.
Security-team message
Security review should be able to map the runtime to existing enterprise controls in operational language: federation, lifecycle, privileged access, segmentation, and traceability. “Supports Entra ID” is not enough; the page needs to show how access is provisioned, approved, elevated, traced, and revoked in practice.
Disaster recovery, resilience, and safe failure model
This needs to show how the runtime behaves when dependencies fail, how human fallback works, and how replay supports recovery. The goal is to prove graceful degradation, not a brittle orchestration engine.
Detect
- Runtime anomalies, queue stalls, evidence lag, replay gaps, and degraded upstream integrations are surfaced early.
- Cross-cloud dependencies and source-system disruption are visible in the control plane instead of being hidden in infrastructure logs.
Contain
- Autonomy can be throttled or downgraded by policy tier when trust conditions degrade.
- Execution pathways can fall back to read-only or human-only modes rather than silently failing or continuing unsafely.
- Supervisor visibility remains available even when automation is reduced.
Recover
- Replay helps reconstruct the incident path, the impacted decisions, and the recovery sequence.
- Human fallback, queue continuity, and evidence integrity remain central to operational restart decisions.
- Post-incident lessons feed governance tuning, resilience certification, and phase progression decisions.
What this answers
- How does bp Sphere fail safely without creating uncontrolled writes or hidden decisions?
- How are incidents replayed and reconstructed for cyber, audit, and operations teams?
- How do human fallback and bounded autonomy work under runtime degradation?
- What is the continuity story when one cloud plane or source system is impaired?
Questions to resolve with bp
- What RTO/RPO and SLA expectations apply to the first pilot domain?
- What DR and failover patterns are already approved across AWS and Azure for this runtime class?
- What crisis-management hooks must integrate with existing resilience workflows?
- What level of degraded operation is acceptable before manual-only fallback becomes mandatory?
Failure scenario that must be demoed
Failure proof surfaces
- The page should show one explicit degraded-mode path instead of implying that resilience exists somewhere in the platform.
- This is where bp sees that replay continuity, human fallback, and policy downgrade are designed into the runtime rather than added after incidents.
Runtime operations center: who runs bp Sphere at 2am?
A boardroom-ready platform story needs an operations answer, not just an architecture answer. This slide explains how the runtime is run, monitored, escalated, and corrected when something goes wrong.
Monitor
- Agent health
- Replay health
- Evidence health
- Queue health
- Policy violations
- Data quality
- Escalation backlog
- Cost posture
Operate
- On-call ownership
- Incident bridge
- Runbooks
- Kill switches
- Fallback activation
- Threshold tuning
Improve
- Root cause patterns
- Override trends
- Failure trends
- Policy friction
- Cost anomalies
- Remediation backlog
Digital workplace, adoption, and workforce operating model
bp will want to understand how work changes, not just how systems connect. This slide should make it credible that analysts, supervisors, and finance leaders move into mission surfaces, inline replay, and human-agent collaboration without adding tool sprawl.
Digital workplace shift
Today
- Inboxes, reports, queues, and system hopping
- Manual prioritization and limited context
- After-the-fact audit reconstruction
- Fragmented collaboration and escalation
Transitional
- Read-only agent assistance
- Replay-backed work queues
- Supervisor-visible escalations
- Policy-aware prioritization and explainability
Future
- Decision theaters instead of dashboard sprawl
- Human-agent collaborative execution
- Enterprise-impact-aware work surfaces
- Institutional learning embedded in daily operations
Adoption and change model
- Start with visible productivity and trust gains rather than autonomous execution.
- Use role-specific mission surfaces for analysts, supervisors, and finance leaders.
- Treat replay, evidence, and explainability as adoption enablers, not technical extras.
- Measure productivity, confidence, override patterns, and operational load reduction as part of rollout governance.
Questions to resolve with bp
- Which user groups should experience the first digital workplace transition surface?
- What training, adoption, and operating-model risks exist for the first rollout?
- What KPIs define meaningful workforce productivity improvement for the pilot?
- How should this align to broader digital workplace and transformation programs already underway?
Hypercare intelligence: how the first production wave is stabilized
The first production wave should not rely on generic project hypercare. It should use runtime intelligence to detect adoption issues, data breakdowns, agent failures, and escalation trends before they damage confidence.
Detect
Adoption issues, data issues, agent failures, policy friction, backlog spikes, and unusual override patterns.
Predict
Which queues will breach, which data sources are becoming unreliable, and which teams are heading toward failure or overload.
Recommend
Corrective actions: retraining, source remediation, policy tuning, threshold adjustment, or fallback activation.
Why this matters
- Because the first 90 days determine whether the platform is trusted or treated as another fragile initiative.
- Because hypercare needs to be intelligence-driven, not only meeting-driven.
- Because bp will judge the platform by how it behaves when adoption, data, or process reality becomes messy.
Questions for bp
- Who owns hypercare decisions across operations, architecture, security, and controllership?
- Which hypercare signals should trigger executive escalation?
- What remediation actions can be automated versus requiring human approval?
Controllership operating model and finance governance alignment
This must explicitly answer how controllership works in the target model. Replay, evidence, policy, approval traceability, close confidence, and treasury exposure need to be framed as part of finance governance, not just runtime plumbing.
Controllership model
What controllers should be able to say
- We can see why a recommendation was made, what policy governed it, and who approved it.
- We can replay how a high-risk decision emerged and what evidence existed at the time.
- We can distinguish read-only intelligence, supervised automation, and controlled write execution clearly.
- We can connect transaction-level interventions to close readiness, exposure, and financial stewardship outcomes.
Questions to resolve with bp controllership
- What are the non-negotiable auditability and replay requirements for finance operations?
- How should close confidence, treasury exposure, and exception posture be surfaced to controllers?
- What dual authorization, evidence retention, and sign-off standards apply?
- What controllership KPIs most clearly prove value in the first phase?
FBT control reality this must map to
- SOX and non-SOX control posture across payables, cash and banking, local close, receivables, tax, and treasury-related processes.
- Unauthorized approval, SOD access, fraud monitoring, and access review concerns already operated by FBT teams.
- Reconciliation, audit assurance, and evidence sufficiency expectations for finance-critical interventions.
- Liquidity, counterparty, and treasury risk implications when a payment, hold, or escalation changes enterprise posture.
What the runtime should eventually expose
- Control-mapping views that show which FBT control families a mission or agentic decision supports.
- Explicit SOD, unauthorized approval, fraud-risk, and reconciliation signals on high-value cases.
- Controller-facing evidence packs that map replay to policy, approval, and audit obligations.
- Treasury and controllership consequence previews attached to sensitive P2P and close decisions.
Next-wave mission agents and enterprise coordination build queue
The next enhancement wave should not be framed as more bots. It should be framed as policy-bound enterprise execution agents operating inside the bp Sphere intelligence fabric. The emphasis is cross-mission orchestration, evidence generation, live operational intelligence, and enterprise decision coordination.
Phase 1 priority agents
- P2P Control Tower Supervisor Agent to coordinate queues, escalations, workload balancing, and SLA jeopardy across AP, procurement, and treasury.
- Enterprise Signal Correlation Agent to connect P2P, Treasury, O2C, FP&A, R2R, and ST&S into replayable causal chains.
- Close Confidence Agent to predict period-end delay, control gaps, and missing evidence for controllership and CFO audiences.
- Enterprise Liquidity Optimization Agent to coordinate cash posture, payment timing, working capital, and treasury actions.
- Evidence and Replay Agent as the signature cross-mission differentiator for immutable evidence packs and provenance.
Why these agents matter
- They elevate the estate from process automation toward enterprise decision operations.
- They strengthen supervisory intelligence, not just task execution.
- They create visible cross-mission coordination rather than isolated domain wins.
- They directly reinforce the control-tower, replay, and policy-runtime claims already shown in the solution narrative.
P2P
Control tower supervision, duplicate intelligence, vendor health and friction, procurement policy drift, and working-capital optimization.
O2C and R2R
Revenue leakage, customer risk and exposure, autonomous dispute resolution, close confidence, journal risk, and reconciliation intelligence.
Treasury, FP&A, and ST&S
Enterprise liquidity optimization, forecast drift, scenario orchestration, supply disruption prediction, margin optimization, and trade exposure intelligence.
ST&S runtime proof now available
The ST&S architecture is no longer only a roadmap claim. The live trading runtime now exposes an autonomy matrix plus a context-graph and decision-ledger surface grounded in Endur-led trade context, shipping state, and source freshness.
How to use it in the room
- Use the matrix to show progressive autonomy and safe first-release scope across trading and shipping.
- Use the ST&S graph to show ontology plus runtime context rather than generic RAG or prompt-only AI.
- Use the decision ledger to answer who recommended what, under which approval boundary, and from which systems.
- Use context quality to address bp concerns about trust, freshness, provenance, and operational safety.
First-wave implementation matrix
The first three agents now need a delivery artifact with agent ID, owner, LaunchPad ID, policy pack, event inputs, replay contract, UI surface, evaluation gate, and delivery phase.
P2P Control Tower Supervisor Agent
LaunchPad ID LP-P2P-CTRL-001. Runtime proof should be live queue reprioritization, replayable escalation rationale, and bounded workload coordination on the analyst mission surface.
Enterprise Signal Correlation and Close Confidence
LaunchPad IDs LP-ENT-CORR-001 and LP-R2R-CLOSE-001. Runtime proof should show cross-mission causal chains, close-confidence change, and controller-facing evidence sufficiency.
Multi-actor operating model and enterprise runtime control plane
This slide must answer how humans and agents coexist, who owns what, and how governance works at runtime. It should make the future-state operating model emotionally and operationally credible rather than abstract.
Human + agent operating model
Enterprise runtime control plane
- Runtime health, event throughput, policy evaluations, escalations, evidence status, and replay availability must be visible as operational intelligence, not just infra telemetry.
- The control plane should support threshold tuning, drift visibility, approval bottlenecks, failure injection, and supervisor intervention from one place.
- This is how bp Sphere becomes a governed enterprise runtime instead of a collection of agents hidden behind screens.
Messier enterprise reality scenarios
- Duplicate supplier master and region-specific workflow variance causing the same operational event to look different across systems.
- Delayed reconciliation and partial ERP coexistence creating close-pressure hotspots that do not map neatly to one source system.
- Conflicting policy versions, hypercare instability, and rollout inconsistency across regions or business units.
- Treasury delay propagation where a blocked approval changes liquidity posture faster than the source teams can see manually.
How the demos should handle reality
- Show one or two messy enterprise scenarios directly, rather than only clean happy-path flows.
- Tie every messy scenario to a specific proof surface: mission runtime, control tower, resilience, replay, or operator console.
- Demonstrate that bp Sphere is useful precisely because the landscape is inconsistent, not because the page pretends it is clean.
Enterprise finance digital twin and strategic intelligence architecture
The narrative should not stop at transaction automation. It should show the long-term direction: invoice, payment, liquidity, supplier trust, and close operations feeding a governed finance digital twin and strategic intelligence runtime.
Local runtime to bp production runtime: the path from proof to controlled enterprise rollout
This is the implementation bridge executives and architects will ask for. It explains how bp Sphere moves from local proof surfaces to controlled enterprise runtime without hand-waving the hard parts away.
Proof runtime
Read-only production observation
Human-approved execution
Scaled operating model
Enterprise execution layer
How the deep dive should run and what it should close with
Session agenda
Best demo moments from one mission surface
Duplicate payment prevention
Signal, orchestration, policy, supervisor escalation, replay, evidence, and measurable protected value.
Early payment discount optimization
A CFO-relevant example showing working-capital intelligence and liquidity-aware decisioning.
Supplier trust escalation
A strong emotional scenario showing cross-functional impact, team routing, and operational criticality.
GR/IR mismatch auto-resolution
The best adaptive-autonomy demonstration: high-confidence recommendation, human approval, and learning loop.
Quarter-close pressure
Shows unresolved invoices, reconciliation backlog, treasury dependency, and close confidence movement.
Policy breach prevention
A governance-first scenario for auditors, controllership, and skeptical architecture leaders.
Questions to resolve with bp
- Which domains and data products are approved for pilot observation first?
- What are the enterprise standards for AI runtime isolation, SIEM forwarding, and privileged execution?
- Who owns runtime certification, policy approval, and autonomy boundary decisions?
- What metrics define a successful phase 1 observe-and-assist pilot?
Risks, assumptions, and dependencies
- Integration readiness and data-product maturity will vary by domain and region.
- Pilot credibility depends on agreed identity, telemetry, and approval patterns up front.
- Governance adoption requires controller, cyber, architecture, and operations sponsorship, not just technology approval.
- Read-only phase 1 scope is the safest commercial and operational entry point; write-back assumptions should stay explicitly bounded.
What we need from bp next
Architecture
Approved integration patterns, target systems for pilot connection, runtime hosting boundaries, and production environment sequencing.
Data
Named data products, stewardship owners, readiness posture, lineage standards, and access approval for the first domains.
Security
Identity model, PAM/JIT expectations, telemetry forwarding requirements, segmentation standards, and privileged execution rules.
Operations
On-call ownership, hypercare model, resilience expectations, pilot KPIs, and governance forums for certification and escalation.
Visible proof 1
Live runtime proof
Signal, orchestration, policy, evidence, replay, and human escalation running on live bp Sphere surfaces.
Visible proof 2
Production swap proof
Explicit mapping from local proof sources to bp production targets, integration contracts, and governance owners.
Visible proof 3
Failure and governance proof
Policy blocking, degraded mode, replay continuity, and human fallback shown as operational behavior rather than slide claims.
The future is governed enterprise execution: replayable, evidence-backed, policy-aware, human-supervised, and credible inside bp’s federated cloud, data, and operational ecosystem. bp Sphere should be presented as the orchestration, trust, and governance fabric that helps bp progressively evolve from fragmented execution into a controllable enterprise decision runtime.