Most enterprises have AI projects. Few have an AI operating model. The difference is the difference between a company that runs experiments and a company that operates AI at scale.
An AI operating model answers four questions that project-level thinking doesn't: Who owns AI decisions? How do AI systems move from idea to production? What governance applies at each stage? And how does the organisation learn and improve its AI capabilities over time?
This guide breaks down what a complete operating model includes and how to build one in phases. For hands-on design support, our AI operating model consulting engagement works through the same diagnose, design, validate, and implement phases with your team.
Why You Need an Operating Model
Without an explicit AI operating model, organisations default to one of two failure modes:
The Two Failure Modes
- Anarchy โ every team builds AI independently. No shared infrastructure, no consistent governance, no reuse. Each project reinvents the wheel, compliance is inconsistent, costs are invisible, and the board has no idea what AI the organisation is actually running.
- Bottleneck โ a central AI team controls everything. Every request goes through a queue. Business units wait months for AI capabilities. Innovation dies. The central team becomes the most resented group in the organisation.
The operating model creates the middle ground: shared infrastructure and governance with distributed execution. The AI equivalent of a platform engineering approach.
The Operating Model Components
1. Organisational Structure
Three models, each with trade-offs:
Centralised AI Team
- Single team owns all AI development and deployment
- Pros: consistent quality, efficient resource use, strong governance
- Cons: slow, becomes a bottleneck, disconnected from business context
- Best for: organisations with fewer than 5 active AI use cases
Federated Model
- Business units have embedded AI capability; central team provides platform and governance
- Pros: faster delivery, business context preserved, scales with the organisation
- Cons: harder to maintain consistency, requires strong platform team
- Best for: organisations with 5-20 AI use cases across multiple business units
Hub and Spoke
- Central Centre of Excellence sets standards and provides advanced capability; spokes in each BU handle day-to-day AI operations
- Pros: balances consistency with agility, enables knowledge sharing
- Cons: complex governance, requires mature organisation
- Best for: large enterprises with 20+ AI use cases
2. AI Lifecycle Management
Every AI system follows a lifecycle. The operating model defines what happens at each stage:
The AI Lifecycle Stages
- Ideation โ use case identification, business case, feasibility assessment. Gate: is this worth pursuing?
- Assessment โ data readiness, technical feasibility, risk classification (EU AI Act), resource requirements. Gate: can we build this responsibly?
- Development โ model selection/training, pipeline development, testing. Gate: does it meet quality and compliance standards?
- Deployment โ production rollout, monitoring setup, user training. Gate: is the operational environment ready?
- Operations โ performance monitoring, drift detection, incident response. Continuous: is it still performing as expected?
- Retirement โ decommissioning, data cleanup, documentation archival. Gate: can we safely turn this off?
Each gate has defined criteria, approvers, and documentation requirements. The rigour scales with risk โ a low-risk internal productivity tool has lighter gates than a high-risk customer-facing decision system.
3. Governance Framework
The governance layer defines who makes what decisions:
- AI Steering Committee โ executive-level. Sets strategy, allocates budget, approves high-risk deployments. Meets quarterly.
- AI Governance Board โ cross-functional (legal, compliance, engineering, business). Reviews risk classifications, policy exceptions, incident escalations. Meets monthly.
- AI Ethics Review โ evaluates fairness, bias, and societal impact for high-risk systems. Triggered by risk classification, not calendar.
- Platform Team โ technical governance. Maintains infrastructure standards, deployment pipelines, security controls. Operates continuously.
4. Platform and Infrastructure
The shared AI platform is the operating model's backbone:
- Model registry โ versioned catalog of all models (internal and third-party) with metadata, approval status, and usage tracking
- Deployment pipelines โ standardised CI/CD for model deployment with automated testing, security scanning, and compliance checks
- Inference infrastructure โ shared model serving with auto-scaling, cost allocation, and multi-tenant isolation
- Observability stack โ unified monitoring for all AI systems: latency, quality, drift, cost, and compliance metrics
- Guardrails engine โ centrally managed content filtering, PII detection, and output validation applied to all deployments
5. Talent and Capability Model
Define the roles, skills, and career paths:
- Core roles โ ML Engineer, Data Engineer, AI Platform Engineer, AI Product Manager, AI Governance Specialist
- Skill matrix โ what capabilities exist today vs. what's needed. Gap analysis drives hiring and training investment
- Career paths โ AI roles need defined progression. Without it, you lose talent to companies that offer it
- Training programme โ continuous upskilling, not one-time onboarding. AI evolves faster than annual training cycles
- External partnerships โ which capabilities do you build internally and which do you source from consulting partners?
6. Financial Model
How AI costs are tracked, allocated, and justified:
- Cost centres โ platform costs (shared infrastructure) vs. project costs (specific AI initiatives)
- Chargeback model โ how business units pay for AI platform usage. Per-inference, per-seat, or allocation-based
- ROI framework โ standardised methodology for measuring AI business value. Without this, every project invents its own ROI story
- Investment governance โ thresholds for AI spending approval. Small experiments should be fast. Large investments need business cases
Kubernetes Recipes
A practical guide for container orchestration and deployment by Grzegorz Stencel & Luca Berton (Apress).
Watch on Skillshare โImplementation Roadmap
You don't build an operating model in one go. Phase it:
Phased Implementation
- Phase 1 (Months 1-3): Foundation โ AI inventory, risk classification of existing systems, basic governance structure, platform team mandate
- Phase 2 (Months 3-6): Platform โ shared infrastructure deployment, deployment pipeline standardisation, monitoring and observability baseline
- Phase 3 (Months 6-9): Governance โ lifecycle management processes, approval workflows, compliance automation, training programme launch
- Phase 4 (Months 9-12): Optimisation โ cost optimisation, advanced capabilities (A/B testing, multi-model routing), cross-BU knowledge sharing, operating model review and adjustment
Measuring Operating Model Maturity
- Level 1 โ Ad hoc: individual projects, no shared infrastructure, no governance
- Level 2 โ Repeatable: some shared tools, basic governance, manual processes
- Level 3 โ Defined: operating model documented, platform in place, governance enforced
- Level 4 โ Managed: metrics-driven, automated compliance, proactive risk management
- Level 5 โ Optimising: continuous improvement, innovation pipeline, industry-leading capabilities
Most regulated enterprises are at Level 1-2. Getting to Level 3 is the critical transition โ it's where AI stops being a series of experiments and becomes an organisational capability. That transition requires an operating model. There's no shortcut.
Back-End Infrastructure: Servers, Secure APIs and Data
Build secure back-end infrastructure from the ground up. In collaboration with Starweaver.
Start on Coursera โRelated Solution
Navigating AI adoption in a regulated environment? Our readiness assessment maps infrastructure, governance, and compliance gaps in 2-3 weeks.
Explore AI Readiness for Regulated Enterprises โ
Luca Berton
