Skip to main content
← All posts·
AI Infrastructure

Designing an AI Center of Excellence: Operating Model, Governance, and Scaling Patterns

How to design and launch an AI Center of Excellence that scales AI adoption across the enterprise. Covers operating models, staffing, governance integration, and common failure patterns to avoid.

Luca Berton14 min read

Every enterprise AI strategy eventually arrives at the same question: how do we scale AI from isolated projects to an organizational capability? The most common answer — and one of the most frequently failed — is the AI Center of Excellence (CoE).

The failure rate of AI CoEs is high. Most fail not because of technology but because of organizational design: wrong mandate, wrong staffing, wrong relationship with the rest of the business. Here's how to design one that works.

CoE Operating Models

Model 1: Centralized CoE

  • Structure: Single team that owns all AI development, deployment, and governance
  • Best for: Early-stage AI adoption (0-5 AI systems in production), small organizations (<5,000 employees), heavily regulated industries where tight control is required
  • Advantage: Consistent quality, strong governance, efficient use of scarce AI talent
  • Risk: Becomes a bottleneck. Business units wait months for the CoE to build what they need. Shadow AI emerges.

Model 2: Federated CoE (Hub and Spoke)

  • Structure: Central team provides platform, governance, and standards. Business unit teams build AI solutions using centrally provided tools and guardrails.
  • Best for: Mid-stage AI adoption (5-50 AI systems), organizations with diverse business units, enterprises needing both speed and governance
  • Advantage: Scales better than centralized. Business units get speed; central team ensures governance.
  • Risk: Quality inconsistency across spokes. Governance becomes advisory rather than enforced.

Model 3: Embedded AI (Distributed)

  • Structure: AI capability embedded in every product and engineering team. Central team provides platform infrastructure only.
  • Best for: Mature AI adoption (50+ AI systems), organizations with strong engineering culture, companies where AI is core to every product
  • Advantage: Maximum speed. AI fully integrated into business operations.
  • Risk: Governance is hard to enforce. Standards fragment. Compliance becomes a challenge.

CoE Functions

Platform Engineering

The CoE builds and operates the ML platform that everyone else uses:

  • ML platform (training, serving, monitoring infrastructure)
  • Feature store and data platform integration
  • Model registry and deployment pipelines
  • Self-service tooling for data scientists and ML engineers

Governance and Compliance

  • AI governance framework design and enforcement
  • Model risk management processes
  • AI Act compliance — risk classification, documentation, conformity assessment support
  • Ethics review for high-risk and sensitive AI use cases
  • Audit support — evidence collection, reporting, regulatory liaison

Enablement and Training

  • AI literacy programs for the broader organization
  • Technical training for data scientists and ML engineers
  • Best practices, templates, and reference architectures
  • Community building — AI community of practice, knowledge sharing, internal conferences

Advisory and Consulting

  • Use case assessment and prioritization support for business units
  • Architecture review for new AI initiatives
  • Vendor evaluation and selection support
  • Embedded support for strategic AI projects
📘 Book

Kubernetes Recipes

A practical guide for container orchestration and deployment by Grzegorz Stencel & Luca Berton (Apress).

Watch on Skillshare

Staffing the CoE

Core Roles

  • CoE Lead / Head of AI: Senior leader with credibility across technology, business, and compliance. This role makes or breaks the CoE.
  • ML Platform Engineers: Build and operate the infrastructure. Kubernetes, MLOps, data engineering skills.
  • AI Governance Lead: Owns governance framework, compliance processes, and regulatory alignment.
  • Senior Data Scientists / ML Engineers: Technical leaders who set quality standards and mentor business unit teams.
  • AI Product Manager: Translates between business needs and technical capabilities. Manages the CoE backlog.
  • AI Ethics Specialist: For organizations deploying high-risk AI — bias testing, fairness assessment, ethical review.

Common Failure Patterns

  1. The Ivory Tower: CoE builds impressive infrastructure that nobody uses because they didn't talk to business users about what they actually need.
  2. The Bottleneck: Every AI request goes through the CoE queue. 6-month wait times. Business units give up or go rogue.
  3. The PowerPoint CoE: Produces frameworks, methodologies, and governance documents — but doesn't build anything or help anyone build anything.
  4. The Science Project: Focuses on technically interesting problems rather than business-valuable solutions. Impressive demos, zero production impact.
  5. No Executive Sponsor: CoE operates without senior leadership backing. First budget cut, it's gone.
🎓 Course with Starweaver

Back-End Infrastructure: Servers, Secure APIs and Data

Build secure back-end infrastructure from the ground up. In collaboration with Starweaver.

Start on Coursera

Success Metrics

  • AI systems in production: Not models built — models deployed and generating business value
  • Time to production: Average time from use case approval to production deployment
  • Business unit satisfaction: NPS or satisfaction score from business units the CoE serves
  • Platform adoption: Percentage of AI projects using CoE platform vs. building their own
  • Governance compliance: Percentage of AI systems compliant with governance framework
  • Business value delivered: Measurable business impact of AI systems supported by the CoE
center of excellence
ai coe
operating model
ai governance
ai scaling
organizational design

Need help applying this in your organization?

Get a free 30-minute assessment with actionable recommendations — whether we work together or not.

Book Your Free AI Platform Assessment

18+ years experience · Ex-Red Hat & Dell · Speaker at KubeCon EU 2026

Luca Berton

Written by

Luca Berton

CEO at Open Empower. 18+ years building enterprise infrastructure at JPMorgan Chase, Red Hat & Dell. Author of 9 technical books. Speaker at Red Hat Summit and KubeCon EU 2026. Instructor on Coursera, Pluralsight & Udemy.

Get more insights like this

Practical AI infrastructure and platform engineering guides — delivered to your inbox.

Subscribe to Newsletter →