LLMs Changed the Game โ But Not the Rules
Traditional MLOps was built for training and deploying predictive models: churn prediction, fraud detection, recommendation engines. LLMs introduce fundamentally different operational challenges: prompt management, hallucination monitoring, cost per token, and the RAG vs fine-tuning decision. But the underlying principles โ reproducibility, monitoring, governance โ still apply.
What's the Same
- Version control โ You still need versioned, reproducible configurations (prompts instead of model weights, but same principle)
- Monitoring โ Models in production need continuous monitoring (drift for ML, hallucination/quality for LLMs)
- Governance โ Audit trails, access control, and compliance requirements don't change because the model architecture changed
- CI/CD โ Automated testing and deployment pipelines are still essential
- Data management โ Training data quality, data lineage, and data governance remain critical
What's Different
| Dimension | Traditional MLOps | LLMOps |
|---|---|---|
| Model | Train from scratch or fine-tune | Use foundation model + prompt/RAG/fine-tune |
| Config | Hyperparameters + feature engineering | Prompts + retrieval config + guardrails |
| Evaluation | Accuracy, F1, AUC (quantitative) | Qualitative + quantitative (human eval, LLM-as-judge) |
| Failure mode | Wrong prediction (measurable) | Hallucination (hard to detect automatically) |
| Cost | Training cost (fixed) + inference (low) | Token cost per request (variable, can be very high) |
| Security | Data poisoning, model extraction | + Prompt injection, jailbreaking, data leakage |
Kubernetes Recipes
A practical guide for container orchestration and deployment by Grzegorz Stencel & Luca Berton (Apress).
Watch on Skillshare โRAG vs Fine-Tuning: The Enterprise Decision
| Factor | RAG | Fine-Tuning |
|---|---|---|
| When to use | Dynamic knowledge, documents, FAQs | Consistent behaviour, domain language, format control |
| Data needed | Documents, knowledge base (no labelling) | Curated input/output examples (100-10K+) |
| Update frequency | Real-time (update index) | Periodic (retrain model) |
| Cost | Higher per-request (longer prompts) | Higher upfront, lower per-request |
| Compliance | Data stays in your infra (retrieval) | Data sent to provider for training (unless self-hosted) |
Enterprise pattern: Most regulated enterprises start with RAG (data stays in your control), then add fine-tuning for specific high-volume use cases where RAG latency or cost is problematic.
LLMOps Infrastructure Stack
- Prompt management: Version control for prompts with A/B testing (LangSmith, Humanloop, custom)
- Vector database: For RAG retrieval (Pinecone, Weaviate, pgvector, Qdrant)
- Guardrails: Input/output validation to prevent prompt injection and data leakage (Guardrails AI, NeMo Guardrails)
- Evaluation: Automated evaluation pipelines using LLM-as-judge + human review for edge cases
- Cost monitoring: Token usage tracking per user/team/use case (prevents surprise bills)
- Observability: Request logging, latency tracking, error rates, hallucination detection (Langfuse, Helicone)
Operationalizing ML Models: MLOps for Scalable AI
Turn ML prototypes into robust, scalable systems using real-world tools. In collaboration with Starweaver.
Start on Coursera โCompliance Implications
- EU AI Act: LLM-powered applications may be high-risk depending on deployment context. The same LLM used for marketing copy (minimal risk) and credit assessment (high-risk) has different compliance obligations.
- GDPR: Prompt logs may contain personal data. Retention, access control, and right-to-erasure apply to LLM interaction logs.
- DORA: If LLMs support critical business functions, they fall under ICT risk management, incident reporting, and resilience testing requirements.
Luca Berton
