| marp | true |
|---|---|
| theme | default |
| paginate | true |
| size | 58140 |
| header | The AI Stack — from infrastructure to application |
| style | section { font-size: 26px; } h1 { font-size: 44px; } h2 { font-size: 34px; } code { font-size: 0.8em; } pre { font-size: 0.7em; line-height: 1.25; } table { font-size: 0.78em; } section.lead h1 { font-size: 52px; } section.lead { text-align: center; } footer, header { color: #888; font-size: 14px; } |
- Why a stack? — a model is one piece
- The 5 layers — infra → models → data → orchestration → application
- The business view — revenue pyramid · 3 types of innovation
- Case studies — chatbot · RAG · the vendor-vs-operator trap
- GenAI for DevOps — incident RCA · observability · IaC · CI/CD risk
- Takeaways
A model is one piece, not the whole system.
To solve a real problem you also need: compute to run it, data to ground it, orchestration to coordinate it, and an application for the user.
Every choice across the stack drives quality · speed · cost · safety.
5 · APPLICATION interfaces & integrations — the user
4 · ORCHESTRATION plan → execute → review — the "brain" (agentic)
3 · DATA sources, pipelines, vector store, RAG
2 · MODELS open/proprietary · size · specialization
1 · INFRASTRUCTURE GPUs — on-premise, cloud, local
We'll go bottom → top, then look at the business picture.
LLMs need AI-specific hardware (GPUs).
Three ways to deploy:
- On-premise — own it; full control, high upfront cost
- Cloud — rent it; scale up/down on demand
- Local — laptop; only the smaller models
The hardware you can access decides what models you can run at all.
Pick along three dimensions:
- Open vs proprietary — control, cost, where it runs
- Size — large (LLM) vs small (SLM)
- Specialization — reasoning, tool-calling, code, language
Use vs build: consume a model, fine-tune it, or (rarely) train it. Lifecycle glue = MLOps.
Need fresh knowledge, not new behavior? Prefer RAG over fine-tuning.
Base models have a knowledge cutoff and don't know your private data.
sources → pipelines (clean/chunk) → embed → vector store → RAG → model
RAG retrieves relevant context and augments the prompt — fresh & private knowledge without retraining.
One prompt in / one answer out = a chat box. Add the loop and it becomes an agent:
plan ──▶ execute (tools) ──▶ review ──▶ (loop)
▲ │
└──────────── memory / state ◀─────────┘
Two kinds: static (reliable workflows) + agentic (autonomous decisions). Fastest-moving layer — agents, MCP.
Where AI meets the user. Two questions:
- Interfaces — text / image / audio · revisions · citations · conversational UI replacing forms
- Integrations — inbound (tools feed AI) & outbound (AI output → tools)
"What the AI does" is orchestration. "What the user sees / where the result lands" is the application layer.
Where is the money — and where is the value?
┌──────────────┐
│ APPS │ fragmented startups
┌─┴──────────────┴─┐
│ ORCHESTRATION │ LangChain ~$16M ARR
┌─┴──────────────────┴─┐
│ DATA │ Scale AI ~$2B
┌─┴──────────────────────┴─┐
│ FOUNDATION MODELS │ OpenAI ~$33B · Anthropic ~$45B*
┌─┴──────────────────────────┴─┐
│ INFRASTRUCTURE / CHIPS │ Nvidia ~$75B/qtr → ~$300B run-rate
└──────────────────────────────┘
Revenue concentrates at the base (Nvidia alone > all model companies). Hype ≠ revenue — orchestration gets the most attention, the least money (LangChain ≈ 1/18,000 of Nvidia).
Clayton Christensen — a lens for any AI product
| Type | Idea | Jobs | AI example |
|---|---|---|---|
| Market-creating | expensive/hard → cheap for all | creates | foundation models for everyone |
| Sustaining | make a good product better | neutral | an AI feature in your app |
| Efficiency | same work, less cost | reduces | coding assistants, automation |
The type isn't the layer — it's how the tech is used.
Klarna × OpenAI: 2.3M chats in month 1 = 700 agents' work, 67% automated, 11 min → <2 min, ~$40M profit (2024). Caveat: cut too deep, rehired in 2025 — AI-first, not AI-only.
ILLUSTRATIVE — 100,000 contacts/mo · human $5 · AI $1 · 67% deflection
AI handles 67,000 × $1 = $67k/mo
vs humans 67,000 × $5 = $335k/mo
net saving ≈ $268k/mo ≈ $3.2M/yr
Players: Sierra · Decagon · Ada · Intercom Fin
Morgan Stanley × OpenAI
- Indexed 350,000 documents (40M words) with RAG
- Before: 30+ min manual search · advisors reached ~20% of knowledge
- After: instant · 98% of teams use it · access 20% → 80%
Frees advisor time for client work; answers grounded in verified sources.
Similar: Glean · Harvey · Hebbia
The pyramid measures who earns money selling AI → base wins.
Case studies measure value from applying AI → lands on the operator's books.
Klarna's $40M isn't an "AI app company" revenue line — it's on Klarna's P&L. The app layer looks thin for vendors, yet is where operators win.
5 · APPLICATION Slack on-call copilot · NL query · Datadog/PagerDuty
4 · ORCHESTRATION agent: query logs/metrics/traces → traverse deps → RCA
3 · DATA vector DB (FAISS/Weaviate): telemetry + incidents + runbooks
2 · MODELS LLM, prompt-engineered/fine-tuned for logs & remediation
1 · INFRASTRUCTURE GPUs to run it
This is why we learned the stack: a production SRE assistant exercises all five layers at once.
An AI agent watches alerts, reads logs / metrics / traces, and proposes a root cause and a fix.
Tools: Cleric · Resolve.ai · Traversal · Rootly · incident.io · Datadog Bits AI
Proven in production:
| Where | Result |
|---|---|
| Traversal @ American Express | 82% RCA accuracy · −32% MTTR |
| Microsoft RCACopilot | ~0.77 RCA accuracy · 4 yrs, 30+ teams |
Vendor "38–90% MTTR" claims are self-reported. Remediation stays human-in-the-loop.
NL querying + anomaly detection on top of your existing tools (Datadog, Prometheus, Grafana, OpenTelemetry):
- Datadog Bits Assistant — query dashboards/logs/traces in plain language
- Grafana Assistant — NL telemetry questions + ML correlation
- Datadog Toto — timeseries foundation model powering anomaly forecasting
Grounding = RAG over telemetry + incident history + runbooks in a vector DB (FAISS / Weaviate), so the LLM reasons over your system, not generic text.
The problem isn't bad code. It's plausible code — code nobody reads line-by-line anymore.
A 200-line Terraform module, AI-written in 30 s. The dev skims it, sees no
syntax errors, runs apply. What goes unchecked:
- Security group with ingress
0.0.0.0/0? - S3 bucket missing a public-access block?
- RDS without encryption-at-rest? · IAM policy with
Resource: "*"? - Secrets hardcoded in
tfvars?
Anti-pattern: AI generates Terraform →
apply→ 💥 (no review, no scan, no policy gate)
Common misconfigurations in AI-generated Terraform (approx.):
| # | Misconfiguration | Seen in | Why |
|---|---|---|---|
| 1 | Security group 0.0.0.0/0 on 22/3389 |
~25% | common demo examples |
| 2 | S3 missing public-access block | ~20% | not blocked by default |
| 3 | RDS not encrypted | ~18% | encryption is opt-in |
| 4 | IAM Resource: "*" |
~15% | convenient in demos |
| 5 | Secrets in tfvars (no SOPS) |
~12% | pattern is in training data |
| 6 | Missing tags | ~30% | tag policy is org-specific |
| 7 | Hardcoded creds in provider block | ~5% | rarer, but critical |
1 · AI GENERATE Claude Code + Terraform MCP
constraint-first, security-first prompts
↓
2 · HUMAN REVIEW read the plan, understand every resource
ask the agent what you don't get — never approve blindly
↓
3 · POLICY GATE tflint → checkov → terraform plan → conftest
→ AI explains the diff → human approves → apply
RULE: checkov not clean (any HIGH) → no apply
L1 is fast but plausible-but-wrong · L2 catches intent, but humans tire · L3 is automated & consistent — it catches what humans miss.
Emerging, less mature than incident response:
- Blast-radius analysis — LLM predicts what a change can break
- Risk scoring of config / schema changes before rollout
- Auto validation & rollback informed by historical outcomes
Honest take: fewer proven products here than in RCA/observability — strong territory for a platform team to build real differentiation.
- AI is a stack, not a model — infra → models → data → orchestration → app
- Money sits at the base; value from applying AI sits with the operator
- Name the innovation type: market-creating / sustaining / efficiency
- Always do the math with both sides — savings and the AI's own cost
- In DevOps, AI drafts — your pipeline verifies
Questions?
Speaker notes & full numbers: README.md and docs/devops-cases.md