Peeking Under the Hood
When you use an AI tool like a chatbot, it can feel a bit like magic. But here is the key insight: that simple interface sits on top of a massive "Jenga tower" of hardware and software dependencies. If any of those lower layers wobble—due to a chip shortage, a cloud outage, or a change in a vendor's policy—your entire AI strategy can feel the impact.
To lead an organization through the AI transformation, you do not need to be a hardware engineer, but you do need to understand how these layers fit together. Think of it this way: understanding the AI stack is like checking the foundation of a building before you decide to add three new floors. Let's walk through the five layers that make modern AI possible.
Layer 1: The Hardware Foundation
At the very bottom are the physical chips that do the heavy lifting. Unlike your laptop's brain (the CPU), which is good at doing one complex thing at a time, AI needs chips that can do thousands of tiny, simple math problems simultaneously. This is why GPUs (Graphics Processing Units) are the gold standard for AI today.
Currently, a single company creates a massive bottleneck. NVIDIA dominates this layerNVIDIA dominates this layer, controlling roughly 80-90% of the market. This creates a "concentration risk"—nearly every AI system your company uses probably depends on this one supplier. To understand the scale of investment here, consider the cost of a single high-end H100 GPUcost of a single high-end H100 GPU, which can range from $25,000 to $40,000.
Layers 2 & 3: Infrastructure and Platforms
The next two layers turn that raw hardware into something your team can actually use. Infrastructure is the cloud environment provided by giants like AWS, Azure, or Google Cloud. This is where your data is physically processed, which raises important governance questions about data residency and whether your information stays within specific borders.
The Platform layer provides the tools and APIs that let developers build AI apps without managing the hardware themselves. This is where you might face "vendor lock-in." Here is what matters: if you build your entire workflow on one specific provider's platform, moving to a competitor later can be expensive and time-consuming. We recommend using an AI Stack Assessment FrameworkAI Stack Assessment Framework to map these dependencies early.
Layers 4 & 5: Models and Applications
Layer 4 is the "intelligence" of the system—the models. Most of the world now uses "foundation models," which are massive systems trained on broad data that can be adapted for many tasks. While some organizations use open-weight models that they host themselves, many rely on closed APIs from companies like OpenAI or Anthropic.
Finally, at the top is the Application layer. This is the chatbot, the resume screener, or the analytics dashboard your employees see. Many of these applications are actually "thin wrappers"—simple interfaces sitting on top of someone else's model and hardware. If the model provider below them changes their rules, your application could change overnight.
Managing the Costs and Risks
As you evaluate your AI portfolio, it is helpful to distinguish between training (the one-time cost to teach the model) and inference (the ongoing cost every time someone uses the model). You might wonder where the energy goes. Surprisingly, inference often consumes more energyinference often consumes more energy in the long run than the initial training. Every query your employees send adds to the bill.
Because the AI stack is so interconnected, simple questions about "buying AI" are rarely simple. We recommend using an Organizational AI Compute Governance ChecklistOrganizational AI Compute Governance Checklist to track your usage and establish budgets for "exploding" cloud costs. By understanding these dependencies, you can move from just "using AI" to governing it with the clarity your organization needs.