The problem isn't the bill. It's that nobody can explain it.
AI spending comes through three different channels: per-user licenses such as Copilot, agent platforms, and custom applications that call AI APIs. Each follows a different pricing model. When costs rise, it is often unclear which team, application, or workload caused the increase.
Start small and focus on your biggest AI cost.
Begin with the largest invoice, an upcoming license renewal, or a fast-growing API account. Review three to six months of data and assign an owner to every major cost. Document who uses it and who manages the budget. A simple spreadsheet is enough to get started.
Match the solution to the type of spend.
Different AI costs need different controls. Unused Copilot licenses may require a usage review before renewal. Custom AI applications need clear ownership, budgets, and spending limits. Using the wrong control wastes time and can disrupt tools people depend on.
Route the AI traffic you control through a central gateway.
A central gateway gives you one place to monitor spending, control usage, cache repeat requests, and direct simpler tasks to lower-cost models. Start with a single application and measure cost and quality before and after the change.
The cheapest model does not always deliver the lowest cost.
A lower-cost model that leads to failed tasks, retries, or manual rework can increase overall expenses. Track the cost per successful outcome rather than the cost per token. That metric shows whether a change actually improves efficiency.
Make cost management part of the process.
Set one simple rule: every AI tool and API key must have an owner, a budget, and a review date. Add a monthly actual-versus-forecast review to catch issues early and prevent the same cost surprises from returning.
This is not simply a billing problem. Enterprise AI is purchased and consumed in multiple ways. An organization may pay per user for Microsoft 365 Copilot, Claude, Gemini, or coding assistants. It may consume credits and connected services through agent platforms. Its own applications may access models through APIs, generating costs across tokens, search, storage, network, and observability. Each model of consumption creates different data, ownership structures, and governance requirements.
AI FinOps has become urgent because AI is no longer a small experimental expense. According to the FinOps Foundation's 2026 State of FinOps, 98% of organizations now manage AI spend, up from 63% in 2025 and 31% in 2024. The same survey identifies AI cost management as the number one capability teams need to develop. AI has moved beyond isolated licenses and pilot projects into applications, agents, and business processes that consume technology continuously.
At the same time, the economics of AI are becoming more complex. While token prices have fallen, providers increasingly charge through a combination of licenses, tokens, credits, cached context, reasoning effort, tools, and agent actions. A single user request may trigger multiple model calls, tool invocations, and retries. As a result, behavior drives consumption, and consumption drives cost. AI FinOps must therefore connect who uses AI, how they use it, what it costs, and whether the outcome is worth that cost. A lower cost per token alone cannot provide that full picture.
A single dashboard doesn't solve the problem. Before organizations can reduce waste or improve forecast accuracy, they need a clear understanding of what they buy, who owns it, how requests flow, and what success looks like for each service. Only then can they implement controls that fit the service, rather than controls that simply appear comprehensive.
At Lingaro, we started by applying these principles to our own AI estate. We mapped spend, assigned ownership, tested controls, and learned which approaches work best across different AI consumption models. That experience, combined with our expertise in AI platforms, FinOps tooling, gateways, and observability, became the foundation for AI Cost Hospital. Today, we help organizations apply the same discipline to their own AI environments.
The AI Cost Hospital framework is built around four stages: Diagnose, Treat, Monitor, and Prevent. First, we identify where spend originates, what drives it, and who owns it. Next, we determine which controls are practical and effective, establish ways to track cost alongside quality and business value, and embed efficiency into everyday governance and engineering processes. This journey helps organizations apply the right governance to each AI consumption model, from user-based licenses and hosted assistants to client-managed applications and APIs.
The first challenge is the mix of commercial models. Pricing varies by product, plan, contract, and configuration. A hosted assistant may be licensed per user, while a managed platform may charge based on the features it uses. Client-managed model APIs typically generate usage-based costs and may also incur charges for search, storage, networking, or observability services.
Each source provides only part of the picture. Reports from Microsoft 365 Copilot and GitHub Copilot can help you review assigned licenses and usage, but usage alone does not demonstrate business value. Microsoft Foundry offers cost analysis, budgets, exports, and tagging, but shared deployments and credentials can still make it difficult to identify which application generated a charge. Copilot Studio costs depend on agent design and feature usage, while connected services may have separate charges. To understand the full cost, you need to combine financial data with clear ownership and measurable business outcomes.
The second challenge is ownership. The invoice sits with finance, the contract with procurement, the cloud account with a platform team, the API credential with a developer, and the business outcome with a product owner. Shared identities and loosely governed experiments weaken these connections. When a cost spike appears in a monthly report, the person best placed to explain it may not even know there is a question.
The final challenge is the pace of change. A pilot becomes a production agent. Assistant adoption grows. An agent begins making more tool calls than its designers expected. Traditional FinOps practices remain important, but AI introduces new considerations such as model quality, behavior, and task success. A lower-cost model is not necessarily the most economical choice if it leads to more failed tasks, retries, or human correction.
| Type of AI spend | Evidence and question | Main management approach | Request-level control |
| Hosted assistants and licenses | Vendor billing and administration reports. Are the right users licensed, and is the product creating enough value? | Review assignments, activity, overlap, user needs, adoption, and renewal options. | Usually limited to controls exposed by the vendor. |
| Managed agent platforms | Platform consumption, environment, agent, and connected-service data. Which agents consume capacity, and what design choices drive it? | Assign owners; review unused agents, capacity, tools, connectors, and supported platform settings. | Platform-specific and dependent on configuration. |
| Client-controlled applications and APIs | Provider billing joined with application and request telemetry. Which application and request type creates the cost, and does the result justify it? | Give applications distinct identities and apply suitable budgets, quotas, model choices, caching, and retry controls at an approved endpoint. | Strong where traffic is deliberately routed through the control point. |
The most important distinction is not the product name. It is who controls the request path. A gateway can govern traffic that is deliberately routed through it, but it cannot replace the model inside a closed, vendor-hosted assistant or observe interactions that the vendor does not expose.
Most organizations begin their AI cost analysis with a symptom rather than a complete picture. An invoice has increased; licenses appear underused, an agent costs more than expected, or no one trusts the forecast. The natural response is to act quickly, but addressing the visible symptom can simply shift the problem elsewhere. Reducing licenses without understanding adoption may remove valuable tools, while switching models without testing quality may lower model costs but increase failed tasks and the need for human correction.
This pattern shaped the framework we call AI Cost Hospital. We start by diagnosing spending patterns and ownership gaps, then apply controls that fit the service. From there, we track whether cost and quality improve together and use those insights to prevent the same issues from recurring. The framework keeps the work practical: understand the problem before making changes, measure the impact of interventions, and turn the results into better day-to-day decisions.
| Diagnose | Treat | Monitor | Prevent |
|
|
|
|
The stages are connected, but they do not follow a rigid sequence. A license renewal may start in Diagnose. A runaway client application may require an immediate Treat action before its attribution model is perfect. Monitoring can reveal new diagnostics questions, while governance should continually prompt the organization to reassess what has changed.
Start with a simple goal: explain your largest AI costs in plain language. What did the organization buy? Who uses it? Who owns it? What caused the bill to increase?
Do not wait for a perfect enterprise-wide inventory. Choose a useful starting point, such as the largest invoice, an upcoming license renewal, or a rapidly growing API account. Review three to six months of available data as a starting point. A spreadsheet is often enough for an initial assessment; while SQL, Power BI, or an existing FinOps platform becomes valuable when the data needs to be refreshed and analyzed regularly.
The information usually comes from a few familiar sources:
| Information needed | Likely source | How to retrieve it |
| Amount paid and charging model | Finance system, invoices, procurement records, contracts | Export invoice lines or purchase order report; record the unit price, commitment, renewal date, and cost center. |
| Assigned and active assistant licenses | Microsoft 365 admin center or the relevant vendor administration portal | Download the license and usage report as CSV, where the product exposes it, or use the vendor API where available. Compare assigned seats with recent activity by team. |
| Cloud and model consumption | Azure Cost Management exports, AWS Cost and Usage Report, Google Cloud Billing export, or model-provider usage reports | Schedule a billing export or download the provider report. Keep subscription, account, project, resource, model, token, and tag fields where available. |
| Application-level API use | An existing API gateway, LiteLLM, application logs, or provider request logs | Export usage by application identity, API key, virtual key, project, model, and environment. Prompt content is not required for basic cost attribution. |
| Business owner and purpose | Service catalogue, CMDB, architecture register, product portfolio, or interviews | Match the account or application ID to a named technical owner, budget owner, and business purpose. Record unknowns rather than guessing. |
Bring the available data into a single, simple list. For each major cost, record the product or service, amount, owner, user group or application, and the source of the evidence. Focus first on the largest or fastest-growing items rather than trying to explain everything at once.
Expect gaps. Shared accounts and API keys may make it difficult to identify which application generated a cost. Tags or ownership information may be missing. License reports may show assignments or activity, but not business value. Invoices and usage reports may also cover different time periods or levels of detail. Record these items as unknown and assign an owner to resolve them rather than filling the gaps with assumptions.
By the end of the Diagnose stage, the team should be able to answer four questions:
What are we paying for?
Which team, user group, or application uses it?
Who owns the budget and the service?
What information is still missing before we can act?
Treatment begins once the organization understands the type of cost it is managing. The goal is not to reduce usage indiscriminately. It is to eliminate avoidable costs while protecting quality, reliability, security, privacy, and adoption.
Central API AI Gateway. Route as much eligible, client-controlled AI traffic as practical through an approved central gateway. This creates a single point for identifying applications, applying authentication, setting quotas and rate limits, tracking usage, and enforcing common policies. Rather than having each team build these controls independently, the organization benefits from a shared governance and control layer.
LiteLLM is an example of a gateway solution. Its proxy provides applications with a common interface across model providers, while virtual keys can separate teams or applications for access control, budgets, rate limits, and spend tracking, depending on how the platform is deployed and configured. This allows applications to integrate once while the platform team manages providers, model aliases, fallbacks, and shared controls centrally.
This does not mean every AI product can or should be routed through a gateway. Hosted assistants and some managed platforms keep request paths within the vendor’s service and must be governed through the controls those products provide. For applications the organization controls, however, gateway coverage offers stronger governance options and a clearer view of API usage.
Smart model routing & caching optimization
Many gateway and model API solutions already support routing based on factors such as availability, rate limits, latency, or cost, along with built-in fallbacks and load balancing. These capabilities can deliver quick improvements without the need for a separate optimization platform. LiteLLM supports several of these patterns and can also expose model aliases or routing groups, allowing applications to use a stable interface without selecting a specific provider for deployment.
More advanced optimization can use organization-defined model bundles, such as a fast bundle for simple tasks, a balanced bundle for general workloads, and a high-capability bundle for complex reasoning. Custom rules or a real-time decision engine can then select the most appropriate bundle based on task type, quality history, latency, cost, availability, and data restrictions. This approach requires robust observability. The decision engine needs reliable signals on requests, outcomes, quality, and cost, along with tested fallbacks and safeguards. Without that evidence, dynamic routing can simply shift cost or quality issues rather than resolve them.
The same control layer can also support caching for stable, repeatable requests while making retries, loops, and unusual consumption patterns easier to identify. Consistent identities, routing records, and cost metadata provide the foundation for testing increasingly advanced rules and confirming that lower costs do not come at the expense of successful outcomes.
Start with one suitable application rather than moving everything at once:
Record its current cost, quality, latency, and successful outcomes.
Give the application a distinct identity and route it through the gateway.
Test one routing, quota, or caching rule on representative work.
Check that cost improved without breaching quality, reliability, privacy, or risk limits.
Keep, adjust, or stop the rule, then use the evidence before expanding coverage.
Monitoring should help people identify meaningful changes, understand them, and take action. The goal is not to collect every possible metric. It is to maintain a shared, explainable view of AI costs, understand what drives them, and spot issues early enough to respond.
AI cost data foundation & reporting
Bring together the main billing records and the identities established during Diagnose and Treat. This may include license reports, cloud and model-provider billing data, gateway usage records, application IDs, service owners, and agreed upon quality and business-outcome measures. Start with cost and attribution metadata rather than raw prompts and responses, following the storage and privacy controls defined during Treat.
The first report should answer practical questions: How much are we spending? Which products, models, teams, or applications drive that cost? How has it changed over time? Use consistent definitions for cost, usage, ownership, and reporting periods so finance, engineering, and product teams work from the same data. Existing finance, analytics, or reporting platforms are often sufficient. A new platform is not automatically required.
Privacy, security, and regulatory review
A gateway and observability layer generates traces and logs that must be stored and governed appropriately. Those records may include prompts, responses, user identifiers, tool inputs, or application data. Employees may also enter personal or non-work information into AI tools, even when policies discourage it. As a result, treating all telemetry as harmless technical data can create unnecessary privacy and security risks.
Plan the data design early, alongside the gateway and routing design. Start with the questions the organization expects to answer. Basic cost allocation may require only application identity, model, token, latency, and cost metadata. More advanced routing optimization may also need task categories, quality outcomes, fallback decisions, and links between related requests. Collect detailed content only when there is a clear need for debugging, evaluation, or compliance, rather than by default.
Store the selected data in an approved environment with encryption, restricted role-based access controls, audit logging, appropriate data residency, and clear retention and deletion policies. Where possible, separate cost metadata from sensitive content, and use redaction, pseudonymization, aggregation, or sampling to reduce exposure. Privacy, security, legal, records management, and AI governance teams should agree on these controls before broad data collection begins.
The GDPR governs the processing of personal data in the EU/EEA. Obligations under the EU AI Act depend on the system, use case, and the organization's role. Logs can support traceability, but logging and cost monitoring alone do not demonstrate compliance. Requirements should still be reviewed on a case-by-case basis.
Observability & alerting
Billing data explains what was charged. Operational data helps explain why. Gateway and application logs can reveal model use, token consumption, latency, errors, retries, loops, cache performance, and unexpected changes in demand. Focus alerts on conditions that require action, such as a sharp increase in costs, repeated failures, or usage outside expected ranges. Each alert should have a clearly defined owner.
Choose tools based on the questions you need to answer and the platforms already in use. LiteLLM or standard gateway logs may be sufficient for basic attribution and alerting. MLflow can be valuable where Databricks or MLflow already supports experimentation and evaluation, while Langfuse may suit teams that need dedicated LLM tracing and investigation capabilities. These are examples rather than required components. Every additional platform also introduces storage, integration, governance, and operational overhead.
AI budget forecasting
A forecast should link expected costs to a small number of understandable drivers. For assistants, these may include license counts, renewal dates, and adoption levels. For APIs, they may include request volume, token consumption, model mix, application growth, and planned releases. Commercial commitments, currency fluctuations, and new projects may also affect future costs.
Use a baseline forecast and a small set of scenarios rather than treating the future as predictable. Compare actual costs with forecasts each month, explain significant variances, and update assumptions when demand, usage patterns, or model choices change. This turns forecasting into an active management tool rather than a number used only for annual budgeting.
A practical operating rhythm is simple:
Alerts highlight cost or usage changes that require immediate action
Monthly reviews compare actual costs with forecasts, explain variances, and assign actions
Quarterly reviews reassess licenses, models, commitments, ownership, and measurement rules
New findings should feed back into the Diagnose stage, creating a continuous cycle of improvement.
Prevention is where cost ownership becomes part of everyday delivery rather than relying on the person who first notices an unexpected invoice. The goal is to make the controls established during Diagnose, Treat, and Monitor repeatable and sustainable as the AI estate grows.
AI cost governance setup
Every new AI tool, agent, API key, application, or model deployment should have a named technical owner, budget owner, business purpose, approved request path, data classification, and review date from the start. A lightweight intake process can capture this information before spending becomes difficult to explain.
Governance should distinguish between policy and technical enforcement. A policy may require client-controlled applications to use an approved gateway, but the organization still needs onboarding processes, identities, logging, bypass detection, budget thresholds, and a clear exception process. Hosted assistants and managed platforms require product-specific controls, such as approved product lists, license assignments, administrative settings, renewal reviews, and data-handling requirements.
The governance model should also define when owners review actual versus forecasted costs, unused licenses, model selections, policy exceptions, and assets that should be retired. Experiments can still move quickly, but they should have a defined budget, a named owner, appropriate data, an expiry date, and a clear decision to scale, revise, or stop.
Developer training: efficient AI tool use
Developers and product teams need practical guidance they can apply when designing and operating AI solutions. They should understand how to use the approved gateway and application identities, select the right model or model bundle, and know which usage and cost signals are recorded and monitored.
Governance can define approved tools and guardrails, but developers still need effective day-to-day practices. With GitHub Copilot, Claude Code, and similar tools, this means providing a clear task and relevant context, breaking larger activities into manageable steps, and asking the agent to verify its results. These habits help reduce unnecessary exploration, repeated tool calls, and rework.
Context should also be managed deliberately. Start a new session when the task changes. Use concise Markdown files, such as PLAN.md or HANDOFF.md, to carry important decisions forward. Store recurring guidance in repository instructions or reusable skills. Trusted MCP servers can provide controlled access to GitHub, documentation, monitoring platforms, and other systems without repeatedly copying large volumes of information into a conversation.
Training should be measured against outcomes rather than attendance. Evaluate completion time, correction effort, model usage, quality, security, and privacy. Use real team workflows and assess whether developers complete useful work faster, with fewer corrections and less unnecessary model or tool usage, while maintaining quality, security, and privacy.
Progress does not require a large scorecard. For each stage, ask one practical question and look for a visible improvement.
| Stage | Practical question | Sign of progress |
| Diagnose | Can we explain the largest AI costs and name their owners? | Less spend is left unknown or unassigned. |
| Treat | Did the chosen action reduce avoidable cost without harming quality or reliability? | The pilot shows a better cost per useful outcome. |
| Monitor | Can owners spot material changes and explain actual cost against forecast? | Alerts lead to action, and forecasts become more accurate. |
| Prevent | Do new AI tools and workloads begin with owners, rules, and review dates? | Fewer costs appear without an owner or approved path. |
Choose a small set of measures that support these questions, such as attributed spend, active license use, forecast variance, gateway coverage, and cost per successful outcome. Establish a baseline before measuring improvement and include the cost of operating the controls themselves.
Start with one area of spend rather than the whole organization:
Choose the target. Select a large invoice, an upcoming license renewal, a managed agent, or a rapidly growing API account.
Inventory the estate. List the relevant tools, licenses, managed platforms, APIs, costs, technical owners, budget owners, business purposes, and evidence sources.
Classify the spend. Separate hosted assistants and licenses, managed agent platforms, and client-controlled applications and APIs so each item receives the most appropriate control.
Mark attribution gaps. Keep missing owners, shared identities, unknown usage, and unallocated costs visible, and assign each gap to an owner.
Check the request path. For eligible client-controlled traffic, identify what uses the approved control point and what bypasses it. Do not assume visibility or coverage across closed vendor-hosted request paths.
Establish a baseline and test one change. Record costs, task success, quality, reliability, latency, security, privacy, adoption, and business value. Then test a single change, such as a license adjustment, improved attribution, revised quotas, routing, caching, retry logic, or a forecasting update.
Assign decisions. Name the owners responsible for budgets, exceptions, forecasts, monthly reviews, renewals, and retirement decisions.
Review the outcome. Compare the results against agreed acceptance criteria, include the operating cost of the controls, and decide whether to expand, adjust, or stop.
Enterprise AI cost management becomes practical when an organization separates different types of spend, assigns clear ownership, and applies controls that match the request paths it can influence. The goal is not simply to reduce model costs. It is to create a stronger relationship between spending, successful outcomes, quality, reliability, security, privacy, adoption, and business value.
Start small. Choose one area of spend, establish ownership, build a baseline, and test one improvement. Measure the outcome, refine what works, and repeat.
Score each dimension independently. The columns are the maturity levels; the rows show what each level means for that dimension.
| Dimension | 1. Ad hoc | 2. Visible | 3. Attributed | 4. Controlled | 5. Managed and optimized |
| Visibility | Invoices appear after spending happens. | Main costs are reported, but in separate views. | Agreed sources include material cost and usage. | Timely views support active decisions. | Data quality and coverage are measured and improved. |
| Attribution | Costs identify only the provider or account. | Some costs link to a team or product. | Most material spend has technical and budget owners. | Important spend links to applications, request types, or user groups. | Unknown attribution is small, monitored, and promptly resolved. |
| Control | Teams manage costs informally. | Some products have local limits or reviews. | Approved paths and basic rules exist. | Suitable controls cover eligible traffic and product decisions. | Controls are tested against outcomes and improved. |
| Forecasting | AI cost is not forecast. | Current spending is projected forward. | Forecasts include known demand and product changes. | Actual versus forecast is reviewed and owned. | Scenarios guide budgets, commitments, and architecture decisions. |
| Value measurement | Value is assumed. | Basic activity or adoption is tracked. | Success and quality are defined for priority use cases. | Cost, quality, reliability, adoption, and outcomes are reviewed together. | Evidence guides investment, optimization, and retirement. |
| Governance | Ownership and review rules are unclear. | Some teams use local rules. | Owners, policies, budgets, and review dates are defined. | Onboarding, enforcement, exceptions, and renewals operate in practice. | Reviews, exceptions, learning, and retirement form a routine cycle. |
Score each dimension from 1 to 5 and add the results: 6-10 Starting; 11-15 Seeing; 16-20 Understanding; 21-25 Controlling; 26-30 Improving.