Home Insights Phi Partners at AWS Community Day Bulgaria 2026 Phi Partners at AWS Community Day Bulgaria 2026 The AI-FinOps Loop: what does it cost to ask an AI agent a question? …More than you might think! A short question can trigger a surprising amount of work behind the scenes, as agents plan an investigation, call tools, review findings and assemble an answer. Each step can potentially add to substantially to the bill. At AWS Community Day Bulgaria 2026, Phi’s Cloud Center of Excellence lead, Ivaylo Vrabchev, explored this relationship in “The AI-FinOps Loop”: how FinOps brings AI costs under control, and how AI can help carry out FinOps itself. His demonstration began with a question familiar to anyone managing cloud expenditure: “Why did our AWS bill jump 18% yesterday?” 99 seconds to answer Ivaylo walked through a five-agent fleet built on Amazon Bedrock AgentCore. A supervisor coordinated four specialists covering anomalies, cost allocation, commitments and usage. Within just 99 seconds, the fleet had traced the spike to a single deployment. An updated data-export job had started routing Amazon S3 traffic through a NAT Gateway, adding $925 a day in data-processing charges. The recommended fix given was to use an S3 gateway endpoint, allowing the traffic to reach S3 without passing through the NAT Gateway. AWS applies no additional charge for using the gateway endpoint. The demonstration showed the appeal of conversational FinOps. A plain-English question produced a specific cause and a recommended action. Then Ivaylo applied the same scrutiny to the fleet’s own bill. Hidden costs The user’s message was just the beginning of the cost. The supervisor had to decide which specialists to involve. Each specialist needed instructions and context, then processed tool results as it investigated. Finally, the supervisor read the specialists’ responses and consolidated their findings. This is token amplification: the volume of text a model processes can grow far beyond the question the user typed. A short exchange at the interface can conceal many model calls and thousands of tokens. The architecture matters too. When specialists return lengthy answers, the supervisor must process them again. As a conversation continues, accumulated context can make subsequent questions more expensive. Understanding that cost requires visibility across the completed task. The price of an individual model call tells only part of the story. Big savings from scheduling In the session’s cost model, the fleet checked 41 AWS accounts every four hours. The underlying billing data refreshed much less frequently, so many runs repeated analysis of information that had not changed. Aligning the schedule with data freshness reduced modelled annual running costs from $104,600 to $40,600, a saving of around 61%. That improvement came from understanding the work and when it needed to happen. Further reductions came from caching stable context, matching model capability to the task and batching work that could wait. Together, these changes brought the modelled annual cost down to $19,100: approximately 82% below the baseline. The same evaluation set produced the same scores and recommendations. The result makes evaluation central to cost control. A cheaper configuration earns its place when it still delivers the quality the task requires. Agent or workflow Conversational agents are useful when someone needs to investigate an unexpected spike, ask follow-up questions or to make a decision before a meeting. Routine sweeps, tagging audits and commitment modelling have different requirements. They can often run on a schedule, with results ready when the team needs them. For these tasks, a defined workflow can coordinate the steps and use AI where analysis is needed. It can also avoid repeatedly paying a supervisor model to decide how familiar work should proceed. Choosing between an agent and a workflow is therefore a cost decision as well as a design decision. The right approach depends on how much judgement the task requires and whether anyone is waiting for the answer. Closing the Loop The principle behind the session was straightforward: establish who and what generated the spend, measure cost per completed task or outcome, then assess value per dollar spent. For a FinOps platform, identifying potential savings is useful, but tracking whether teams implement the recommendations matters too. Savings identified and savings realised can be very different. This is where the loop closes. Findings and feedback improve future investigations. Evaluation tests whether shorter prompts or cheaper models still perform adequately. The platform’s own expenditure becomes another cost to explain and govern. AI can help teams understand their cloud bill. FinOps helps establish whether that assistance is worth what it costs. If your AI spend is growing faster than your ability to explain it, Phi’s Cloud Practice can help. Get in touch. Contact | Phi Partners Related news The problem isn’t always headcount: why adding people isn’t the same as adding capability Blog 28 May 2026 Meet our Sales Consultants: Romain Houllier, Junior Partner – Head of Sales, France & Benelux Blog 29 Apr 2026 From GenAI Enablement to Enterprise Enablement: The CTO Function Reconsidered Insights 19 Mar 2026 Celebrating Experience at Phi Blog 11 Mar 2026