← Back to case studies
A Taiwan-born AI startup scaling across borders (de-identified)Electronics & tech

AI usage went from an end-of-month surprise to a cost curve you can see — and trim

How an enterprise's in-house AI Agent, running across many models, used ATP Petrichor's project-level governance to consolidate scattered token usage into a single billing window — cost attributable, waste visible, budgets under control.

100%

of model calls converged to one governance gateway

1

consolidated bill replacing multi-vendor reconciliation

0

token spend without budget attribution

This case continues the in-house AI Agent rollout: once the same agent was running across departments and models every day, the real challenge shifted from "can we use it" to "can we afford it and control it."

The challenge: the more useful the agent, the more illegible the bill

Once the AI agent took over meeting notes, knowledge Q&A, and content generation, it was calling GPT, Claude, Gemini, and DeepSeek back and forth all day. As usage grew fast, cost lost focus:

  • Usage wasn't attributable: the monthly total was clear, but which function or department spent it was not
  • Waste hid in the dark: lightweight tasks running on flagship models, dead test keys still billing — no one could see it, so no one fixed it
  • Budget was an end-of-month result, not a variable you could steer: overruns surfaced only when the invoice arrived

This was never a "models are expensive" problem. It was the absence of a governance plane that makes usage legible and points out the waste.

The solution: converge every call onto an attributable gateway

We migrated all of this agent's model calls onto ATP Petrichor, applying an "organization → workspace → project" hierarchy so usage carries attribution from the very first request:

  1. Split projects by function: meeting automation, knowledge base, and content generation each became a separate Project — usage naturally separated and billed on its own
  2. Authorize models at the project level: a task can only reach the model tier it's allowed — lightweight work can't run on a flagship model, and waste is blocked at the source
  3. Log every request: model, token count, and the function behind it are queryable in real time, so anomalies surface the same week rather than at month-end

ATP's role in this case

Governance dimensionBeforeAfter (ATP Petrichor)
Usage attributionCompany-wide total onlyAttributed to each project and function
Model selectionLeft to developer disciplineProject-level authorization; permission is the boundary
Billing windowReconciliation across vendorsOne platform, one consolidated bill
Cost controlKnown at month-endLive quota dashboard, alerts before overrun

Outcomes

  • Waste became visible and trimmable: mismatched models and idle keys, once flagged, were converged — spend flows to the calls that actually create value
  • Cost became predictable: each function's token spend is a line on a live dashboard, not a month-end surprise
  • Governance without slowing iteration: engineers keep building; permissions and quotas take effect at the platform layer — safer to use, not hands tied

Real cost reduction was never about which model you pick — it's about giving every unit of usage a destination and making every ounce of waste visible. We proved this governance plane on our own AI Agent first, and it can turn your enterprise's AI usage into a cost curve you can actually read. Explore ATP Petrichor →

Your industry deserves a case study like this

A free one-hour assessment maps your use cases and data, and shows how intelligence can flow into your daily operations.

Book an assessment
Token Governance for an In-House AI Agent: Turning Usage Waste into Predictable Cost|Case studies — Horizon AI