There’s a pattern that shows up almost every time an organization scales generative AI workloads on Amazon Bedrock past a single team: at some point, someone asks a simple question — “which service spent what, on which model, last month?” — and nobody can answer it with real confidence. Not because the data is missing. CloudWatch metrics and Cost and Usage Reports are sitting right there. It’s because cost attribution usually gets bolted on after model access has already been opened up broadly, and by then you’re reconstructing intent from usage logs instead of reading it off a clean signal.
The more I’ve worked through this, the more I think the interesting insight isn’t really about cost at all. It’s that cost attribution failures are almost always a symptom of an access-control design decision made earlier, upstream, often without anyone framing it as one.
The default shape of Bedrock access
Out of the box, Bedrock makes it easy for any role with bedrock:InvokeModel to call any foundation model it has access to, with no built-in per-caller cost signal beyond “someone in this account used this model.” In an environment with 15-plus services sharing a handful of AWS accounts — a completely normal shape for any org past its first year of cloud maturity — that means the model layer starts out cost-transparent to AWS and opaque to everyone else.
The natural response is to reach for tagging conventions, or to build a reconciliation job that tries to reverse-engineer per-service usage from CloudTrail after the fact. Both are reasonable-sounding FinOps moves. But they’re solving the problem at the reporting layer, when the actual gap opened up one layer down, at the access layer, before a single invocation was made.
Where the signal actually gets created
This is what made per-service Application Inference Profiles click for me as the real lever, rather than just another provisioning task. Instead of every service sharing a generic model endpoint, each service gets its own inference profile — its own attributable identity in front of the model call. Doing this consistently across the full roster of consuming services (not just the obvious high-spend ones) means every invocation now carries an identity that maps cleanly through CloudWatch and into CUR → Athena → QuickSight, with no manual reconciliation step in between.
The insight worth sitting with here: attribution isn’t something you compute after the fact — it’s something you either design into the call path, or you don’t. Inference profiles are what make the difference between a cost model and a cost estimate.

bedrock governance pipeline
The path a Bedrock call actually takes: gated at merge time, given an attributable identity, boxed in by policy, and only then allowed to reach the model — with the cost signal falling out the other end.
Enforcement is what makes attribution durable
None of this holds up on its own, though, because a governed path is only meaningful if it’s the only path. This is where Service Control Policies do quiet but important work — denying direct foundation-model invocation outside the inference-profile path, applied at the OU level so it isn’t something an individual account owner can opt out of. Permission boundaries on developer and SSO roles close the other common gap: the “I’ll just call it directly for a quick test” habit, or the IAM drift that accumulates in any multi-account org simply because permissions are additive by default and nobody goes back to prune them.
What’s interesting about this pairing is that neither control is really “about” cost. They’re both access-boundary decisions. The cost clarity is a downstream effect of getting the boundary right.
The exceptions are where the model gets tested
Not everything fits neatly into Bedrock’s native model catalog. Bringing in models through a path like Bedrock Mantle — for something like Gemma-class models — tends to come with a genuinely different billing shape: project-based billing instead of per-invocation, short-lived token issuance instead of static credentials. It would be simpler to fold this into the existing IAM structure and move on, but that’s usually where the attribution model quietly breaks — because the new integration doesn’t actually behave like the old ones, even though it’s tempting to pretend it does.
Giving it its own IAM namespace and its own line in the attribution pipeline costs more up front, but it’s a useful test of whether the governance model generalizes or whether it only worked for the cases you happened to anticipate.
Enforcement lives where the code ships
The part that ties everything together, in practice, is that none of this survives normal engineering velocity unless it’s checked where changes actually merge. A governance pattern documented in Confluence is a description of intent. The version that holds is the one wired into the pipeline — in our case, Jenkins with Semgrep checks confirming that a new service requesting Bedrock access comes with a properly scoped inference profile, rather than falling back to direct invocation. Catching a gap in a quarterly cost review means you’ve already absorbed a quarter of ungoverned spend before you even knew there was a question to ask.
Conclusion:
Bedrock cost attribution isn’t really a reporting problem — it’s a governance problem. When service identity is designed into the model access path through inference profiles, enforced with IAM and SCPs, and validated in the deployment pipeline, cost attribution becomes a natural outcome rather than a reconciliation exercise.
The goal isn’t simply to know where AI spend went last month. It’s to build the system so you never have to guess.