Tool Registry Design for Multi-Agent Systems
Governance and discovery infrastructure matter more than the directory itself.

A tool registry fails in production for a specific reason: it has plenty of entries but no authority behind them. That is an architectural gap, not a data entry problem, and closing it requires a different kind of infrastructure than most teams start with.
Most early implementations solved one problem and assumed they had solved two. Connectivity covers how an agent calls an external capability at all, while governance covers who may call what, when, with what authority, and under what audit trail. These get conflated constantly, but they need separate infrastructure answers. A directory solves connectivity. One agent calls one tool in a controlled environment, everyone nods, and the system ships.
Real governance asks for more than a list. It asks for policy enforcement on who may invoke what. It asks for lifecycle control over what is actually approved to run, an audit trail of every action taken, and health signals that stop agents from calling tools that are degraded or quietly deprecated. A directory has none of these mechanisms, because it was never built to have them.
The gap turns structural the moment a system grows past a handful of agents. Nobody designed the directory to say no to anything.
That timing made 2026 the year enterprise registry tooling arrived in force. That discovery, repeated across enough companies, is what forced the shift from directory thinking to governance thinking. The rest of this piece walks through what that shift demands at the architecture level.
What the schema must carry for governance
Governance starts with the record itself. A tool entry that captures only a name, an endpoint, and a short description cannot support any of the controls described above, because none of those fields say who is accountable, what state the tool is in, or what it is actually allowed to do. The schema has to carry provenance, protocol, lifecycle state, and authorization surface as first-class fields rather than notes bolted on later.
Publisher identity comes first. Without it, there's no one to ask when a tool misbehaves.
Protocol declaration matters just as much. A record has to state what interface the tool exposes, whether that's MCP, A2A, or a custom schema, so agents can invoke it correctly and the registry can validate that the tool actually conforms to what it claims.
Capability manifest rounds out what the tool exposes. For MCP tools specifically, that means declaring the three core primitives: tools, which are actions agents can invoke; resources, which are data agents can read; and prompts, which are templates for structured interactions. Leaving any of these undeclared means an agent discovers the gap only at call time, which is the worst possible moment to discover it.
Lifecycle state tells the system where a record sits in the approval pipeline, and this field has to be explicit and machine-readable rather than implied. Deprecation and retirement flags close the loop, signaling to agents and orchestrators that a given record should stop receiving calls.
AWS Agent Registry, which reached general availability in August 2026 as part of Amazon Bedrock AgentCore, structures its records around exactly this logic. DEPRECATED is the terminal state, and that's a design choice, not an oversight.
Once agents are built against a registry schema, changing that schema breaks every caller depending on it. The schema functions as a public API surface, and it has to be versioned with the same discipline anyone would apply to a public API, because the cost of getting it wrong is a cascade of broken agents, not a bug report.
How the approval lifecycle turns schema into enforced policy
A lifecycle field in a schema does nothing by itself. Add a state called "approved" to a record; nothing changes in practice unless something actually enforces the transition into and out of that state. A registry becomes a governance layer at the exact point where state transitions require human or automated sign-off, not before.
The pipeline typically runs through four stages, and each one functions as a gate rather than a formality. A tool starts in draft, registered but not yet discoverable by any agent, which prevents accidental invocation of something nobody has reviewed. It moves to pending approval, visible to curators but still not callable, which is the human-in-the-loop moment where someone actually looks at what's being proposed. It becomes approved once it clears review, discoverable and invocable by agents that have the right access scope. Eventually it reaches deprecated or retired, where the record stays in the catalog for audit purposes but stops being routable, so any agent still holding a reference to it gets a clear signal to go re-resolve that reference rather than failing silently.
Consider what happens without this enforcement. Coordinated rollouts and fast rollbacks only work if the registry is treated as the single source of truth for which version is approved at any given moment. Without that discipline, the registry becomes one more place that might be lying to you.
AWS Agent Registry operationalizes the review step using Amazon EventBridge, which notifies curators the moment a record enters the approval queue. Different tool categories can also route to different approvers: a tool that writes to a production database should clear a stricter gate than a tool that only reads from a public API, and the lifecycle mechanism is what makes that distinction enforceable rather than aspirational.
Governance scope doesn't have to stop at the project boundary, either. Put together, this means governance can extend to the organizational boundary rather than staying trapped inside whichever team happened to stand the registry up first.
How agents discover tools at runtime without hardcoded dependencies
Everything covered so far happens on the write path: registering tools, approving them, tracking their versions. None of it matters if agents bypass the registry at call time by hardcoding the tools they depend on. Hardcoded endpoints defeat the entire governance model, because approval state, version, and health all live in the registry, and a hardcoded reference ignores all three.
AWS Agent Registry exposes itself as a remote MCP endpoint. An AI agent can query the registry directly to find other agents or tools. That opens up a coordination pattern where one agent finds and delegates to another entirely through the catalog, with no peer address hardcoded anywhere in either agent's code.
Google Cloud takes a related but distinct approach. That discovery model deliberately resembles a package registry, the kind of interface engineers already know from pulling dependencies in any modern language ecosystem, which lowers the learning curve considerably.
Access to those discovery results is governed through IAM roles on Google Cloud, specifically the Agent Registry API Viewer role for users, while agent-to-agent communication is controlled separately through Identity-Aware Proxy egress policies. An agent that lacks the right role doesn't get an error message. It simply doesn't see the records it isn't authorized to call, which makes the access control invisible to the agent itself while still being enforced at the infrastructure layer.
This is where the design pays off. Once one agent can discover another through the registry and invoke it through a known protocol, the catalog becomes the coordination substrate for the whole system rather than a lookup tool. Research on LLM-based multi-agent systems describes this pattern directly: a Tool-Agent Registry functioning as a centralized repository for dynamic discovery, selection, and use, which improves output quality while reducing both model workload and latency. But that same capability raises an obvious question. If any approved agent can discover any other agent through the registry, what actually stops it from reaching tools it has no business touching?
Enforcing scoping and access control at the registry layer
That question has one defensible answer: scoping has to live in the registry, not in the agent. Access control written into an agent's prompt or its code is a convention, not a boundary, and conventions are things a sufficiently capable or sufficiently pressured agent can ignore, misread, or route around without even trying to.
Telling an agent, in its system prompt, to only call tools in the billing namespace is an instruction, and instructions degrade. But why take that bet on a production system handling financial data or customer records?
Enforcement at the registry layer sidesteps the bet. The registry returns only the records a calling agent's identity is authorized to see, so the agent never even receives a reference to a tool it can't call. The constraint doesn't need the agent's cooperation, because the agent was never shown the option.
This gets more complicated under a stateless protocol. The July 28, 2026 revision to MCP removed protocol-level session tracking, making MCP stateless at the protocol level. Protocol version and client capabilities now travel in a _meta parameter attached to every request, and client identity is recommended but not required. That means the registry has to validate identity on every single call. It can no longer lean on session state, because session state no longer exists.
It's worth being honest about where this breaks down. Project-and-region scoping is sound architecture for a single cloud. It is not the architecture of a registry spanning multiple clouds, and teams running agents across providers need to plan for that gap directly, ahead of time, rather than discovering it in the middle of an audit when it's far more expensive to fix.
Health signaling when a tool the registry knows about stops working
Schema, approval, and scoping all share a hidden assumption: that an approved tool is a working tool. That assumption doesn't always hold, and health signaling is what covers the distance between what the registry's records say and what's actually happening at the moment of invocation.
Picture the failure mode. An agent resolves a tool reference through the registry. The tool is marked approved. But the underlying service is degraded or down. The agent proceeds anyway, because nothing told it not to, and the failure appears as a workflow error rather than a routing error, which makes it far slower to trace back to the real cause.
There's a sharper version of this problem: tracking what an agent has already done. Without tracking that side-effect state, retries turn dangerous: an agent that can't tell whether an API call succeeded before a network timeout cut it off might just repeat the action, and if that action was destructive, repeating it is not a minor inconvenience.
Health signaling, properly built, has to cover three layers of state. Side-effect state tracks what the agent has actually done out in the world; losing track of it is the most dangerous failure of the three.
Memory stores give agents the context to make sense of health signals rather than just reacting to them. Episodic memory logs past tool calls and their outcomes. Together, these give an agent enough history to recognize a degraded tool for what it is, instead of repeating a pattern that already failed once.
The audit trail as proof that governance happened
None of the previous sections mean much without a way to prove, after the fact, that they actually happened. A registry without an audit trail is a system of claims about what it enforces rather than evidence that it enforced anything. The audit trail has to capture every action taken on the registry: registration, state transitions, invocation, deprecation, and retirement, each one tied to the identity of the actor and a timestamp.
AWS Agent Registry routes control-plane actions, things like registration, state transitions, and deprecation, through AWS CloudTrail by default. Data-plane actions, like record discovery and MCP invocation, need a separately configured trail with data event logging turned on explicitly. A complete audit log doesn't just appear on its own. It has to be set up as an infrastructure artifact, deliberately, rather than left to whatever application-level logging a developer happened to write and might just as easily have left out.
This matters most at the moment an auditor asks a specific question: which agents had access to a given tool during a specific window? The answer has to come from an immutable, infrastructure-level log. Application logs written by the agents themselves don't clear that bar, because the agents are the subject of the audit, not a neutral party reporting on it.
A recent, concrete example makes the stakes clear. AWS Agent Registry moved to the agent-registry namespace starting August 6, 2026, reaching general availability on August 31, 2026, and replacing the public-preview bedrock-agentcore path. Any team that didn't track that migration in its own registry records would have ended up with workflows silently pointing at a deprecated path, with no clear signal that anything had changed. That's precisely the class of untracked change an audit trail exists to catch, and it's a far better way to find out about a broken path than watching a production workflow fail without explanation.
A registry built on these principles in a running agent infrastructure
An audit trail, once the pieces are put together, turns every one of those mechanisms into something provable rather than assumed.
What ties these together is where the enforcement actually sits. Every one of these controls lives at the registry layer, not inside individual agents, and that's the architectural decision the whole piece has been circling. An agent that behaves well because its prompt tells it to is relying on cooperation. An agent that behaves well because the registry never showed it an unauthorized tool, never gave it a stale reference, and never let it run an unapproved version is relying on infrastructure. The second kind scales. The first kind relies on cooperation that holds only until the system grows from a demo into something a business actually depends on, and then it fails.


