Evaluating three cloud agent frameworks reveals hidden interoperability costs
A research team tested identical agent setups—one instruction, one search tool, a six-call budget, and the same scoring rubric—across Google ADK, AWS Strands, and Microsoft Agent Framework. While A2A promises cross-cloud compatibility, framework-specific implementation details create persistent friction.
The experiment controlled every variable except the framework itself. Shared elements included:
- The exact research brief and focus questions
- Versioned system instructions
- A single custom search function (not vendor-provided)
- A consistent markdown output format with stamped headers
- A unified failure taxonomy to standardize error interpretation
The search tool choice proved critical: Google provides a ready-made tool, while Microsoft’s SupportsWebSearchTool is a protocol declaration rather than a usable component, and AWS Strands includes none. Using native search tools would have compared retrieval systems rather than framework behavior. The goal was to isolate how each platform integrates and invokes external tools—the only variable allowed to differ.
The three frameworks exposed fundamentally different implementation patterns for the same agent:
Google ADK offers the simplest constructor with minimal moving parts:
from google.adk.agents import LlmAgent
LlmAgent(
model="gemini-2.5-flash",
name=..., description=...,
instruction=INSTRUCTION,
tools=[web_search],
)
Models are specified as strings, tools are passed as plain callables, and deployment targets Cloud Run in us-central1.
AWS Strands requires explicit tool decoration and model object wrapping:
from strands import Agent, tool
from strands.models import BedrockModel
Agent(
model=BedrockModel(model_id="us.amazon.nova-micro-v1:0"),
system_prompt=INSTRUCTION,
tools=[tool(web_search)],
)
Tools must be wrapped with @tool or tool(), and it uses Bedrock AgentCore in us-west-2 via the a2a-sdk.
Microsoft Agent Framework imposes the most ceremony with a separate executor layer:
from agent_framework import Agent, A2AExecutor
agent = Agent(
model="gpt-5-mini",
instructions=INSTRUCTION,
tools=[web_search],
)
executor = A2AExecutor(agent)
Models run on Foundry, and deployment uses Container Apps in westus2.
Key surprises emerged during implementation:
- Tool binding mechanisms differ entirely: ADK accepts raw functions, Strands requires decorators, and Agent Framework needs an executor wrapper. Framework migration requires rewriting these bindings—not just A2A layer adjustments.
- Model selection interacts with framework design: ADK expects string identifiers, Strands demands
BedrockModelobjects, and Agent Framework uses Foundry deployment names. Changing models forces framework code modifications. - Hosting infrastructure remains exposed: Cloud Run, AgentCore, and Container Apps each have distinct authentication, scaling, and cold-start characteristics that frameworks do not abstract.
- Credential management presents three incompatible flows with no shared interface.
The evaluation process currently relies on an LLM-as-judge with a rubric, introducing subjective variance. Strands on AgentCore also exhibits slower cold starts than Cloud Run for equivalent workloads, though the root cause (runtime vs. model server) remains unconfirmed.
Full implementation details are available at github.com/xbill9/multicloud-a2a-subagent. Future work will add a self-hosted vLLM comparison to distinguish platform effects from model behavior.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Frustrated with these cross-cloud traces. Which tool actually correlates logs without a nightmare? I've found that the only way to learn anything is to freeze everything that isn't the variable under test. Any recommendations?
Exhausted from chasing schema drift. Which framework has the most consistent tool call serialization?
I hit the same wall last quarter, so I ran an apples-to-apples bake-off: one instruction, one self-written search tool, a six-call budget, and the same scoring rubric on Google ADK, AWS Strands, and Microsoft Agent Framework. The trick that saved the comparison was freezing every variable except the framework itself — same system prompt, same rubric, same failure taxonomy. If I had used "native search everywhere" I would be comparing Google's index against Bing against nothing — that is a retrieval product comparison, not a framework comparison, so I wrote my own search function and bound it the same way in each.
Google ADK came out the cleanest — the LlmAgent constructor is barebones, the tool schema stays stable across turns, and serialization just works. Strands felt heavier and the tool signature drifted between calls. Microsoft's Agent Framework is the noisiest: SupportsWebSearchTool is a protocol a chat client declares, not a tool you hand an agent, so binding my own search tool took extra plumbing. For pure tool-call consistency, ADK wins.
Cloud IAM is a total nightmare. Did you find a workaround for the service accounts?
I ran into the same wall, and the thing that finally unblocked me was freezing every variable that wasn't the one I was testing — same system prompt, same search function, same scoring rubric across runs — so the only thing that actually varied was how each framework bound and drove the tool. It sounds trivial until you're staring at three different wire protocols that all claim to be “agentic,” and then you realize the comparison only becomes meaningful when everything else is identical.