Skip to content

Agent Frameworks & SDKs Arena

Google ADK vs AutoGen

Google ADK wins · 238 (15 drawn)

Agenticness — how well agents can access and operate the productAgenticness

How well agents can access and operate the product

Agent access

  1. ai-native userPoint an agent at llms.txt or agent-oriented docs

    weight 2 · round to Google ADK
    Google ADKdisputedcontradicted3/10

    Docs claim 'AI-aware developer resources' and coding-assistant integration (google-adk-docs-9), suggesting agent-oriented documentation exists, but direct probes for llms.txt and markdown-rendered docs both return 404 (google-adk-probe-1, google-adk-probe-2), and no OpenAPI/machine-readable spec is discoverable (google-adk-probe-3), contradicting the claim that an agent can straightforwardly consume these docs. Missing for 10: a working llms.txt or agent-readable doc endpoint, confirmation that the 'AI-aware resources' are actually machine-fetchable rather than just a marketing phrase.

    • [claimed-docs] ADK is designed to be written by both humans and AI. Connect your favorite coding assistant to our ADK developer Skills and AI-aware develop…
    • [probe] PROBE llms.txt: HTTP 404 at https://google.github.io/llms.txt
    • [probe] PROBE docs-md: HTTP 404 at https://google.github.io/adk-docs/get-started/.md
    • [probe] PROBE openapi: all candidate paths 404 (https://google.github.io/openapi.json, https://google.github.io/swagger.json, https://google.github.…
    AutoGennone0/10

    Probes explicitly show no llms.txt (404) and no markdown-formatted docs endpoint (404) or OpenAPI spec, so there is no evidence AutoGen exposes agent-oriented docs formats; all other evidence is standard human-readable documentation.

    • [probe] PROBE llms.txt: HTTP 404 at https://microsoft.github.io/llms.txt
    • [probe] PROBE docs-md: HTTP 404 at https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/index.html.md
    • [probe] PROBE openapi: all candidate paths 404 (https://microsoft.github.io/openapi.json, https://microsoft.github.io/swagger.json, https://microsof…
  2. ai-native userRun the product headlessly / in CI for automation

    weight 2 · round to Google ADK
    Google ADKfullclaimed8/10

    ADK provides a CLI (`adk run`, `adk web`, `adk eval`, `adk deploy docker`) that supports headless invocation and scripted evaluation, plus containerized deployment for CI/production pipelines. missing for 10: explicit CI pipeline examples (e.g. GitHub Actions), independent third-party confirmation of headless CI usage.

    • [github] adk run path/to/my_agent
    • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
    • [github] adk eval \ samples_for_testing/hello_world \ samples_for_testing/hello_world/hello_world_eval_set_001.evalset.json
    • [github] adk deploy docker --with_ui <agent-folder>
    • [claimed-docs] You can manually package your Agent into a container image and then run it in any environment that supports container images.
    • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session).
    AutoGenpartialclaimed6/10

    AutoGen ships as a pip-installable Python library (autogen-agentchat) with agents/teams fully scriptable and exportable to plain Python code, which implies it can run headlessly in CI pipelines without any GUI dependency; AutoGen Studio's 'export and run teams in python code' and Docker execution further support automatable, non-interactive runs. However, there is no explicit CI/CD documentation, GitHub Actions example, or headless-mode flag demonstrated in the evidence pack. missing for 10: explicit CI/automation docs or examples, confirmation of non-interactive/headless flags, independent report of someone running it in a CI pipeline.

    • [github] pip install -U "autogen-agentchat" "autogen-ext[openai]"
    • [claimed-docs] pip install -U "autogen-agentchat"
    • [claimed-docs] Export and run teams in python code
    • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
    • [claimed-docs] Serialize Components: Serialize and deserialize components
  3. ai-native userPlug MCP servers into this product so it can use their tools

    weight 3 · round to Google ADK
    Google ADKfullclaimed8/10

    Docs explicitly state an ADK agent can act as an MCP client and use tools provided by external MCP servers, directly matching the story. missing for 10: independent/hands-on corroboration beyond first-party docs, and more detail on multi-server configuration or auth handling.

    • [claimed-docs] An ADK agent can act as an MCP client and use tools provided by external MCP servers.
    • [claimed-docs] Exposing ADK Tools via an MCP Server: How to build an MCP server that wraps ADK tools, making them accessible to any MCP client.
    • [claimed-docs] How to build an MCP server that wraps ADK tools, making them accessible to any MCP client.
    AutoGenpartialclaimed6/10

    AutoGen's GitHub docs explicitly show creating an agent that uses the Playwright MCP server, confirming MCP server tool integration is supported, but the evidence pack lacks first-party documentation detailing a general-purpose MCP client/adapter API, configuration options, or broader ecosystem support beyond this single example. Missing for 10: dedicated MCP integration documentation, examples with multiple/varied MCP servers, independent hands-on corroboration of MCP tool usage reliability.

    • [github] Create a web browsing assistant agent that uses the Playwright MCP server.
  4. ai-native userUse an official CLI

    weight 2 · round to Google ADK
    Google ADKfullclaimed8/10

    ADK ships an official CLI (`adk run`, `adk web`, `adk eval`, `adk deploy docker`) documented in the GitHub repo with concrete command examples, plus docs reference an "Agents CLI" for scaffolding/build/test/deploy workflows tailored to AI-native/agentic use. Missing for 10: independent third-party hands-on review of the CLI's AI-native ergonomics beyond first-party docs/repo.

    • [github] adk run path/to/my_agent
    • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
    • [github] adk eval \ samples_for_testing/hello_world \ samples_for_testing/hello_world/hello_world_eval_set_001.evalset.json
    • [github] adk deploy docker --with_ui <agent-folder>
    • [claimed-docs] Go from idea to coded ADK agent in minutes. Use your favorite AI-enabled developer environment to scaffold, build, test, evaluate, and deplo…
    • [claimed-docs] Migrate existing agents and workflows to ADK with Agents CLI.
    AutoGennone0/10

    AutoGen ships a Python package (pip install), a Studio GUI, and library APIs, but no evidence of an official standalone CLI tool for AI-native workflows.

    • ai-native userDrive the product through a documented public API

      weight 3 · round to AutoGen
      Google ADKpartialprobed6/10

      ADK is a Python framework/CLI (adk run, adk web, adk eval, adk deploy) with documented programmatic APIs for building and driving agents, plus MCP client/server support, but there is no evidence of a formal public REST/OpenAPI-style API surface — probes for openapi/swagger specs and llms.txt all 404. missing for 10: a documented public HTTP/OpenAPI API spec, independent third-party confirmation of programmatic drivability beyond first-party docs.

      • [claimed-docs] Create your first Python ADK agent in minutes.
      • [github] adk run path/to/my_agent
      • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
      • [github] adk eval \ samples_for_testing/hello_world \ samples_for_testing/hello_world/hello_world_eval_set_001.evalset.json
      • [github] adk deploy docker --with_ui <agent-folder>
      • [probe] PROBE openapi: all candidate paths 404 (https://google.github.io/openapi.json, https://google.github.io/swagger.json, https://google.github.…
      • [probe] PROBE llms.txt: HTTP 404 at https://google.github.io/llms.txt
      AutoGenfullprobed7/10

      AutoGen is a Python library/framework with an extensively documented public API (AssistantAgent, GroupChat variants, tool integration, memory, serialization) that AI-native users can directly script against via pip-installed packages. missing for 10: no formal OpenAPI/REST spec (probe shows 404s), no independent third-party corroboration of API stability beyond community sentiment.

      • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
      • [claimed-docs] RoundRobinGroupChat is a simple yet effective team configuration where all agents share the same context and take turns responding in a roun…
      • [claimed-docs] The UserProxyAgent is a special built-in agent that acts as a proxy for a user to provide feedback to the team.
      • [claimed-docs] Create your own agents with custom behaviors
      • [claimed-docs] Multi-agent coordination through a shared context and centralized, customizable selector
      • [claimed-docs] SelectorGroupChat: A team that selects the next speaker using a ChatCompletion model after each message.
      • [claimed-docs] Swarm: A team that uses HandoffMessage to signal transitions between agents.
      • [claimed-docs] GraphFlow: Multi-agent workflows through a directed graph of agents.
      • [github] pip install -U "autogen-agentchat" "autogen-ext[openai]"
      • [github] You can use `AgentTool` to create a basic multi-agent orchestration setup.
      • [probe] PROBE openapi: all candidate paths 404 (https://microsoft.github.io/openapi.json, https://microsoft.github.io/swagger.json, https://microsof…
    • ai-native userIssue scoped/least-privilege API credentials for an agent

      weight 2 · round drawn
      Google ADKnone0/10

      No evidence in the pack describes issuing scoped or least-privilege API credentials/tokens for agents; the docs cover tools, MCP, workflows, deployment, and evaluation but nothing about credential scoping or permission management for agent identities.

        AutoGennone0/10

        No evidence in the pack mentions scoped/least-privilege API credential issuance or any credential-management/permissioning system for agents; AutoGen's docs focus on agent orchestration, teams, and tools, not credential scoping.

        • ai-native userBuild against official SDKs

          weight 2 · round to Google ADK
          Google ADKfullclaimed9/10

          Google ADK is itself an official Python SDK/framework with extensive first-party documentation, code examples, CLI tooling (adk run/web/eval/deploy), and a public GitHub repo, giving AI-native developers a fully documented, official SDK to build against. Minor gap — missing for 10: independent third-party corroboration beyond vendor docs/repo, and llms.txt/OpenAPI probes returned 404s suggesting some machine-readable doc surfaces are incomplete.

          • [claimed-docs] Create your first Python ADK agent in minutes.
          • [claimed-docs] Building an agent with just a model, instructions, and tools is a great place to start for most developers.
          • [claimed-docs] agent = Agent( name="researcher", model="gemini-flash-latest", instruction="You help users research topics thoroughly.", too…
          • [github] Agent Config: Build agents without code.
          • [github] adk run path/to/my_agent
          • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
          • [claimed-docs] ADK is designed to be written by both humans and AI. Connect your favorite coding assistant to our ADK developer Skills and AI-aware develop…
          AutoGenfullprobed8/10

          AutoGen ships official Python SDKs (autogen-agentchat, autogen-ext) with pip install instructions, documented core classes (AssistantAgent, UserProxyAgent, teams, tools), and GitHub-hosted source, giving AI-native developers a genuine first-party SDK to build against. Missing for 10: no evidence of official SDKs in other languages, no machine-readable API reference (openapi probes 404), and no independent third-party validation of SDK stability/versioning beyond community sentiment.

          • [github] pip install -U "autogen-agentchat" "autogen-ext[openai]"
          • [claimed-docs] pip install -U "autogen-agentchat"
          • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
          • [claimed-docs] The UserProxyAgent is a special built-in agent that acts as a proxy for a user to provide feedback to the team.
          • [claimed-docs] Create your own agents with custom behaviors
          • [claimed-docs] How to migrate from AutoGen 0.2.x to 0.4.x.
          • [probe] PROBE openapi: all candidate paths 404 (https://microsoft.github.io/openapi.json, https://microsoft.github.io/swagger.json, https://microsof…
        • ai-native userSubscribe to events via webhooks

          weight 2 · round drawn
          Google ADKnone0/10

          ADK's evidence shows only in-process callbacks/hooks for agent execution lifecycle, not an external webhook subscription mechanism; no docs mention registering webhook URLs or event push notifications. Missing for 10: any webhook registration API, outbound event delivery docs, or third-party confirmation of webhook support.

          • [claimed-docs] Callbacks: Hook into specific events during an agent's execution lifecycle to add logging, monitoring, or custom side-effects without alteri…
          AutoGennone0/10

          No evidence in the pack mentions webhooks or event subscription mechanisms; AutoGen's docs focus on agent orchestration, teams, and Studio UI with no webhook API or subscription feature described.

          Agentic features

          1. ai-native userSet up automations that run autonomously in the background

            weight 2 · round drawn
            Google ADKpartialclaimed6/10

            ADK supports deployable, auto-scaling agent runtimes (Cloud Run, GKE, Agent Runtime) and workflow orchestration with retries, state, and scheduling-like execution (fan-out/fan-in, loops), enabling agents to run unattended once deployed. However, evidence does not show explicit scheduling/triggers (e.g., cron-like autonomous kick-off) or a dedicated 'background automation' mode distinct from deployment. missing for 10: explicit trigger/schedule mechanism for autonomous background runs, independent evidence of long-running unattended operation, and confirmation of persistent background execution outside a deploy/response cycle.

            • [claimed-docs] Agent Runtime is a fully managed auto-scaling service on Google Cloud specifically designed for deploying, managing, and scaling AI agents b…
            • [claimed-docs] Cloud Run is a managed auto-scaling compute platform on Google Cloud that enables you to run your agent as a container-based application.
            • [claimed-docs] GKE is a good option if you need more control over the deployment as well as for running Open Models.
            • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
            • [github] A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan-out/fan-in, loops…
            • [claimed-docs] In ADK, any agent application that has more than one agent or executable Node is considered a workflow.

            AutoGen supports building agent teams (RoundRobinGroupChat, SelectorGroupChat, Swarm, GraphFlow) that execute autonomously without human intervention unless a UserProxyAgent is added, and AutoGen Studio lets you export teams to run as Docker containers or set up endpoints, which enables non-interactive/background execution. However, there is no explicit documentation of scheduling, triggers, persistent background daemons, or always-on automation management — the framework is oriented toward orchestrated agent conversations/workflows rather than dedicated 'set-and-forget' background automation tooling. Missing for 10: scheduling/trigger mechanisms, persistent background service management, and independent evidence of long-running unattended automations in production.

            • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
            • [claimed-docs] The UserProxyAgent is a special built-in agent that acts as a proxy for a user to provide feedback to the team.
            • [claimed-docs] Multi-agent coordination through a shared context and centralized, customizable selector
            • [claimed-docs] SelectorGroupChat: A team that selects the next speaker using a ChatCompletion model after each message.
            • [claimed-docs] Swarm: A team that uses HandoffMessage to signal transitions between agents.
            • [claimed-docs] GraphFlow: Multi-agent workflows through a directed graph of agents.
            • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
            • [community] i noticed autogen creates a new docker container each time code is executed by agents (default behaviour, can be turned off), so it's safe a…
          2. ai-native userDelegate tasks to a built-in AI assistant inside the product

            weight 3 · round to AutoGen
            Google ADKnone0/10

            ADK is a framework for building agents that developers run themselves; the docs explicitly describe connecting *external* coding assistants (e.g., 'Connect your favorite coding assistant to our ADK developer Skills') rather than shipping a built-in AI assistant that end-users delegate tasks to inside the product itself. No evidence shows ADK embedding its own persistent assistant persona for task delegation.

            • [claimed-docs] ADK is designed to be written by both humans and AI. Connect your favorite coding assistant to our ADK developer Skills and AI-aware develop…
            • [claimed-docs] Go from idea to coded ADK agent in minutes. Use your favorite AI-enabled developer environment to scaffold, build, test, evaluate, and deplo…
            AutoGenpartialclaimed6/10

            AutoGen ships a built-in AssistantAgent (LLM + tool use) and a UserProxyAgent that lets a human delegate/oversee tasks, plus AutoGen Studio gives an interactive UI to run agent teams without writing full code. However, this is fundamentally a developer framework requiring setup/config rather than a ready-made assistant embedded in an end-user product experience. Missing for 10: evidence of a zero-setup, end-user-facing 'assistant' experience (vs. code/Studio configuration), and independent hands-on confirmation of ease of delegation.

            • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
            • [claimed-docs] The UserProxyAgent is a special built-in agent that acts as a proxy for a user to provide feedback to the team.
            • [claimed-docs] Interactive environment for testing and running agent teams
            • [claimed-docs] A visual interface for creating agent teams through declarative specification (JSON) or drag-and-drop
          3. ai-native userOperate the product with natural-language commands

            weight 2 · round to AutoGen
            Google ADKpartialclaimed5/10

            ADK docs claim it is designed to be built and operated via AI coding assistants (Agent Config for no-code agent building, 'Agents CLI' for AI-enabled dev environments to scaffold/build/test/deploy) which supports some natural-language-driven operation, but the primary operating surface is a traditional CLI (adk run/web/eval/deploy) and Python code, not direct NL commands to the tool itself. Missing for 10: concrete example of natural-language command controlling ADK end-to-end, independent/hands-on confirmation that Agent Config or coding-assistant integration works as a full NL interface.

            • [claimed-docs] ADK is designed to be written by both humans and AI. Connect your favorite coding assistant to our ADK developer Skills and AI-aware develop…
            • [claimed-docs] Go from idea to coded ADK agent in minutes. Use your favorite AI-enabled developer environment to scaffold, build, test, evaluate, and deplo…
            • [github] Agent Config: Build agents without code.
            • [github] Agent Config: Build agents without code. Check out the Agent Config feature.
            • [github] Build agents without code. Check out the Agent Config feature.
            AutoGenpartialclaimed6/10

            AutoGen's agents (AssistantAgent, UserProxyAgent) communicate and are steered via natural-language messages, and UserProxyAgent explicitly lets a human give feedback in natural language during human-in-the-loop workflows; AutoGen Studio also provides an interactive environment for running/testing teams. However, this evidence centers on agent-to-agent conversation and a GUI/declarative builder rather than a dedicated natural-language 'command' interface for the whole product, and there's no first-party doc or hands-on example showing a user simply typing commands to control the system end-to-end. Missing for 10: explicit documentation of a natural-language command/control layer for the overall product (vs. per-agent chat), and independent hands-on confirmation that NL commands reliably drive product behavior.

            • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
            • [claimed-docs] The UserProxyAgent is a special built-in agent that acts as a proxy for a user to provide feedback to the team.
            • [claimed-docs] Interactive environment for testing and running agent teams
            • [claimed-docs] A visual interface for creating agent teams through declarative specification (JSON) or drag-and-drop

          Api quality

          1. ai-native userExplore an interactive API reference with runnable examples

            weight 2 · round drawn
            Google ADKnone0/10

            The evidence pack shows standard docs, code snippets, and CLI examples, but no interactive/runnable API reference (e.g., a Swagger/OpenAPI explorer or live code sandbox); probes for openapi.json and similar endpoints explicitly returned 404s.

            • [probe] PROBE openapi: all candidate paths 404 (https://google.github.io/openapi.json, https://google.github.io/swagger.json, https://google.github.…
            • [probe] PROBE docs-md: HTTP 404 at https://google.github.io/adk-docs/get-started/.md
            • [claimed-docs] agent = Agent( name="researcher", model="gemini-flash-latest", instruction="You help users research topics thoroughly.", too…
            AutoGennone0/10

            AutoGen ships conventional static docs and tutorials (autogen-docs-1..22) but no evidence of an interactive, runnable API reference (e.g., embedded live code execution, Jupyter-style sandbox tied to reference pages); probes confirm no llms.txt, no markdown-served docs, and no OpenAPI spec (autogen-probe-1,2,3), indicating the docs are not AI-native/interactive in the way the story describes.

            • [probe] PROBE llms.txt: HTTP 404 at https://microsoft.github.io/llms.txt
            • [probe] PROBE docs-md: HTTP 404 at https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/index.html.md
            • [probe] PROBE openapi: all candidate paths 404 (https://microsoft.github.io/openapi.json, https://microsoft.github.io/swagger.json, https://microsof…
            • [claimed-docs] How to migrate from AutoGen 0.2.x to 0.4.x.
            • [claimed-docs] Migration Guide: How to migrate from AutoGen 0.2.x to 0.4.x.
          2. ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)

            weight 2 · round drawn
            Google ADKnone0/10

            No evidence of a downloadable OpenAPI/machine-readable spec for ADK; explicit probes for openapi.json/swagger.json and llms.txt all return 404, indicating no such spec is published.

            • [probe] PROBE llms.txt: HTTP 404 at https://google.github.io/llms.txt
            • [probe] PROBE docs-md: HTTP 404 at https://google.github.io/adk-docs/get-started/.md
            • [probe] PROBE openapi: all candidate paths 404 (https://google.github.io/openapi.json, https://google.github.io/swagger.json, https://google.github.…
            AutoGennone0/10

            Probes for OpenAPI/Swagger endpoints and llms.txt all returned 404, and no documentation mentions a machine-readable API spec for AutoGen; AutoGen is a Python framework/library, not a hosted API service, but a downloadable spec is still a fair question for its SDK surface and none is provided.

            • [probe] PROBE llms.txt: HTTP 404 at https://microsoft.github.io/llms.txt
            • [probe] PROBE docs-md: HTTP 404 at https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/index.html.md
            • [probe] PROBE openapi: all candidate paths 404 (https://microsoft.github.io/openapi.json, https://microsoft.github.io/swagger.json, https://microsof…
          3. ai-native userTest against a sandbox environment without touching production data

            weight 1 · round to Google ADK
            Google ADKpartialclaimed5/10

            ADK supports local dev/test workflows (adk run, adk web, adk eval, local evaluation with test files and eval sets) that inherently run against a local/dev environment rather than production, and offline/disconnected deployment is mentioned. However, there's no explicit documentation of a dedicated 'sandbox' environment or data isolation guarantee distinct from production. missing for 10: explicit sandbox/staging environment docs, explicit statement that test runs are isolated from production data/state, independent confirmation of this isolation.

            • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session).
            • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session). It's most eff…
            • [claimed-docs] Expected Intermediate Tool Use Trajectory: The tool calls we expect the agent to make in order to respond correctly to the user query.
            • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
            • [github] adk eval \ samples_for_testing/hello_world \ samples_for_testing/hello_world/hello_world_eval_set_001.evalset.json
            • [claimed-docs] This is a good option if you prefer to run offline or disconnected, or otherwise in a system that has no connection to Google Cloud.

            AutoGen isolates code execution in per-run Docker containers by default (autogen-comm-2), which provides a sandbox for agent-executed code and reduces risk to the host system, and AutoGen Studio offers an 'interactive environment for testing and running agent teams' (autogen-docs-7) and can 'run teams in a docker container' (autogen-docs-17). However, there is no explicit documentation of a dedicated sandbox vs production-data separation, test data isolation, or staging environment concept for AI-native testing workflows. Missing for 10: explicit sandbox/production data separation, dedicated test-environment tooling, first-party guidance on avoiding production data during agent testing.

            • [community] i noticed autogen creates a new docker container each time code is executed by agents (default behaviour, can be turned off), so it's safe a…
            • [claimed-docs] Interactive environment for testing and running agent teams
            • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
          4. ai-native userRely on versioned APIs with a documented deprecation policy

            weight 2 · round to AutoGen
            Google ADKnone0/10

            No evidence in the pack mentions API versioning schemes or a documented deprecation policy for ADK; probes for OpenAPI specs and related docs all returned 404s. missing for 10: versioning scheme documentation, deprecation policy, changelog/migration guides, semantic versioning commitments.

            • [probe] PROBE llms.txt: HTTP 404 at https://google.github.io/llms.txt
            • [probe] PROBE docs-md: HTTP 404 at https://google.github.io/adk-docs/get-started/.md
            • [probe] PROBE openapi: all candidate paths 404 (https://google.github.io/openapi.json, https://google.github.io/swagger.json, https://google.github.…
            AutoGenpartialclaimed3/10

            AutoGen provides a migration guide for moving from 0.2.x to 0.4.x, showing awareness of versioning and breaking changes, but there is no documented deprecation policy, semver commitments, or API stability guarantees in the evidence. missing for 10: explicit deprecation policy statement, semantic versioning guarantees, timelines for deprecating old APIs.

            • [claimed-docs] How to migrate from AutoGen 0.2.x to 0.4.x.
            • [claimed-docs] Migration Guide: How to migrate from AutoGen 0.2.x to 0.4.x.

          Agents tools — stories about agents tools in this arenaAgents tools

          Stories about agents tools in this arena

          Agent authoring

          1. developerDefine an agent with typed custom tools in a few lines of code

            weight 3 · round drawn
            Google ADKpartialclaimed6/10

            The docs show a concrete few-line example of defining an Agent with a model, instructions, and a tools list (google-adk-docs-22), and other docs confirm tools are a core, optional component of agent definition (google-adk-docs-2, google-adk-docs-13). However, the evidence never shows a custom Python tool function with type hints/typed parameters being defined and passed in — only a prebuilt tool (google_search) is used in the example. Missing for 10: an explicit example of writing a custom typed tool function, and documentation of automatic schema/type inference from function signatures.

            • [claimed-docs] agent = Agent( name="researcher", model="gemini-flash-latest", instruction="You help users research topics thoroughly.", too…
            • [claimed-docs] Building an agent with just a model, instructions, and tools is a great place to start for most developers.
            • [claimed-docs] The basic components of an Agent are an artificial intelligence (AI) model, task instructions, and optionally, a set of tools to be used by …
            AutoGenpartialclaimed6/10

            AutoGen's AssistantAgent supports tool use and custom agent creation, and AgentTool is documented for building agents with tools, indicating typed tool integration is possible in relatively few lines. However, the evidence pack lacks a concrete code example showing typed tool schemas (e.g., function signatures/Pydantic typing) wired directly into an agent definition. missing for 10: a first-party minimal code snippet demonstrating typed tool definition and attachment to an agent, independent hands-on confirmation of the 'few lines of code' ergonomics.

            • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
            • [claimed-docs] Create your own agents with custom behaviors
            • [github] You can use `AgentTool` to create a basic multi-agent orchestration setup.

          Ai buildability

          1. ai-native userHave a coding agent scaffold a new agent project from an official CLI or template in one command

            weight 2 · round to Google ADK
            Google ADKfullclaimed8/10

            ADK docs explicitly advertise an official 'Agents CLI' to scaffold, build, test, evaluate, and deploy agents in minutes, and the GitHub README shows concrete one-line commands (adk run, adk web, adk deploy) plus a no-code 'Agent Config' template feature for scaffolding agents. This directly matches the ai-native scaffolding story via an official CLI/template workflow. Missing for 10: independent/hands-on confirmation of the one-command scaffold experience beyond first-party docs.

            • [claimed-docs] Go from idea to coded ADK agent in minutes. Use your favorite AI-enabled developer environment to scaffold, build, test, evaluate, and deplo…
            • [claimed-docs] Migrate existing agents and workflows to ADK with Agents CLI.
            • [github] adk run path/to/my_agent
            • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
            • [github] Agent Config: Build agents without code. Check out the Agent Config feature.
            • [github] Build agents without code. Check out the Agent Config feature.
            AutoGennone0/10

            AutoGen provides pip install and a Python library/AutoGen Studio for building agents, but there is no evidence of a scaffolding CLI or project template command for one-command project generation. missing for 10: an official CLI scaffold/init command, project template generation, evidence of one-command bootstrap workflow.

            • ai-native userRun the framework's example agents headlessly from a terminal so an agent can verify what it just built

              weight 2 · round to Google ADK
              Google ADKfullclaimed7/10

              ADK provides a documented CLI (`adk run path/to/my_agent`) to run agents headlessly from a terminal, plus `adk eval` for automated verification of agent behavior against eval sets, matching the 'verify what it just built' use case for an ai-native/agentic workflow. Missing for 10: explicit confirmation that shipped 'example agents' (vs. user-authored ones) work with this flow, and independent/hands-on corroboration beyond the official repo docs.

              • [github] adk run path/to/my_agent
              • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
              • [github] adk eval \ samples_for_testing/hello_world \ samples_for_testing/hello_world/hello_world_eval_set_001.evalset.json
              • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session).

              AutoGen agents are plain Python objects installed via pip and run as scripts, and community evidence confirms code execution happens in isolated Docker containers by default, which supports a terminal/headless workflow. However, there is no documented CLI or explicit 'run example agents headlessly to verify a build' feature — examples are shown as notebooks, and AutoGen Studio (the interactive runner) is UI-first with only python-code export, not a described headless verification loop. missing for 10: a documented CLI/headless example-runner, explicit self-verification workflow, independent confirmation of headless terminal use for build-verification purposes.

              • [github] pip install -U "autogen-agentchat" "autogen-ext[openai]"
              • [community] i noticed autogen creates a new docker container each time code is executed by agents (default behaviour, can be turned off), so it's safe a…
              • [community] FWIW the 'group research' and 'chess' examples from the notebooks folder in their repo have been the best for explaining the utility of this…
              • [claimed-docs] Export and run teams in python code
              • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
            • ai-native userRely on strict typing and schema validation so a coding agent catches its own mistakes at build time

              weight 2 · round drawn
              Google ADKnone0/10

              The evidence pack covers ADK's agent orchestration, deployment, and evaluation features, but contains no mention of strict typing, schema validation, or build-time error detection for tool/agent definitions — the evaluation features described (docs-20, docs-21, docs-25) are runtime test-set based, not compile/build-time type checks.

                AutoGennone0/10

                The evidence pack covers AutoGen's multi-agent orchestration, teams, memory, and studio UI, but contains no mention of strict typing, schema validation, or build-time error catching for a coding agent's own output. This is a fair question for an agent framework (e.g., via Pydantic-typed messages or structured outputs) but no such capability is documented here.

                Automation depth — how much of the product can run unattendedAutomation depth

                How much of the product can run unattended

                1. ai-native userPerform bulk operations across many items at once

                  weight 2 · round to Google ADK
                  Google ADKpartialclaimed3/10

                  ADK's Workflow Runtime offers fan-out/fan-in and loop constructs that could be used by developers to build bulk-item processing pipelines, but there is no documented built-in 'bulk operations' feature or example for end users acting across many items at once. Missing for 10: explicit bulk-operation tooling/UI, documented examples of processing many items in one call, and evidence of end-user (not just developer-framework) bulk workflows.

                  • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
                  • [claimed-docs] you can use the ADK development framework to expand them into workflows, which allow you to combine and orchestrate multiple agents and code…
                  AutoGennone0/10

                  The evidence pack covers AutoGen's multi-agent orchestration (teams, group chats, graph flows) but contains no mention of bulk/batch processing of many items at once (e.g., batch task queues, mass data operations). This is a fair axis for an automation framework, but no documentation or community evidence shows such a capability.

                  • ai-native userDefine rules that trigger actions automatically on events

                    weight 3 · round to Google ADK
                    Google ADKfullclaimed7/10

                    ADK explicitly supports event-driven automation via Callbacks ("Hook into specific events during an agent's execution lifecycle... without altering core agent logic") and a Workflow Runtime graph engine with routing, retry, fan-out/fan-in and dynamic nodes for triggering actions on execution events, matching the story of defining rules that fire on events. missing for 10: independent/hands-on evidence of callback-triggered rules in production use, and more detail on condition-based rule syntax beyond docs summaries.

                    • [claimed-docs] Callbacks: Hook into specific events during an agent's execution lifecycle to add logging, monitoring, or custom side-effects without alteri…
                    • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
                    • [github] A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan-out/fan-in, loops…
                    • [claimed-docs] In ADK, any agent application that has more than one agent or executable Node is considered a workflow.
                    AutoGenpartialclaimed4/10

                    AutoGen provides some conditional/event-driven orchestration primitives — Swarm's HandoffMessage triggers transitions between agents, and GraphFlow defines directed-graph workflows that route execution based on conditions — which can be used to build reactive, event-triggered behavior. However, there is no dedicated declarative 'rule' definition system (e.g., event-condition-action rules, triggers/webhooks) documented; the automation is implemented via developer code (agents, handoffs, selectors) rather than a rules engine an AI-native user configures directly. Missing for 10: explicit rule/trigger definition API or UI, event-listener/webhook mechanism, and independent evidence of this being used for automated event-driven actions outside code-defined agent handoffs.

                    • [claimed-docs] Swarm: A team that uses HandoffMessage to signal transitions between agents.
                    • [claimed-docs] Multi-agent workflows through a directed graph of agents.
                    • [claimed-docs] GraphFlow: Multi-agent workflows through a directed graph of agents.
                    • [claimed-docs] SelectorGroupChat: A team that selects the next speaker using a ChatCompletion model after each message.
                  • ai-native userSchedule recurring jobs or workflows

                    weight 2 · round drawn
                    Google ADKnone0/10

                    The evidence covers agent/workflow orchestration, deployment targets (Cloud Run, GKE, Agent Runtime), and evaluation, but nothing describes scheduling, cron-like triggers, or recurring execution of jobs/workflows. Absence of evidence for this applicable automation-depth capability yields 'none'.

                      AutoGennone0/10

                      No evidence of a scheduler, cron-like trigger, or recurring workflow execution feature; AutoGen's docs focus on agent teams, orchestration patterns (RoundRobin, Selector, Swarm, GraphFlow) and AutoGen Studio, none of which mention scheduling or recurrence.

                      • ai-native userVersion, review, and roll back my automations

                        weight 1 · round drawn
                        Google ADKnone0/10

                        ADK is a framework for building agents (code, workflows, tools, deployment) but the evidence pack shows no version control, review, or rollback mechanism for automations themselves — no changelog/versioning UI, no approval/review workflow for agent definitions, no rollback feature. Agent code could theoretically be tracked via external git, but ADK itself provides no such capability in the evidence. Missing for 10: any versioning system, review/approval workflow, or rollback capability for automations.

                          AutoGennone0/10

                          AutoGen docs mention serializing/deserializing components and exporting team configs as JSON/Python, but there is no evidence of built-in versioning, review workflows, or rollback of automations/agent teams. Missing for 10: version history tracking, diff/review UI, rollback mechanism, audit trail for changes.

                          • [claimed-docs] Serialize Components: Serialize and deserialize components
                          • [claimed-docs] Export and run teams in python code
                          • [claimed-docs] A visual interface for creating agent teams through declarative specification (JSON) or drag-and-drop

                        Deployment portability — stories about deployment portability in this arenaDeployment portability

                        Stories about deployment portability in this arena

                        Deployment

                        1. engineering-leadDeploy an agent to a managed runtime and call it as an API endpoint

                          weight 2 · round to Google ADK
                          Google ADKfullclaimed8/10

                          ADK docs explicitly describe deploying agents to a fully managed, auto-scaling Agent Engine/Agent Runtime on Google Cloud, plus alternative managed options like Cloud Run and GKE, with the stated purpose being to make the agent 'accessed, queried, and used in production' as an API endpoint. Missing for 10: no explicit hands-on/independent confirmation of the API contract (e.g., request/response schema) or third-party verification of endpoint behavior beyond first-party docs.

                          • [claimed-docs] Agent Runtime is a fully managed auto-scaling service on Google Cloud specifically designed for deploying, managing, and scaling AI agents b…
                          • [claimed-docs] Cloud Run is a managed auto-scaling compute platform on Google Cloud that enables you to run your agent as a container-based application.
                          • [claimed-docs] GKE is a good option if you need more control over the deployment as well as for running Open Models.
                          • [claimed-docs] Once you've built and tested your agent using ADK, the next step is to deploy it so it can be accessed, queried, and used in production
                          • [claimed-docs] You can manually package your Agent into a container image and then run it in any environment that supports container images.
                          AutoGenpartialclaimed4/10

                          AutoGen Studio docs mention the ability to 'setup and test endpoints based on a team configuration' and 'run teams in a docker container,' which shows some API-endpoint and containerization support, but this is self-managed docker, not a Microsoft-managed runtime/PaaS. Missing for 10: evidence of an actual managed/hosted runtime service, deployment guides, or cloud endpoint provisioning beyond local docker export.

                          • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
                          • [claimed-docs] A visual interface for creating agent teams through declarative specification (JSON) or drag-and-drop
                        2. engineering-leadRun my agents entirely on my own infrastructure with no dependence on the vendor's platform

                          weight 2 · round drawn
                          Google ADKfullclaimed8/10

                          ADK is an open-source framework (google/adk-python) that supports running agents locally via `adk run`/`adk web`, packaging into containers with `adk deploy docker`, deploying to any container-supporting environment, and explicitly documents an offline/disconnected mode with no Google Cloud connection required, alongside optional managed services like Agent Runtime/Cloud Run/GKE. missing for 10: independent/hands-on confirmation of fully vendor-free operation (e.g., third-party report of running ADK completely offline with non-Google models) and clarity on whether any telemetry/model calls still phone home by default.

                          • [claimed-docs] This is a good option if you prefer to run offline or disconnected, or otherwise in a system that has no connection to Google Cloud.
                          • [claimed-docs] You can manually package your Agent into a container image and then run it in any environment that supports container images.
                          • [github] adk deploy docker --with_ui <agent-folder>
                          • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
                          • [claimed-docs] ADK can work with almost any generative AI model. The framework provides easy access to Gemini as well as other leading models, and we provi…
                          • [claimed-docs] GKE is a good option if you need more control over the deployment as well as for running Open Models.
                          AutoGenfullcommunity8/10

                          AutoGen is an open-source Python framework (pip installable) that runs locally, supports Docker-based code execution, and has no required vendor SaaS backend for agent execution; AutoGen Studio can also run teams in a docker container fully self-hosted. missing for 10: explicit vendor statement on air-gapped/offline deployment and independent case studies of fully on-prem production use.

                          • [github] pip install -U "autogen-agentchat" "autogen-ext[openai]"
                          • [claimed-docs] pip install -U "autogen-agentchat"
                          • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
                          • [community] i noticed autogen creates a new docker container each time code is executed by agents (default behaviour, can be turned off), so it's safe a…

                        Portability

                        1. developerSwap the underlying LLM provider or model without rewriting my agent

                          weight 3 · round to Google ADK
                          Google ADKfullclaimed8/10

                          Docs explicitly state ADK works with 'almost any generative AI model,' providing adapters for Gemini and many other model providers including locally running models, and the agent definition just takes a `model` string parameter (e.g., 'gemini-flash-latest'), implying swapping providers is a config change rather than a rewrite. Missing for 10: independent/hands-on confirmation that switching providers requires no code changes to agent logic, and no explicit list of supported non-Gemini providers with concrete migration examples.

                          • [claimed-docs] ADK can work with almost any generative AI model. The framework provides easy access to Gemini as well as other leading models, and we provi…
                          • [claimed-docs] agent = Agent( name="researcher", model="gemini-flash-latest", instruction="You help users research topics thoroughly.", too…
                          • [claimed-docs] The basic components of an Agent are an artificial intelligence (AI) model, task instructions, and optionally, a set of tools to be used by …

                          AutoGen's model client abstraction (autogen-ext[openai] and similar model-client packages) and AssistantAgent design imply pluggable LLM providers, and community mentions confirm pointing agents at different LLMs (e.g., GPT-4 and others) without rearchitecting agent logic. However, the evidence pack lacks explicit documentation of a unified model-client interface listing multiple supported providers or a concrete swap example. Missing for 10: explicit docs enumerating supported model providers/backends, a documented config-only swap example, and independent confirmation of zero-code-change provider switching.

                          • [github] pip install -U "autogen-agentchat" "autogen-ext[openai]"
                          • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
                          • [community] However from his examples (and his own admission) it seems that AutoGen isn't benefitting from full GPT4-level performance even tho he's poi…

                        Evals observability — stories about evals observability in this arenaEvals observability

                        Stories about evals observability in this arena

                        Evals

                        1. engineering-leadScore agent quality with built-in evals and run them as part of CI

                          weight 2 · round to Google ADK
                          Google ADKfullclaimed8/10

                          ADK ships a first-party evaluation framework with groundtruth and rubric-based metrics, expected tool-use trajectories, evalset.json test files, and a documented CLI command (`adk eval <agent> <evalset>`) that can be scripted/invoked headlessly, which is exactly the shape needed for CI integration. Missing for 10: explicit CI/CD pipeline documentation (e.g., a GitHub Actions example) and independent/third-party corroboration of running adk eval in CI.

                          • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session).
                          • [claimed-docs] ADK provides both groundtruth based and rubric based tool use evaluation metrics.
                          • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session). It's most eff…
                          • [claimed-docs] Expected Intermediate Tool Use Trajectory: The tool calls we expect the agent to make in order to respond correctly to the user query.
                          • [github] adk eval \ samples_for_testing/hello_world \ samples_for_testing/hello_world/hello_world_eval_set_001.evalset.json
                          AutoGennone0/10

                          No evidence of built-in evaluation/scoring tools or CI integration for agent quality; the pack only covers agent/team constructs, logging, and AutoGen Studio, none of which address evals or CI test scoring.

                          Testing

                          1. developerUnit-test agents with mocked models and tools

                            weight 2 · round to Google ADK
                            Google ADKpartialclaimed5/10

                            ADK docs describe a test-file based evaluation approach explicitly described as 'a form of unit testing' for single agent-model interactions, with expected tool-use trajectories and groundtruth/rubric metrics plus an `adk eval` CLI — but none of this evidence explicitly describes mocking models or tools (e.g., swapping in fake LLM responses or stub tool implementations) for isolated unit tests. Missing for 10: explicit mocked-model/mocked-tool test fixtures or APIs, independent/hands-on confirmation of mocking support, and unit-test framework integration examples (e.g., pytest with mock objects).

                            • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session).
                            • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session). It's most eff…
                            • [claimed-docs] Expected Intermediate Tool Use Trajectory: The tool calls we expect the agent to make in order to respond correctly to the user query.
                            • [claimed-docs] ADK provides both groundtruth based and rubric based tool use evaluation metrics.
                            • [github] adk eval \ samples_for_testing/hello_world \ samples_for_testing/hello_world/hello_world_eval_set_001.evalset.json
                            AutoGennone0/10

                            No evidence pack item describes mocking models/tools or unit-testing utilities for AutoGen agents; docs cover building agents, teams, memory, logging, and serialization but nothing about test harnesses or mock LLM/tool clients.

                            Tracing

                            1. developerTrace every LLM call and tool invocation of an agent run in an observability UI

                              weight 3 · round to Google ADK
                              Google ADKpartialclaimed5/10

                              ADK ships a built-in development Web UI explicitly for testing, evaluating, and debugging agents, and provides callbacks to hook into execution lifecycle events for logging/monitoring, which together imply some run-level visibility into tool and model calls. However, the evidence never explicitly describes a trace view showing each LLM call and tool invocation of a run, nor mentions integration with tracing standards (e.g., OpenTelemetry) or a dedicated observability dashboard beyond the dev/eval UI. Missing for 10: explicit documentation of per-call tracing UI, tool-invocation-level trace inspection, and any third-party/hands-on confirmation of this granularity.

                              • [github] A built-in development UI to help you test, evaluate, debug, and showcase your agent(s).
                              • [github] Web UI (supports multi-agent directories or pointing directly to a single agent folder)
                              • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
                              • [claimed-docs] Callbacks: Hook into specific events during an agent's execution lifecycle to add logging, monitoring, or custom side-effects without alteri…
                              AutoGenpartialclaimed4/10

                              Docs confirm AutoGen has a logging/tracing feature for 'traces and internal messages' and a separate AutoGen Studio UI for building/testing/running teams, but no evidence explicitly ties these together into a UI that visualizes per-call LLM/tool traces for a given run. Missing for 10: explicit documentation or screenshots of an observability/tracing UI showing individual LLM calls and tool invocations, and any independent corroboration of this capability.

                              • [claimed-docs] Logging: Log traces and internal messages
                              • [claimed-docs] Interactive environment for testing and running agent teams
                              • [claimed-docs] A visual interface for creating agent teams through declarative specification (JSON) or drag-and-drop
                              • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container

                            Guardrails safety — stories about guardrails safety in this arenaGuardrails safety

                            Stories about guardrails safety in this arena

                            Guardrails

                            1. developerAttach input/output guardrails that validate, transform, or block unsafe content

                              weight 3 · round to Google ADK
                              Google ADKpartialclaimed5/10

                              ADK exposes general extensibility hooks—Callbacks to intercept execution events for custom logic/side-effects, Plugins for pre-packaged behaviors, and a Tool Confirmation (HITL) flow that can guard tool execution—which developers could use to build input/output guardrails, but there is no dedicated 'guardrails' feature, built-in content-safety/validation API, or example showing blocking/transforming unsafe content end-to-end. Missing for 10: explicit guardrail/validation API or moderation integration, documented examples of blocking/transforming unsafe input or output, and any third-party/community confirmation of this pattern in practice.

                              • [claimed-docs] Callbacks: Hook into specific events during an agent's execution lifecycle to add logging, monitoring, or custom side-effects without alteri…
                              • [claimed-docs] Plugins: Integrate complex, pre-packaged behaviors and third-party services directly into your agent's workflow.
                              • [github] Tool Confirmation: A tool confirmation flow (HITL) that can guard tool execution with explicit confirmation and custom input.
                              • [github] A tool confirmation flow (HITL) that can guard tool execution with explicit confirmation and custom input.
                              AutoGennone0/10

                              No evidence of built-in input/output guardrails, content validation/transformation, or blocking mechanisms; docs cover agents, teams, memory, logging, serialization but nothing on safety/guardrail features. Community discussion touches on code-execution sandboxing (docker) but not content guardrails. missing for 10: any documentation of guardrail/validation hooks, content moderation APIs, or examples of blocking/transforming unsafe outputs.

                              • [claimed-docs] Create your own agents with custom behaviors
                              • [claimed-docs] Add memory capabilities to your agents
                              • [claimed-docs] Logging: Log traces and internal messages
                              • [community] i noticed autogen creates a new docker container each time code is executed by agents (default behaviour, can be turned off), so it's safe a…
                            2. engineering-leadRestrict what an agent may do with fine-grained tool permissions and sandboxed execution

                              weight 2 · round to AutoGen
                              Google ADKpartialclaimed4/10

                              ADK provides a Tool Confirmation (HITL) flow that can gate tool execution with explicit confirmation/custom input, plus callbacks/plugins hooks to intercept agent actions, giving some control over agent behavior. However there is no evidence of fine-grained per-tool permission policies or an actual sandboxed execution environment for code/tool runs. Missing for 10: explicit sandboxing of tool/code execution, a permissions/ACL system scoping tool access, and independent verification of these guardrails in practice.

                              • [github] Tool Confirmation: A tool confirmation flow (HITL) that can guard tool execution with explicit confirmation and custom input.
                              • [github] A tool confirmation flow (HITL) that can guard tool execution with explicit confirmation and custom input.
                              • [claimed-docs] Callbacks: Hook into specific events during an agent's execution lifecycle to add logging, monitoring, or custom side-effects without alteri…
                              • [claimed-docs] Plugins: Integrate complex, pre-packaged behaviors and third-party services directly into your agent's workflow.

                              Community evidence confirms code execution runs in ephemeral Docker containers by default (sandboxed execution), and AutoGen Studio can run teams in a docker container, giving real sandboxing support corroborated by hands-on users. However, there is no documented fine-grained, per-tool permission system (e.g., allow/deny lists, scoped capabilities) for agents — only that agents 'have the ability to use tools' and can create custom agents. Missing for 10: explicit fine-grained tool-permission/allow-list mechanism, first-party docs on restricting specific tool access per agent, and independent verification of sandbox robustness beyond one HN thread noting it 'can be turned off'.

                              • [community] i noticed autogen creates a new docker container each time code is executed by agents (default behaviour, can be turned off), so it's safe a…
                              • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
                              • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
                              • [claimed-docs] Create your own agents with custom behaviors

                            Human in the loop — stories about human in the loop in this arenaHuman in the loop

                            Stories about human in the loop in this arena

                            Approval flows

                            1. developerPause an agent mid-run for human input or approval and resume with the human's decision

                              weight 3 · round to Google ADK
                              Google ADKfullclaimed8/10

                              ADK explicitly documents a Tool Confirmation flow described as HITL that can 'guard tool execution with explicit confirmation and custom input,' plus a Workflow Runtime and Task API both explicitly listing human-in-the-loop support with state management for pausing and resuming execution. This directly matches pausing mid-run for human approval and resuming with the decision, though missing for 10: a concrete end-to-end code example showing pause/resume state persistence and independent third-party corroboration beyond vendor GitHub README claims.

                              • [github] Tool Confirmation: A tool confirmation flow (HITL) that can guard tool execution with explicit confirmation and custom input.
                              • [github] A tool confirmation flow (HITL) that can guard tool execution with explicit confirmation and custom input.
                              • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
                              • [github] A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan-out/fan-in, loops…
                              • [github] Task API: Structured agent-to-agent delegation with multi-turn task mode, single-turn controlled output, mixed delegation patterns, human-in…
                              • [github] Structured agent-to-agent delegation with multi-turn task mode, single-turn controlled output, mixed delegation patterns, human-in-the-loop,…
                              AutoGenpartialclaimed6/10

                              AutoGen has a documented UserProxyAgent that provides human-in-the-loop feedback to the team, supporting pausing for human input, but the evidence pack doesn't detail explicit pause/resume with persisted state or approval gating mid-run (e.g., checkpoint/resume semantics). missing for 10: explicit resume-from-interrupt mechanics, evidence of approval gating on specific actions, independent hands-on confirmation of pause/resume behavior.

                              • [claimed-docs] The UserProxyAgent is a special built-in agent that acts as a proxy for a user to provide feedback to the team.
                            2. engineering-leadRequire human approval before specific sensitive tool calls execute

                              weight 2 · round to Google ADK
                              Google ADKfullclaimed8/10

                              ADK explicitly documents a 'Tool Confirmation' HITL flow that guards tool execution with explicit confirmation and custom input, plus broader human-in-the-loop support in its workflow/task orchestration engines, directly matching the story of requiring approval before sensitive tool calls execute. Missing for 10: no independent/hands-on validation or detailed walkthrough of configuring per-tool approval policies beyond the feature summary.

                              • [github] Tool Confirmation: A tool confirmation flow (HITL) that can guard tool execution with explicit confirmation and custom input.
                              • [github] A tool confirmation flow (HITL) that can guard tool execution with explicit confirmation and custom input.
                              • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
                              • [github] A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan-out/fan-in, loops…
                              • [github] Task API: Structured agent-to-agent delegation with multi-turn task mode, single-turn controlled output, mixed delegation patterns, human-in…
                              • [github] Structured agent-to-agent delegation with multi-turn task mode, single-turn controlled output, mixed delegation patterns, human-in-the-loop,…
                              AutoGenpartialclaimed5/10

                              AutoGen's UserProxyAgent is documented as enabling human-in-the-loop feedback within a team, which could be used to pause and approve steps, but the evidence never shows a mechanism to gate specific sensitive tool calls (e.g., per-tool approval hooks) rather than general conversational feedback. missing for 10: documentation of tool-call-level approval/interrupt hooks, examples of selectively requiring approval only for sensitive tools, and independent confirmation this works as described.

                              • [claimed-docs] The UserProxyAgent is a special built-in agent that acts as a proxy for a user to provide feedback to the team.
                              • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.

                            Memory context — stories about memory context in this arenaMemory context

                            Stories about memory context in this arena

                            Memory

                            1. developerTrim, summarize, or filter conversation history to keep an agent inside its context window

                              weight 2 · round to Google ADK
                              Google ADKfullclaimed7/10

                              ADK docs explicitly state it "automatically filters irrelevant events, summarizes older conversational turns, lazy-loads artifacts, and tracks token usage," directly addressing trimming/summarizing/filtering to manage context window, reinforced by mention of designing for AI context window limits. Missing for 10: no code example/API reference showing how a developer configures or customizes this summarization/filtering behavior, and no independent/hands-on corroboration beyond first-party docs.

                              • [claimed-docs] ADK automatically filters irrelevant events, summarizes older conversational turns, lazy-loads artifacts, and tracks token usage.
                              • [claimed-docs] Use prebuilt or custom Agent Skills to extend agent capabilities in a way that works efficiently inside AI context window limits.
                              AutoGennone0/10

                              The evidence pack mentions general 'Memory' and 'Logging' capabilities but never documents any mechanism for trimming, summarizing, or filtering conversation history to manage context window size — no mention of context buffering, truncation, or summarization utilities.

                            2. developerGive agents long-term memory that persists across sessions and threads

                              weight 2 · round to AutoGen
                              Google ADKpartialclaimed3/10

                              The docs mention session-based interactions and automatic context management (filtering irrelevant events, summarizing older turns, tracking token usage) but there is no explicit evidence of a dedicated long-term memory service or store that persists agent knowledge across separate sessions/threads. missing for 10: explicit memory/session-store API docs, cross-session persistence guarantees, first-party examples of retrieving memory in a new thread.

                              • [claimed-docs] ADK automatically filters irrelevant events, summarizes older conversational turns, lazy-loads artifacts, and tracks token usage.
                              • [claimed-docs] This approach involves creating individual test files, each representing a single, simple agent-model interaction (a session).
                              AutoGenpartialclaimed5/10

                              AutoGen's docs advertise a 'Memory' component ('Add memory capabilities to your agents') as part of AgentChat, indicating some support for giving agents memory, but the evidence pack gives no detail on how this memory persists across sessions/threads (e.g., storage backend, serialization, retrieval across conversations). missing for 10: explicit documentation of session/thread-persistent memory implementation, examples of memory surviving across separate runs, independent/hands-on confirmation.

                            Openness — open source, data portability, and self-hosting storiesOpenness

                            Open source, data portability, and self-hosting stories

                            1. ai-native userDo everything through the API that I can do in the UI

                              weight 2 · round to AutoGen
                              Google ADKpartialprobed5/10

                              ADK is primarily a code-first Python framework where agents are built and orchestrated programmatically (Agent(), Workflow Runtime, Task API), and the CLI (adk run/web/eval/deploy) exposes most dev-loop actions including the same UI functions, suggesting reasonable parity between programmatic/CLI and the built-in dev UI. However, there's no evidence of a documented REST/OpenAPI API for driving the dev UI's specific features programmatically, and probes show no OpenAPI spec or llms.txt discoverability. missing for 10: explicit API/CLI parity documentation for every dev-UI feature (debug, evaluate, showcase), a published OpenAPI/REST spec, and confirmation that UI-only actions (e.g. visual debugging, showcase mode) are fully scriptable.

                              • [github] A built-in development UI to help you test, evaluate, debug, and showcase your agent(s).
                              • [github] Web UI (supports multi-agent directories or pointing directly to a single agent folder)
                              • [github] adk run path/to/my_agent
                              • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
                              • [github] adk eval \ samples_for_testing/hello_world \ samples_for_testing/hello_world/hello_world_eval_set_001.evalset.json
                              • [github] adk deploy docker --with_ui <agent-folder>
                              • [probe] PROBE openapi: all candidate paths 404 (https://google.github.io/openapi.json, https://google.github.io/swagger.json, https://google.github.…
                              • [probe] PROBE llms.txt: HTTP 404 at https://google.github.io/llms.txt
                              AutoGenpartialclaimed6/10

                              AutoGen's core framework is code/API-first, and AutoGenStudio (the visual UI) explicitly supports exporting team configurations to Python code and setting up API endpoints from a team config, indicating parity between UI-built and API-driven workflows (autogen-docs-5, autogen-docs-17). However there's no evidence confirming every UI feature (e.g. community component gallery/hub, drag-and-drop specifics) has a documented equivalent API path, and no independent corroboration of full parity. Missing for 10: explicit 1:1 mapping of all AutoGenStudio UI features (gallery/hub, docker run options) to API calls, and independent/hands-on confirmation of parity.

                              • [claimed-docs] A visual interface for creating agent teams through declarative specification (JSON) or drag-and-drop
                              • [claimed-docs] Export and run teams in python code
                              • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
                              • [claimed-docs] Central hub for discovering and importing community-created components
                            2. ai-native userRead the product's source under an open license

                              weight 2 · round to Google ADK
                              Google ADKfullclaimed7/10

                              The evidence repeatedly links to the public GitHub repository https://github.com/google/adk-python, which hosts the full source code and CLI (adk run, adk web, adk eval, adk deploy) that AI-native users can read and inspect directly. Missing for 10: an explicit citation of the license file/type (e.g., Apache-2.0) confirming the open-license terms, and independent third-party confirmation of licensing.

                              • [github] Agent Config: Build agents without code.
                              • [github] A built-in development UI to help you test, evaluate, debug, and showcase your agent(s).
                              • [github] adk run path/to/my_agent
                              • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
                              AutoGenpartialclaimed6/10

                              The evidence pack shows AutoGen is distributed via a public GitHub repository (autogen-gh-1,2,3) and pip-installable packages, indicating its source is publicly readable, but no evidence explicitly states or confirms an open-source license (e.g., MIT/Apache) in the pack. missing for 10: explicit license file/badge citation, confirmation of open-license terms, independent verification of license compliance.

                              • [github] pip install -U "autogen-agentchat" "autogen-ext[openai]"
                              • [github] Create a web browsing assistant agent that uses the Playwright MCP server.
                              • [github] You can use `AgentTool` to create a basic multi-agent orchestration setup.
                            3. ai-native userSelf-host the core product

                              weight 3 · round drawn
                              Google ADKfullclaimed8/10

                              ADK is an open-source framework (github.com/google/adk-python) that can be run entirely locally via `adk run`/`adk web`, packaged into containers, and deployed offline/disconnected from Google Cloud, evidencing full self-hosting capability without requiring the vendor's managed service. Missing for 10: no independent third-party report confirming a full self-hosted production deployment, and no explicit self-hosted infra requirements/scaling guidance beyond container packaging.

                              • [claimed-docs] This is a good option if you prefer to run offline or disconnected, or otherwise in a system that has no connection to Google Cloud.
                              • [claimed-docs] You can manually package your Agent into a container image and then run it in any environment that supports container images.
                              • [github] adk run path/to/my_agent
                              • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
                              • [github] adk deploy docker --with_ui <agent-folder>
                              • [claimed-docs] GKE is a good option if you need more control over the deployment as well as for running Open Models.
                              AutoGenfullcommunity8/10

                              AutoGen is an open-source Python framework installed via pip (autogen-agentchat) and available on GitHub, meaning the core agent/team runtime is self-hosted by design; AutoGen Studio can also be run locally or in a Docker container. Missing for 10: dedicated self-hosting/deployment guide or infrastructure requirements documentation beyond pip install and docker run mentions.

                              • [github] pip install -U "autogen-agentchat" "autogen-ext[openai]"
                              • [claimed-docs] pip install -U "autogen-agentchat"
                              • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
                              • [community] i noticed autogen creates a new docker container each time code is executed by agents (default behaviour, can be turned off), so it's safe a…

                            Orchestration multi agent — stories about orchestration multi agent in this arenaOrchestration multi agent

                            Stories about orchestration multi agent in this arena

                            Multi agent

                            1. developerOrchestrate multiple agents — handoffs, subagents, or crews — inside one workflow

                              weight 3 · round drawn
                              Google ADKfullclaimed9/10

                              ADK explicitly supports multi-agent orchestration: workflows are defined as any application with more than one agent/node, with a graph-based Workflow Runtime supporting routing, fan-out/fan-in, loops, nested workflows, and a Task API for structured agent-to-agent delegation including multi-turn task mode and mixed delegation patterns; the CLI/Web UI explicitly supports multi-agent directories. missing for 10: independent third-party hands-on validation of complex multi-agent orchestration at scale.

                              • [claimed-docs] you can use the ADK development framework to expand them into workflows, which allow you to combine and orchestrate multiple agents and code…
                              • [claimed-docs] In ADK, any agent application that has more than one agent or executable Node is considered a workflow.
                              • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
                              • [github] Task API: Structured agent-to-agent delegation with multi-turn task mode, single-turn controlled output, mixed delegation patterns, human-in…
                              • [github] Structured agent-to-agent delegation with multi-turn task mode, single-turn controlled output, mixed delegation patterns, human-in-the-loop,…
                              • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
                              AutoGenfullcommunity9/10

                              AutoGen provides extensive first-party documentation of multi-agent orchestration patterns: RoundRobinGroupChat, SelectorGroupChat, Swarm (handoff-based), and GraphFlow (directed-graph workflows), plus AgentTool for nesting agents as tools, all within a single workflow. Community feedback corroborates real-world use of multi-agent conversation control. Missing for 10: independent hands-on benchmark of complex multi-agent handoff reliability beyond the HN thread's mixed performance comments.

                              • [claimed-docs] Multi-agent coordination through a shared context and centralized, customizable selector
                              • [claimed-docs] Multi-agent coordination through a shared context and localized, tool-based selector
                              • [claimed-docs] Multi-agent workflows through a directed graph of agents.
                              • [claimed-docs] SelectorGroupChat: A team that selects the next speaker using a ChatCompletion model after each message.
                              • [claimed-docs] Swarm: A team that uses HandoffMessage to signal transitions between agents.
                              • [claimed-docs] GraphFlow: Multi-agent workflows through a directed graph of agents.
                              • [github] You can use `AgentTool` to create a basic multi-agent orchestration setup.
                              • [community] Have been working with this and very impressed so far - it's a step ahead of LangChain agents and seems to be receiving more attention/devel…
                              • [community] The breakthrough I've had is realizing how important it is to control the conversation between agents. Just like in our work environments an…

                            Workflow control

                            1. developerCompose agents into an explicit graph or workflow with branching, loops, and parallel steps

                              weight 2 · round to Google ADK
                              Google ADKfullclaimed9/10

                              ADK provides a dedicated graph-based Workflow Runtime with explicit support for routing, fan-out/fan-in (parallel), loops, retry, nested workflows, and dynamic nodes, plus structured Task API for agent delegation and workflow nodes—directly matching branching/loops/parallel composition; docs also describe 'graph-based architectures with explicit execution paths.' Missing for 10: independent/hands-on third-party validation beyond vendor docs and GitHub README.

                              • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
                              • [github] A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan-out/fan-in, loops…
                              • [github] Task API: Structured agent-to-agent delegation with multi-turn task mode, single-turn controlled output, mixed delegation patterns, human-in…
                              • [github] Structured agent-to-agent delegation with multi-turn task mode, single-turn controlled output, mixed delegation patterns, human-in-the-loop,…
                              • [claimed-docs] Weave deterministic code with adaptive AI reasoning. Orchestrate complex tasks through structured, graph-based architectures, with explicit …
                              • [claimed-docs] In ADK, any agent application that has more than one agent or executable Node is considered a workflow.
                              AutoGenpartialclaimed6/10

                              AutoGen explicitly ships GraphFlow, described as enabling 'multi-agent workflows through a directed graph of agents,' which directly supports the story's core ask of graph-based orchestration alongside other team patterns (RoundRobin, Selector, Swarm) for different coordination styles. However, the evidence pack never details how branching conditions, loop constructs, or parallel step execution are configured within GraphFlow, so the specific mechanics of the story are only partially substantiated. Missing for 10: explicit documentation/examples of conditional branching syntax, loop/cycle handling, and parallel step execution within GraphFlow, plus independent hands-on confirmation of these features working as described.

                              • [claimed-docs] Multi-agent workflows through a directed graph of agents.
                              • [claimed-docs] GraphFlow: Multi-agent workflows through a directed graph of agents.
                              • [claimed-docs] SelectorGroupChat: A team that selects the next speaker using a ChatCompletion model after each message.
                              • [claimed-docs] Swarm: A team that uses HandoffMessage to signal transitions between agents.
                              • [claimed-docs] RoundRobinGroupChat is a simple yet effective team configuration where all agents share the same context and take turns responding in a roun…

                            Privacy posture — data-handling and privacy storiesPrivacy posture

                            Data-handling and privacy stories

                            1. ai-native userControl data retention and deletion

                              weight 2 · round drawn
                              Google ADKnone0/10

                              The evidence describes ADK as a self-hosted/deployable agent framework (Cloud Run, GKE, offline/disconnected deployment) but contains no documentation of explicit data retention policies, session/state deletion APIs, or user-facing controls for purging stored data. missing for 10: explicit retention/deletion controls, session data lifecycle docs, any privacy/compliance statements about stored artifacts or memory.

                              • [claimed-docs] This is a good option if you prefer to run offline or disconnected, or otherwise in a system that has no connection to Google Cloud.
                              • [claimed-docs] ADK automatically filters irrelevant events, summarizes older conversational turns, lazy-loads artifacts, and tracks token usage.
                              AutoGennone0/10

                              AutoGen is an open-source, self-hosted framework, so data retention/deletion would be determined by the user's own infrastructure, but no evidence pack item discusses any built-in retention policy, data deletion controls, or configuration for purging stored conversation/memory data.

                              • ai-native userOpt out of telemetry and usage tracking

                                weight 2 · round drawn
                                Google ADKnone0/10

                                No evidence in the pack addresses telemetry collection or an opt-out mechanism for ADK; the docs cover agent building, deployment, evaluation, and workflows but never mention usage tracking or privacy controls. This is a fair axis for a developer framework/SDK, but absence of evidence means it counts as none. missing for 10: any mention of telemetry collection, an opt-out flag/env var, or a privacy policy describing data tracking.

                                  AutoGennone0/10

                                  No evidence in the pack mentions telemetry, usage tracking, or an opt-out mechanism for AutoGen; as an open-source, self-hosted framework this axis plausibly applies but is unaddressed.

                                  State durability — stories about state durability in this arenaState durability

                                  Stories about state durability in this arena

                                  Durable state

                                  1. developerCheckpoint agent state so a run can resume exactly where it left off after a crash or restart

                                    weight 3 · round to AutoGen
                                    Google ADKnone0/10

                                    Evidence only mentions generic 'state management' as one feature in the workflow runtime engine, with no documentation of session/state persistence, checkpointing, or resuming an agent run after a crash or restart. Missing for 10: explicit checkpoint/save-state API, resume-from-crash mechanism, persistence backend documentation, and any hands-on confirmation of durable resumption.

                                    • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
                                    • [github] A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan-out/fan-in, loops…
                                    • [claimed-docs] ADK automatically filters irrelevant events, summarizes older conversational turns, lazy-loads artifacts, and tracks token usage.
                                    AutoGenpartialclaimed3/10

                                    Docs mention a 'Serialize Components' feature for serializing/deserializing components and a separate logging/tracing feature, which are the building blocks for state persistence, but there is no explicit documentation of a checkpoint/resume workflow after a crash or restart. missing for 10: explicit checkpoint/resume API or tutorial, crash-recovery guarantees, independent confirmation that resumed runs continue exactly where they left off.

                                    • [claimed-docs] Serialize Components: Serialize and deserialize components
                                    • [claimed-docs] Logging: Log traces and internal messages
                                  2. engineering-leadRun long-lived agents durably across process restarts and deploys, natively or via durable-execution integrations

                                    weight 2 · round to Google ADK
                                    Google ADKpartialclaimed3/10

                                    ADK's Workflow Runtime mentions 'state management' and 'retry' in its graph-based execution engine, and deployment docs describe scalable hosting (Agent Runtime, Cloud Run), but there is no explicit evidence of session/state persistence surviving process restarts or redeploys, nor any named durable-execution integration (e.g., Temporal, Cloud Workflows checkpointing). Missing for 10: documented durable state store or checkpoint/resume mechanism, explicit claim of surviving restarts/redeploys, and any third-party durable-execution integration.

                                    • [github] Workflow Runtime: A graph-based execution engine for composing deterministic execution flows for agentic apps, with support for routing, fan…
                                    • [claimed-docs] Agent Runtime is a fully managed auto-scaling service on Google Cloud specifically designed for deploying, managing, and scaling AI agents b…
                                    • [claimed-docs] Cloud Run is a managed auto-scaling compute platform on Google Cloud that enables you to run your agent as a container-based application.
                                    • [claimed-docs] ADK automatically filters irrelevant events, summarizes older conversational turns, lazy-loads artifacts, and tracks token usage.
                                    AutoGennone0/10

                                    Evidence shows serialization of components and logging/tracing, but there is no mention of durable execution, process-restart recovery, checkpointing/resume across deploys, or integrations with durable-execution engines (e.g., Temporal). missing for 10: durable-execution runtime or integration, state checkpoint/resume across restarts, deploy-survival guarantees, any documentation or example of long-lived agent persistence.

                                    • [claimed-docs] Serialize Components: Serialize and deserialize components
                                    • [claimed-docs] Logging: Log traces and internal messages

                                  Streaming output — stories about streaming output in this arenaStreaming output

                                  Stories about streaming output in this arena

                                  Streaming

                                  1. developerStream tokens and intermediate agent events (tool calls, steps) to my UI in real time

                                    weight 3 · round to Google ADK
                                    Google ADKpartialclaimed4/10

                                    The evidence shows a built-in Web/dev UI (`adk web`) for testing/debugging agents and a Callbacks mechanism to hook into execution-lifecycle events (tool calls, steps), which implies some visibility into intermediate agent activity, but nothing explicitly documents token-level streaming to a custom UI (no mention of SSE/websocket/streaming API). missing for 10: explicit documentation of real-time token streaming API/protocol, evidence of streaming tool-call/step events to an arbitrary UI beyond the built-in dev UI, independent confirmation of streaming behavior.

                                    • [github] A built-in development UI to help you test, evaluate, debug, and showcase your agent(s).
                                    • [github] Web UI (supports multi-agent directories or pointing directly to a single agent folder)
                                    • [github] adk run path/to/my_agent # Web UI (supports multi-agent directories or pointing directly to a single agent folder) adk web path/to/agents_d…
                                    • [claimed-docs] Callbacks: Hook into specific events during an agent's execution lifecycle to add logging, monitoring, or custom side-effects without alteri…
                                    AutoGennone0/10

                                    The evidence pack mentions logging of traces/internal messages and an AutoGen Studio interactive environment, but nowhere describes token-level streaming or real-time emission of intermediate agent/tool events to a UI. Missing for 10: any mention of a streaming API (e.g., token/event streaming methods), UI integration examples, or independent confirmation of real-time event delivery.

                                    • [claimed-docs] Logging: Log traces and internal messages
                                    • [claimed-docs] Interactive environment for testing and running agent teams

                                  Structured output

                                  1. developerGet schema-validated structured output from an agent, with automatic retries when validation fails

                                    weight 3 · round drawn
                                    Google ADKnone0/10

                                    No evidence in the pack mentions schema-validated structured output (e.g., Pydantic output_schema) or automatic retry-on-validation-failure behavior for ADK agents; the evidence covers agent setup, tools, workflows, deployment, and evaluation but not structured output validation. Missing for 10: any mention of output schema enforcement, structured output configuration, or validation-retry mechanism.

                                      AutoGennone0/10

                                      No evidence in the pack mentions schema-validated structured output or automatic retry-on-validation-failure behavior for AutoGen agents; docs cover agents, teams, tools, and orchestration but not structured output validation. Missing for 10: any mention of structured output schemas (e.g., Pydantic models), validation error handling, or retry logic tied to output parsing.

                                      Not comparable on these axes

                                      1. ai-native userConnect an agent via an official MCP server

                                        weight 3 · not comparable
                                        Google ADKpartialclaimed6/10

                                        ADK's official docs explicitly document how to expose ADK tools via an MCP server ('build an MCP server that wraps ADK tools, making them accessible to any MCP client'), showing the framework supports the server side of MCP, not just being an MCP client. However, this is a build-your-own-server guide rather than a turnkey, pre-hosted official MCP endpoint, so it requires developer setup work. Missing for 10: a ready-made hosted/official MCP server endpoint, independent hands-on confirmation that the generated server works reliably with third-party MCP clients.

                                        • [claimed-docs] Exposing ADK Tools via an MCP Server: How to build an MCP server that wraps ADK tools, making them accessible to any MCP client.
                                        • [claimed-docs] How to build an MCP server that wraps ADK tools, making them accessible to any MCP client.
                                        • [claimed-docs] An ADK agent can act as an MCP client and use tools provided by external MCP servers.
                                        AutoGenn/a

                                        AutoGen is an agent framework/client, and the story concerns serving tools via an official MCP server (agent-as-server role). The only relevant evidence (autogen-gh-2) shows AutoGen agents connecting to an external Playwright MCP server, which is client-side usage and does not make the server axis applicable per the agent-role exception.

                                        • [github] Create a web browsing assistant agent that uses the Playwright MCP server.
                                      2. ai-native userGet AI-generated insights and suggestions from my data inside the product

                                        weight 2 · not comparable
                                        Google ADKn/a

                                        Google ADK is a developer framework/SDK for building agent applications, not an end-user product with a data surface that itself surfaces AI-generated insights to a user; the evidence is entirely about developer tooling (agent definitions, workflows, deployment, evaluation), not about a product feature that analyzes 'my data' and surfaces insights within an application UI.

                                          AutoGenpartialclaimed4/10

                                          AutoGen provides building blocks (AssistantAgent with tool use, memory, multi-agent teams, AutoGen Studio for building/testing agents) that a developer could use to construct data-analysis or insight-generating agents, but there is no evidence of a built-in, turnkey feature that ingests a user's data and surfaces AI-generated insights inside the product itself — it remains a framework requiring custom agent construction. Missing for 10: a documented out-of-the-box 'analyze my data / dashboard insights' feature, evidence of automatic data ingestion, and independent hands-on confirmation of insight quality on real datasets.

                                          • [claimed-docs] AssistantAgent is a built-in agent that uses a language model and has the ability to use tools.
                                          • [claimed-docs] Add memory capabilities to your agents
                                          • [claimed-docs] Interactive environment for testing and running agent teams
                                          • [claimed-docs] A visual interface for creating agent teams through declarative specification (JSON) or drag-and-drop
                                          • [github] You can use `AgentTool` to create a basic multi-agent orchestration setup.
                                        • ai-native userExport all of my data in open formats and leave

                                          weight 3 · not comparable
                                          Google ADKn/a

                                          Google ADK is an open-source, locally-run agent-building framework where agent code/configs are files developers own directly (in their own repos), not a hosted service that stores user data centrally requiring an 'export and leave' capability. The data-export/lock-in axis is designed for SaaS platforms holding user data hostage, which does not match ADK's dev-framework category.

                                            AutoGenpartialclaimed4/10

                                            AutoGen Studio supports exporting team configurations as JSON or Python code and the framework has component serialization/deserialization features, which are open, portable formats. However, there is no evidence of exporting broader user data (conversation history, memory stores, logs) in a comprehensive open-format package, and no documentation of a full account/data 'leave' export process. Missing for 10: evidence of exporting full conversation/memory history, a documented data-portability/export-all workflow, and independent confirmation of format openness beyond configs.

                                            • [claimed-docs] A visual interface for creating agent teams through declarative specification (JSON) or drag-and-drop
                                            • [claimed-docs] Export and run teams in python code
                                            • [claimed-docs] Export and run teams in python code * Setup and test endpoints based on a team configuration * Run teams in a docker container
                                            • [claimed-docs] Serialize Components: Serialize and deserialize components
                                          • ai-native userChoose where my data is stored (region/residency)

                                            weight 2 · not comparable
                                            Google ADKnone0/10

                                            ADK is a framework that can be deployed via Cloud Run, GKE, or self-hosted/offline (google-adk-docs-7, google-adk-docs-14, google-adk-docs-19), which implies developers control infrastructure location, but there is no explicit documentation about data residency, region selection, or storage location controls for agent data.

                                            • [claimed-docs] This is a good option if you prefer to run offline or disconnected, or otherwise in a system that has no connection to Google Cloud.
                                            • [claimed-docs] Cloud Run is a managed auto-scaling compute platform on Google Cloud that enables you to run your agent as a container-based application.
                                            • [claimed-docs] GKE is a good option if you need more control over the deployment as well as for running Open Models.
                                            AutoGenn/a

                                            AutoGen is a self-hosted, open-source multi-agent framework/library, not a hosted SaaS that stores user data — deployment location and data residency are entirely determined by the user's own infrastructure, not a product feature to select. No evidence pack content addresses region/residency selection, and the axis is a category error for a framework with no first-party data storage.

                                            • ai-native userPrevent my data from being used to train AI models

                                              weight 3 · not comparable
                                              Google ADKn/a

                                              Google ADK is an open-source developer framework for building agents, run locally or self-hosted, not a hosted AI service with a data-training policy to opt out of; this privacy-posture question about model-training data usage is a category error for a framework/SDK.

                                                AutoGenn/a

                                                AutoGen is an open-source multi-agent framework you self-host, not a hosted AI service with a data-training policy for user data; there's no vendor relationship where 'my data used for training' applies (users bring their own LLM API keys/providers). This axis is a category error for a framework rather than a hosted product with its own training policy.