Skip to content

Feature request: emit structured evidence bundles for actions and failures (text + json content) #341

Description

@KazuCocoa

(Wrote this doc with LLM's help)

I'd like to find better evidence for each action. So, current response is text only basis. If we gives more better usage in CI, probably it makes sense to add more context for each action such as what kind of locator the tool actually used. This would help for the future CI usage etc.

This assumption could be wrong if current info was sufficient though.

Body

Problem

Appium MCP is positioned as an MCP server for mobile automation and recent releases show active development, but public-facing descriptions emphasize actions, session control, and generation more than evidence-oriented execution artifacts. For automation at scale, actions are not enough; every run needs replayable and diagnosable output.

Today, when a tool call fails (e.g., tap_element, type_text, or future wait_for_element), users often see only a high-level error. It’s difficult to answer questions like:

  • Which locator was actually used?
  • Which element did we resolve (by WebDriver/Appium element ID)?
  • What platform/app/context/screen was active at the time?
  • What did the UI look like before and after the action?
  • Did the app navigate, crash, or stay on the same screen?
  • Was it a timeout, a locator issue, or something else?

This makes failures hard to debug (especially in CI) and keeps Appium MCP closer to an “interactive operator surface” than an auditable automation platform.


Proposal

For every “meaningful” tool call (starting with core action tools and later waits/asserts), Appium MCP should return a structured evidence bundle alongside a compact summary.

To keep token usage under control while still providing rich evidence, I’m proposing a dual-content pattern:

  • A short text summary for the LLM
  • A structured json object for tools/CI/logs

EvidenceBundle shape (example)

type EvidenceBundle = {
  id: string;
  toolName: string;
  args: any;

  locator?: {
    strategy: string;  // e.g., "accessibilityId", "id", "xpath"
    value: string;     // e.g., "login_button"
  };

  element?: {
    webdriverId?: string; // Appium/WebDriver element id
    logicalId?: string;   // optional: test-oriented ID like "LoginScreen.LoginButton"
  };

  context?: {
    platform: "android" | "ios";
    appPackageOrBundle?: string;
    activityOrScreen?: string;
    contextName?: string; // e.g., "NATIVE_APP", "WEBVIEW_com.example"
  };

  screenshots?: {
    beforePath?: string; // file paths, not inline base64
    afterPath?: string;
  };

  hierarchy?: {
    pageSourcePath?: string;     // optional: UI hierarchy snapshot
    accessibilityPath?: string;  // optional: accessibility tree snapshot
  };

  timing: {
    startedAt: string;   // ISO timestamp
    finishedAt: string;  // ISO timestamp
    durationMs: number;
  };

  error?: {
    code: string;    // e.g., "ELEMENT_NOT_FOUND", "TIMEOUT", "CONTEXT_NOT_AVAILABLE", "ACTION_FAILED"
    message: string; // human-readable summary
  };

  evidenceRef?: string; // path or ID for the persisted bundle
};

This adds element-level information (both WebDriver element ID and an optional logical/test ID), which makes it easier to correlate with driver logs and test models.

Return format (text + json)

Instead of stuffing everything into text, the server would return:

return {
  content: [
    {
      type: "json",
      mimeType: "application/vnd.appium.evidence+json",
      data: {
        id: evidence.id,
        toolName: "tap_element",
        locator: evidence.locator,
        element: evidence.element,
        context: evidence.context,
        timing: evidence.timing,
        error: evidence.error,
        evidenceRef: evidenceJsonPath // e.g., runs/.../evidence.json
      }
    },
    {
      type: "text",
      text: !evidence.error
        ? `tap_element succeeded (locator=${evidence.locator?.strategy}:${evidence.locator?.value}, evidenceId=${evidence.id})`
        : `tap_element failed (code=${evidence.error.code}, locator=${evidence.locator?.strategy}:${evidence.locator?.value}, evidenceId=${evidence.id})`
    }
  ]
};

This keeps the LLM-facing text small, while providing a structured json blob for tools and CI systems to consume.


Token-efficiency considerations

To avoid unnecessary token usage:

  • The text summary should remain compact:
    • Short sentence + error code + evidence ID
  • Heavy data should not be embedded in text:
    • Screenshots saved as files (before.png, after.png)
    • Full hierarchy snapshots saved as files
  • The json evidence should reference heavy artifacts by path/ID, not inline base64 blobs.

Example for tap_element:

const summary = evidence.error
  ? `tap_element failed (code=${evidence.error.code}, evidenceId=${evidence.id})`
  : `tap_element succeeded (evidenceId=${evidence.id})`;

return {
  content: [
    {
      type: "json",
      mimeType: "application/vnd.appium.evidence+json",
      data: {
        id: evidence.id,
        locator: evidence.locator,
        element: evidence.element,
        context: evidence.context,
        timing: evidence.timing,
        error: evidence.error,
        evidenceRef: evidenceJsonPath
      }
    },
    {
      type: "text",
      text: summary
    }
  ]
};

Example flow (narrative)

For a tap_element tool call:

  1. Before:
    • Capture context (platform/app/contextName)
    • Capture optional before.png (saved to disk)
  2. Action:
    • Resolve locator (strategy + selector)
    • Call findElement and record the WebDriver element ID in evidence.element.webdriverId
    • Optionally map to a logical/test ID (e.g., LoginScreen.LoginButton)
    • Attempt click
    • Normalize any failure into { code, message }
  3. After:
    • Capture optional after.png
    • Record timing
    • Persist evidence.json to disk (if configured)
  4. Return:
    • json evidence content (lightweight subset + evidenceRef)
    • text summary with code + evidence ID

Expected benefit

This change would:

  • Make Appium MCP much easier to debug:
    • Every action has a reproducible, structured trace, including which element was actually targeted.
  • Provide a clear CI story:
    • Failures come with artifacts (screenshots, element IDs, context, timing) as a matter of course.
  • Improve agent behavior:
    • Agents can check error codes, element IDs, and evidence references instead of guessing.
  • Preserve token efficiency:
    • LLMs see a small summary; heavy evidence stays in separate files or structured json data.
  • Move Appium MCP closer to a full-fledged test runner for agents, not just a thin tool bridge.

You can paste this as your Issue 2, or tweak field names if you want to align with your internal naming conventions.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions