(Wrote this doc with LLM's help)
I'd like to find better evidence for each action. So, current response is text only basis. If we gives more better usage in CI, probably it makes sense to add more context for each action such as what kind of locator the tool actually used. This would help for the future CI usage etc.
This assumption could be wrong if current info was sufficient though.
Body
Problem
Appium MCP is positioned as an MCP server for mobile automation and recent releases show active development, but public-facing descriptions emphasize actions, session control, and generation more than evidence-oriented execution artifacts. For automation at scale, actions are not enough; every run needs replayable and diagnosable output.
Today, when a tool call fails (e.g., tap_element, type_text, or future wait_for_element), users often see only a high-level error. It’s difficult to answer questions like:
- Which locator was actually used?
- Which element did we resolve (by WebDriver/Appium element ID)?
- What platform/app/context/screen was active at the time?
- What did the UI look like before and after the action?
- Did the app navigate, crash, or stay on the same screen?
- Was it a timeout, a locator issue, or something else?
This makes failures hard to debug (especially in CI) and keeps Appium MCP closer to an “interactive operator surface” than an auditable automation platform.
Proposal
For every “meaningful” tool call (starting with core action tools and later waits/asserts), Appium MCP should return a structured evidence bundle alongside a compact summary.
To keep token usage under control while still providing rich evidence, I’m proposing a dual-content pattern:
- A short
text summary for the LLM
- A structured
json object for tools/CI/logs
EvidenceBundle shape (example)
type EvidenceBundle = {
id: string;
toolName: string;
args: any;
locator?: {
strategy: string; // e.g., "accessibilityId", "id", "xpath"
value: string; // e.g., "login_button"
};
element?: {
webdriverId?: string; // Appium/WebDriver element id
logicalId?: string; // optional: test-oriented ID like "LoginScreen.LoginButton"
};
context?: {
platform: "android" | "ios";
appPackageOrBundle?: string;
activityOrScreen?: string;
contextName?: string; // e.g., "NATIVE_APP", "WEBVIEW_com.example"
};
screenshots?: {
beforePath?: string; // file paths, not inline base64
afterPath?: string;
};
hierarchy?: {
pageSourcePath?: string; // optional: UI hierarchy snapshot
accessibilityPath?: string; // optional: accessibility tree snapshot
};
timing: {
startedAt: string; // ISO timestamp
finishedAt: string; // ISO timestamp
durationMs: number;
};
error?: {
code: string; // e.g., "ELEMENT_NOT_FOUND", "TIMEOUT", "CONTEXT_NOT_AVAILABLE", "ACTION_FAILED"
message: string; // human-readable summary
};
evidenceRef?: string; // path or ID for the persisted bundle
};
This adds element-level information (both WebDriver element ID and an optional logical/test ID), which makes it easier to correlate with driver logs and test models.
Return format (text + json)
Instead of stuffing everything into text, the server would return:
return {
content: [
{
type: "json",
mimeType: "application/vnd.appium.evidence+json",
data: {
id: evidence.id,
toolName: "tap_element",
locator: evidence.locator,
element: evidence.element,
context: evidence.context,
timing: evidence.timing,
error: evidence.error,
evidenceRef: evidenceJsonPath // e.g., runs/.../evidence.json
}
},
{
type: "text",
text: !evidence.error
? `tap_element succeeded (locator=${evidence.locator?.strategy}:${evidence.locator?.value}, evidenceId=${evidence.id})`
: `tap_element failed (code=${evidence.error.code}, locator=${evidence.locator?.strategy}:${evidence.locator?.value}, evidenceId=${evidence.id})`
}
]
};
This keeps the LLM-facing text small, while providing a structured json blob for tools and CI systems to consume.
Token-efficiency considerations
To avoid unnecessary token usage:
- The
text summary should remain compact:
- Short sentence + error code + evidence ID
- Heavy data should not be embedded in
text:
- Screenshots saved as files (
before.png, after.png)
- Full hierarchy snapshots saved as files
- The
json evidence should reference heavy artifacts by path/ID, not inline base64 blobs.
Example for tap_element:
const summary = evidence.error
? `tap_element failed (code=${evidence.error.code}, evidenceId=${evidence.id})`
: `tap_element succeeded (evidenceId=${evidence.id})`;
return {
content: [
{
type: "json",
mimeType: "application/vnd.appium.evidence+json",
data: {
id: evidence.id,
locator: evidence.locator,
element: evidence.element,
context: evidence.context,
timing: evidence.timing,
error: evidence.error,
evidenceRef: evidenceJsonPath
}
},
{
type: "text",
text: summary
}
]
};
Example flow (narrative)
For a tap_element tool call:
- Before:
- Capture context (platform/app/contextName)
- Capture optional
before.png (saved to disk)
- Action:
- Resolve locator (strategy + selector)
- Call
findElement and record the WebDriver element ID in evidence.element.webdriverId
- Optionally map to a logical/test ID (e.g.,
LoginScreen.LoginButton)
- Attempt
click
- Normalize any failure into
{ code, message }
- After:
- Capture optional
after.png
- Record timing
- Persist
evidence.json to disk (if configured)
- Return:
json evidence content (lightweight subset + evidenceRef)
text summary with code + evidence ID
Expected benefit
This change would:
- Make Appium MCP much easier to debug:
- Every action has a reproducible, structured trace, including which element was actually targeted.
- Provide a clear CI story:
- Failures come with artifacts (screenshots, element IDs, context, timing) as a matter of course.
- Improve agent behavior:
- Agents can check error codes, element IDs, and evidence references instead of guessing.
- Preserve token efficiency:
- LLMs see a small summary; heavy evidence stays in separate files or structured
json data.
- Move Appium MCP closer to a full-fledged test runner for agents, not just a thin tool bridge.
You can paste this as your Issue 2, or tweak field names if you want to align with your internal naming conventions.
(Wrote this doc with LLM's help)
I'd like to find better evidence for each action. So, current response is text only basis. If we gives more better usage in CI, probably it makes sense to add more context for each action such as what kind of locator the tool actually used. This would help for the future CI usage etc.
This assumption could be wrong if current info was sufficient though.
Body
Problem
Appium MCP is positioned as an MCP server for mobile automation and recent releases show active development, but public-facing descriptions emphasize actions, session control, and generation more than evidence-oriented execution artifacts. For automation at scale, actions are not enough; every run needs replayable and diagnosable output.
Today, when a tool call fails (e.g.,
tap_element,type_text, or futurewait_for_element), users often see only a high-level error. It’s difficult to answer questions like:This makes failures hard to debug (especially in CI) and keeps Appium MCP closer to an “interactive operator surface” than an auditable automation platform.
Proposal
For every “meaningful” tool call (starting with core action tools and later waits/asserts), Appium MCP should return a structured evidence bundle alongside a compact summary.
To keep token usage under control while still providing rich evidence, I’m proposing a dual-content pattern:
textsummary for the LLMjsonobject for tools/CI/logsEvidenceBundle shape (example)
This adds element-level information (both WebDriver element ID and an optional logical/test ID), which makes it easier to correlate with driver logs and test models.
Return format (text + json)
Instead of stuffing everything into
text, the server would return:This keeps the LLM-facing text small, while providing a structured
jsonblob for tools and CI systems to consume.Token-efficiency considerations
To avoid unnecessary token usage:
textsummary should remain compact:text:before.png,after.png)jsonevidence should reference heavy artifacts by path/ID, not inline base64 blobs.Example for
tap_element:Example flow (narrative)
For a
tap_elementtool call:before.png(saved to disk)findElementand record the WebDriver element ID inevidence.element.webdriverIdLoginScreen.LoginButton)click{ code, message }after.pngevidence.jsonto disk (if configured)jsonevidence content (lightweight subset +evidenceRef)textsummary with code + evidence IDExpected benefit
This change would:
jsondata.You can paste this as your Issue 2, or tweak field names if you want to align with your internal naming conventions.