Skip to content

fix(openhands-automation): stop steering agents away from deterministic scripts - #686

Merged
enyst merged 6 commits into
mainfrom
fix/automation-skill-preset-balance
Oct 4, 2026
Merged

enyst merged 6 commits into
mainfrom
fix/automation-skill-preset-balance

Conversation

@enyst

@enyst enyst commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

HUMAN:
The agent is too prone to create automations on prompt, when deterministic ones would be more appropriate. Adjust the skill to show them side by side better.


AGENT:

Why

The skill's rule 0 tells the agent to prefer a no-LLM custom script for deterministic tasks, but the rest of the text pulls toward the prompt preset: the prompt examples model deterministic tasks, the custom-automation reference opens with a "do not use unless explicitly requested" warning, there is no complete no-LLM example in the body, and custom scripts carry an extra consent step that presets do not. An agent follows the weight of the document, not the rule.

Summary

  • Replace prompt-preset examples that model deterministic tasks (hourly health check, run tests on push, weekly temp-file cleanup) with tasks that need judgment: Sentry error triage, bug-ticket reproduction, weekly dependency review, flagging risky pushes to main.
  • Delete the stop-sign warning at the top of references/custom-automation.md and list deterministic tasks as a first-class reason to use the custom path.
  • Add a complete no-LLM worked example (hourly health check with a Slack alert: script, tarball, upload, create) next to the prompt-preset examples. Its get_secret matches the helper on main (fix(openhands-automation): read secrets from the OpenHands API on Cloud and Enterprise runs #709), so it reads secrets from the agent server on a local Agent Canvas and from the OpenHands API on Cloud and Enterprise.
  • Collapse rules 2 to 4 into one symmetric rule: present prompt preset, plugin preset, and custom script side by side and build what the user picks. The existing "ready to deploy?" confirmation applies to every path. Matching one-sided sentences in "Choosing the Right Preset" and "Reference Files" are rewritten, the (see rule 5) pointer is updated to rule 4, and a duplicated references/security.md bullet is removed.
  • Update the skill's README.md to match: the prompt preset is no longer marked "(recommended)", and custom scripts are listed next to it as the fit for deterministic tasks instead of "for advanced users".
  • Regenerate skills/index.js.

Issue Number

Fixes #685

How to Test

uv sync --group test
uv run python scripts/sync_extensions.py --check
uv run pytest -q tests/test_skills_catalog.py tests/test_catalogs.py tests/test_sync_extensions.py tests/test_skill_plugin_loading.py tests/test_sdk_loading.py

Run on the current head: the sync check passes and the suite reports 189 passed. node scripts/build-skills-catalog.mjs produces no further diff.

The example script was extracted from SKILL.md and run directly:

  • python3 -m py_compile main.py succeeds.
  • With URL pointed at a page that returns 200, it prints https://example.com/ -> 200 and exits 0.
  • With the unreachable default URL and no run environment, it takes the failure path, prints the error to stderr, and exits 1.
  • get_secret was called against a local stub HTTP server in both modes. With OPENHANDS_CLOUD_API_URL and SANDBOX_ID=sb1 set it requested /api/v1/sandboxes/sb1/settings/secrets/SLACK_BOT_TOKEN; with AGENT_SERVER_URL set it requested /api/settings/secrets/SLACK_BOT_TOKEN. Both sent X-Session-API-Key and returned the secret.

Not tested: deploying the example as a real automation on Cloud or a local Agent Canvas, and the Slack post itself.

To check the behavioral change, load the skill into a conversation and ask for "post a reminder to Slack every 5 minutes". The agent should present the custom-script option next to the presets with a cost note, rather than defaulting to the prompt preset or asking for extra permission to write a script.

Video/Screenshots

Not applicable; Markdown and generated catalog changes only.

Notes

Text only; no change to the automation API or to what the presets do. The regenerated skills/index.js diff is the updated content string for this one skill.

The sync-sdk-skill check fails on this head, but it checks skills/openhands-sdk/SKILL.md against the live docs index, which this PR does not touch; it is failing on other open PRs as well.

🤖 Generated with Claude Code

…ic scripts

The skill's rule 0 tells agents to prefer a no-LLM custom script for
deterministic tasks, but the rest of the text pulled the other way:

- Prompt-preset examples modelled deterministic tasks (hourly health
  check, run tests on push, weekly temp-file cleanup). Replace them with
  tasks that need judgment: error triage, bug reproduction, dependency
  review, flagging risky pushes.
- references/custom-automation.md opened with "Do NOT use this reference
  unless the user has explicitly requested a custom automation", which
  contradicted rule 0 and the Reference Files section. Remove it and list
  deterministic tasks as a first-class reason to use the custom path.
- Add a complete no-LLM worked example (hourly health check with Slack
  alert) next to the prompt-preset examples, so both paths are equally
  concrete.
- Make consent symmetric: present the options side by side and build the
  one the user picks. The "ready to deploy?" confirmation applies to
  every path; custom scripts no longer need an extra permission step.

Also drops a duplicated references/security.md bullet and regenerates
the skills catalog.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

@all-hands-bot all-hands-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review was posted by an AI agent (OpenHands).

Summary

The change does what the linked issue (#685, ready-for-dev, priority:low) asks for: the deterministic prompt examples are replaced with agent-shaped ones (Sentry triage, bug reproduction, dependency review, risky-push review), the stop-sign warning at the top of references/custom-automation.md is gone and deterministic tasks are now a first-class reason to use the custom path, a complete no-LLM worked example is added, and rules 2-4 are collapsed into one symmetric rule. I verified in the workspace that skills/index.js is regenerated (node scripts/build-skills-catalog.mjs produces no further diff), uv run python scripts/sync_extensions.py --check passes, and the focused suite is green (179 passed). All check runs on the head are green. Scope is correct for this repository: it is a skill-content change only, with no API or preset behavior change.

Two non-blocking but real defects remain, both caused by incomplete follow-through on this same edit.

1. Dangling rule cross-reference after the renumbering (inline)

Collapsing rules 2-4 shifts the webhook-reachability rule from 5 to 4, but the in-body pointer near Polling as a Webhook Alternative still reads (see rule 5). There is no rule 5 any more, so an agent following that pointer lands nowhere. Anchored inline at the renumbered rule.

2. README.md was not updated, contrary to the issue's acceptance criteria and the repo convention

skills/openhands-automation/README.md is untouched by this PR (git diff HEAD~1 HEAD -- .../README.md is empty). It still frames the prompt preset as (recommended) in the Features list and describes custom automations as being For advanced users who need full control - exactly the imbalance this PR removes from SKILL.md. The issue's triage acceptance criteria explicitly require the README to stop presenting the prompt preset as the recommended default, and AGENTS.md states: "If you change a skill's behavior or scope, update its README.md (if present) accordingly." The README also still calls /preset/prompt "(recommended)" in prose while SKILL.md moved to an even-handed framing, so the two artifacts now disagree. This is the one acceptance criterion the PR does not meet; adding a short README edit closes it.

Everything else I checked holds up: the new example script compiles, the references/custom-automation.md "when to use" list names deterministic tasks first, the duplicated references/security.md bullet is removed, and no remaining preset example matches a rule 0 instant-recognition pattern (the residual "Generate a weekly status report" / "Daily Report" string is generic enough to be judgment-based and is pre-existing).

🔄 CHANGES REQUESTED

Comment thread skills/openhands-automation/SKILL.md
Address review feedback on #686:

- SKILL.md: the polling section pointed at rule 5, which became rule 4
  when rules 2-4 were collapsed.
- README.md: stop presenting the prompt preset as the recommended default
  and custom automations as advanced-only; list custom scripts alongside
  prompt-based creation and point to the no-LLM worked example.
- Regenerate skills/index.js.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@enyst
enyst requested a review from all-hands-bot October 4, 2026 11:51

@all-hands-bot all-hands-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review was posted by an AI agent (OpenHands).

Summary

The skill-text rebalance meets issue #685's acceptance criteria: deterministic prompt examples are replaced with agent-shaped ones, the "do not use unless explicitly requested" warning is gone from references/custom-automation.md with deterministic tasks now a first-class reason to use the custom path, a complete no-LLM worked example is added, rules 2-4 are collapsed into one symmetric rule, and the duplicated references/security.md bullet is removed. I verified in the workspace that skills/index.js is regenerated (scripts/sync_extensions.py --check passes), the focused suite is green (179 passed), and the referenced anchors (#deterministic-script-no-llm, #sdk-based-scripts, #state-persistence-kv-store) all resolve.

One artifact on the head must not merge.

1. .beads/issues.jsonl is internal task-tracker data added to a public repo (blocking)

The second commit (b588ba6) adds .beads/issues.jsonl with 21 records of a private issue tracker. It is not part of the stated scope (the PR body says "Markdown and generated catalog changes only"), does not exist on main or in the first commit (166b757), and appears nowhere else in the repository. AGENTS.md states the primary artifacts here are Markdown skill definitions and plugin configs, and forbids committing internal data. The records name unreleased internal work and private repositories - for example "Private enyst/automations has metadata-only snapshots...", "export the authenticated enyst Cloud definitions to private enyst/automations main", and "read built-in github_token instead of uploaded ENYST_GH_TOKEN". This repository is public and @openhands/extensions publishes to npm with ./skills/* and ./plugins/* subpaths, so this content would ship. Remove .beads/issues.jsonl from the branch, and add .beads/ to .gitignore if the tracker is used locally. Flagged inline.

2. Non-blocking: the branch is stale against main

The head merges cleanly in content terms, but the branch base predates a large amount of main, which has since reworked the No-LLM Script Helpers (get_secret now branches for OPENHANDS_CLOUD_API_URL/SANDBOX_ID so it works on Cloud). After rebasing, update the new worked example's duplicated helpers to match and regenerate skills/index.js (it currently conflicts).

🔄 CHANGES REQUESTED

Comment thread .beads/issues.jsonl Outdated
enyst and others added 3 commits October 4, 2026 14:06
Resolve the conflict in the generated skills/index.js by regenerating it
with scripts/build-skills-catalog.mjs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…on Cloud

The worked example added by this PR carried the pre-#709 get_secret, which
only reads from the agent server. Match the helper now on main so the
example also fetches secrets from the OpenHands API on Cloud and
Enterprise runs. Regenerate skills/index.js.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
b588ba6 picked up .beads/issues.jsonl through the local bd pre-commit
hook, which stages that file on every commit. It is local task-tracker
data and does not belong in this repository.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@enyst
enyst requested a review from all-hands-bot October 4, 2026 12:09
@OpenHands OpenHands deleted a comment from openhands-ai Bot Oct 4, 2026

@all-hands-bot all-hands-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review was posted by an AI agent (OpenHands).

Summary

The skill-text rebalance now meets issue #685 (enhancement, ready-for-dev, priority:low): the deterministic prompt examples are replaced with agent-shaped ones (Sentry triage, bug reproduction, dependency review, risky-push review), the "do not use unless explicitly requested" warning is gone from references/custom-automation.md with deterministic tasks listed first as a reason to use the custom path, a complete no-LLM worked example is added, rules 2-4 are collapsed into one symmetric rule, the duplicated references/security.md bullet is removed, and README.md no longer frames the prompt preset as the recommended default.

I verified on the current head that scripts/sync_extensions.py --check passes, the focused suite is green (189 passed), the references/custom-automation.md anchors (#deterministic-script-no-llm, #sdk-based-scripts, #state-persistence-kv-store) all resolve, and the new example script compiles. The .beads/issues.jsonl tracker file that an earlier revision added by mistake is no longer on this head, so that blocker is resolved. No code-level defects found.

Two required checks are red on this head and mergeStateStatus is BLOCKED:

1. Validate PR description fails - the PR body does not follow the current template (action needed by the author)

The head's PR body has no HUMAN: or AGENT: markers; the first visible line is - [ ] A human has tested these changes.. The template in .github/pull_request_template.md now requires HUMAN:, a short human-written note of at least 20 characters, then AGENT:, and the check (.github/scripts/check_pr_description.py) fails on all three. This is not a code issue, but it is a repository gate that blocks the merge and only the author can fix it (the human attestation must not be written by an agent). Restoring the HUMAN: / AGENT: markers with a human note clears the check.

2. sync-sdk-skill fails - unrelated to this PR

This check runs scripts/sync_openhands_sdk_skill.py --check against skills/openhands-sdk/SKILL.md, which this PR does not touch (it is not in the PR diff). It is a time-dependent regeneration check that also fails independently of these changes. It does not reflect a defect in this PR, but it does leave the required status rollup red, so it needs to be resolved (regenerate the SDK skill) before merge.

🔄 CHANGES REQUESTED

Pick up the SDK skill sync fix from #725.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@enyst
enyst requested a review from all-hands-bot October 4, 2026 14:18

@all-hands-bot all-hands-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review was posted by an AI agent (OpenHands).

Summary

The skill-text rebalance meets issue #685 (enhancement, ready-for-dev, priority:low). The change is confined to skills/openhands-automation/ and the generated skills/index.js, which is the correct repository and scope for a skill-content fix; no API, schema, or preset behavior changes.

What I verified on this head (3adaead):

  • The deterministic prompt examples are replaced with agent-shaped ones (Sentry triage, bug-ticket reproduction, weekly dependency review, risky-push review); none of the remaining preset examples matches a rule 0 instant-recognition pattern.
  • references/custom-automation.md no longer opens with the "Do NOT use this reference unless the user has explicitly requested" warning, and its "when to use" list now names deterministic/no-LLM tasks first.
  • A complete no-LLM worked example (script, tarball, upload, create) is present, its embedded script compiles (py_compile), and its get_secret/fire_callback helpers match the canonical helpers on main (including the Cloud OPENHANDS_CLOUD_API_URL/SANDBOX_ID branch).
  • Rules 2-4 are collapsed into one symmetric rule, and the (see rule 5) pointer now reads (see rule 4).
  • The duplicated references/security.md bullet is removed (one remains), and README.md no longer marks the prompt preset "(recommended)" or frames custom automation as only for advanced users.
  • All anchors resolve: #deterministic-script-no-llm, #sdk-based-scripts, and #state-persistence-kv-store all exist in references/custom-automation.md.
  • scripts/sync_extensions.py --check passes, node scripts/build-skills-catalog.mjs produces no further diff, and the focused suite is green (189 passed).
  • All required checks on this head are green (sync-extensions, sync-sdk-skill, test, check, validate-claude-code, package, Validate PR description). The .beads/issues.jsonl tracker file that an earlier revision added by mistake is gone, and the PR body now follows the HUMAN:/AGENT: template.

No material bugs, security problems, or design flaws found. The two blocking items from the previous revision (the accidental tracker file and the PR-description gate) are both resolved.

✅ APPROVED

@enyst
enyst merged commit 3bfeaf0 into main Oct 4, 2026
15 checks passed
@enyst
enyst deleted the fix/automation-skill-preset-balance branch October 4, 2026 14:24
@openhands-release-bot openhands-release-bot Bot added the released: v0.28.0 Shipped in v0.28.0 label Oct 5, 2026
@openhands-release-bot

Copy link
Copy Markdown
Contributor

🚀 Released in v0.28.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

released: v0.28.0 Shipped in v0.28.0 type: fix A bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

openhands-automation skill steers agents toward prompt presets even for deterministic tasks

2 participants