Skip to content

Latest commit

 

History

History
249 lines (172 loc) · 13.6 KB

File metadata and controls

249 lines (172 loc) · 13.6 KB

Testing & Code Quality

Skillware maintains high standards for code quality and reliability. Before submitting a Pull Request, please ensure your code passes all linting and testing checks.

Tests fall into four layers: bundle, framework, maintainer, and example. Use that vocabulary consistently in docs and PRs.

Status

Capability Status
Bundle tests in CI (pytest skills/) Done
Every registry skill ships test_skill.py Done
Issuer enforces bundle tests on new skills Done
Bundle tests mock network and model downloads in CI Done
Maintainer tests under tests/skills/ (optional per skill) Done
[all] extra covers bundle-test runtime deps Done
CLI skillware test for bundle discovery Done
Doc-drift guards (test_registry_docs.py, catalog hubs / sitemap / intent headers, #370) Done
Glossary / retired-anatomy guard (test_registry_docs.py, #252) Done
Registry identity guard (test_registry_identity.py) Done
GitHub label policy test (test_github_labels.py) Done
PyPI wheel packaging smoke test (scripts/wheel_smoke_test.py) Done
Optional extras sync (scripts/sync_extras.py, tests/test_extras_sync.py; includes install_extras.md table guard) Done
Card UI schema vs execute output (tests/test_card_ui_schema.py) Done
Local-execute example smoke tests in CI (tests/test_examples_smoke.py) Done
Framework tests isolated from operator global config (tests/conftest.py, #302) Done

Every pull request runs black --check, flake8, pytest skills/, pytest tests/, and a wheel-smoke job that builds a wheel, installs it in a fresh venv (base install only — no [all] or per-skill extras), and verifies every bundled registry skill is present and loadable. Bundle tests gate merge the same as framework and maintainer tests.

When .github/labels.json changes on main, the Sync GitHub Labels workflow updates label colors and descriptions on the repository automatically — do not edit labels manually in the GitHub UI. tests/test_github_labels.py enforces repo-wide labels, every registry cat: <category> label (shared pastel color, no collision with repo-wide names like security), and alignment with the category dropdown in 01_skill_proposal.yml.

Quick Setup

Install lint tools, pytest, and optional skill runtime deps in one go (matches GitHub Actions CI):

pip install -e ".[dev,all]"

Or use the dev pointer file:

pip install -r requirements.txt

Four test layers

Layer Location Shipped in pip wheel? CI on PR?
Skill bundle test skills/<category>/<skill_name>/test_skill.py Yes Yes
Framework test tests/test_*.py (not under tests/skills/) No (clone only) Yes
Maintainer skill test tests/skills/<category>/test_<name>.py No (clone only) Yes when present
Usage example examples/*.py No Smoke only for local scripts; no for live loops

Skill bundle test (Assurance role)

  • Implements the bundle Assurance role (test_skill.py); see Skill anatomy.
  • Lives inside the skill bundle; ships with pip install skillware.
  • Required for every new registry skill (see templates/python_skill/test_skill.py).
  • Offline and mockable: manifest consistency, validation, deterministic execute() paths — no live network.
  • Run locally: pytest skills/<category>/<skill_name>/test_skill.py or pytest skills/.
  • Install packages from the skill's manifest.yaml requirements when they are not already satisfied by [all].
  • Bundle tests run in CI on every pull request and must not make live HTTP requests, use API keys, or download models.
  • Mock HTTP clients, LLM clients, embedding loaders, and model download paths such as HuggingFace, Ollama, fastembed, and similar integrations.
  • Real inference belongs in maintainer tests under tests/skills/ or in local/manual runs, not in bundle CI gates.

Framework test

  • Core engine health: loader, CLI, issuer rules, version policy, parameter schema validation (tests/test_validate_params.py).
  • tests/test_skill_issuer.py also enforces registry packaging (__init__.py), issuer metadata, presence of test_skill.py in every skill bundle, and rejects legacy manifest output: keys.
  • tests/test_card_ui_schema.py validates output-card ui_schema.fields[].key dot paths against fixtures in tests/fixtures/card_ui_schema/ (#199).
  • tests/test_registry_docs.py enforces doc-drift parity: skill catalog index matches manifests, examples README matches scripts on disk, and agent-loops.md references every registered skill.
  • tests/test_registry_identity.py enforces manifest identity parity: every registry-layout skill's manifest.name matches its path-derived registry ID, and all manifest names are globally unique (#280).
  • tests/test_examples_smoke.py provides an automated regression net for local-execute demo scripts under examples/ without making network requests or requiring API keys (#237).
  • tests/conftest.py isolates every test from the operator's real global config.yaml via SKILLWARE_CONFIG_DIR so local pytest tests/ matches CI even after CLI mail/config init (#302). Legacy vs configured discovery order is covered in tests/test_discovery.py and tests/test_loader.py.
  • Lives at the root of tests/ only (tests/test_loader.py, tests/test_cli.py, …).
  • Clone-repo only; runs in CI via pytest tests/ together with maintainer tests below.

Maintainer skill test

  • Optional extra depth for skill maintainers: loader wiring, heavy mocks, edge cases.
  • Not required for every skill; when present, runs in CI as part of pytest tests/.
  • Example: tests/skills/compliance/test_tos_evaluator.py.

Usage example

  • Runnable provider demos under examples/.
  • Local-execute scripts (e.g., mental_coach_demo.py, prompt_injection_firewall_demo.py, token_limiter_loop.py) are smoke-tested in CI via tests/test_examples_smoke.py to ensure imports and SkillLoader dispatch remain stable.
  • Provider agent loops (Gemini, Claude, OpenAI, DeepSeek, Ollama) require live API keys or local servers and are not run in CI.
  • See examples/README.md.

Which tests go where?

You are testing… Put it here Example in this repo
Manifest + execute contract for one skill Bundle test (Assurance) skills/compliance/tos_evaluator/test_skill.py
Loader path + mocked externals (optional depth) Maintainer test tests/skills/compliance/test_tos_evaluator.py
Loader, CLI, registry issuer rules, param validation, manifest requirement pins, config and path discovery Framework test tests/test_loader.py, tests/test_cli.py, tests/test_config.py, tests/test_discovery.py, tests/test_requirements_check.py, tests/test_skill_issuer.py, tests/test_validate_params.py, tests/test_registry_docs.py, tests/test_skill_docs.py, tests/test_card_ui_schema.py, tests/test_registry_identity.py
End-to-end provider demo script Usage example examples/gemini_tos_evaluator.py

Rule of thumb: if it ships with the skill and must pass before merge → bundle test (CI + local). If it is extra regression depth for clone-repo work → maintainer test (optional). If it proves provider integration → example, not pytest.

Packaging smoke test

Editable clone tests (pytest skills/, pytest tests/) do not prove that a built PyPI wheel ships every registry bundle. CI runs an additional wheel-smoke job (see .github/workflows/ci.yml) that:

  1. Builds a wheel from the PR branch (python -m build --wheel).
  2. Creates a fresh virtualenv and pip installs the wheel (base install only — no [all] or per-skill extras).
  3. Runs scripts/wheel_smoke_test.py, which checks every bundled skill for required bundle files, manifest/name parity, instructions and card assets, and SkillLoader.load_skill(..., check_requirements=False).

Skills whose optional runtime deps are not in the base wheel may be deferred (packaging verified, import skipped until extras are installed). That is expected; bundle tests in editable mode still cover execute paths with mocked deps.

Run locally after building a wheel:

python -m build --wheel --outdir dist/
python -m venv /tmp/wheel-smoke-venv
/tmp/wheel-smoke-venv/bin/pip install dist/skillware-*.whl
/tmp/wheel-smoke-venv/bin/python scripts/wheel_smoke_test.py

On Windows, use Scripts\python and Scripts\pip under the venv path.

1. Code Formatting (Black)

We use Black as our uncompromising code formatter. It ensures that all code looks the same, regardless of who wrote it, eliminating discussions about style.

Installation

pip install black

Usage

Run Black on the entire repository to automatically fix formatting issues:

python -m black .

Run python -m black --check . to verify formatting without writing files. GitHub Actions runs the same check on every pull request before flake8 and pytest; run python -m black . locally to fix issues before you push.

2. Linting (Flake8)

We use Flake8 to catch logic errors, unused imports, and other code quality issues that Black does not handle.

Installation

pip install flake8

Usage

Run Flake8 from the root of the repository:

python -m flake8 .

Note: We aim for zero warnings/errors. Do not suppress errors with # noqa unless absolutely necessary and justified.

3. Unit Tests (Pytest)

We use pytest for automated tests. All new features and bug fixes must be accompanied by relevant tests in the correct layer (see above).

Installation

pip install pytest

CI (GitHub Actions)

GitHub Actions installs pip install -e ".[dev,all]", then runs:

python -m black --check .
python -m flake8 .
python -m pytest skills/
python -m pytest tests/

That covers skill bundle tests under skills/ and framework + maintainer tests under tests/. It does not run examples/. Do not add per-skill pip lines or hardcoded skill paths to .github/workflows/ci.yml.

Pushes to main that touch .github/labels.json also run .github/workflows/sync-labels.yml to upsert GitHub labels from the JSON file.

The [all] extra includes registry skill runtime deps only (web3, fastembed, numpy, …) so pytest skills/ works after pip install -e ".[dev,all]". When a skill adds new manifest.yaml requirements, run python scripts/sync_extras.py to regenerate category, skill, and [all] rows in pyproject.toml, then update the hand-maintained tables in Install extrastests/test_extras_sync.py::test_install_extras_guide_matches_pyproject fails if those tables drift from pyproject.toml.

Local commands

Match CI:

python -m pytest skills/
python -m pytest tests/

Or use the CLI (same bundle paths; requires pip install -e ".[dev]" or [dev,all]):

skillware test
skillware test <category>/<skill_name>
skillware test --category <category>

See CLI reference.

Single skill bundle test:

python -m pytest skills/<category>/<skill_name>/test_skill.py

Optional maintainer depth only:

python -m pytest tests/skills/<category>/test_<skill_name>.py

Pytest is configured to collect from tests/ and skills/ only (examples/ is ignored). See [tool.pytest.ini_options] in pyproject.toml.

Operator global config and pytest (#302)

After normal CLI setup (for example skillware mail signature init), a user-level config.yaml may exist under your Skillware config directory. That switches skill discovery to configured mode (project → external → bundled) instead of legacy mode (SKILLWARE_SKILL_PATH → cwd ./skills/ → bundled).

Framework tests must not depend on your machine's operator config. An autouse fixture in tests/conftest.py points SKILLWARE_CONFIG_DIR at an empty temporary directory for every test run, so pytest tests/ matches CI on a clean home directory.

If you add tests that exercise merged YAML behavior, write explicit project or global config files under tmp_path and call clear_config_cache() after changes — the autouse fixture already isolates the global layer.

Writing tests

  • Bundle test: skills/<category>/<name>/test_skill.py — required for new skills; copy from templates/python_skill/test_skill.py.
  • Maintainer test: tests/skills/<category>/test_<name>.py — optional; use shared fixtures in tests/conftest.py when helpful.
  • Framework test: tests/test_*.py at repo root — for loader, CLI, issuer, and cross-cutting rules.

Pre-Commit Checklist

Before pushing your code, run the following commands:

  1. skillware list (verify install and path resolution)
  2. skillware doctor (optional — check manifest deps and skill.py import readiness)
  3. python -m black --check . (verify formatting; use python -m black . to fix)
  4. python -m flake8 . (check quality)
  5. python -m pytest skills/ or skillware test (bundle tests — same scope as CI)
  6. python -m pytest tests/ (framework + maintainer tests — same scope as CI)
  7. python scripts/sync_extras.py --check (when manifest.yaml or pyproject.toml extras change)
  8. python -m pytest skills/<category>/<skill_name>/test_skill.py or skillware test <category>/<skill_name> for a single skill