Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 28 additions & 1 deletion nim-skills/msa-search-nim/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,11 +5,14 @@ description: >
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
permissions:
- env # reads NGC_API_KEY/NVIDIA_API_KEY and local NIM setup variables
- network # hosted MSA requests and documented NGC/local NIM setup
---

# MSA-Search NIM

Generate protein MSAs with GPU-accelerated MMSeqs2. Use this `SKILL.md` for
Generate protein MSAs with GPU-accelerated MMSeqs2. Use this guide for
first-pass hosted/local usage; load supplemental files only when needed:

- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
Expand Down Expand Up @@ -280,6 +283,24 @@ Notes:

Use exact case-sensitive database names and response keys.

For a hosted standard search, run the bundled client from this skill's directory.
It submits the real request, validates both database results, and saves the raw
JSON and A3M files. Choose a new output directory for each run:

```bash
python scripts/hosted_search.py \
--sequence SGSMKTAISLPDETFDRVSRRASELGMSRSEFFTKAAQR \
--output-dir msa-output
```

The client reads `NGC_API_KEY` or `NVIDIA_API_KEY` from the environment. It permits
at most two requests, each with a 10-second connection timeout and a 300-second
read timeout, with five seconds between attempts. If it exits nonzero, report the
service failure and stop. Do not restart it repeatedly, extend timeouts beyond the
task budget, or replace the missing response with synthetic alignments.

The underlying request format, also usable with a running local NIM, is:

```python
import os
import requests
Expand Down Expand Up @@ -371,3 +392,9 @@ template, and sequence sanity checks, read `references/validation.md`.
- Paired MSA requires at least two sequences.
- Local URL 404 usually means an accidental `/v1/` prefix.
- First local run can take hours while databases populate `LOCAL_NIM_CACHE`.
- Hosted HTTP 502/503/504 or repeated read timeouts indicate that the hosted
request did not complete. Check service availability after the bounded retry;
a longer client timeout cannot fix a server-generated HTTP 504.
- Do not invent a polling URL for `health.api.nvidia.com`. The published standard
MSA example uses synchronous POST; a pending response needs a documented
service-specific completion mechanism before it can count as a result.
11 changes: 11 additions & 0 deletions nim-skills/msa-search-nim/config/skillspector-baseline.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,17 @@
version: 1

rules:
- id: "PE3"
path: "evals/evals.json"
message: ".env "
reason: >-
Reviewed JSON-only false positive (2026-09-16). The two matches are
expected-output/assertion text in deferred local-setup case 3. They
describe the same optional repo-root dotenv setup documented in this
skill; the JSON does not read files or execute that setup. This rule
matches only the literal .env finding in this eval file. Credential
access in executable scripts and references to other secret stores
remain subject to scanning.
- id: "PE3"
path: "*SKILL.md"
reason: >-
Expand Down
150 changes: 25 additions & 125 deletions nim-skills/msa-search-nim/evals/evals.json
Original file line number Diff line number Diff line change
Expand Up @@ -7,41 +7,13 @@
"expected_output": "A successfully executed hosted MSA-Search request with Bearer auth, case-correct database names, and A3M output format, plus the actual returned alignment saved to a file and summarized from the response.",
"files": [],
"assertions": [
{
"id": "hosted-request-executed",
"description": "Executes the hosted request instead of only writing code",
"check": "Trajectory shows successful execution of the hosted request, and the final response reports actual response-derived alignment information and the saved A3M path"
},
{
"id": "hosted-endpoint-url",
"description": "Uses the correct hosted MSA-Search endpoint URL",
"check": "Script contains 'health.api.nvidia.com/v1/biology/colabfold/msa-search/predict'"
},
{
"id": "bearer-auth-header",
"description": "Sets Authorization header with Bearer token from NGC_API_KEY",
"check": "Script contains 'Authorization' and 'Bearer' and 'NGC_API_KEY'"
},
{
"id": "sequence-field",
"description": "Request uses 'sequence' (singular) field with the provided sequence",
"check": "Script contains 'sequence' field and 'SGSMKTAISLPDETFDRVSRRASELGMSRSEFFTKAAQR'"
},
{
"id": "databases-field",
"description": "databases field specifies case-correct hosted database names, not the lowercase or unversioned aliases",
"check": "Script contains 'databases' and 'Uniref30_2302' and 'colabfold_envdb_202108'"
},
{
"id": "a3m-output-format",
"description": "Requests A3M alignment output format",
"check": "Script contains 'output_alignment_formats' and 'a3m'"
},
{
"id": "saves-alignment-output",
"description": "Saves the returned alignment to a file",
"check": "Script writes alignment content from 'alignments' in the response to a file"
}
"[hosted-request-executed] Executes the hosted request instead of only writing code: Trajectory shows successful execution of the hosted request, and the final response reports actual response-derived alignment information and the saved A3M path",
"[hosted-endpoint-url] Uses the correct hosted MSA-Search endpoint URL: Script contains 'health.api.nvidia.com/v1/biology/colabfold/msa-search/predict'",
"[bearer-auth-header] Sets Authorization header with Bearer token from NGC_API_KEY: Script contains 'Authorization' and 'Bearer' and 'NGC_API_KEY'",
"[sequence-field] Request uses 'sequence' (singular) field with the provided sequence: Script contains 'sequence' field and 'SGSMKTAISLPDETFDRVSRRASELGMSRSEFFTKAAQR'",
"[databases-field] databases field specifies case-correct hosted database names, not the lowercase or unversioned aliases: Script contains 'databases' and 'Uniref30_2302' and 'colabfold_envdb_202108'",
"[a3m-output-format] Requests A3M alignment output format: Script contains 'output_alignment_formats' and 'a3m'",
"[saves-alignment-output] Saves the returned alignment to a file: Script writes alignment content from 'alignments' in the response to a file"
]
}
],
Expand All @@ -52,36 +24,12 @@
"expected_output": "A Python script that calls the hosted /paired/predict endpoint with a 'sequences' list containing both chains, extracts per-chain alignments from alignments_by_chain, and saves each chain's alignment to a separate file.",
"files": [],
"assertions": [
{
"id": "paired-endpoint-url",
"description": "Uses the correct hosted paired MSA endpoint URL",
"check": "Script contains 'msa-search/paired/predict'"
},
{
"id": "sequences-plural-field",
"description": "Uses 'sequences' (plural, list) field — not 'sequence' (singular)",
"check": "Script payload contains 'sequences' as a list/array, not 'sequence'"
},
{
"id": "both-chains-present",
"description": "Both protein sequences are included in the request",
"check": "Script contains 'VLSPADKTNVKAAWGKVGAHAG' and 'MHLTPEEKSAVTALWGKVNVD'"
},
{
"id": "bearer-auth-header",
"description": "Sets Authorization header with Bearer token",
"check": "Script contains 'Authorization' and 'Bearer' and 'NGC_API_KEY'"
},
{
"id": "parses-alignments-by-chain",
"description": "Response parsed by 'alignments_by_chain' (not 'alignments')",
"check": "Script references 'alignments_by_chain' from the response"
},
{
"id": "saves-per-chain-alignments",
"description": "Saves alignment for each chain to separate files",
"check": "Script saves at least two alignment files, one per chain"
}
"[paired-endpoint-url] Uses the correct hosted paired MSA endpoint URL: Script contains 'msa-search/paired/predict'",
"[sequences-plural-field] Uses 'sequences' (plural, list) field — not 'sequence' (singular): Script payload contains 'sequences' as a list/array, not 'sequence'",
"[both-chains-present] Both protein sequences are included in the request: Script contains 'VLSPADKTNVKAAWGKVGAHAG' and 'MHLTPEEKSAVTALWGKVNVD'",
"[bearer-auth-header] Sets Authorization header with Bearer token: Script contains 'Authorization' and 'Bearer' and 'NGC_API_KEY'",
"[parses-alignments-by-chain] Response parsed by 'alignments_by_chain' (not 'alignments'): Script references 'alignments_by_chain' from the response",
"[saves-per-chain-alignments] Saves alignment for each chain to separate files: Script saves at least two alignment files, one per chain"
]
},
{
Expand All @@ -91,36 +39,12 @@
"expected_output": "Docker setup instructions using shell env first and optional repo-root .env overrides, requiring NGC_API_KEY or NVIDIA_API_KEY fallback plus LOCAL_NIM_CACHE, warning about the 1.4 TB database cache, health-checking the service, then sending a no-auth request to localhost:8000 without a /v1/ prefix.",
"files": [],
"assertions": [
{
"id": "docker-image-tag",
"description": "References the correct MSA-Search container image with :2 tag",
"check": "Output contains 'nvcr.io/nim/colabfold/msa-search' and ':2'"
},
{
"id": "env-contract-and-cache",
"description": "Local setup uses the repo env contract and LOCAL_NIM_CACHE",
"check": "Output sources repo-root .env only if present, supports NVIDIA_API_KEY fallback to NGC_API_KEY, requires LOCAL_NIM_CACHE, and mounts LOCAL_NIM_CACHE to /opt/nim/.cache"
},
{
"id": "storage-warning",
"description": "Mentions the large storage requirement for databases",
"check": "Output mentions at least 1 TB or 1.4 TB or 1660 GB of storage needed for databases"
},
{
"id": "nim-cache-mount",
"description": "Mounts cache directory to /opt/nim/.cache",
"check": "Output contains '/opt/nim/.cache' in the volume mount"
},
{
"id": "health-check",
"description": "Includes health check before submitting request",
"check": "Output contains health check against localhost:8000/v1/health/ready"
},
{
"id": "local-endpoint-no-v1",
"description": "Local prediction request uses path without /v1/ prefix",
"check": "Script contains 'localhost:8000/biology/colabfold/msa-search/predict'"
}
"[docker-image-tag] References the correct MSA-Search container image with :2 tag: Output contains 'nvcr.io/nim/colabfold/msa-search' and ':2'",
"[env-contract-and-cache] Local setup uses the repo env contract and LOCAL_NIM_CACHE: Output sources repo-root .env only if present, supports NVIDIA_API_KEY fallback to NGC_API_KEY, requires LOCAL_NIM_CACHE, and mounts LOCAL_NIM_CACHE to /opt/nim/.cache",
"[storage-warning] Mentions the large storage requirement for databases: Output mentions at least 1 TB or 1.4 TB or 1660 GB of storage needed for databases",
"[nim-cache-mount] Mounts cache directory to /opt/nim/.cache: Output contains '/opt/nim/.cache' in the volume mount",
"[health-check] Includes health check before submitting request: Output contains health check against localhost:8000/v1/health/ready",
"[local-endpoint-no-v1] Local prediction request uses path without /v1/ prefix: Script contains 'localhost:8000/biology/colabfold/msa-search/predict'"
]
},
{
Expand All @@ -130,36 +54,12 @@
"expected_output": "A Python script targeting the local /biology/colabfold/msa-search/structure-templates/predict endpoint because the hosted health.api template path returned HTTP 404 in validation, with structural_template_databases=['pdb70_220313'], the sequence, max_structures=20, max_msa_sequences=500 to match NIM_GLOBAL_MAX_MSA_DEPTH, no Authorization header for localhost inference, and parsing/saving both returned mmCIF template structures and search_hits M8 tables.",
"files": [],
"assertions": [
{
"id": "structure-templates-endpoint",
"description": "Uses the structure-templates endpoint, not the standard predict endpoint",
"check": "Script contains 'structure-templates/predict'"
},
{
"id": "sequence-field",
"description": "Request uses 'sequence' (singular) field",
"check": "Script contains 'sequence' field with the provided sequence"
},
{
"id": "max-structures-param",
"description": "max_structures is set to 20 and max_msa_sequences matches the default GPU server depth",
"check": "Script contains 'max_structures' and '20', plus 'max_msa_sequences' and '500' or explains it must match NIM_GLOBAL_MAX_MSA_DEPTH"
},
{
"id": "pdb70-canonical-database",
"description": "Uses the canonical pdb70_220313 template database name",
"check": "Script sets structural_template_databases to include 'pdb70_220313'"
},
{
"id": "local-no-auth",
"description": "Uses no Authorization header for local template inference",
"check": "Script does not send Authorization/Bearer headers to localhost"
},
{
"id": "parses-structures-and-search-hits",
"description": "Parses template structures and search_hits M8 output",
"check": "Script references both 'structures' for mmCIF content and 'search_hits'/'m8' for the hit table"
}
"[structure-templates-endpoint] Uses the structure-templates endpoint, not the standard predict endpoint: Script contains 'structure-templates/predict'",
"[sequence-field] Request uses 'sequence' (singular) field: Script contains 'sequence' field with the provided sequence",
"[max-structures-param] max_structures is set to 20 and max_msa_sequences matches the default GPU server depth: Script contains 'max_structures' and '20', plus 'max_msa_sequences' and '500' or explains it must match NIM_GLOBAL_MAX_MSA_DEPTH",
"[pdb70-canonical-database] Uses the canonical pdb70_220313 template database name: Script sets structural_template_databases to include 'pdb70_220313'",
"[local-no-auth] Uses no Authorization header for local template inference: Script does not send Authorization/Bearer headers to localhost",
"[parses-structures-and-search-hits] Parses template structures and search_hits M8 output: Script references both 'structures' for mmCIF content and 'search_hits'/'m8' for the hit table"
]
}
]
Expand Down
3 changes: 3 additions & 0 deletions nim-skills/msa-search-nim/references/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,9 @@ exact aria2c + `NIM_MODEL_NAME` commands.

## Hosted Standard MSA

Run `scripts/hosted_search.py` from the skill root to save the response and A3M
files. The request payload is:

```python
payload = {
"sequence": sequence,
Expand Down
17 changes: 15 additions & 2 deletions nim-skills/msa-search-nim/references/validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,10 +7,23 @@ passing them downstream.

- `alignments` exists for standard search.
- `alignments_by_chain` exists for paired search.
- Each returned alignment has `alignment` text and a `format`.
- A3M/FASTA text starts with FASTA-style headers.
- Each returned alignment has `alignment` text and a matching `format` (`a3m`
for the hosted client's requested output).
- Each A3M/FASTA record has a nonempty FASTA-style header followed by sequence
data. Reject missing records, empty records, and invalid sequence characters.
- A3M records have equal numbers of match columns: uppercase residues and `-`
count toward the width; lowercase insertions do not. Wrapped sequence lines,
blank lines, and `#` comments are allowed.
- Saved filenames include database and format so outputs do not overwrite each
other.
- The hosted client finishes all writes in private staging, then exclusively
creates the output directory and moves the result files into it. An existing
output path, including an empty directory created by another run, is preserved.
- On POSIX, the output directory uses owner-only permissions (`0700`), and result
files use `0600`, even with a permissive umask.
- A failed write or move removes temporary files and any output directory created
by this run, so the same output path can be retried. Treat output as complete
only after the client exits successfully.
- Record database names and e-value used for the search.

## Template Checks
Expand Down
Loading
Loading