← skillsmith scanner · all guides
skillsmith is an open-source security scanner for Claude Agent Skills
(SKILL.md files). Paste a skill or point it at a GitHub URL and get
a risk verdict with plain-language findings in seconds. There is also a REST
API, an MCP server, a behavioral sandbox, and a Python CLI
(skillsmith-scanner on PyPI).
Yes. Sign-in is free and includes 5 scans/day. Paid tiers exist for heavier use (Pro: 100 scans/day; pay-per-use single scans from $0.02). Payments are handled via USDC on Solana — no credit card required. Scanning one file by hand will never cost anything.
Yes. skillsmith is a static heuristic scanner. It does not execute the skill. A clean verdict means no known suspicious pattern matched — it is not a guarantee of safety. Likewise, flagged findings can be false positives. Always read short files yourself before running them; that is why every finding ships with a human-readable explanation.
The security score is a number from 0 to 100 where higher means safer. It is derived from the risk score (total weight of all detected findings) using the formula:
security_score = max(0, 100 − risk_score × 4)
Each finding has a weight (3–10). A skill with no findings scores 100 (clean). A skill with 25 cumulative risk points scores 0 (high risk). The score is shown on scan results and in the trends graph to help you track whether a skill is getting safer or riskier over time.
You can see your skill's score history by clicking the sparkline on the scan results page.
What is the behavioral sandbox?The behavioral sandbox goes beyond pattern matching: an AI analyst inside an isolated container simulates what an agent following the skill would actually do — step-by-step actions, capabilities, indicators of compromise, and a 0–10 severity score. It catches attacks that never match any string pattern, such as paraphrased exfiltration instructions. The skill is never executed against real systems.
Often yes. Every scanned skill gets a "Skill-DNA" fingerprint (simhash).
Near-duplicate variants with tiny cosmetic changes stay within Hamming
distance ≤12 of known scans and appear via
/api/similar. Rug-pull watching (/api/watch) covers
the other trick: content that changes after you vetted it.
A scan stores only the hash, verdict metadata and — if you opt in with "publish" — the full content. Private scans keep your skill text private; near-duplicate search masks names of unpublished skills. Clean-scanned skills appear in the public safe-skills registry and feed without their source code unless published.
Add the official GitHub Action (or call the API directly) to lint and scan
skills in your repository on every push — see the
API docs and the
repository README.
The CLI works offline: pip install skillsmith-scanner, then
skillsmith scan ./SKILL.md.
The static engine checks prompt-injection phrasing including paraphrases,
roleplay jailbreaks (including restriction-free personas), guideline overrides (including qualified forms like
"ignore the safety guidelines"; linter-style usage stays clean), prompt-extraction attempts (phrasings that try to make
the agent reveal its system prompt or hidden rules) and text hidden in the
description frontmatter, credential and
secret snooping (in code and in prose instructions), hardcoded secret material (embedded PEM private keys, AWS/GitHub/Google/Slack API key literals; documentation examples stay clean), exfiltration instructions, concealment of agent operations (hidden
logging, invisible actions, fake success reports), dangerous code patterns
(eval, pickle, raw sockets, destructive shell commands), and URLs
that carry credential-looking query parameters. It also normalizes common
obfuscation: zero-width characters (both as word separators and hidden inside
words), RTL/bidi direction overrides, fullwidth characters, combining marks,
Cyrillic and Greek homoglyph look-alikes, invisible control, format, and space-like characters, and base64-encoded payloads (including
UTF-16) are decoded, normalized, or folded before scanning — hiding
instructions in reversed, disguised, or encoded text does not work.
Note on languages: the phrase-based detection is anchored on English wording. Injection instructions written entirely in other languages may not be flagged by the static engine; the behavioral sandbox is the language-independent layer. We document this limit openly rather than claim full multilingual coverage.
For behavior that only becomes clear at runtime, use the behavioral sandbox.
Yes — MIT-licensed, including pattern categories inspired by NVIDIA's SkillSpector research (Apache-2.0, credited in-file). Audit the scanner itself at GitHub; we eat our own dog food by scanning every contributed change.
The text field accepts up to 100,000 characters (about 100 KB).
For larger files, use the url field with a
github.com blob link
or a raw.githubusercontent.com URL;
the scanner fetches the file server-side and applies the same 200 KB fetch cap.
Both modes produce identical risk verdicts.
Behavioral sandbox has the same 100 KB limit on submitted text; URL-mode is not supported there (the sandbox reads the file inside an isolated container, not the scanner process).
Each account has daily limits that reset at 00:00 UTC:
/api/public_scan: 200 lookups/day/IP (soft cap)/api/report): 20 reports/day/keyPay-per-use single scans ($0.02) add to your scan quota on top of the tier limits. Credits never expire.
Yes. When you re-scan the same SKILL.md content (same sha256 hash), the
result includes a trend field with:
direction (improved / declined /
unchanged), delta (score difference vs. the
previous scan), previous_security_score, and a
history array of the last 10 scores.
The web UI shows this as an arrow indicator (↑ improved, ↓ declined,
→ unchanged) and a mini sparkline of the score history. The
explanation field translates each raw pattern hit into
plain-language advice.
Please open an issue at GitHub Issues — include the SKILL.md (redacted if needed) and the reported verdict. False positives directly improve the engine weights.
This page documents the detection logic, threat model, and design decisions behind skillsmith. If you are evaluating the scanner for a high-stakes deployment, the answers below should help you understand exactly what the tool does and does not catch.
Every scan runs through four detection stages, in order: frontmatter parse (YAML safe_load, with any parse error flagged), body normalize (NFC, strip zero-width/bidi, fold homoglyphs), base64 decode pass (with UTF-16 fallback, then re-scan), and pattern match against four weighted lists (prompt-injection, code-pattern, paraphrase, dropper). The final risk score is the sum of matched weights, bucketed into clean / low / medium / high / critical.
skillsmith is designed to defend against: (a) an attacker who publishes a malicious skill hoping agents will install it; (b) a maintainer whose previously-clean skill is compromised via supply-chain attack; (c) a benign-looking skill that contains a hidden second-stage prompt in a comment, frontmatter field, or base64 blob. The scanner is not designed to defend against a malicious agent — for that, you need runtime sandboxing.
Every pattern has a weight, and verdict thresholds are tuned to minimize false negatives on known-bad samples while keeping false positive rates acceptable for typical skill content. A low-weight pattern alone will not push a clean skill into "high" — you need multiple matches or one high-weight match. The trade-off is intentional: better to ask a human to look at a borderline result than to miss a real attack.