Frequently Asked Questions

skillsmith.ch · Claude Agent Skill security

What is skillsmith?

skillsmith is an open-source security scanner for Claude Agent Skills (SKILL.md files). Paste a skill or point it at a GitHub URL and get a risk verdict with plain-language findings in seconds. There is also a REST API, an MCP server, a behavioral sandbox, and a Python CLI (skillsmith-scanner on PyPI).

Is skillsmith free?

Yes. Sign-in is free and includes 5 scans/day. Paid tiers exist for heavier use (Pro: 100 scans/day; pay-per-use single scans from $0.02). Payments are handled via USDC on Solana — no credit card required. Scanning one file by hand will never cost anything.

Can a "clean" result be wrong?

Yes. skillsmith is a static heuristic scanner. It does not execute the skill. A clean verdict means no known suspicious pattern matched — it is not a guarantee of safety. Likewise, flagged findings can be false positives. Always read short files yourself before running them; that is why every finding ships with a human-readable explanation.

How is the security score calculated?

The security score is a number from 0 to 100 where higher means safer. It is derived from the risk score (total weight of all detected findings) using the formula:

security_score = max(0, 100 − risk_score × 4)

Each finding has a weight (3–10). A skill with no findings scores 100 (clean). A skill with 25 cumulative risk points scores 0 (high risk). The score is shown on scan results and in the trends graph to help you track whether a skill is getting safer or riskier over time.

You can see your skill's score history by clicking the sparkline on the scan results page.

What is the behavioral sandbox?

The behavioral sandbox goes beyond pattern matching: an AI analyst inside an isolated container simulates what an agent following the skill would actually do — step-by-step actions, capabilities, indicators of compromise, and a 0–10 severity score. It catches attacks that never match any string pattern, such as paraphrased exfiltration instructions. The skill is never executed against real systems.

Someone renamed a malicious skill — can you detect that?

Often yes. Every scanned skill gets a "Skill-DNA" fingerprint (simhash). Near-duplicate variants with tiny cosmetic changes stay within Hamming distance ≤12 of known scans and appear via /api/similar. Rug-pull watching (/api/watch) covers the other trick: content that changes after you vetted it.

Do you store my skills publicly?

A scan stores only the hash, verdict metadata and — if you opt in with "publish" — the full content. Private scans keep your skill text private; near-duplicate search masks names of unpublished skills. Clean-scanned skills appear in the public safe-skills registry and feed without their source code unless published.

How do I use skillsmith in CI?

Add the official GitHub Action (or call the API directly) to lint and scan skills in your repository on every push — see the API docs and the repository README. The CLI works offline: pip install skillsmith-scanner, then skillsmith scan ./SKILL.md.

What kinds of attacks does the scanner detect?

The static engine checks prompt-injection phrasing including paraphrases, roleplay jailbreaks (including restriction-free personas), guideline overrides (including qualified forms like "ignore the safety guidelines"; linter-style usage stays clean), prompt-extraction attempts (phrasings that try to make the agent reveal its system prompt or hidden rules) and text hidden in the description frontmatter, credential and secret snooping (in code and in prose instructions), hardcoded secret material (embedded PEM private keys, AWS/GitHub/Google/Slack API key literals; documentation examples stay clean), exfiltration instructions, concealment of agent operations (hidden logging, invisible actions, fake success reports), dangerous code patterns (eval, pickle, raw sockets, destructive shell commands), and URLs that carry credential-looking query parameters. It also normalizes common obfuscation: zero-width characters (both as word separators and hidden inside words), RTL/bidi direction overrides, fullwidth characters, combining marks, Cyrillic and Greek homoglyph look-alikes, invisible control, format, and space-like characters, and base64-encoded payloads (including UTF-16) are decoded, normalized, or folded before scanning — hiding instructions in reversed, disguised, or encoded text does not work.

Note on languages: the phrase-based detection is anchored on English wording. Injection instructions written entirely in other languages may not be flagged by the static engine; the behavioral sandbox is the language-independent layer. We document this limit openly rather than claim full multilingual coverage.

For behavior that only becomes clear at runtime, use the behavioral sandbox.

Is the detection engine open source?

Yes — MIT-licensed, including pattern categories inspired by NVIDIA's SkillSpector research (Apache-2.0, credited in-file). Audit the scanner itself at GitHub; we eat our own dog food by scanning every contributed change.

Is there a size limit on skills I can scan?

The text field accepts up to 100,000 characters (about 100 KB). For larger files, use the url field with a github.com blob link or a raw.githubusercontent.com URL; the scanner fetches the file server-side and applies the same 200 KB fetch cap. Both modes produce identical risk verdicts.

Behavioral sandbox has the same 100 KB limit on submitted text; URL-mode is not supported there (the sandbox reads the file inside an isolated container, not the scanner process).

What are the rate limits?

Each account has daily limits that reset at 00:00 UTC:

Pay-per-use single scans ($0.02) add to your scan quota on top of the tier limits. Credits never expire.

Does skillsmith track changes over time?

Yes. When you re-scan the same SKILL.md content (same sha256 hash), the result includes a trend field with: direction (improved / declined / unchanged), delta (score difference vs. the previous scan), previous_security_score, and a history array of the last 10 scores.

The web UI shows this as an arrow indicator (↑ improved, ↓ declined, → unchanged) and a mini sparkline of the score history. The explanation field translates each raw pattern hit into plain-language advice.

I found a bug or a false positive.

Please open an issue at GitHub Issues — include the SKILL.md (redacted if needed) and the reported verdict. False positives directly improve the engine weights.

Behind the Scanner: How skillsmith Works

This page documents the detection logic, threat model, and design decisions behind skillsmith. If you are evaluating the scanner for a high-stakes deployment, the answers below should help you understand exactly what the tool does and does not catch.

Detection Pipeline

Every scan runs through four detection stages, in order: frontmatter parse (YAML safe_load, with any parse error flagged), body normalize (NFC, strip zero-width/bidi, fold homoglyphs), base64 decode pass (with UTF-16 fallback, then re-scan), and pattern match against four weighted lists (prompt-injection, code-pattern, paraphrase, dropper). The final risk score is the sum of matched weights, bucketed into clean / low / medium / high / critical.

Threat Model

skillsmith is designed to defend against: (a) an attacker who publishes a malicious skill hoping agents will install it; (b) a maintainer whose previously-clean skill is compromised via supply-chain attack; (c) a benign-looking skill that contains a hidden second-stage prompt in a comment, frontmatter field, or base64 blob. The scanner is not designed to defend against a malicious agent — for that, you need runtime sandboxing.

What the Scanner Does Not Catch

Calibration and False Positives

Every pattern has a weight, and verdict thresholds are tuned to minimize false negatives on known-bad samples while keeping false positive rates acceptable for typical skill content. A low-weight pattern alone will not push a clean skill into "high" — you need multiple matches or one high-weight match. The trade-off is intentional: better to ask a human to look at a borderline result than to miss a real attack.