How the scan works
SkillsVetted downloads a skill's files from GitHub at a specific commit and checks every line against a set of rules for known risky patterns. Nothing is executed.
The grade
Every skill starts at 100. Each finding subtracts points by severity: critical 35, high 15, medium 6, low 2. Only the two most severe hits of any single rule count, so one noisy pattern can't sink a score on its own.
| Grade | Score | Meaning |
|---|---|---|
| A | 90–100 | No significant risks found |
| B | 75–89 | Minor patterns worth a glance |
| C | 55–74 | Review the findings before installing |
| D | 35–54 | Several risky patterns |
| F | 0–34, or any critical finding | Do not install without a careful review |
Findings inside Markdown files are usually documentation, so most rules drop one severity level there. Rules aimed at the AI itself, like prompt injection and hidden characters, keep full severity everywhere, because Markdown is exactly where those attacks live.
What we check
Hidden characters
Zero-width spaces, bidirectional overrides and Unicode tag characters that hide text from people but not from models.
Prompt injection
- High Tries to override earlier instructions. Text telling the AI to ignore or replace its existing instructions is a classic prompt-injection pattern.
- High Asks the AI to hide actions from the user. Instructions to keep actions secret from the user are a strong sign of malicious intent.
- Medium Asks the AI to skip permission or confirmation. Skills should not instruct the agent to bypass approval prompts or act without asking.
- Medium Attempts to redefine the AI's role or safety rules. Phrases like 'you are now' or 'developer mode' are used to jailbreak models out of their normal behavior.
Secrets & credentials
- High Reads sensitive credential files. References to SSH keys, cloud credentials or keychains. A skill rarely has a legitimate reason to touch these.
- Critical Accesses browser cookies, passwords or wallets. Paths to browser profile data or crypto wallets are typical of credential-stealing malware.
- Medium Reads secret environment variables. Reads variables named like API keys, tokens or passwords, or dumps the whole environment. Check where the values go.
- Medium Reads .env files. .env files usually contain API keys and passwords.
Network access
- High Contacts a host commonly used for data exfiltration. Webhook relays, paste sites and tunnels are frequently used to send stolen data out.
- Medium Sends data to a remote server. Code that POSTs or uploads data. Check what is being sent and where.
- Critical Downloads and runs remote code. Piping a download straight into a shell or interpreter runs code that was never reviewed.
Code execution
- Low Runs shell commands. Many legitimate skills run commands. Check that the commands match what the skill says it does.
- Medium Evaluates code dynamically. eval/exec on strings makes it impossible to know what will run by reading the code.
Obfuscation
- Critical Decodes hidden data and executes it. Decoding a base64/hex blob and running it is a standard way to hide malicious code.
- Medium Contains a long encoded blob. Very long base64 or hex strings can hide code or data. Legitimate uses include embedded images or fonts.
Destructive commands
- High Deletes files recursively at a dangerous path. Recursive deletion of the home directory, root or wildcard paths can wipe data.
- Critical Formats or overwrites disks. Commands like mkfs or dd to a device can destroy entire drives.
- Medium Weakens permissions or escalates privileges. sudo or world-writable permissions are rarely needed by a skill.
Persistence
- High Modifies startup files or schedules tasks. Writing to shell profiles, cron, LaunchAgents or systemd lets code keep running after the skill finishes.
- High Changes AI agent settings or hooks. Editing agent settings, hooks or MCP config can silently grant broader permissions in future sessions.
Install scripts
- Medium Installs packages from unpinned or remote sources. Installing from git URLs or arbitrary indexes can pull in unreviewed code.
Structure and maintenance
- Binary or compiled files, which can't be reviewed by reading them.
- npm install hooks that run code automatically.
- Missing SKILL.md or frontmatter, archived repositories, no license, and no updates in over a year.
What a scan can't tell you
Pattern matching finds known techniques. It can't judge intent, follow logic across files, or see what a script downloads at runtime. A clean result lowers risk; it does not prove a skill is safe. For anything that will touch sensitive data, get a full human review.