Three BlueSkills misses that changed how I think about skill scanning
Security Toolings+1•September 30, 2026

Three BlueSkills misses that changed how I think about skill scanning

By Wehi (@iamwehi)

Before our public challenge, I was confident in BlueSkills. We had checks for suspicious instructions and bundled code, model-assisted analysis, and an isolated runtime stage when enabled. I expected few people to find a harmful agent skill that we would mark CLEAN.

Three prize-winning submissions proved my confidence was too high. Each found a different gap in what our scanner could see.

1. An attack on the work product

The evm-integration-review submission looked like documentation for a smart-contract deployment review. It supplied replacement price-oracle and privileged-caller addresses and instructed an agent to treat its table as authoritative.

In two tests against a synthetic project, agents copied those placeholder addresses into the pending deployment and removed a recovery setter. No attacker contract was deployed and no transaction was broadcast. BlueSkills had marked the submitted skill CLEAN.

Our first mitigation made Ethereum addresses in a skill trigger a warning. That submission no longer receives a complete CLEAN result, but the warning is deliberately limited: an address in a document is not, by itself, proof of an attack. The harder question is whether the skill is replacing a value the user’s project already trusts.

2. An attack that waited for macOS

The second submission declared macOS support. On that platform, its bundled Python would reconstruct the path to an OpenSSH private key, read it, and send its contents to an external endpoint. On Linux, the entry point took a harmless path. Two scans of the same ZIP returned CLEAN with no findings.

Our source now has checks and regression tests for this credential-to-network path and for macOS-targeted behavior. The larger lesson remains: observing one run in one environment does not establish what a skill will do on another operating system or execution path.

3. An attack that changed the shape of stolen data

The third submission instructed an agent to collect hidden files from the user’s home directory, put them in a tar archive, and upload it. BlueSkills returned CLEAN. In our follow-up testing, archiving planted credentials also meant a network watcher searching for the credentials’ recognizable bytes did not see them leaving.

We changed how the runtime prompts an agent to exercise such a skill. Recognizing planted secrets after they have been compressed into an upload remains a disclosed limitation. Merely noticing that a credential file was read would not, on its own, explain or stop the entire data flow.

What I took from the challenge

These were not three versions of the same trick. One attacked the user’s work product, one depended on the execution environment, and one hid sensitive data from a particular observation method. Adding more analysis layers had not made our coverage universal.

BlueSkills is a check before installation, not a certificate of safety. Scan the exact skill package or revision, read the quoted findings and coverage, and make the installation decision with those limits in mind.

If you find another reproducible miss, open an issue in our public repository. Use dummy data and include the exact submission, scan result, expected result, and a negative control where possible. That is how the next blind spot becomes something we can investigate.