Agent Skills

PrivacyScrubber PII Masker

Zero-Trust Data Sanitization (ZTDS) local PII and secrets scrubber for secure LLM pipelines.

Install

Install and configure the MCP from https://github.com/moxno/privacyscrubber-mcp now. Follow the repository's installation instructions, ask me for anything you can't complete yourself, and verify its tools load.
README

@privacyscrubber/mcp-server

CISO-Approved Zero-Trust PII & Secrets Redaction MCP Server for Cursor, Windsurf, and Claude Desktop. Locally scrubs PII, secrets, credentials, and custom regex rules from files and text contexts before they reach remote LLM providers to prevent API leaks and ensure HIPAA/SOC 2 compliance at the developer endpoint.

⭐ Support Zero-Trust Open Source: If PrivacyScrubber protects your API keys and code from leaks, please Star this repository or run gh repo star moxno/privacyscrubber-mcp in your terminal!


πŸ”’ Zero-Trust Data Flow

All sensitive parameters, identifiers, and variables are intercepted locally inside your machine's RAM. They are replaced by tokens (e.g. [EMAIL_1]) before being sent to the AI. Once the AI responds, the tokens are safely swapped back to original values in your local context.

[Raw Input / Files] ──> [MCP sanitize_text] ──> [Masked Tokens] ──> [LLM API]
                               β”‚                                       β”‚
                        (In-Memory Map)                             (Result)
                               β”‚                                       β”‚
[Original Output] <─── [MCP reveal_text] <β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸš€ Installation

1. Install via Smithery

To automatically configure and run with your preferred client, install using Smithery:

npx -y @smithery/cli install @privacyscrubber/mcp-server --write-to-clients

2. Instant Run with NPX

Run the server directly without local installation:

npx -y @privacyscrubber/mcp-server

3. Programmatic Node.js / TypeScript SDK (Lightweight Presidio Alternative)

Need direct, in-memory zero-trust PII sanitization in your backend microservice, Next.js app, or RAG vector pipeline rather than an MCP server? Use our official zero-dependency SDK:

npm install @privacyscrubber/sdk
import OpenAI from 'openai';
import { wrapOpenAI } from '@privacyscrubber/sdk';

// Transparently masks PII before sending to LLM and rehydrates responses:
const openai = wrapOpenAI(new OpenAI({ apiKey: process.env.OPENAI_API_KEY }));

// Outbound prompt is sanitized in local RAM before leaving your machine:
// "Schedule a call with [NAME_1] at [EMAIL_1] regarding API key [AWS_KEY_1]."
const completion = await openai.chat.completions.create({
  model: 'gpt-4o',
  messages: [{ role: 'user', content: 'Schedule a call with Alice Smith at alice@acme.com with AKIAIOSFODNN7EXAMPLE.' }]
});

// Incoming LLM answer is automatically rehydrated with "Alice Smith (alice@acme.com)":
console.log(completion.choices[0].message.content);

⚑ Why @privacyscrubber/sdk vs Microsoft Presidio?

Microsoft Presidio is the Python standard, but deploying it in a Node.js / TypeScript stack requires running heavy Python microservices, Docker containers, and 500MB+ spaCy NLP models with 35–120ms latency. @privacyscrubber/sdk runs 100% in-process with zero dependencies:

Feature / Metric@privacyscrubber/sdkMicrosoft PresidioAWS Comprehend / Google Cloud DLP
Runtime & Dependencies0 dependencies (~150KB)Python + Docker + spaCy (~500MB)Heavy Cloud SDKs
Execution LatencySub-millisecond in RAM (<1ms)35–120ms HTTP/gRPC roundtrip180–400ms external cloud roundtrip
Docker / Sidecar NeededNone (Pure in-process Node/WASM)Mandatory Docker containerVPC endpoints & IAM configurations
Network Egress0 Bytes (100% Air-gapped)Local internal network hopFull unencrypted payload to cloud
Data Loss & Reversibility0 Loss (RAM-reversible via restore())Manual token vault configurationIrreversible masking / hashing
OpenAI / LangChain 1-LinerBuilt-in (wrapOpenAI, wrapAiStream)Complex custom pipeline glueCustom proxy architecture
DevOps Secrets InterceptionBuilt-in (AWS, JWT, DB URIs, GitHub PAT)Custom regex recognizers neededCloud-specific classifiers
Streaming RehydrationBuilt-in (wrapAiStream, TransformStream)Buffering / chunk split failuresNot supported in real-time streams

πŸ‘‰ View @privacyscrubber/sdk on NPM | Full documentation, Express middleware, LangChain transforms, and 25 compliance profiles.

πŸ€– AI Orchestrator Recipes: LangChain, LlamaIndex, CrewAI & AutoGPT

If you are building autonomous agents, RAG vector pipelines, or backend services rather than single-user IDE prompts, use @privacyscrubber/sdk to sanitize data in-memory:

1. LangChain.js (LCEL & RAG Document Ingestion)
import { ChatOpenAI } from '@langchain/openai';
import { createLangChainTransform, createDocumentTransformer } from '@privacyscrubber/sdk';

// A. RAG Pre-Ingestion: sanitize documents before embedding into vector stores
const docTransformer = createDocumentTransformer({
  profile: 'General',
  sensitiveMetadataKeys: ['account_owner', 'submitter_email'],
  attachTelemetry: true
});
const sanitizedDocs = await docTransformer.transformDocuments(rawDocs);

// B. LCEL Runtime Chains: sanitize prompts and restore responses in local RAM
const transform = createLangChainTransform({ defaultProfile: 'Dev' });
const { scrubbedText, tokenMap } = transform.preprocess(
  "Deploying database with user admin and pwd postgresql://user:SecretPass123@db.internal:5432/prod"
);
const model = new ChatOpenAI({ model: 'gpt-4o' });
const response = await model.invoke(scrubbedText);
const finalOutput = transform.postprocess(response.content, tokenMap);
2. LlamaIndex.TS & Universal Vector DB Ingestion (Chroma, Pinecone, Qdrant)
import { createVectorIngestionGuard, createLlamaIndexTransform } from '@privacyscrubber/sdk';

// Initialize in-memory Vector Ingestion Guard (<1ms per batch, zero network egress)
const guard = createVectorIngestionGuard({
  profile: 'Finance',
  sensitiveMetadataKeys: ['contractor_email', 'billing_contact']
});

// A. Sanitize Chroma columnar batches ({ ids, documents, metadatas })
const cleanChroma = guard.sanitizeRecords(chromaBatch);

// B. Sanitize Pinecone / Qdrant record arrays ({ id, text/payload, metadata })
const cleanRecords = guard.sanitizeRecords(pineconeRecords);

// C. Align search query tokens with sanitized vector space before embedding
const { query: alignedQuery } = guard.sanitizeQuery("Search user john@example.com records");

// D. Transparent Vector Store Proxy: auto-sanitizes on write, re-hydrates on query
const guardedStore = guard.wrapVectorStore(nativeVectorStore);

Complete runnable recipes are included inside the npm package under @privacyscrubber/sdk/examples/ (or run npx @privacyscrubber/sdk).

3. CrewAI & Multi-Agent Swarms (Sidecar Mode / Python Interop)
# Run lightweight zero-dependency local daemon (<1ms in-memory, 0 external network calls):
npx @privacyscrubber/sdk sidecar --port 3100
# In your Python CrewAI / AutoGen Agent tools:
import requests

def sanitize_agent_input(prompt: str) -> tuple[str, dict]:
    res = requests.post("http://127.0.0.1:3100/sanitize", json={"text": prompt, "profile": "Dev"}).json()
    return res["scrubbedText"], res["tokenMap"]

def restore_agent_output(ai_response: str, token_map: dict) -> str:
    res = requests.post("http://127.0.0.1:3100/restore", json={"text": ai_response, "tokenMap": token_map}).json()
    return res["restoredText"]
4. AutoGPT / Multi-Turn Autonomous Agents
import { PrivacyScrubberEngine, createGuardedTools } from '@privacyscrubber/sdk';

const engine = new PrivacyScrubberEngine({ defaultProfile: 'Dev' });
// Equips agent with tools that intercept commands, file reads, and git diffs before model exposure:
const guardedTools = createGuardedTools(engine);

🏒 Enterprise & Production Licensing:

  • Self-Serve Developer SDK ($299/mo flat or $2,990/yr): Unlimited internal backend nodes, microservices, and RAG pipelines. Developer SDK Specs & Pricing
  • PrivacyScrubber TEAMS ($99/mo flat): Unlimited team seats, centralized policy enforcement, encrypted session handoff. Deploy TEAMS
  • Enterprise Air-Gapped License: On-premise source code distribution, zero-network custom models. Contact Enterprise

πŸ›‘οΈ Architecture & Security Deep-Dive (Zero-Trust vs Cloud DLP)

When AI IDEs (Cursor, Claude Desktop, Windsurf) connect to model providers, developer credentials, database connection strings, and internal customer PII are at continuous risk of prompt exfiltration. The @privacyscrubber/mcp-server enforces four immutable architectural guarantees:

  1. Stdio Air-Gapped Transport: The MCP server communicates exclusively over local standard input/output (stdio) child processes spawned by your IDE. It opens zero external listening ports and initiates zero remote network requests.
  2. Volatile RAM-Only Token Map: The mapping table between synthetic tokens ([AWS_KEY_1], [EMAIL_1]) and raw cleartext is maintained exclusively in ephemeral node memory and is wiped the moment your IDE session closes.
  3. Deterministic AST Lookarounds vs Cloud Proxy Overhead: Unlike cloud DLP gateways (Nightfall, Skyflow) that add 200–400ms latency and transmit unencrypted code to third parties, @privacyscrubber/mcp-server runs locally in <2ms with zero data egress.
  4. Permanent Standards & Academic Validation:

βš™οΈ Client Integrations

Claude Desktop

Add this to your Claude Desktop config file:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %APPDATA%\Claude\claude_desktop_config.json
{
  "mcpServers": {
    "privacyscrubber": {
      "command": "npx",
      "args": ["-y", "@privacyscrubber/mcp-server"],
      "env": {
        "PRIVACYSCRUBBER_KEY": "YOUR_OPTIONAL_PRO_LICENSE_KEY"
      }
    }
  }
}

Cursor / Windsurf

  1. Navigate to Settings -> Features -> MCP.
  2. Add new MCP server:
    • Name: privacyscrubber
    • Type: command
    • Command: npx -y @privacyscrubber/mcp-server
  3. Optional: Set PRIVACYSCRUBBER_KEY as an environment variable in your system shell.

Cline / Roo Code

Add to cline_mcp_settings.json:

{
  "mcpServers": {
    "privacyscrubber": {
      "command": "npx",
      "args": ["-y", "@privacyscrubber/mcp-server"],
      "env": {
        "PRIVACYSCRUBBER_KEY": "YOUR_OPTIONAL_PRO_LICENSE_KEY"
      }
    }
  }
}

Claude Code CLI

Add directly from your terminal:

claude mcp add privacyscrubber -- npx -y @privacyscrubber/mcp-server

πŸ› οΈ Provided Tools & JSON-RPC Specifications

1. sanitize_text

Redacts PII, secrets, API keys, and credentials from a text block and populates the volatile local replacement mapping.

  • Arguments:
    • text (string, required): The raw content or logs to sanitize.
    • profile (string, optional): Gated industry detection profile (e.g., 'General', 'Dev', 'Medical', 'Legal', 'Compliance'). Defaults to 'General'.
  • JSON-RPC Call Example:
    {
      "method": "tools/call",
      "params": {
        "name": "sanitize_text",
        "arguments": {
          "text": "Contact me at dev-key-1234 or jane.doe@company.com",
          "profile": "General"
        }
      }
    }
    
  • Response Example:
    {
      "content": [
        {
          "type": "text",
          "text": "Contact me at [SECRET_1] or [EMAIL_1]"
        }
      ]
    }
    

2. reveal_text

Detokenizes the AI response back to the original values locally.

  • Arguments:
    • text (string, required): The response from the LLM containing tokenized placeholders.
  • JSON-RPC Call Example:
    {
      "method": "tools/call",
      "params": {
        "name": "reveal_text",
        "arguments": {
          "text": "Please reach out to [EMAIL_1] regarding the update."
        }
      }
    }
    
  • Response Example:
    {
      "content": [
        {
          "type": "text",
          "text": "Please reach out to jane.doe@company.com regarding the update."
        }
      ]
    }
    

3. sanitize_file

Reads a local file, extracts text, sanitizes it, and returns the redacted template for LLM analysis.

  • Supported Formats: Plain text (source code, logs, CSV, JSON, markdown) and Microsoft Word (.docx) documents.
  • Arguments:
    • filePath (string, required): Absolute file path to read and sanitize.
    • profile (string, optional): The industry detection profile.

4. guard_exec (Command Execution Firewall)

Safely executes terminal commands in an isolated child process, masking stdout/stderr PII, database credentials, and API keys in local RAM before passing them to the AI agent. Includes a CISO audit receipt in stderr.

  • Arguments:
    • command (string, required): The shell command to execute (e.g. cat .env, docker logs web, git diff).
    • cwd (string, optional): Working directory.
    • profile (string, optional): Detection profile (defaults to Dev).
    • timeout_ms (number, optional): Timeout in ms (defaults to 15000).

5. guard_read_file (Credential-Masking File Reader)

Reads files (.env, configs, source code, database dumps) and tokenizes all passwords, JWTs, and PII in volatile memory, returning safe redacted content for AI reasoning.

  • Arguments:
    • file_path (string, required): Path to file.
    • profile (string, optional): Detection profile (defaults to Dev).
    • max_lines (number, optional): Line cap for large files (defaults to 500).

6. guard_git_diff (Pre-Commit & Diff Sanitizer)

Inspects staged (--cached) or unstaged repository diffs, redacting any newly introduced secrets or PII in local RAM before AI code review or commit message generation.

  • Arguments:
    • staged (boolean, optional): If true, inspects staged changes (git diff --cached). Defaults to false.
    • cwd (string, optional): Working directory.
    • profile (string, optional): Detection profile (defaults to Dev).

7. guard_apply_patch (Safe Patch Applicator)

Reverses token placeholders ([API_KEY_1], [SECRET_1]) in AI-generated code or text by looking up the local RAM session map, creating a .bak backup, and writing authentic cleartext directly to disk. The remote LLM never sees real secrets.

  • Arguments:
    • file_path (string, required): Path to target file.
    • content (string, required): Content containing tokens to restore on disk.
    • create_backup (boolean, optional): Backup existing file before write (defaults to true).

8. create_agent_rules (1-Click Agent Rule Injection)

Automatically scaffolds CISO-grade Zero-Trust directives into .cursorrules, .windsurfrules, CLAUDE.md, .github/copilot-instructions.md, or .clinerules.

  • Arguments:
    • agent_types (array, optional): ["all"], ["cursor"], ["windsurf"], ["claude_code"], ["copilot"], ["cline"]. Defaults to ["all"].
    • workspace_dir (string, optional): Workspace directory.

9. check_status

Returns a visual dashboard showing your current tier, session request count, active profiles, and upgrade instructions. Use it at any time to check your license status or get setup help.

  • Arguments: (none required)
  • JSON-RPC Call Example:
    {
      "method": "tools/call",
      "params": { "name": "check_status", "arguments": {} }
    }
    
  • Response Example (Free Tier):
    ╔══════════════════════════════════════════════════╗
    β•‘       PrivacyScrubber MCP Server v2.2.4          β•‘
    ╠══════════════════════════════════════════════════╣
    β•‘  πŸ”“ Tier: FREE                                   β•‘
    β•‘  πŸ“Š Session requests: 5                          β•‘
    β•‘  πŸ“ Input size limit: 15,000 characters/request  β•‘
    ╠══════════════════════════════════════════════════╣
    β•‘  🏷️  Profiles: General only β€” PRO unlocks 25 more β•‘
    β•‘  πŸ“‹ Custom rules: πŸ”’ Locked β€” requires PRO       β•‘
    ╠══════════════════════════════════════════════════╣
    β•‘  πŸ’³ Upgrade to PRO β€” $110 Lifetime               β•‘
    β•‘     https://privacyscrubber.com/pricing?utm_source=npm&utm_medium=readme&utm_campaign=mcp_server          β•‘
    ╠══════════════════════════════════════════════════╣
    β•‘  After purchase, add your key to MCP config:     β•‘
    β•‘  "PRIVACYSCRUBBER_KEY": "<your-key-here>"        β•‘
    β•‘  Full setup guide:                               β•‘
    β•‘  https://privacyscrubber.com/pii-mcp/?utm_source=npm&utm_medium=readme&utm_campaign=mcp_server          β•‘
    β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
    

πŸ›‘οΈ Standalone CLI: ps-guard

PrivacyScrubber bundles ps-guard for Unix pipe, pre-commit, and agentic workflows:

# Pipe any output through RAM redaction
cat .env | npx ps-guard --profile dev

# Execute commands through the ZTDS safety wrapper
npx ps-guard -- npm test

# Review git diff with secrets redacted
npx ps-guard --diff --staged

# Generate rules for all AI IDEs (.cursorrules, .windsurfrules, CLAUDE.md, copilot, cline)
npx ps-guard --rules

🌐 Browser Extension & Web Client

Looking for real-time protection directly inside your web browser?

πŸ“„ License & Commercial Upgrade

By default, the server runs under the Free Tier (restricted to 15,000 characters per request and the basic General PII profile). To unlock 30 specialized engineering, medical, legal, and financial PII profiles, as well as team-wide custom rules, you can purchase a commercial license.

Feature Comparison

FeatureFree TierPRO TierTEAMS Tier
Volatile Tokenizationβœ… Yesβœ… Yesβœ… Yes
Standard PII Maskingβœ… Yesβœ… Yesβœ… Yes
Max Character Length15,000 chars♾️ Unlimited♾️ Unlimited
Industry ProfilesGeneral Only30 Profiles30 Profiles
Custom Regex Rules❌ Locked♾️ Unlimited♾️ Unlimited
Team Rules Sync (GPO)❌ No❌ Noβœ… Yes (Shared Link)
Licensing Cost$0$110 Lifetime$99/mo Flat Rate

πŸ‘‰ Acquire a PRO / TEAMS License Key at privacyscrubber.com/pricing


πŸ” After Purchase: Activate PRO in Your MCP Client

After purchasing a PRO license at privacyscrubber.com/pricing, you will receive a license key. Add it to your MCP client config as an environment variable: PRIVACYSCRUBBER_KEY.

Claude Desktop

Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "privacyscrubber": {
      "command": "npx",
      "args": ["-y", "@privacyscrubber/mcp-server"],
      "env": {
        "PRIVACYSCRUBBER_KEY": "YOUR_LICENSE_KEY_HERE"
      }
    }
  }
}

Restart Claude Desktop after saving.

Cursor

  1. Go to Settings β†’ Features β†’ MCP Servers.
  2. Find privacyscrubber and click Edit.
  3. Add the environment variable: PRIVACYSCRUBBER_KEY=YOUR_LICENSE_KEY_HERE.
  4. Restart Cursor.

Alternatively, export it system-wide so all tools pick it up:

# macOS / Linux β€” add to ~/.zshrc or ~/.bashrc
export PRIVACYSCRUBBER_KEY="YOUR_LICENSE_KEY_HERE"

Windsurf

Edit ~/.codeium/windsurf/mcp_config.json:

{
  "mcpServers": {
    "privacyscrubber": {
      "command": "npx",
      "args": ["-y", "@privacyscrubber/mcp-server"],
      "env": {
        "PRIVACYSCRUBBER_KEY": "YOUR_LICENSE_KEY_HERE"
      }
    }
  }
}

Verify Activation

After adding the key, ask your AI agent to call check_status:

Use the check_status tool from PrivacyScrubber MCP

πŸ”— Ecosystem & Production Architecture Guides


πŸ“š Academic Foundations & Regulatory Verification

PrivacyScrubber and the Zero-Trust Data Sanitization (ZTDS) protocol are backed by published scientific, clinical, and legal treatises:

Repository / ArchiveDOI / IdentifierFocus AreaRegulatory & Compliance Scope
IETF Specificationdraft-sibiryakov-ztds-protocolThe ZTDS Protocol for Frontier AI Ingestion (Internet-Draft)Global AI Privacy, Zero-Egress Architecture
Zenodo / CERN10.5281/zenodo.22058770Zero-Trust Data Sanitization (ZTDS) Protocol FoundationCross-Border AI Privacy, ISO 27001 A.8.11
OSF (Center for Open Science)10.17605/OSF.IO/5BYJFEmpirical Latency Benchmark & Memory Profiling (<2ms RAM)Performance vs Cloud DLP Proxies
SSRN / ElsevierSSRN ID: 7335581Enterprise Generative AI GovernanceEU AI Act, UK GDPR, US State Privacy
medRxiv (Cold Spring Harbor)MEDRXIV/2026/361661Multi-Center Clinical Trial De-IdentificationHIPAA Safe Harbor Section 164.514(b)
Law Archive / OSFLawArchive ID: 4wc86Preserving Attorney-Client Privilege in AI WorkflowsABA Model Rules & Legal Ethics

Citing PrivacyScrubber in Research & Audits

@software{sibiryakov2026privacyscrubber,
  author = {Sibiryakov, Ilya},
  title = {PrivacyScrubber: Zero-Trust Data Sanitization (ZTDS) Engine & MCP Server},
  year = {2026},
  publisher = {Zenodo},
  doi = {10.5281/zenodo.22058770},
  url = {https://github.com/moxno/privacyscrubber-mcp}
}

πŸ“„ License

MIT Β© Ilya Sibiryakov (BrandMeWeb)


βš–οΈ Intellectual Property & Virtual Patent Marking

The Zero-Trust Data Sanitization (ZTDS) architecture, in-memory deterministic tokenization, cryptographic session handoff, and stdio execution methods implemented in this package are proprietary technology of Ilya Sibiryakov (BrandMeWeb) and are protected under Patent Pending status:

  • Patent Office: State of Israel Ministry of Justice, Patent Office (ILPO)
  • Application Number: 331905 (Tracking ID: 94221)
  • Filing / Priority Date: September 14, 2026 (Paris Convention Art. 4 & 35 U.S.C. Β§ 119 Priority)
  • Official Title: SYSTEM AND METHOD FOR CLIENT-SIDE ZERO-TRUST DATA SANITIZATION AND CRYPTOGRAPHIC SESSION HANDOFF IN ARTIFICIAL INTELLIGENCE WORKFLOWS
  • Virtual Patent Marking: privacyscrubber.com/patents/ in accordance with 35 U.S.C. Β§ 287(a).

🌐 Internet Engineering Task Force (IETF) Specification

Search skills and MCP servers

Search across 31,816 skills and MCPs