Agent Skills

misata

Synthetic data that hits the numbers you declare, exactly. Multi-table with verified foreign-key integrity, deterministic, no model in the data path. Python + MCP server. In simple terms, a powerful demo data generator for sales/demos/seed data.

Install

uvx misata
README.md
Misata

Misata

The relational synthetic data engine that satisfies exact business outcomes.

Generate complete multi-table databases with foreign keys that resolve, math that reconciles, and domain-authentic prose with zero Lorem Ipsum. From a single sentence, YAML schema, or live database.

PyPI version Python versions CI License Open in Colab Paper smithery badge Misata Studio

Prefer a visual interface? Try Misata Studio to design schemas on an interactive canvas and generate datasets directly in your browser.


The Paradigm Shift

Most synthetic data tools take existing data and imitate it. But in modern engineering, you rarely have clean data to start withβ€”or you need data designed around a specific target outcome:

  • "Monthly revenue rises from $50k to $200k with a Q3 slump."
  • "Fraud rate starts at 1% in Q1 and climbs to 6% by Q4."
  • "Every customer's total_spent strictly equals the sum of their order line items."

Misata works in reverse: you declare the outcome, and Misata solves for the individual rows that hit it to $0.00 error, guaranteed by a closed-form Gamma conditional-sum mechanism (arXiv:2606.08736). No machine learning models, no real data required, and zero hallucinated foreign keys.


⚑ 30-Second Quickstart

Install via pip:

pip install misata

CLI

Generate a complete relational dataset from a single description:

misata generate \
  --story "Brazilian fintech with R$ payments, CPF verification, and 3% fraud" \
  --rows 1000 \
  --output-dir ./demo_data

Outputs clean CSVs and an oracle_report.json verifying zero foreign key orphans, constraint satisfaction, and statistical fidelity.

Python

import misata

# Generate multi-table DataFrames from plain English
data = misata.generate("SaaS startup with 500 users, monthly subscriptions, and 12% churn")

users_df = data["users"]
subscriptions_df = data["subscriptions"]

Enrich Existing Data in 1 Line

Replace boring or blank columns in an existing DataFrame with domain-authentic prose:

import misata
import pandas as pd

df = pd.read_csv("support_cases.csv")
# Automatically detects semantic columns (subjects, resolution notes, memos, error traces)
df_enriched = misata.enrich_text(df, seed=42)

πŸš€ The 6 Unique Capabilities of Misata

What separates Misata from legacy libraries like Faker and imitation models like SDV:

1. Exact Outcome Conformance ($0.00 Error)

Declare the aggregate outcome (revenue curve, churn rate, seasonal surge, default curve), and Misata generates micro-rows whose monthly or annual sums match your target to $0.00 error. While off-the-shelf synthesizers miss aggregate targets by 74–86%, Misata hits them provably (read the research paper).

2. Instant Sandbox Oracle for AI Coding Agents

Misata includes a native Model Context Protocol (MCP) server. Cursor, Claude Code, and Windsurf can call create_sandbox to spin up an isolated SQLite database seeded with realistic relational data in 2 seconds. AI agents can test SQL queries and test code against real tables instead of guessing schema.

pip install "misata[mcp]" && misata mcp install --client all

3. Topological DAG Relational Integrity (0 Orphan Foreign Keys)

Parent-child relationships across 10+ tables are resolved via topological rank ordering. Self-referential hierarchies, composite unique constraints, and temporal causality (signup_date <= order_date <= payment_date <= refund_date) are strictly enforced.

4. Cross-Vertical Text Realism (Zero Lorem Ipsum)

No Latin placeholder gibberish. Combinatorial microtext generators provide authentic text across 15+ real-world industries:

  • E-Commerce: Category-conditioned product descriptions (electronics, apparel, home, beauty), authentic return reasons, delivery notes.
  • B2B SaaS: Support ticket subjects, multi-line issue descriptions, agent resolution notes, competitor churn reasons.
  • FinTech: Bank statement descriptors, ACH/wire remittance memos, AML audit overrides.
  • Healthcare: Clinical SOAP progress notes, chief complaints, discharge instructions, medical dosage schedules (TID, QID, PRN).
  • Customer Reviews: Sentiment calibrated provably to 1-to-5 star ratings.

5. Vectorized Speed (100x–500x Faster Than Faker)

While Faker iterates row-by-row in pure Python (5,000–15,000 rows/s), Misata uses vectorized NumPy array operations. It generates 500,000 to 16,000,000 rows/second on a single CPU core.

6. Introspect & Seed Live Databases (misata seed)

Point Misata directly at a PostgreSQL, MySQL, or SQLite database. Misata introspects foreign keys, check constraints, enums, and column types, generates a topological insertion plan, and seeds production-like data directly back into your database without writing schema code.


🎯 Dozens of Real-World Use Cases

Misata is built for engineers, testers, data teams, and founders across dozens of everyday workloads:

πŸ› οΈ Data Engineering & ETL Pipelines

  • Known-Answer Pipeline Testing: Declare exact KPI targets (e.g. $1.2M Q4 revenue), generate synthetic raw tables, and verify your dbt, Spark, or SQL transforms return the exact expected figure.
  • High-Throughput Stress Testing: Churn 10,000,000+ rows in seconds to test partition boundaries, shuffle performance, and warehouse scaling.
  • Schema Migration Dry Runs: Rehearse destructive column backfills and foreign key additions against realistic data before deploying to production.
  • Deterministic CI/CD Fixtures: Reproducible seeds ensure test assertions never flake across continuous integration runs.

πŸ€– AI Coding Agents & LLM Development

  • Agent SQL Sandboxes: Provide Cursor, Claude Code, and Windsurf with isolated, pre-seeded databases to validate SQL queries without production access.
  • Text-to-SQL Benchmark Generation: Build complex relational schema benches with diverse joins to evaluate fine-tuned coding models.
  • Evaluation Database Packs (Evalpacks): Create verified eval datasets with independent DuckDB answer keys where the ground truth cannot be wrong.

πŸ—„οΈ Application Development & Database Seeding

  • Local & Staging Environment Seeding: Seed staging Postgres, MySQL, or SQLite databases with realistic customer histories and 0 broken foreign keys (misata seed).
  • ORM Model Fixtures: Generate rich fixtures matching your Prisma schema or SQLAlchemy declarative models.
  • Multi-Tenant Isolation Verification: Test row-level security (RLS) and tenant isolation rules without cross-tenant key leakage.
  • Incremental Data Growth (generate_diff): Add 5,000 new rows to an existing dataset while auto-offsetting IDs and preserving referential integrity.

πŸ“Š BI, Product Demos & Sales Engineering

  • Board-Ready Dashboard Demos: Populate Tableau, PowerBI, and Metabase dashboards with convincing seasonality (Black Friday spikes, summer dips) rather than flat random noise.
  • Sales Engineering Prototypes: Demo customer-facing analytics with authentic company names, human names, and transaction histories with zero PII exposure.
  • Feature Previews: Preview upcoming charts, cohorts, and metrics before production customer data accumulates.

πŸ’³ FinTech, Banking & Payments

  • Double-Entry Ledger Balancing: Generate accounting transactions where total debits strictly equal credits across every ledger account.
  • Credit Risk & Loan Tapes: Calibrate delinquency curves, credit score distributions, and default rates for credit portfolio testing.
  • AML & Fraud Detection Testing: Inforce exact fraud incidence rates (e.g. 2.4%) with authentic transaction memos and AML audit trails.
  • Payment Remittance: Simulate SWIFT, ACH, and card transactions with valid routing numbers, CVVs, and statement descriptors.

πŸ₯ Healthcare & Clinical Informatics

  • HIPAA Safe-Harbor Synthetic Cohorts: Generate realistic patient populations, vital signs, and encounter histories with zero PHI liability.
  • Clinical NLP Model Evaluation: Evaluate healthcare LLMs against authentic SOAP notes, chief complaints, and discharge summaries.
  • Ward & Scheduling Simulation: Simulate hospital appointment grids with realistic 15-minute intervals, business hours, and weekend dips.

πŸ“¦ E-Commerce & Supply Chain Logistics

  • Multi-Category Catalog Modeling: Produce realistic item specs and descriptions conditioned on category (electronics, apparel, home, industrial).
  • Return & Refund Workflows: Simulate return logistics with authentic return reasons, restocking milestones, and customer refund dates.
  • Route & Fleet Optimization: Compute realistic routes with Haversine distance calculations and valid secondary addresses (Apt, Suite, Bldg).

πŸ›‘οΈ Cybersecurity & IT Infrastructure

  • Network Intrusion Datasets: Generate netflow logs, port scans, and DDoS traffic patterns for security tool benchmarking.
  • System Exception & Error Analysis: Populate observability dashboards with realistic deadlocks, HTTP 504 timeouts, and connection pool exhaustion logs.
  • Compliance Audit Logging: Simulate SOC2/HIPAA access logs with documented managerial access override justifications.

πŸ”¬ Machine Learning & Statistical Research

  • Synthetic Twins from CSV (misata.mimic): Clone distributions and correlations from sensitive CSVs without copying a single original row.
  • Hierarchical Cluster Modeling (ICC): Generate multi-site data with specified Intraclass Correlation Coefficients for mixed-effects regression.
  • Time-Series Autocorrelation (AR1): Generate longitudinal entity trajectories that maintain realistic temporal memory.

⚑ Why You Should Never Use Faker Again

Faker was built over a decade ago for single-attribute mock values. For modern applications, it introduces critical failure modes:

Problem in 2026 Faker Reality Misata 0.9.6.60 Advantage
Relational Topology βœ— 0 concept of databases or FKs; manual glue code required βœ“ Strict topological DAG; 0 orphan FKs guaranteed
Cross-Column Coherence βœ— Incoherent (e.g. "Male" name, mismatched email, invalid city) βœ“ Coherent identities, addresses, and causality
Text Realism βœ— 2,000-year-old Latin "Lorem Ipsum" or robotic templates βœ“ 15+ domain microtext pools (SOAP notes, tickets, memos)
Mathematical Consistency βœ— Violates basic accounting (price * qty != total) βœ“ Exact mathematical formulas and balanced ledgers
Performance βœ— ~10k rows/s (single-threaded Python loops) βœ“ 500k to 16M rows/s (Vectorized NumPy engine)
Aggregate Targets βœ— Impossible (uniform random noise) βœ“ Exact closed-form outcome conformance ($0.00 error)
Database Seeding βœ— Manual SQL scripts or ORM boilerplate βœ“ One-command introspection and seeding (misata seed)

Read the complete Faker vs SDV vs Misata Guide for full benchmarks and code comparisons.


πŸ› οΈ Eight Ways to Generate Data

Misata fits whatever workflow you already use:

Input Mode Best For Learn More
1. Plain English Story Rapid prototyping, zero configuration Story Guide
2. YAML Schema-as-Code Committing versioned data definitions to git YAML Guide
3. Live Database Seeding Introspecting and populating Postgres, MySQL, SQLite Database Seeding Guide
4. Python Dict Schema Programmatic in-memory generation in Python scripts Dict Schema Guide
5. dbt Project Schemas Generating fixtures directly from schema.yml dbt Seeding Guide
6. Prisma Schema Next.js and Node.js developers seeding full-stack apps Prisma Guide
7. Multi-Provider LLMs Groq, OpenAI, Claude, Gemini, or Ollama-driven schemas LLM Guide
8. Incremental Growth Appending rows with offset IDs and preserved FKs Incremental Guide

🌐 20+ Built-in Industry Domains

Generate domain-complete schemas with tuned statistical distributions out of the box:

SaaS Β· E-Commerce Β· FinTech Β· Healthcare Β· Logistics Β· Credit Risk Β· HR & People Β· Streaming Media Β· Insurance Β· CRM & Sales Β· Food Delivery Β· Travel & Hospitality Β· Gaming Β· Crypto & DeFi Β· Predictive Maintenance Β· Network Intrusion Β· Islamic Finance Β· EdTech Β· Real Estate Β· Contact Centers Β· Manufacturing SPC

See the Complete Domain Catalog.


⚑ Performance

Measured on standard Apple M-series hardware (single CPU core, no GPU):

Workload Row Count Generation Time Throughput
Single table (lognormal distribution) 1,000,000 0.06 s ~16M rows/s
Star schema (5 tables, 4 FK dependencies) 1,055,030 1.54 s ~687k rows/s
Multi-table enterprise database 100,000 0.42 s ~240k rows/s

πŸ“š Documentation Index

For in-depth guides, API references, and architecture deep dives:


πŸ“„ Research & Citation

The closed-form exact-outcome conformance engine is formalised in arXiv preprint 2606.08736:

@article{rasin2026declarative,
  title   = {Declarative Outcome-Conformant Synthesis: Exact, Closed-Form
             Specification Satisfaction and a Conformance Benchmark},
  author  = {Rasin, Muhammed},
  year    = {2026},
  url     = {https://arxiv.org/abs/2606.08736v1}
}

🀝 Contributing & Development Status

Misata is currently under massive, rapid development to push synthetic realism to its absolute limit: expanding real-world domain knowledge, deepening seed pool vocabularies, elevating textual column realism, and advancing statistical fidelity across every industry vertical.

Contributions from domain experts, data engineers, and researchers are warmly welcomed! Whether you want to:

  • Enrich Text & Vocabulary Pools: Add authentic seeds and grammar rules for specialized domains in misata/vocab_seeds.py and misata/microtext.py.
  • Contribute a Domain Capsule: Expand built-in industry templates (healthcare, legal, banking, engineering, supply chain).
  • Advance Statistical Fidelity: Improve multi-variate copulas, time-series dynamics, or outcome-curve solvers.
  • Report Edge Cases & Realism Flaws: Open an issue or discussion whenever generated values don't look 100% human-authentic.
git clone https://github.com/rasinmuhammed/misata
cd Misata
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest -q

Misata is open-source under the MIT License.

Search skills and MCP servers

Fuzzy search across 23,137 skills and servers