data-analysis
Analyze spreadsheet data, generate insights, create visualizations, and build reports from Excel/CSV data.
Install
npx skills add https://github.com/claude-office-skills/skills --skill data-analysisSKILL.md
Data Analysis Assistant
Analyze data in spreadsheets, uncover insights, and create compelling visualizations.
Overview
This skill helps you:
- Understand and explore your data
- Perform statistical analysis
- Generate insights and recommendations
- Create charts and visualizations
- Write formulas and queries
How to Use
Getting Started
- Share your spreadsheet or data file
- Describe what you want to analyze
- Get insights, formulas, or visualizations
Analysis Types
Exploratory Analysis
"What patterns do you see in this data?"
"Give me an overview of this dataset"
"What are the key statistics?"
Specific Questions
"What was the total revenue by region?"
"Which products had the highest growth?"
"Is there a correlation between X and Y?"
Visualization Requests
"Create a chart showing sales trends"
"Make a comparison chart of Q1 vs Q2"
"Show the distribution of customer ages"
Output Formats
Data Overview
## Dataset Overview
**Rows**: 1,234
**Columns**: 15
**Date Range**: Jan 2025 - Dec 2025
### Column Summary
| Column | Type | Non-null | Unique | Sample Values |
|--------|------|----------|--------|---------------|
| date | Date | 100% | 365 | 2025-01-01 |
| revenue | Number | 98% | 890 | $1,234.56 |
| region | Text | 100% | 5 | North, South |
### Data Quality Issues
- [X] rows have missing values in [column]
- [Y] potential duplicates detected
Statistical Analysis
## Statistical Summary
### [Metric Name]
- **Mean**: X
- **Median**: Y
- **Std Dev**: Z
- **Min/Max**: A / B
### Key Findings
1. [Finding with statistical support]
2. [Finding with statistical support]
### Recommendations
- [Action based on analysis]
Insight Report
## Analysis Report: [Topic]
### Executive Summary
[2-3 sentence overview of key findings]
### Key Metrics
| Metric | Value | Change |
|--------|-------|--------|
| Total Revenue | $X | +Y% |
| Avg Order Value | $Z | -W% |
### Trends
1. **[Trend 1]**: [Description with data]
2. **[Trend 2]**: [Description with data]
### Recommendations
1. [Actionable recommendation]
2. [Actionable recommendation]
Common Analysis Workflows
Sales Analysis
1. "Show total sales by month"
2. "Which products are top performers?"
3. "What's the customer segment breakdown?"
4. "Compare this year vs last year"
5. "Forecast next quarter based on trends"
Customer Analysis
1. "What's the customer distribution by segment?"
2. "Calculate customer lifetime value"
3. "Which customers are at risk of churning?"
4. "What's the acquisition cost vs LTV ratio?"
Financial Analysis
1. "Calculate profit margins by product"
2. "What's the expense breakdown?"
3. "Show cash flow trends"
4. "Compare budget vs actual"
Formula Generation
Request Formulas
"Write a formula to calculate year-over-year growth"
"Create a VLOOKUP to match customer data"
"Make a dynamic sum based on criteria"
Formula Output
## Formula: [Purpose]
### Excel/Google Sheets
```excel
=SUMIFS(Sales[Amount], Sales[Region], "North", Sales[Date], ">="&DATE(2025,1,1))
Explanation
SUMIFS: Sums values meeting multiple criteria- First argument: Column to sum
- Subsequent pairs: Criteria column + criteria value
Usage
Place in cell [X] where you want the result.
## Visualization Recommendations
### Choose the Right Chart
| Data Type | Best Chart |
|-----------|------------|
| Trends over time | Line chart |
| Part of whole | Pie/Donut chart |
| Comparison | Bar chart |
| Distribution | Histogram |
| Correlation | Scatter plot |
| Geographic | Map chart |
### Chart Specifications
```markdown
## Recommended Chart: [Type]
**Data Series**:
- X-axis: [Column] (e.g., Date)
- Y-axis: [Column] (e.g., Revenue)
- Series: [Column] (e.g., Region)
**Formatting**:
- Title: "[Descriptive title]"
- Colors: Use consistent color scheme
- Labels: Show values on data points
**Chart Description**:
[What this chart shows and why it's useful]
Advanced Analysis
Pivot Table Design
## Pivot Table: [Purpose]
**Rows**: [Field 1], [Field 2]
**Columns**: [Field 3]
**Values**: SUM of [Field 4], AVG of [Field 5]
**Filters**: [Field 6]
Expected Output:
| Region | Q1 | Q2 | Q3 | Q4 | Total |
|--------|----|----|----|----|-------|
| North | $X | $X | $X | $X | $X |
| South | $X | $X | $X | $X | $X |
Cohort Analysis
## Cohort Analysis
**Cohort Definition**: Customers grouped by [first purchase month]
**Metric**: [Retention rate / Revenue / etc.]
**Time Period**: [12 months]
| Cohort | M0 | M1 | M2 | M3 | ... |
|--------|-----|-----|-----|-----|-----|
| Jan 25 | 100%| 45% | 32% | 28% | ... |
| Feb 25 | 100%| 48% | 35% | 30% | ... |
Best Practices
For Better Analysis
- Clean data first: Handle missing values, duplicates
- Define metrics clearly: What exactly are you measuring?
- Consider context: Industry benchmarks, seasonality
- Validate findings: Cross-check with other data sources
For Better Visualizations
- Keep it simple: One main message per chart
- Label clearly: Title, axes, legend
- Use appropriate scale: Don't truncate misleadingly
- Consider colorblind users: Use patterns or distinct colors
Limitations
- Cannot directly execute code on your data
- Large datasets may need sampling
- Complex statistical models need specialized tools
- Real-time data requires live connections
- Cannot guarantee 100% accuracy on OCR'd data
Related skills
researchmattpocock575KInvestigate a question against high-trust primary sources and capture the findings as a Markdown file in the repo. Use when the user wants a topic researched, docs or API facts gathered, or reading legwork delegated to a background agent.paper-context-resolverlllllllama451KRigor Paper Context helper for README-first deep learning repo reproduction. Use only when the README and repository files leave a narrow reproduction-critical gap and the task is to resolve a specific paper detail such as dataset split, preprocessing, evaluation protocol, checkpoint mapping, or runtime assumption from primary paper sources while recording conflicts. Do not use for general paper summary, repo scanning, environment setup, command execution, title-only paper lookup, or replacing Renv-and-assets-bootstraplllllllama450KRigor Setup skill for README-first deep learning repo reproduction. Use when the task is specifically to prepare a conservative conda-first environment, checkpoint and dataset path assumptions, cache location hints, and setup notes before any run on a README-documented repository. Do not use for repo scanning, full orchestration, paper interpretation, final run reporting, or generic environment setup that is not tied to a specific reproduction target.ai-research-explorelllllllama311KRigor Explore compatible skill slug for meaningful and potentially novel deep learning research candidates. Use when the researcher has chosen the task family, dataset, benchmark, evaluation method, provided SOTA references, and wants candidate-only exploration on top of `current_research` with auditable repo understanding, idea gating, fair comparison, and governed experiments written to `explore_outputs/`. Do not use for README-first trusted reproduction, open-ended direction finding, narrow c