experiment-code
Write ML experiment code with iterative improvement. Generate training/evaluation pipelines, debug errors, and optimize results through code reflection. Use when implementing experiments for a research paper.
Install
npx skills add https://github.com/lingzhi227/agent-research-skills --skill experiment-codeSKILL.md
Experiment Code
Generate and iteratively improve ML experiment code for research papers.
Input
$0— Task:generate,improve,debug,plot$1— Research plan, idea description, or error message
References
- Experiment prompts and patterns:
~/.claude/skills/experiment-code/references/experiment-prompts.md - Code patterns (error handling, repair, hill-climbing):
~/.claude/skills/experiment-code/references/code-patterns.md
Action: generate
Generate initial experiment code following this structure:
- Plan experiments first — List all runs needed (hyperparameter sweeps, ablations, baselines)
- Write self-contained code — All code in project directory, no external imports from reference repos
- Include proper logging — Save results to JSON, print intermediate metrics
- Generate figures — At minimum Figure_1.png and Figure_2.png
Mandatory Structure
project/
├── experiment.py # Main experiment script
├── plot.py # Visualization script
├── notes.txt # Experiment descriptions and results
├── run_1/ # Results from run 1
│ └── final_info.json
├── run_2/
└── ...
Constraints
- No placeholder code (
pass,...,raise NotImplementedError) - Must use actual datasets (not toy data unless explicitly requested)
- PyTorch or scikit-learn preferred (no TensorFlow/Keras)
- Each run uses:
python experiment.py --out_dir=run_i
Action: improve
Improve existing experiment code:
- Read current code and results
- Reflect on what worked and what didn't
- Apply targeted edits (prefer small edits over full rewrites)
- Re-run and compare scores
- Keep the best-performing code variant
Action: debug
Fix experiment code errors:
- Read the error message (truncate to last 1500 chars if very long)
- Identify the root cause
- Apply minimal fix
- Up to 4 retry attempts before changing approach
Action: plot
Generate publication-quality plots from experiment results:
- Read all
run_*/final_info.jsonfiles - Generate comparison plots with proper labels
- Use the figure-generation skill for styling
Rules
- Always plan experiments before writing code
- After each run, document results in notes.txt
- Include print statements explaining what results show
- Method MUST not get 0% accuracy — verify accuracy calculations
- Use seeds for reproducibility
- Before each experiment include a print statement explaining exactly what the results are meant to show
Related Skills
- Upstream: experiment-design, algorithm-design
- Downstream: data-analysis, backward-traceability
- See also: code-debugging, paper-to-code
Related skills
researchmattpocock575KInvestigate a question against high-trust primary sources and capture the findings as a Markdown file in the repo. Use when the user wants a topic researched, docs or API facts gathered, or reading legwork delegated to a background agent.paper-context-resolverlllllllama451KRigor Paper Context helper for README-first deep learning repo reproduction. Use only when the README and repository files leave a narrow reproduction-critical gap and the task is to resolve a specific paper detail such as dataset split, preprocessing, evaluation protocol, checkpoint mapping, or runtime assumption from primary paper sources while recording conflicts. Do not use for general paper summary, repo scanning, environment setup, command execution, title-only paper lookup, or replacing Renv-and-assets-bootstraplllllllama450KRigor Setup skill for README-first deep learning repo reproduction. Use when the task is specifically to prepare a conservative conda-first environment, checkpoint and dataset path assumptions, cache location hints, and setup notes before any run on a README-documented repository. Do not use for repo scanning, full orchestration, paper interpretation, final run reporting, or generic environment setup that is not tied to a specific reproduction target.ai-research-explorelllllllama311KRigor Explore compatible skill slug for meaningful and potentially novel deep learning research candidates. Use when the researcher has chosen the task family, dataset, benchmark, evaluation method, provided SOTA references, and wants candidate-only exploration on top of `current_research` with auditable repo understanding, idea gating, fair comparison, and governed experiments written to `explore_outputs/`. Do not use for README-first trusted reproduction, open-ended direction finding, narrow c