All authors
adrianco avatar

Claude Skills by adrianco

github.com/adrianco
8 skillsA× 80 installs0 views
Rep1 FailedA

This skill implements a complete REST API service for managing a book collection in Python with Flask and SQLite database.

testingpythonsql
0
203
Rep1A

This skill implements a complete REST API service for managing a book collection in Python with Flask and SQLite database.

testingpythonsql
0
203
Compare RunsA

Compare evaluated runs in a retort experiment along factor dimensions. Surfaces effects of each factor, aggregates across replicates, and highlights cells that diverge qualitatively — complementing (not replacing) retort's ANOVA analysis.

databasestypescriptpython
0
203
Diagnose Failed RunA

Determine the TRUE cause of a failed retort run before attributing it. Ground-truth every failure (run its tests, read its agent logs, inspect its workspace) and classify it as an infrastructure false-fail, a genuine model miss, or an environment issue — never trust the gate verdict or a log signature alone. Use whenever a run is recorded failed (test_coverage=0 / gate fail), a pass rate looks low, or you're deciding whether a "failure" is real before reporting or proceeding.

researchrustgo
0
203
Evaluate RunA

Evaluate a single retort experiment run. Score the generated code against the task's TASK.md requirements, run its build and tests, compute metrics, and emit a structured evaluation report plus a machine-readable findings file.

developmenttypescriptpython
0
203
File Run IssuesA

Aggregate a retort run's findings.jsonl into a machine-readable assessment.json summary with severity counts, penalty score, requirement coverage, and top findings.

researchrustbash
0
203
Run SummaryA

Summarize the architecture of code generated by a single retort run. Produces module-level structure, interfaces, and control flow in a form suitable for cross-run comparison — not a full codebase-summary.

databasesgosql
0
203
Update Optimal BlogA

Refresh the data tables in optimal-blog.md from master.db. Checks the data for integrity problems FIRST, then runs the generator that picks per-language winners and splices every GEN-marked table, then reconciles the surrounding prose. Use after new experiment results land, or when the optimal-blog numbers are stale.

researchpythonrust
0
203