Apply hugging face datasets in reproducible local data workflows with version-aware APIs, explicit assumptions, and validation. Use when the user chooses hugging face datasets or its strengths fit the task.
Pro scans all 2 files and shows the line behind each finding
Scanned 9/26/2026
npx -y skills add sandbaseai/sandbase-skills --skill huggingface-datasets --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Huggingface Datasets?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/sandbaseai-huggingface-datasets)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: huggingface-datasets
description: "Apply hugging face datasets in reproducible local data workflows with version-aware APIs, explicit assumptions, and validation. Use when the user chooses hugging face datasets or its strengths fit the task."
---
# Hugging Face Datasets
Use this Skill to produce a bounded, verifiable Hugging Face Datasets outcome. Preserve the user's chosen stack, source material, and authorization boundaries.
Read [the SandBase API map](references/sandbase-api-map.md) only when the task genuinely needs an external data source or generative model.
## Workflow
1. Inspect the available files, runtime, versions, inputs, and existing conventions before deciding what to change.
2. Restate the requested outcome, constraints, acceptance checks, and any assumption that could change the result.
3. Produce the smallest complete implementation, analysis, or artifact that satisfies those checks.
4. Verify the real output with appropriate tests, previews, calculations, or source comparison; do not infer success from file creation alone.
5. Return the deliverable, evidence of validation, material assumptions, and unresolved limitations.
## Quality gates
- Inspect shapes, types, units, missing values, sampling, target leakage, and train/test boundaries before modeling or transformation.
- Pin or record relevant library versions, random seeds, parameters, and environment assumptions for reproducibility.
- Validate against a baseline or independent calculation and report diagnostics, uncertainty, failure modes, and resource use.
## Focus checks
- Verify dataset card, license, configuration, split, feature schema, revision, streaming and cache behavior, and avoid assuming viewer samples represent the full dataset.
## SandBase boundary
Keep the core Hugging Face Datasets work local. Use SandBase only for an explicitly requested external dataset or model inference step that is not part of the local analysis.
1. Call `sandbase_discover` with a short capability query.
2. Call `sandbase_inspect` for viable candidates and compare the live schema, coverage, limits, output, execution mode, and price.
3. Prefer a dedicated tool or API the user already has. Send only the minimum necessary data.
4. Before any paid call, show the endpoint, important arguments, current unit price, call count, and total estimate or uncertainty, then obtain confirmation.
5. Use `sandbase_account` before an approved multi-call batch and call `sandbase_run` only with current schema-defined arguments.
6. Poll asynchronous work with `sandbase_run_get` using the same run ID; never resubmit merely because it is pending.
7. Use `sandbase_runs` only to recover status or reconcile observed cost.
If SandBase is unavailable, continue with local work and authorized sources when possible. Do not silently switch providers, fabricate external results, or claim a generation or retrieval succeeded.
## Handoff
Provide the completed artifact or findings, concise reproduction steps, checks actually run, source or asset provenance, SandBase endpoint and run IDs when used, observed cost when available, and any follow-up that still requires user action.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!