Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food. Use when the user wants to benchmark on CUB-200-2011, NABirds, iNatbirds, Food-101, FoodX-251, Food-2K, or asks about evaluating this task. Reports Top-1 k-NN accuracy.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill pas-dataset-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Pas Dataset Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-pas-dataset-eval)More formats (shields.io, HTML) on the badges page.
---
name: pas-dataset-eval
description: Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food. Use when the user wants to benchmark on CUB-200-2011, NABirds, iNatbirds, Food-101, FoodX-251, Food-2K, or asks about evaluating this task. Reports Top-1 k-NN accuracy.
metadata:
skill_kind: dataset_eval
source_arxiv: 2407.03463
bibtex_key: rodriguez2024precisionatscale
confidence: high
---
# pas-dataset-eval
> Precision at Scale: Domain-Specific Datasets On-Demand — Jesús M Rodríguez-de-Vera et al. (2024) (arXiv:2407.03463, 2024)
## What this evaluates
Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food.
## Datasets
- **CUB-200-2011** — total ?; splits: test (-1)
- **NABirds** — total ?; splits: test (-1)
- **iNatbirds** — total ?; splits: test (-1)
- **Food-101** — total ?; splits: test (-1)
- **FoodX-251** — total ?; splits: test (-1)
- **Food-2K** — total ?; splits: test (-1)
## Metrics
- `Top-1 k-NN accuracy` **(primary)** — range: percent
- Accuracy of a k-nearest neighbors classifier on frozen backbone features. Predicts the majority class among the k closest training examples in feature space.
- `Linear accuracy` — range: percent
- Accuracy of a linear probe trained on frozen backbone features. A linear classifier is trained on the extracted features and evaluated on the test set.
- `mAP` — range: percent
- Mean Average Precision for object detection tasks, measuring the precision-recall curve across confidence thresholds.
## Input / output format
**Input**: RGB images from the target domain (birds or food) along with their ground-truth class labels or bounding boxes.
**Output**: Class predictions (for classification) or bounding box coordinates and class scores (for object detection).
## Scoring recipe
```python
def evaluate_knn(features, labels, k=5):
dists = np.linalg.norm(test_features[:, None] - features[None, :], axis=2)
topk = np.argsort(dists, axis=1)[:, :k]
preds = [np.bincount(labels[idx]).argmax() for idx in topk]
return np.mean(np.array(preds) == test_labels)
def evaluate_linear(features, labels, epochs=50):
clf = LogisticRegression(max_iter=epochs)
clf.fit(features, labels)
return clf.score(test_features, test_labels)
```
## Common pitfalls
- The paper evaluates both k-NN and linear probing on frozen features, but does not specify the value of k for k-NN or the exact optimizer/learning rate for the linear probe in the main text (referenced in supplementary).
- Data leakage is mitigated by filtering images resembling test sets, but the exact similarity threshold (0.45) and duplicate removal method (SSCD) must be carefully applied to avoid inflating downstream performance.
- Pretraining uses MoCo v3 for ViTs and NNCLR for ResNets with different epoch counts (300 vs 500); mixing these setups without adhering to the specified protocols will yield non-comparable results.
## Evidence (verbatim from paper)
> We report Top-1 k-NN and Linear accuracies for all datasets. As can be seen in Table [1], PaS datasets prove to be better pretrainers even at the same scale.
## Citation
```bibtex
@misc{rodriguez2024precisionatscale,
title={Precision at Scale: Domain-Specific Datasets On-Demand},
author={Jesús M Rodríguez-de-Vera et al. (2024)},
year={2024},
note={arXiv:2407.03463}
}
```
- arXiv: 2407.03463
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!