Evaluates a model's ability to detect all instances of a person matching a natural language description in an image, including handling multiple instances and correctly rejecting cases where the described person is absent. Use when the user wants to benchmark on HumanRef, or asks about evaluating this task. Reports DensityF1 Score.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill humanref-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Humanref Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-humanref-eval)More formats (shields.io, HTML) on the badges page.
---
name: humanref-eval
description: Evaluates a model's ability to detect all instances of a person matching a natural language description in an image, including handling multiple instances and correctly rejecting cases where the described person is absent. Use when the user wants to benchmark on HumanRef, or asks about evaluating this task. Reports DensityF1 Score.
metadata:
skill_kind: dataset_eval
source_arxiv: 2503.08507
bibtex_key: jiang2025referringtoanyperson
confidence: high
---
# humanref-eval
> Referring to Any Person — Qing Jiang et al. (2025) (arXiv:2503.08507, 2025)
## What this evaluates
Evaluates a model's ability to detect all instances of a person matching a natural language description in an image, including handling multiple instances and correctly rejecting cases where the described person is absent.
## Datasets
- **HumanRef** — total ?; splits: test (-1); repo https://github.com/IDEA-Research/RexSeek
## Metrics
- `Precision` — range: [0, 1]
- Fraction of predicted bounding boxes that correctly match a ground truth instance (IoU > threshold).
- `Recall` — range: [0, 1]
- Fraction of ground truth instances that are successfully matched by at least one predicted bounding box.
- `DensityF1 Score` **(primary)** — range: [0, 1]
- Average F1 score across instances weighted by a density penalty factor: DensityF1 = (1/N) * Σ [2*(P_i*R_i)/(P_i+R_i) * D_i], where D_i = min(1.0, GT_Count_i / Predicted_Count_i). Penalizes over-detection.
- `Rejection Score` — range: percent
- Percentage of referring expressions for which the model predicts zero bounding boxes, measuring hallucination/rejection capability.
## Input / output format
**Input**: An image and a natural language referring expression describing one or more persons.
**Output**: One or more bounding boxes (or a single point for point-output models) indicating the location(s) of the referred person(s).
## Scoring recipe
```python
def compute_metrics(predictions, gold, iou_thresholds=range(0.5, 0.96, 0.05)):
ious = compute_iou(predictions, gold)
matched = ious.max(axis=0) > iou_thresholds
precision = matched.sum() / max(len(predictions), 1)
recall = matched.sum() / max(len(gold), 1)
f1 = 2 * precision * recall / (precision + recall + 1e-8)
density_penalty = min(1.0, len(gold) / max(len(predictions), 1))
density_f1 = f1 * density_penalty
return precision, recall, density_f1
```
## Common pitfalls
- Models trained on single-instance datasets (e.g., RefCOCO) severely drop in recall when referring to multiple persons.
- Point-in-mask evaluation is less strict than IoU-based evaluation, making direct comparisons unfair.
- Models often hallucinate bounding boxes even when the described person is absent, leading to low rejection scores.
## Evidence (verbatim from paper)
> We evaluate the referring task using Precision, Recall, and DensityF1 Score. Given a referring expression, the model predicts one or more bounding boxes, and a prediction is considered correct if its IoU with any ground truth box exceeds a predefined threshold. Following the evaluation protocol in COCO[[37]], we report the average performance across IoU thresholds from 0.5 to 0.95 in increments of 0.05. For the rejection subset, we calculate the number of referring expressions that the model does not predict any boxes and divide it by the number of total expressions.
## Citation
```bibtex
@misc{jiang2025referringtoanyperson,
title={Referring to Any Person},
author={Qing Jiang et al. (2025)},
year={2025},
note={arXiv:2503.08507}
}
```
- arXiv: 2503.08507
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!