Evaluates the fairness of federated learning models across heterogeneous client distributions. It probes whether server-level aggregation masks persistent unfairness at the individual client level by measuring demographic disparity and equal opportunity difference on bias-heterogeneous tabular datasets. Use when the user wants to benchmark on attribute-silo, value-silo, attribute-device, value-device, or asks about evaluating this task. Reports DD.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill feda4fair-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Feda4fair Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-feda4fair-eval)More formats (shields.io, HTML) on the badges page.
---
name: feda4fair-eval
description: Evaluates the fairness of federated learning models across heterogeneous client distributions. It probes whether server-level aggregation masks persistent unfairness at the individual client level by measuring demographic disparity and equal opportunity difference on bias-heterogeneous tabular datasets. Use when the user wants to benchmark on attribute-silo, value-silo, attribute-device, value-device, or asks about evaluating this task. Reports DD.
metadata:
skill_kind: dataset_eval
source_arxiv: 2506.21095
bibtex_key: heilmann2025feda4fair
confidence: high
---
# feda4fair-eval
> FeDa4Fair: Client-Level Federated Datasets for Fairness Evaluation — Heilmann et al. (2025) (arXiv:2506.21095, 2025)
## What this evaluates
Evaluates the fairness of federated learning models across heterogeneous client distributions. It probes whether server-level aggregation masks persistent unfairness at the individual client level by measuring demographic disparity and equal opportunity difference on bias-heterogeneous tabular datasets.
## Datasets
- **attribute-silo** — total ?; splits: clients (51); repo https://anonymous.4open.science/r/FeDa4Fair-AEB3/
- **value-silo** — total ?; splits: clients (51); repo https://anonymous.4open.science/r/FeDa4Fair-AEB3/
- **attribute-device** — total ?; splits: clients (100); repo https://anonymous.4open.science/r/FeDa4Fair-AEB3/
- **value-device** — total ?; splits: clients (111); repo https://anonymous.4open.science/r/FeDa4Fair-AEB3/
## Metrics
- `DD` **(primary)** — range: [0, 1]
- Demographic Disparity (DD) measures the absolute difference in positive prediction rates across sensitive attribute groups: max_g P(Ŷ=1|A=g) - min_g P(Ŷ=1|A=g). A dataset is considered biased if both LR and XGBoost models exhibit the maximum DD for the same attribute/value and this value exceeds 0.09.
- `EOD` — range: [0, 1]
- Equal Opportunity Difference (EOD) measures the difference in true positive rates across sensitive groups: max_g P(Ŷ=1|Y=1,A=g) - min_g P(Ŷ=1|Y=1,A=g). Reported for completeness alongside DD.
## Input / output format
**Input**: Tabular features and target labels partitioned across multiple FL clients, with sensitive attributes (RACE, SEX) provided for bias analysis.
**Output**: Per-client model predictions, aggregated to compute client-level fairness metrics (DD and EOD) across sensitive attribute groups.
## Scoring recipe
```python
def compute_dd(preds, sensitive):
groups = np.unique(sensitive)
rates = [np.mean(preds[sensitive == g]) for g in groups]
return max(rates) - min(rates)
def compute_eod(preds, labels, sensitive):
groups = np.unique(sensitive)
tprs = [np.mean(preds[sensitive == g] == labels[sensitive == g]) for g in groups]
return max(tprs) - min(tprs)
# Evaluate per client, then aggregate or report per-client values
client_dd = {cid: compute_dd(preds[cid], sensitive[cid]) for cid in clients}
client_eod = {cid: compute_eod(preds[cid], labels[cid], sensitive[cid]) for cid in clients}
```
## Common pitfalls
- Relying solely on global/server-level fairness metrics, which the benchmark demonstrates can completely mask severe client-level unfairness.
- Using a single ML model to quantify bias; the protocol requires consistency across both Logistic Regression and XGBoost to confirm a dataset is truly biased.
- Ignoring the explicit bias threshold (DD > 0.09); datasets are only valid for bias evaluation if both models exceed this threshold for the same attribute/value.
## Evidence (verbatim from paper)
> We measured DD as our primary fairness metric, but also reported EOD results for completeness. We consider a dataset biased toward a specific attribute or attribute value if both models exhibited the maximum DD value for the same attribute/value, and the minimum of these maximum DD values exceeded $0.09$.
## Citation
```bibtex
@misc{heilmann2025feda4fair,
title={FeDa4Fair: Client-Level Federated Datasets for Fairness Evaluation},
author={Heilmann et al. (2025)},
year={2025},
note={arXiv:2506.21095}
}
```
- arXiv: 2506.21095
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!