Evaluates the ability to predict click-through, conversion, and effective conversion rates in a large-scale e-commerce recommender system. It specifically probes how well models handle cascade delayed feedback, sample selection bias, and data sparsity when predicting user purchase and refund behaviors. Use when the user wants to benchmark on Alibaba Production Dataset, or asks about evaluating this task. Reports AUC.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill ecvr-prediction-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Ecvr Prediction Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-ecvr-prediction-eval)More formats (shields.io, HTML) on the badges page.
---
name: ecvr-prediction-eval
description: Evaluates the ability to predict click-through, conversion, and effective conversion rates in a large-scale e-commerce recommender system. It specifically probes how well models handle cascade delayed feedback, sample selection bias, and data sparsity when predicting user purchase and refund behaviors. Use when the user wants to benchmark on Alibaba Production Dataset, or asks about evaluating this task. Reports AUC.
metadata:
skill_kind: dataset_eval
source_arxiv: 2308.04768
bibtex_key: zhao2023entire
confidence: high
---
# ecvr-prediction-eval
> Entire Space Cascade Delayed Feedback Modeling for Effective Conversion Rate Prediction — Zhao et al. (2023) (arXiv:2308.04768, 2023)
## What this evaluates
Evaluates the ability to predict click-through, conversion, and effective conversion rates in a large-scale e-commerce recommender system. It specifically probes how well models handle cascade delayed feedback, sample selection bias, and data sparsity when predicting user purchase and refund behaviors.
## Datasets
- **Alibaba Production Dataset** — total 560000000; splits: train (-1), test (-1)
## Metrics
- `AUC` **(primary)** — range: [0, 1]
- Area under the Receiver Operating Characteristic curve. Indicates the probability that a positive sample is ranked higher than a negative one.
- `PR-AUC` — range: [0, 1]
- Area under the Precision-Recall curve. More sensitive than AUC when negative samples significantly outnumber positive ones.
- `RI-AUC` — range: percent
- Relative Improvement in AUC: (AUC_model - AUC_base) / (AUC_oracle - AUC_base). Approaches 100% as performance improves relative to the base and oracle models.
- `RI-PR-AUC` — range: percent
- Relative Improvement in PR-AUC: (PR-AUC_model - PR-AUC_base) / (PR-AUC_oracle - PR-AUC_base). Approaches 100% as performance improves relative to the base and oracle models.
## Input / output format
**Input**: Click traffic log instances containing 209 features (user and item attributes). Each instance represents a user-item interaction with a timestamp.
**Output**: Probability scores: p_cvr (conversion rate), p_rfr (refund rate), and p_ecvr (effective conversion rate).
## Scoring recipe
```python
def compute_ri_auc(model_auc, base_auc, oracle_auc):
return (model_auc - base_auc) / (oracle_auc - base_auc)
def compute_ri_pr_auc(model_pr_auc, base_pr_auc, oracle_pr_auc):
return (model_pr_auc - base_pr_auc) / (oracle_pr_auc - base_pr_auc)
# Standard metrics
auc = roc_auc_score(y_true, y_pred)
pr_auc = average_precision_score(y_true, y_pred)
# Relative improvements
ri_auc = compute_ri_auc(auc, base_auc, oracle_auc)
ri_pr_auc = compute_ri_pr_auc(pr_auc, base_pr_auc, oracle_pr_auc)
```
## Common pitfalls
- The attribution window for labels is strictly 3 days for conversion/refund and 5 days for ECVR; ignoring this leads to label leakage or stale labels.
- The test set consists of only 1 day of data, partitioned into 10 folds for averaging; results are not cross-dated across different time periods.
- RI-AUC and RI-PR-AUC are normalized relative to a specific 'Base' and 'Oracle' model, not absolute improvements, so they cannot be compared across different experimental setups.
## Evidence (verbatim from paper)
> Following previous works [5, 24], we adopt two widely used metrics for evaluating offline experimental results. The first metric is the area under the ROC curve (AUC) which indicates the probability that a positive sample is ranked higher than a negative one (0.1% improvement on AUC in industrial datasets is deemed as significant [15, 35]). The second metric is the area under the precision-recall curve (PR-AUC), which is more sensitive than AUC in scenarios where negative samples significantly outnumber positive ones. Besides, We also report the relative improvement of methods over the base model in each task on AUC and PR-AUC [5]. Taking the relative improvement of ECAD on AUC (denoted as RI-AUC) in CVR prediction as the example, the RI-AUC_ECDA can be computed as follows: RI-AUC_ECDA = (AUC_ECDA - AUC_CVR-Base) / (AUC_ESMM-Oracle - AUC_CVR-Base)
## Citation
```bibtex
@misc{zhao2023entire,
title={Entire Space Cascade Delayed Feedback Modeling for Effective Conversion Rate Prediction},
author={Zhao et al. (2023)},
year={2023},
note={arXiv:2308.04768}
}
```
- arXiv: 2308.04768
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!