Evaluates zero-shot language-queried audio source separation on diverse environmental sounds using natural language captions. The benchmark tests isolation of a target sound from a concatenated background mixture. Use when the user wants to benchmark on Clotho v2, or asks about evaluating this task. Reports SDRi.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill clotho-sep-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Clotho Sep Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-clotho-sep-eval)More formats (shields.io, HTML) on the badges page.
---
name: clotho-sep-eval
description: Evaluates zero-shot language-queried audio source separation on diverse environmental sounds using natural language captions. The benchmark tests isolation of a target sound from a concatenated background mixture. Use when the user wants to benchmark on Clotho v2, or asks about evaluating this task. Reports SDRi.
metadata:
skill_kind: dataset_eval
source_arxiv: 2308.05037
bibtex_key: liu2023separate
confidence: high
---
# clotho-sep-eval
> Separate Anything You Describe — Liu et al. (2023) (arXiv:2308.05037, 2023)
## What this evaluates
Evaluates zero-shot language-queried audio source separation on diverse environmental sounds using natural language captions. The benchmark tests isolation of a target sound from a concatenated background mixture.
## Datasets
- **Clotho v2** — total 5225; splits: test (5225)
## Metrics
- `SDRi` **(primary)** — range: dB
- Signal-to-distortion ratio improvement, calculated as the difference between the SDR of the separated output and the original mixture.
- `SI-SDR` — range: dB
- Scale-invariant signal-to-distortion ratio, evaluates separation quality independent of amplitude scaling between prediction and target.
## Input / output format
**Input**: A 15-30 second audio mixture (SNR 0 dB) formed by concatenating and truncating two clips, paired with one of five human-annotated captions.
**Output**: Separated audio waveform corresponding to the target sound described by the caption.
## Scoring recipe
```python
def score(pred, gold, mixture):
sdr_out = compute_sdr(pred, gold)
sdr_mix = compute_sdr(mixture, gold)
sdri = sdr_out - sdr_mix
si_sdr = compute_si_sdr(pred, gold)
return sdri, si_sdr
```
## Common pitfalls
- Mixtures are created by concatenating two clips and truncating to the target length, not simple additive mixing.
- Each target clip is mixed with five different background clips, yielding 5 variations per original evaluation clip.
## Evidence (verbatim from paper)
> The Clotho v2 evaluation set includes 1045 audio clips, each provided with five human-annotated captions. ... This procedure culminates in a total of 5225 mixtures for evaluation. ... We utilize signal-to-distortion ratio improvement (SDRi) and scale-invariant SDR (SI-SDR) to evaluate the performance of sound separation systems.
## Citation
```bibtex
@misc{liu2023separate,
title={Separate Anything You Describe},
author={Liu et al. (2023)},
year={2023},
note={arXiv:2308.05037}
}
```
- arXiv: 2308.05037
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!