Compute the MulticlassF1Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassF1Score, or asks how to score with MulticlassF1Score.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill multiclassf1score --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Multiclassf1score?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-multiclassf1score)More formats (shields.io, HTML) on the badges page.
---
name: multiclassf1score
description: Compute the MulticlassF1Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassF1Score, or asks how to score with MulticlassF1Score.
metadata:
skill_kind: metric
source_lib: torchmetrics
import_path: torchmetrics.classification.MulticlassF1Score
source: library_introspection
---
# multiclassf1score
> Metric `MulticlassF1Score` from `torchmetrics` (torchmetrics.classification.MulticlassF1Score)
## When to invoke this skill
The user has predictions + ground truth and asks to evaluate with MulticlassF1Score, or
mentions `torchmetrics.classification.MulticlassF1Score` directly, or wants the standard torchmetrics implementation.
## Reference signature
```python
from torchmetrics.classification import MulticlassF1Score
# MulticlassF1Score(num_classes: int, top_k: int = 1, average: Optional[Literal['micro', 'macro', 'weighted', 'none']] = 'macro', multidim_average: Literal['global', 'samplewise'] = 'global', ignore_index: Optional[int] = None, validate_args: bool = True, zero_division: float = 0, **kwargs: Any) -> None
```
## Library docstring
```
Compute F-1 score for multiclass tasks.
.. math::
F_{1} = 2\frac{\text{precision} * \text{recall}}{(\text{precision}) + \text{recall}}
The metric is only proper defined when :math:`\text{TP} + \text{FP} \neq 0 \wedge \text{TP} + \text{FN} \neq 0`
where :math:`\text{TP}`, :math:`\text{FP}` and :math:`\text{FN}` represent the number of true positives, false
positives and false negatives respectively. If this case is encountered for any class, the metric for that class
will be set to `zero_division` (0 or 1, default is 0) and the overall metric may therefore be affected in turn.
As input to ``forward`` and ``update`` the metric accepts the following input:
- ``preds`` (:class:`~torch.Tensor`): An int tensor of shape ``(N, ...)`` or float tensor of shape ``(N, C, ..)``.
If preds is a floating point we apply ``torch.argmax`` along the ``C`` dimension to automatically convert
probabilities/logits into an int tensor.
- ``target`` (:class:`~torch.Tensor`): An int tensor of shape ``(N, ...)``
As output to ``forward`` and ``compute`` the metric returns the following output:
- ``mcf1s`` (:class:`~torch.Tensor`): A tensor whose returned shape depends on the ``average`` and
``multidim_average`` arguments:
- If ``multidim_average`` is set to ``global``:
- If ``average='micro'/'macro'/'weighted'``, the output will be a scalar tensor
- If ``average=None/'none'``, the shape will be ``(C,)``
- If ``multidim_average`` is set to ``samplewise``:
- If ``average='micro'/'macro'/'weighted'``, the shape will be ``(N,)``
- If ``average=None/'none'``, the shape will be ``(N, C)``
If ``multidim_average`` is set to ``samplewise`` we expect at least one additional dimension ``...`` to be present,
which the reduction will then be applied over instead of the sample dimension ``N``.
Args:
preds: Tensor with predictions
target: Tensor with true labels
num_classes: Integer specifying the number of classes
average:
Defines the reduction that is applied over labels. Should be one of the following:
- ``micro``: Sum statistics over all labels
- ``macro``: Calculate statistics for each label and average the
```
## Quick recipe
```python
import torchmetrics.classification as _m
score = _m.MulticlassF1Score(y_true, y_pred)
```
## Don'ts
- Don't reimplement when the library version handles edge cases (NaN, ties, empty inputs) better than a hand-rolled formula.
- Always check the library version's argument order — sklearn is `(y_true, y_pred)` while torchmetrics is `(preds, target)`.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!