This benchmark evaluates a model's ability to generate culturally accurate and coherent explanations for Vietnamese visual question answering. It probes both linguistic fluency and the model's capacity to ground visual evidence in domain-specific cultural knowledge through structured, stepwise reasoning. Use when the user wants to benchmark on Vietnamese VQA dataset, or asks about evaluating this task. Reports Cultural Accuracy.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill vietmeagent-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Vietmeagent Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-vietmeagent-eval)More formats (shields.io, HTML) on the badges page.
---
name: vietmeagent-eval
description: This benchmark evaluates a model's ability to generate culturally accurate and coherent explanations for Vietnamese visual question answering. It probes both linguistic fluency and the model's capacity to ground visual evidence in domain-specific cultural knowledge through structured, stepwise reasoning. Use when the user wants to benchmark on Vietnamese VQA dataset, or asks about evaluating this task. Reports Cultural Accuracy.
metadata:
skill_kind: dataset_eval
source_arxiv: 2511.09058
bibtex_key: nguyen2025vietmeagent
confidence: high
---
# vietmeagent-eval
> VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering — Nguyen et al. (2025) (arXiv:2511.09058, 2025)
## What this evaluates
This benchmark evaluates a model's ability to generate culturally accurate and coherent explanations for Vietnamese visual question answering. It probes both linguistic fluency and the model's capacity to ground visual evidence in domain-specific cultural knowledge through structured, stepwise reasoning.
## Datasets
- **Vietnamese VQA dataset** — total ?; splits: (unstated)
## Metrics
- `BLEU-4` — range: [0, 1]
- Standard 4-gram precision metric with a brevity penalty to penalize overly short generations. It measures surface-level lexical overlap between the predicted and reference rationales.
- `METEOR` — range: [0, 1]
- Harmonic mean of unigram precision and recall, incorporating stemming and synonymy matching to better capture semantic equivalence than exact n-gram overlap.
- `ROUGE-L` — range: [0, 1]
- F-score based on the longest common subsequence (LCS) between the predicted and reference texts, capturing sentence-level fluency and structural alignment.
- `Cultural Accuracy` **(primary)** — range: [0, 1]
- Measures correctness against expert-validated cultural facts. It is computed as the proportion of generated explanations that correctly align with verified cultural knowledge entries for the depicted visual content.
- `Explanation Quality` — range: [0, 1]
- Assesses the coherence and grounding of multimodal rationales. It evaluates whether the generated stepwise reasoning logically connects visual evidence to cultural semantics without hallucination or disjointed claims.
## Input / output format
**Input**: An image and a Vietnamese question requiring visual identification and cultural reasoning.
**Output**: A textual, stepwise explanation linking visual evidence to cultural semantics, followed by the final answer.
## Scoring recipe
```python
def compute_metrics(predictions, gold):
bleu4 = compute_bleu(predictions, gold, n=4)
meteor = compute_meteor(predictions, gold)
rouge_l = compute_rouge_l(predictions, gold)
# Domain-specific metrics rely on expert validation
cultural_acc = sum(1 for p in predictions if is_culturally_valid(p)) / len(predictions)
expl_quality = compute_coherence_grounding(predictions, gold)
return {
'BLEU-4': bleu4,
'METEOR': meteor,
'ROUGE-L': rouge_l,
'Cultural Accuracy': cultural_acc,
'Explanation Quality': expl_quality
}
```
## Common pitfalls
- Cultural Accuracy depends entirely on expert-validated facts; models without access to the curated knowledge base will systematically underperform regardless of visual capability.
- Explanation Quality evaluates coherence and grounding, which can be highly subjective without a strict annotation rubric or inter-annotator agreement reporting.
- BLEU-4 may overestimate generation quality by rewarding fluent but culturally incorrect or hallucinated rationales.
## Evidence (verbatim from paper)
> We evaluate language quality with BLEU-4 *[papineni2002bleu]*, complement it with METEOR *[banerjee2005meteor]* and ROUGE-L *[lin2004rouge]*, and introduce two domain-oriented criteria: Cultural Accuracy, which measures correctness against expert-validated cultural facts, and Explanation Quality, which assesses the coherence and grounding of multimodal rationales.
## Citation
```bibtex
@misc{nguyen2025vietmeagent,
title={VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering},
author={Nguyen et al. (2025)},
year={2025},
note={arXiv:2511.09058}
}
```
- arXiv: 2511.09058
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!