Evaluates large language models' ability to generate correct, compilable COBOL code from natural language specifications, and to translate bidirectionally between COBOL and Java. It probes functional correctness, compilation reliability, and practical utility for legacy system modernization. Use when the user wants to benchmark on COBOLEval, COBOLCodeBench, COBOL-JavaTrans, or asks about evaluating this task. Reports Pass@1.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill cobol-codegen-translation-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Cobol Codegen Translation Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-cobol-codegen-translation-eval)More formats (shields.io, HTML) on the badges page.
---
name: cobol-codegen-translation-eval
description: Evaluates large language models' ability to generate correct, compilable COBOL code from natural language specifications, and to translate bidirectionally between COBOL and Java. It probes functional correctness, compilation reliability, and practical utility for legacy system modernization. Use when the user wants to benchmark on COBOLEval, COBOLCodeBench, COBOL-JavaTrans, or asks about evaluating this task. Reports Pass@1.
metadata:
skill_kind: dataset_eval
source_arxiv: 2604.03986
bibtex_key: dau2026cobolcoder
confidence: high
---
# cobol-codegen-translation-eval
> COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation — Dau et al. (2026) (arXiv:2604.03986, 2026)
## What this evaluates
Evaluates large language models' ability to generate correct, compilable COBOL code from natural language specifications, and to translate bidirectionally between COBOL and Java. It probes functional correctness, compilation reliability, and practical utility for legacy system modernization.
## Datasets
- **COBOLEval** — total 146; splits: test (-1)
- **COBOLCodeBench** — total 46; splits: test (-1)
- **COBOL-JavaTrans** — total 143; splits: test (-1)
## Metrics
- `Compilation Success Rate (CSR)` — range: [0, 1]
- Proportion of generated solutions that successfully compile using the language-specific compiler (GnuCOBOL 2.0.0 for COBOL, javac 17.0.18 for Java).
- `Pass@1` **(primary)** — range: [0, 1]
- Percentage of tasks for which the model's first generated solution passes all provided unit tests.
## Input / output format
**Input**: For code generation: a COBOL function signature and a natural language specification. For translation: source code (COBOL or Java) paired with the corresponding task specification or reference implementation.
**Output**: A single code block containing the generated COBOL program or the translated code (Java for COBOL-to-Java, COBOL for Java-to-COBOL).
## Scoring recipe
```python
def compute_csr(predictions, compilers):
compiled = sum(1 for code in predictions if compile(code, compilers[code.lang]) == 0)
return compiled / len(predictions)
def compute_pass_at_1(predictions, test_cases):
passed = sum(1 for code, tests in zip(predictions, test_cases) if run_tests(code, tests))
return passed / len(predictions)
```
## Common pitfalls
- Evaluations are strictly zero-shot with temperature=0.0 and n=1; results must be averaged over three independent runs to account for LLM randomness.
- Compilation success is highly sensitive to compiler versions; the paper mandates GnuCOBOL 2.0.0 for COBOL and javac 17.0.18 for Java.
- Not all HumanEval tasks can be implemented in COBOL; COBOL-JavaTrans only includes 143 of the original 164 tasks that compile and pass tests.
## Evidence (verbatim from paper)
> We assess the correctness of both generated and translated code using two metrics, including the Compilation Success Rate (CSR) and Pass@1. CSR: It measures the proportion of generated solutions that compile successfully. We use the corresponding compiler for different languages (e.g., GnuCOBOL—version 2.0.0 for COBOL and javac—version 17.0.18 for Java). Pass@1: Pass@1 evaluates functional correctness by measuring the percentage of tasks for which the model’s first generated solution passes all test cases.
## Citation
```bibtex
@misc{dau2026cobolcoder,
title={COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation},
author={Dau et al. (2026)},
year={2026},
note={arXiv:2604.03986}
}
```
- arXiv: 2604.03986
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!