**arXiv ID:** 2410.20297 **Authors:** Daniel C. Ruiz, John Sell **Published:** 2024-10-27T00:39:24Z **Abstract:** In recent years, the widespread adoption of Large Language Models (LLMs) has sparked interest in their potential for application within the military domain. However, the current generation of LLMs demonstrate sub-optimal performance on Army use cases, due to the prevalence of domain-specific vocabulary and jargon. In order to fully leverage LLMs in-domain, many organizations have ...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill finetuning-and-evaluating-opensource-large-language-models-for-the-army-domain --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Finetuning And Evaluating Opensource Large Language Models For The Army Domain?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-finetuning-and-evaluating-opensource-large-languag)More formats (shields.io, HTML) on the badges page.
# Fine-Tuning and Evaluating Open-Source Large Language Models for the Army Domain
**arXiv ID:** 2410.20297
**Authors:** Daniel C. Ruiz, John Sell
**Published:** 2024-10-27T00:39:24Z
**Abstract:**
In recent years, the widespread adoption of Large Language Models (LLMs) has sparked interest in their potential for application within the military domain. However, the current generation of LLMs demonstrate sub-optimal performance on Army use cases, due to the prevalence of domain-specific vocabulary and jargon. In order to fully leverage LLMs in-domain, many organizations have turned to fine-tuning to circumvent the prohibitive costs involved in training new LLMs from scratch. In light of this trend, we explore the viability of adapting open-source LLMs for usage in the Army domain in order to address their existing lack of domain-specificity. Our investigations have resulted in the creation of three distinct generations of TRACLM, a family of LLMs fine-tuned by The Research and Analysis Center (TRAC), Army Futures Command (AFC). Through continuous refinement of our training pipeline, each successive iteration of TRACLM displayed improved capabilities when applied to Army tasks and use cases. Furthermore, throughout our fine-tuning experiments, we recognized the need for an evaluation framework that objectively quantifies the Army domain-specific knowledge of LLMs. To address this, we developed MilBench, an extensible software framework that efficiently evaluates the Army knowledge of a given LLM using tasks derived from doctrine and assessments. We share preliminary results, models, methods, and recommendations on the creation of TRACLM and MilBench. Our work significantly informs the development of LLM technology across the DoD and augments senior leader decisions with respect to artificial intelligence integration.
## Skill Description
This skill is generated from the arXiv paper: Fine-Tuning and Evaluating Open-Source Large Language Models for the Army Domain (2410.20297).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2410.20297](http://arxiv.org/abs/2410.20297v1)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!