llamafile by Mozilla bundles open-source LLMs into a single portable executable that runs locally on macOS, Windows, Linux, and BSD with zero installation. It combines llama.cpp inference with Cosmopolitan Libc to collapse model weights, server, and runtime into one file.
Scanned 6/8/2026
Install via CLI
openskills install agentskillexchange/skills---
name: "llamafile Single-File LLM Distribution and Runner by Mozilla"
slug: "llamafile-single-file-llm-runner-mozilla"
description: "llamafile by Mozilla bundles open-source LLMs into a single portable executable that runs locally on macOS, Windows, Linux, and BSD with zero installation. It combines llama.cpp inference with Cosmopolitan Libc to collapse model weights, server, and runtime into one file."
github_stars: 24134
verification: "security_reviewed"
source: "https://github.com/mozilla-ai/llamafile"
category: "Developer Tools"
framework: "Multi-Framework"
tool_ecosystem:
github_repo: "mozilla-ai/llamafile"
github_stars: 24134
---
# llamafile Single-File LLM Distribution and Runner by Mozilla
llamafile by Mozilla bundles open-source LLMs into a single portable executable that runs locally on macOS, Windows, Linux, and BSD with zero installation. It combines llama.cpp inference with Cosmopolitan Libc to collapse model weights, server, and runtime into one file.
## Installation
Basic usage or getting-started notes:
- **llamafile lets you distribute and run LLMs with a single file.**
- show which version of the server they have been bundled with ([0.9.* example](https://huggingface.co/mozilla-ai/llava-v1.5-7b-llamafile), [0.10.* example](https://huggingface.co/mozilla-ai/llamafile_0.10)), so you wil...
- Download and run your first llamafile in minutes:
- Source: https://github.com/mozilla-ai/llamafile
- Extracted from upstream docs: https://raw.githubusercontent.com/mozilla-ai/llamafile/HEAD/README.md
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/llamafile-single-file-llm-runner-mozilla/)
No comments yet. Be the first to comment!
Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, wenyan-lite, wenyan-full, wenyan-ultra. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.