Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.
Scanned 9/12/2026
Install to Claude Code
npx -y skills add thedixitjain/the-mega-skill-library --skill picking-a-format --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Picking A Format?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/thedixitjain-picking-a-format)More formats (shields.io, HTML) on the badges page.
---
name: picking-a-format
description: "Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair."
category: general-purpose
source_repo: hashgraph-online/awesome-codex-plugins
source_path: "plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/picking-a-format/SKILL.md"
source_url: https://github.com/hashgraph-online/awesome-codex-plugins/blob/HEAD/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/picking-a-format/SKILL.md
---
# Picking a format
Kreuzberg has two orthogonal format knobs. Get them right up front and the
downstream code stays simple.
| Knob | What it controls | Values | Default |
| ------------------- | ------------------------------------------------- | -------------------------------------- | ---------------- |
| `--format` | How the CLI prints the result | `text`, `json` | `text` (`extract`), `json` (`batch`) |
| `--content-format` | How extracted content is rendered inside `result` | `plain`, `markdown`, `djot`, `html` | `plain` |
| `--token-reduction` | Strip whitespace / boilerplate for LLM contexts | `off`, `light`, `moderate`, `aggressive` | `off` |
`--format json` always returns the full `ExtractionResult` (content +
metadata + tables + images). `--format text` prints just `content`.
`--content-format` is what shows up inside that `content` field.
## Decision tree
```text
Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│ --format text --content-format markdown
├── Vector store / RAG indexer
│ --format json --content-format markdown
│ (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│ --format json --content-format plain
│ (cleanest text + structured metadata)
├── Human review / archival
│ --format text --content-format markdown
├── HTML re-rendering / web display
│ --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│ --format json --content-format djot
└── Token-budget-constrained pipeline
--format text --content-format plain
(drops markup; add --token-reduction moderate for further savings)
```
## Examples
Feed a PDF directly into an LLM:
```bash
kreuzberg extract paper.pdf --content-format markdown
```
Index a corpus into a RAG store with tables and headings preserved:
```bash
kreuzberg batch docs/*.pdf --format json --content-format markdown \
| jq -c '.[] | {path: .metadata.path, content: .content, tables: .tables}'
```
Strip a file to bare text for a token-tight summarizer:
```bash
kreuzberg extract long.pdf \
--content-format plain \
--token-reduction moderate
```
Pull metadata only, ignore content:
```bash
kreuzberg extract file.pdf --format json | jq '.metadata'
```
## When in doubt
- **Default to `markdown`** as the content format. It is the best
compromise across LLMs, RAG, and human review, and Kreuzberg has the
most faithful renderer for it.
- Reach for `plain` only when downstream cannot tolerate any markup.
- Reach for `djot` only if you're already in a djot/pandoc pipeline.
- Reach for `html` only when re-rendering for the web.
## Token-reduction (orthogonal)
`--token-reduction` collapses whitespace, strips repeated headers/footers,
and trims boilerplate. It composes with any `--content-format`:
- `off` (default), `light`, `moderate`, `aggressive`, `maximum`.
Use `moderate` as a safe starting point for LLM context windows. `maximum`
is lossy — verify before relying on it.
See `references/cli-reference.md` for the full flag set and
`references/configuration.md` for the equivalent `output_format` and
`token_reduction` keys in `kreuzberg.toml`.
---
**Source:** [`hashgraph-online/awesome-codex-plugins`](https://github.com/hashgraph-online/awesome-codex-plugins) → `plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/picking-a-format/SKILL.md`
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!