Skip to content
Back to skills

Cassandra Sstable Format Parsing

ASecurity

Guide parsing of Cassandra 5.0+ SSTable components (Data.db, Index.db, Statistics.db, Summary.db, TOC) with compression support (LZ4, Snappy, Deflate). Use when working with SSTable files, binary format parsing, hex dumps, compression issues, offset calculations, BTI index, partition layout, or debugging parsing errors.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 27, 2026
developmentrustjavabashdebugging

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned September 27, 2026

npx -y skills add David-Li0406/meta-skill-evloving --skill cassandra-sstable-format-parsing --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Cassandra Sstable Format Parsing?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Cassandra Sstable Format Parsing
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/david-li0406-cassandra-sstable-format-parsing/badge)](https://www.skillsdirectory.com/skills/david-li0406-cassandra-sstable-format-parsing)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: Cassandra SSTable Format Parsing
description: Guide parsing of Cassandra 5.0+ SSTable components (Data.db, Index.db, Statistics.db, Summary.db, TOC) with compression support (LZ4, Snappy, Deflate). Use when working with SSTable files, binary format parsing, hex dumps, compression issues, offset calculations, BTI index, partition layout, or debugging parsing errors.
allowed-tools: Read, Grep, Glob
---

# Cassandra SSTable Format Parsing

This skill helps with parsing and understanding Cassandra 5.0+ SSTable file formats.

## When to Use This Skill

- Parsing Data.db, Index.db, Statistics.db files
- Debugging binary format mismatches
- Analyzing hex dumps of SSTable data
- Working with compression (LZ4, Snappy, Deflate)
- Investigating offset calculation errors
- Understanding BTI (Big Table Index) format
- Validating partition boundaries

## Key SSTable Components

### Data.db
Contains the actual row data with:
- Partition headers
- Row data (clustering + cells)
- Compression blocks
- Checksums

### Index.db
Contains partition index entries with:
- BTI (Big Table Index) format in Cassandra 5.0+
- Partition key → file offset mapping
- Promoted index entries

### Statistics.db
Contains serialization metadata:
- Encoding stats (min/max timestamps, TTLs)
- Column definitions
- Schema information
- Compression parameters

### Summary.db
Contains sampling of index entries for faster lookups

## Format References

**Primary Source of Truth**: `docs/sstables-definitive-guide/`

Key chapters:
- **Ch.5**: Data.db Format - Row layout, flags, V5CompressedLegacy
- **Ch.6**: Index.db and Summary.db - Partition lookups
- **Ch.9**: CompressionInfo.db - Compression metadata, chunking
- **Ch.17**: BTI Formats - Trie-based indexes
- **Appendix B**: Encoding Cheat Sheet - VInt, cell flags
- **Appendix F**: Known Limitations - What doesn't work yet

## Common Debugging Techniques

### Hex Dump Analysis
When debugging parsing errors:

1. **Extract hex at specific offset**:
   ```bash
   hexdump -C Data.db -s <offset> -n 64
   ```

2. **Compare with expected format**:
   - Check magic bytes (if applicable)
   - Verify VInt encoding
   - Validate flag bytes

3. **Look for patterns**:
   - Repeated byte sequences may indicate arrays/collections
   - All zeros may indicate padding
   - Non-zero high bytes suggest multi-byte integers

### Offset Validation
Track byte consumption at each parsing stage:
- Clustering prefix (may be 0 bytes)
- Row sizes (2 VInts)
- Liveness info (conditional)
- Deletion info (conditional)
- Column bitmap (conditional)
- Cell data

### Zero-Copy Considerations
When implementing parsers:
- Use `Bytes` crate for buffer sharing
- Avoid copying large cell values
- Keep references to original buffer
- Use byte slices not owned Vecs

## Integration with Rust Code

Current implementation in `cqlite-core/src/storage/sstable/reader/parsing/`:
- `v5_compressed_legacy.rs` - Main V5 format parser (1997 lines)
- Uses zero-copy patterns with `Bytes`
- Handles compression transparently

## PRD Alignment

**Supports Milestone M1** (Core Reading Library):
- 100% Cassandra 5 SSTable format support
- All compression formats (LZ4, Snappy, Deflate)
- Zero-copy deserialization
- Memory target: <128MB for large files

## Quick Reference

### Flag Bytes (Row)
- `0x01`: HAS_IS_MARKER
- `0x02`: HAS_ALL_COLUMNS (inverted - 0x20 means all present)
- `0x04`: HAS_TIMESTAMP
- `0x08`: HAS_TTL
- `0x10`: HAS_DELETION
- `0x20`: HAS_ALL_COLUMNS (flag set = all columns present)
- `0x40`: IS_STATIC
- `0x80`: EXTENSION_FLAG

### VInt Encoding
Variable-length integer encoding:
- First byte indicates length
- Subsequent bytes contain value
- Used for row sizes, timestamps, offsets

## Next Steps

When parser encounters issues:
1. Log byte offsets at each stage
2. Compare against Java source (UnfilteredSerializer.java)
3. Validate against sstabledump output
4. Check compression block boundaries
5. Verify delta encoding calculations

Files in this skill

  • SKILL.md3.9 KB
  • cassandra5-format-reference.md5.2 KB
  • compression-formats.md4.7 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…