Designing around token limits, memory, and conversation persistence.
Scanned 5/28/2026
Install via CLI
openskills install Owl-Listener/ai-design-skills---
name: context-window-design
description: Designing around token limits, memory, and conversation persistence.
---
# Context Window Design
Every AI model has a finite context window. Designing within this constraint — and designing the user experience around it — is a core skill for AI product design.
## The Context Window as a Design Material
The context window is not just a technical limitation. It's a design material:
- **What goes in**: System prompts, conversation history, retrieved documents, tool results, user preferences
- **What gets dropped**: Older messages, less relevant context, verbose instructions
- **What the user sees**: The conversation as presented may differ from what the model actually processes
Designers must understand context window allocation to design reliable experiences.
## Memory and Persistence
Users expect AI to remember. Design for different memory horizons:
- **Within-conversation memory**: What was said earlier in this chat. Usually handled by the context window itself.
- **Cross-conversation memory**: Preferences, past decisions, ongoing projects. Requires explicit memory systems.
- **Shared memory**: Context shared across multiple users or agents. Requires careful privacy design.
## Strategies for Limited Context
- **Summarisation**: Compress earlier conversation into summaries to free up tokens
- **Retrieval-augmented generation**: Pull in relevant context on demand rather than keeping everything loaded
- **Priority ordering**: Put the most important context closest to the prompt (recency bias in attention)
- **User-controlled context**: Let users pin, remove, or prioritise what the AI remembers
- **Graceful degradation**: When context is lost, acknowledge it rather than hallucinating continuity
## Design Artefacts
- Context budget allocations (how many tokens for system prompt, history, retrieval, etc.)
- Memory architecture diagrams showing what persists and what's ephemeral
- Context overflow UX flows (what happens when the window fills up)
- User-facing memory controls specification
No comments yet. Be the first to comment!
Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, wenyan-lite, wenyan-full, wenyan-ultra. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.
Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...
**Complete production-ready guide for Google Gemini embeddings API** This skill provides comprehensive coverage of the `gemini-embedding-001` model for generating text embeddings, including SDK usage, REST API patterns, batch processing, RAG integration with Cloudflare Vectorize, and advanced use cases like semantic search and document clustering. ---
Interview, source-challenge, verify, save, and ADR-gate fuzzy coding requests into Codex-ready implementation specs. Use when a feature, bugfix, refactor, migration, repo-wide change, or architecture task needs user-verified requirements, source-backed decisions, durable architecture decisions, acceptance criteria, validation commands, rollout notes, saved spec/ADR files, and a Codex execution prompt. Do not use when already fully specified or when the user wants direct implementation now.
Use when a repo needs CodeGraph plus ast-grep for Codex MCP setup, exploration, impact analysis, structural search, or safe refactor planning.