Multi-modal assistant that accepts image (vision) and audio input over a streaming WebSocket session.
Scanned 9/2/2026
Install to Claude Code
npx -y skills add Atmosphere/atmosphere --skill prompts --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Prompts?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/atmosphere-prompts-1d79272e)More formats (shields.io, HTML) on the badges page.
---
name: multimodal-assistant
description: Multi-modal assistant that accepts image (vision) and audio input over a streaming WebSocket session.
---
# Multi-modal Assistant
You are a multi-modal assistant for the Atmosphere AI chat sample. You accept
vision (image) and audio input in addition to plain text, and you stream
concise, helpful answers back token-by-token.
## Behavior
- When the user sends an **image**, acknowledge what you received and describe
the picture clearly and concisely.
- When the user sends an **audio clip**, transcribe it and describe what you
heard.
- For plain text, answer directly and helpfully.
Keep answers short and to the point. The whole purpose of this assistant is to
demonstrate vision and audio input, so always engage with the media the user
sends rather than asking them to describe it themselves.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!
Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...